How to choose an AI development company: what to ask, what to verify, what to walk away from
Every agency now sells AI. Few can tell you what happens after the demo. Here is how to separate teams that ship AI into production from teams that ship slideware.
Working on something like this?
Get an estimate- A demo proves a model can do something once. Production proves it can do it reliably, safely, and cheaply.
- Ask how they will measure quality before you ask what they will build.
- The unglamorous work, evaluation, data access, monitoring, and cost control, is what you are really buying.
- Start with one workflow and a success number. Walk away from anyone selling a platform on day one.
Two years ago, finding a team that could build with large language models was the hard part. Now nearly every software company lists AI among its services, and the hard part is telling them apart. The difference rarely shows in a demonstration, because a demonstration is exactly the thing everyone can now produce in a week.
What separates a capable partner is what happens on the days after: when the model is confidently wrong, when the cost per request climbs, when the data you need sits behind a system nobody has documented. This guide covers what to look for, what to ask, and what should make you stop the conversation.
First, decide what you are actually buying
“AI development” covers very different work, and the right partner differs for each.
Many AI projects are, at heart, ordinary software projects with one unusual component. If the integration, security, and data work is most of the effort, a strong engineering partner with real AI experience usually serves you better than a specialist shop that treats everything else as an afterthought.
What to evaluate beyond the demo
How they measure quality
Ask what a good answer looks like and how they will know the system produces one. A serious team will propose an evaluation set built from your real cases, a target accuracy, and a way to catch regressions when the model or prompt changes. A team that cannot describe this will judge quality by feel, and so will your users.
How they handle being wrong
Models fail in plausible ways. Ask what happens when the system is unsure, when it is wrong, and when it is asked something outside its remit. Good answers involve confidence thresholds, human review for high-stakes steps, citations back to source material, and clear fallbacks.
How they treat your data
Where does it go, who can see it, is it used for training, and how is access controlled? Look for permission-aware retrieval, so the assistant never shows someone a document they could not open directly. Ask where the model is hosted and what the contract says.
How they control cost
A prototype that costs pennies can cost real money at volume. Ask about caching, model selection by task, token budgets, and how they will report spend. Cost per completed task is the number that decides whether the project survives its first budget review.
How they take it to production
Monitoring, logging, versioning of prompts and models, rollback, and an owner after launch. This is the gap that kills most pilots, and we cover it in moving an AI pilot to production.
For a worked example of that last step, see how an enterprise AI assistant moved from pilot to production.
Red flags in a proposal
- No success metric. “Improve efficiency with AI” is a hope, not a deliverable.
- A platform before a use case. If the first thing on the table is a multi-year roadmap, they are selling their capacity, not your outcome.
- Accuracy promised in advance. Nobody can guarantee a figure before seeing your data. Credible teams commit to measuring it.
- Vague answers on data and security. If the data flow is hand-waved in the sales call, it will be worse in the contract.
- Only demos, no production references. Ask for something running in front of real users, and for what went wrong with it.
- One model, always. The best choice varies by task, and a team that cannot say when not to use a large model has not run enough of them.
- Nothing about people. Adoption, review workflows, and change management decide whether anyone uses it.
Questions worth asking in the first call
- What would you measure to decide this worked, and how will we build the test set?
- Show us something in production. What failed, and what did you change?
- How do you keep the assistant from exposing data a user should not see?
- What will this cost per completed task at ten times today's volume?
- What happens when the underlying model is deprecated or changes behaviour?
- Who owns the prompts, evaluation sets, and integrations when we finish?
- Which part of this would you advise us not to use AI for?
Start small, with a number attached
The lowest-risk first engagement is one workflow, one measurable target, and a short timeline. Choose something repetitive and well understood, where a wrong answer is caught cheaply, and where you already know what it costs today. Define the target before anyone writes code, for example the share of cases handled without escalation, or the time taken per document, and agree what happens if it is not reached.
Resist the pull to start with the most ambitious idea. A narrow success that reaches production teaches your organisation more than a broad prototype that stalls. The same discipline applies as in any build, which is why scoping a first release tightly matters here as much as anywhere.
Where governance fits
Larger organisations will need policy on acceptable use, data handling, human oversight, and vendor risk before anything touches customers. It is far cheaper to settle this alongside the first project than to discover it afterwards. If you have not yet decided where AI belongs in your business, an AI readiness and governance engagement is a sensible first step, ahead of any build.
Build, buy, or partner
Not every AI need requires a custom build. Off-the-shelf tools cover many common cases faster and cheaper, and the honest advice is sometimes to buy. Our framework for build versus buy applies directly: build when the capability is a differentiator or must integrate deeply with your own data and systems, and buy when it is commodity.
If you want a straight read on whether your idea is a good AI candidate, what it would take, and what we would not recommend, see our artificial intelligence services or get in touch.
Working through this on a real project?
Tell us what you are building. You will get a scoped estimate and an architecture you own, not a capability deck.

