New: CentriCall AI voice agents that answer, qualify, and book around the clock
Choosing a Partner5 min read

How to choose an AI development company: what to ask, what to verify, what to walk away from

Every agency now sells AI. Few can tell you what happens after the demo. Here is how to separate teams that ship AI into production from teams that ship slideware.

The short version
  • A demo proves a model can do something once. Production proves it can do it reliably, safely, and cheaply.
  • Ask how they will measure quality before you ask what they will build.
  • The unglamorous work, evaluation, data access, monitoring, and cost control, is what you are really buying.
  • Start with one workflow and a success number. Walk away from anyone selling a platform on day one.

Two years ago, finding a team that could build with large language models was the hard part. Now nearly every software company lists AI among its services, and the hard part is telling them apart. The difference rarely shows in a demonstration, because a demonstration is exactly the thing everyone can now produce in a week.

What separates a capable partner is what happens on the days after: when the model is confidently wrong, when the cost per request climbs, when the data you need sits behind a system nobody has documented. This guide covers what to look for, what to ask, and what should make you stop the conversation.

First, decide what you are actually buying

“AI development” covers very different work, and the right partner differs for each.

You needWhat it involvesWhere to look
An assistant or agent that does work inside your systemsRetrieval, tool use, permissions, evaluation, guardrailsAI agents and automation, enterprise AI assistants
Extraction from documents such as invoices, contracts, and formsParsing, validation, human review, accuracy measurementDocument intelligence
A decision on where to start, with governance in placeUse-case ranking, data readiness, risk and policyAI readiness and governance
An AI feature in an existing productAPI integration, cost control, fallbacks, monitoringA software engineering team that happens to know AI

Many AI projects are, at heart, ordinary software projects with one unusual component. If the integration, security, and data work is most of the effort, a strong engineering partner with real AI experience usually serves you better than a specialist shop that treats everything else as an afterthought.

What to evaluate beyond the demo

1

How they measure quality

Ask what a good answer looks like and how they will know the system produces one. A serious team will propose an evaluation set built from your real cases, a target accuracy, and a way to catch regressions when the model or prompt changes. A team that cannot describe this will judge quality by feel, and so will your users.

2

How they handle being wrong

Models fail in plausible ways. Ask what happens when the system is unsure, when it is wrong, and when it is asked something outside its remit. Good answers involve confidence thresholds, human review for high-stakes steps, citations back to source material, and clear fallbacks.

3

How they treat your data

Where does it go, who can see it, is it used for training, and how is access controlled? Look for permission-aware retrieval, so the assistant never shows someone a document they could not open directly. Ask where the model is hosted and what the contract says.

4

How they control cost

A prototype that costs pennies can cost real money at volume. Ask about caching, model selection by task, token budgets, and how they will report spend. Cost per completed task is the number that decides whether the project survives its first budget review.

5

How they take it to production

Monitoring, logging, versioning of prompts and models, rollback, and an owner after launch. This is the gap that kills most pilots, and we cover it in moving an AI pilot to production.

For a worked example of that last step, see how an enterprise AI assistant moved from pilot to production.

Red flags in a proposal

  • No success metric. “Improve efficiency with AI” is a hope, not a deliverable.
  • A platform before a use case. If the first thing on the table is a multi-year roadmap, they are selling their capacity, not your outcome.
  • Accuracy promised in advance. Nobody can guarantee a figure before seeing your data. Credible teams commit to measuring it.
  • Vague answers on data and security. If the data flow is hand-waved in the sales call, it will be worse in the contract.
  • Only demos, no production references. Ask for something running in front of real users, and for what went wrong with it.
  • One model, always. The best choice varies by task, and a team that cannot say when not to use a large model has not run enough of them.
  • Nothing about people. Adoption, review workflows, and change management decide whether anyone uses it.

Questions worth asking in the first call

Take these into the room
  • What would you measure to decide this worked, and how will we build the test set?
  • Show us something in production. What failed, and what did you change?
  • How do you keep the assistant from exposing data a user should not see?
  • What will this cost per completed task at ten times today's volume?
  • What happens when the underlying model is deprecated or changes behaviour?
  • Who owns the prompts, evaluation sets, and integrations when we finish?
  • Which part of this would you advise us not to use AI for?

Start small, with a number attached

The lowest-risk first engagement is one workflow, one measurable target, and a short timeline. Choose something repetitive and well understood, where a wrong answer is caught cheaply, and where you already know what it costs today. Define the target before anyone writes code, for example the share of cases handled without escalation, or the time taken per document, and agree what happens if it is not reached.

Resist the pull to start with the most ambitious idea. A narrow success that reaches production teaches your organisation more than a broad prototype that stalls. The same discipline applies as in any build, which is why scoping a first release tightly matters here as much as anywhere.

Where governance fits

Larger organisations will need policy on acceptable use, data handling, human oversight, and vendor risk before anything touches customers. It is far cheaper to settle this alongside the first project than to discover it afterwards. If you have not yet decided where AI belongs in your business, an AI readiness and governance engagement is a sensible first step, ahead of any build.

Build, buy, or partner

Not every AI need requires a custom build. Off-the-shelf tools cover many common cases faster and cheaper, and the honest advice is sometimes to buy. Our framework for build versus buy applies directly: build when the capability is a differentiator or must integrate deeply with your own data and systems, and buy when it is commodity.

If you want a straight read on whether your idea is a good AI candidate, what it would take, and what we would not recommend, see our artificial intelligence services or get in touch.

Working through this on a real project?

Tell us what you are building. You will get a scoped estimate and an architecture you own, not a capability deck.

Common questions

Evaluate how they measure quality, handle wrong answers, protect your data, control cost, and take systems to production, rather than how impressive their demo is. Ask to see something running with real users, ask what failed, and be cautious of anyone proposing a platform or promising a specific accuracy before seeing your data.
A working system in production with a defined success metric, an evaluation set you own, monitoring, documentation, and a clear owner afterwards. A prototype or a demonstration alone is not a deliverable, because it does not tell you how the system behaves on real, messy inputs.
It varies widely with scope, the quality of your data, integrations, and the accuracy and safety bar. Two costs are often overlooked: evaluation and monitoring during build, and the ongoing per-request cost once the system is used at volume. Ask any vendor for a range with assumptions, plus an estimate of running cost per completed task.
In-house suits organisations that will keep building AI as a core capability and can attract the talent. An agency suits a first project, a time-bound build, or a need for a mixed team across engineering, data, and security. Many organisations start with a partner and transfer ownership once the system is stable.
No defined success metric, a platform pitched before a use case, accuracy guaranteed in advance, vague answers about data and security, no production references, and no mention of review workflows or user adoption. A vendor who cannot name a task where AI is the wrong tool deserves scepticism.
It can be, with the right controls. Ask where data is processed and stored, whether it is used for model training, how access is permissioned, and what the contract says about retention. Permission-aware retrieval, so users only see what they are already allowed to see, is a baseline for any assistant working over internal documents.
Pick one repetitive, well-understood workflow where mistakes are caught cheaply, set a measurable target before building, and keep the first phase short. A narrow result that reaches production usually teaches more, and builds more internal confidence, than a broad prototype that stalls.
Consulting helps you decide where AI fits, which use cases to prioritise, and what governance you need. Development builds and runs the systems. Many projects benefit from a short readiness phase first, followed by a build, ideally with continuity of people between the two.