A fintech MVP, kickoff to paid pilot in 12 weeks
12 weeks from kickoff to the first paying pilot customer
Everyone liked the demo. Nobody would put it in front of a customer, because there was no evaluation set — so “better” was an opinion held by whoever had spoken last.
correct-and-cited on a 400-ticket evaluation set, from 61% at the start
weeks from the embed starting to general availability, after six months in pilot
cross-tenant retrieval defects found in pre-release testing, all before a customer saw one
Figures are as reported by the client over the period named in the body below, and were not independently audited by us.
Working on something like this?
Get an estimateThe assistant had been in pilot for six months. Everyone liked the demo. Nobody would put it in front of a customer, and the reasons were never written down in a form anyone could act on.
It answered the questions it had been demoed with and invented answers to the long tail. There was no evaluation set, so every proposed improvement was assessed by someone trying twenty questions they had thought of themselves, which measures the questions rather than the assistant.
Four hundred real support tickets, each with the answer a competent support engineer would have given and the document it should have cited. Scored for correct-and-cited rather than for plausibility. This was the entire deliverable of the first month, and it was the change that made every later change decidable.
Tenant isolation enforced where the documents are stored, not as a predicate on the query that someone can forget to apply. A filter is a thing you remember; an index boundary is a thing you cannot bypass. Three cross-tenant defects were found in pre-release testing, and all three would have passed a filter-based design's own tests.
Not prompt instructions. If the retrieved context does not support an answer, the assistant says so and offers the ticket path — and that behaviour has its own tests and its own line in the specification, so it survives a model change.
Under the retention policy the vendor already had. This is what legal had been asking for and what nobody had framed as an engineering requirement: a complaint about an answer given in March has to be reconstructable in June.
Two of our engineers, on the client's SDLC, branching model, code review, and change control. Not a parallel team producing a handover — the handover in this model is continuous, and there is no week where knowledge transfers in a meeting.
Rolled out by customer cohort rather than by date, starting with the tenants whose document sets were largest and messiest, on the principle that the hardest cohort tells you the most and the easiest cohort tells you nothing.
Because until there was an evaluation set, a model change was indistinguishable from a mood. The team had already spent six months making changes that could not be assessed, and the correct read of that is not that the changes were bad — several were good — but that nobody could tell, so none of them could be defended.
Once the 400 tickets existed, the sequence became ordinary engineering: measure, change one thing, measure again. Correct-and-cited went from 61% to 89% over ten weeks, and roughly two-thirds of that came from retrieval and chunking rather than from the model or the prompt. That distribution is typical and it is not what the pilot had been spending its time on.
correct-and-cited on the 400-ticket evaluation set, from 61%
from embed to general availability, after six months in pilot
fewer tier-one tickets in the first quarter after GA
The 61% to 89% is measured on one evaluation set built from one product's support history. It says the assistant improved against the questions its own customers actually ask. It does not transfer to your product, and a vendor quoting a percentage without naming the set is quoting nothing.
The three cross-tenant defects are the number we would rather a buyer asked about. They were found because the evaluation set included deliberately adversarial retrieval cases, not because someone spotted them while using the product. Without that, the first person to find them would have been a customer.
An embedded team on an AI feature is a good fit under specific conditions, and a poor one outside them.
The capability is AI development services, the model is engineers embedded in your team, and the platform context is a SaaS platform build. If you are in the same place this vendor was, getting an AI pilot into production and multi-tenant architecture choices are the two pieces that matter most.
A demo is easy and a deployment is not. Build the evaluation set first, and the arguments stop being about how it felt on a Tuesday.
“For six months we argued about whether it felt better. Four weeks in we were arguing about a number instead, and those arguments finished.”
Tell us what is not working. You will get a scoped estimate and an architecture you own, not a capability deck.
12 weeks from kickoff to the first paying pilot customer
2 days to make a pricing change, down from roughly six weeks
4 days median referral to first appointment, down from eleven, over two quarters