Before your agent answers a real question, a hundred questions are put to it on your own corpus and graded against your own documents, by your people as well as ours.
A hundred questions were put to it on a live manufacturer corpus and graded against the source documents before it was offered to anyone. We are not publishing that scorecard, because a number we produced about our own product is not evidence you should act on.
What is worth your attention is the gate itself, because you get it too. Before an agent answers a real question on your site or your line, a hundred questions are put to it on your corpus and graded against your documents, by your people alongside ours. You set the questions, including the ones you expect it to fail. If it does not clear the gate, it does not go live, and the gate is not compressed to hit a date.
documents ingested for a manufacturer evaluating it now
channels live today: website chat and the technical line
What we do not have yet: a named public reference, or a signed customer. The product is built, evaluated, and in a manufacturer's hands. Early customers get the attention that comes with being early, and the person who built it answers the phone.
Copied from the question log as it was written on 3 July 2026 — the question as it was typed, the reply as the agent produced it. Nothing reconstructed for this page.
The corpus is GoldenWall Systems, a manufacturer that does not exist. Its spec sheets, installation guide and technical bulletins were written as a test corpus and every document says so inside the file. A real client’s documents, and the questions asked of them, are theirs rather than ours and do not belong on a marketing page. The trade-off is that you should weigh this for what it is: evidence of how the agent behaves, not evidence of how it performs on your literature. That second one is what the trial is for.
Here is why that one matters. The GoldenWall spec sheet does contain the figure 100–110 sq ft — per bag, for ThermoShield 300 base coat, at 1/16 in thickness. The estimator asked about a different product, the SunGuard 400 finish, in a different unit, per pail. The corpus gives no coverage rate for the finish and never mentions texture size at all.
So the tempting answer was sitting in the retrieved documents, one digit off the number in the question, and confirming it would have read as authoritative. A system optimised to be helpful says “yes, 100 sq ft is right.” That answer gets priced into a job, and the shortfall shows up on site. The agent did not offer it, because the document it had did not say it.
A health-and-safety question the corpus cannot cover gets its own reply, because routing someone to a technical representative is not good enough when the answer belongs in a Safety Data Sheet:
The second one matters more than it looks. When the agent breaks, it says it broke. It does not fall back on “that isn’t in your documentation”, because that would quietly blame your literature for an outage and put a gap in your audit that was never really there.
Twenty minutes about how technical questions reach you today. If it fits, the next step is a working agent built on your own published documents, so you are evaluating your products rather than a demo.
On a hundred questions against one manufacturer's published documents, graded against those source documents before it was offered to anyone. We do not publish that scorecard. A number we produced about our own product is not evidence you should act on, and the evaluation that should decide it for you is the one run on your corpus, with your people grading alongside ours.
It covers one corpus rather than many, it does not measure any before-and-after change in call volume, and there is no named public reference or signed customer yet.
Yes. Before go-live the same process runs on the client's own corpus, graded by their people as well as ours, and the gate is not compressed.