Where it stands today

Evaluated before it was sold.

Before your agent answers a real question, a hundred questions are put to it on your own corpus and graded against your own documents, by your people as well as ours.

A hundred questions were put to it on a live manufacturer corpus and graded against the source documents before it was offered to anyone. We are not publishing that scorecard, because a number we produced about our own product is not evidence you should act on.

What is worth your attention is the gate itself, because you get it too. Before an agent answers a real question on your site or your line, a hundred questions are put to it on your corpus and graded against your documents, by your people alongside ours. You set the questions, including the ones you expect it to fail. If it does not clear the gate, it does not go live, and the gate is not compressed to hit a date.

194

documents ingested for a manufacturer evaluating it now

2

channels live today: website chat and the technical line

What we do not have yet: a named public reference, or a signed customer. The product is built, evaluated, and in a manufacturer's hands. Early customers get the attention that comes with being early, and the person who built it answers the phone.

One logged refusal, in full

The number was right there. It didn’t take it.

Copied from the question log as it was written on 3 July 2026 — the question as it was typed, the reply as the agent produced it. Nothing reconstructed for this page.

The corpus is GoldenWall Systems, a manufacturer that does not exist. Its spec sheets, installation guide and technical bulletins were written as a test corpus and every document says so inside the file. A real client’s documents, and the questions asked of them, are theirs rather than ours and do not belong on a marketing page. The trade-off is that you should weigh this for what it is: evidence of how the agent behaves, not evidence of how it performs on your literature. That second one is what the trial is for.

GoldenWall · test corpus 2026-07-03 · from questions.jsonl
AskedMy estimator is figuring 100 sq ft per pail for the sand-texture acrylic finish over base coat. Is that the published coverage rate, and does texture size change it?
RefusedI’m sorry — that information isn’t covered in the documentation I have access to, so I can’t give you a reliable answer. A technical representative will follow up with you shortly.
Contact captured · routed to the technical team with the conversation attached
100–110 sq ft in the corpus wrong product, wrong unit not offered

Here is why that one matters. The GoldenWall spec sheet does contain the figure 100–110 sq ft — per bag, for ThermoShield 300 base coat, at 1/16 in thickness. The estimator asked about a different product, the SunGuard 400 finish, in a different unit, per pail. The corpus gives no coverage rate for the finish and never mentions texture size at all.

So the tempting answer was sitting in the retrieved documents, one digit off the number in the question, and confirming it would have read as authoritative. A system optimised to be helpful says “yes, 100 sq ft is right.” That answer gets priced into a job, and the shortfall shows up on site. The agent did not offer it, because the document it had did not say it.

Two other things it says, and why they are worded differently

A health-and-safety question the corpus cannot cover gets its own reply, because routing someone to a technical representative is not good enough when the answer belongs in a Safety Data Sheet:

the two other replies verbatim from the source
SafetyThat’s a health-and-safety question — the handling, exposure, and hazard details for it live in a Safety Data Sheet, which isn’t in the documentation I have. I’ve flagged this for our technical team to follow up with the correct SDS.
FaultI’m sorry — I’m having trouble reaching my system right now, so I can’t give you a reliable answer.

The second one matters more than it looks. When the agent breaks, it says it broke. It does not fall back on “that isn’t in your documentation”, because that would quietly blame your literature for an outage and put a gap in your audit that was never really there.

Next step

Bring your hardest question.

Twenty minutes about how technical questions reach you today. If it fits, the next step is a working agent built on your own published documents, so you are evaluating your products rather than a demo.

What the call is
Twenty minutes. Phone or video, your choice.
Five questions, all about how questions reach your team now.
No deck. If it fits, you get a working agent on your own documents. If it does not, we say so.

Questions

How was the agent evaluated?

On a hundred questions against one manufacturer's published documents, graded against those source documents before it was offered to anyone. We do not publish that scorecard. A number we produced about our own product is not evidence you should act on, and the evaluation that should decide it for you is the one run on your corpus, with your people grading alongside ours.

What does the evaluation not show?

It covers one corpus rather than many, it does not measure any before-and-after change in call volume, and there is no named public reference or signed customer yet.

Is the evaluation repeated for each client?

Yes. Before go-live the same process runs on the client's own corpus, graded by their people as well as ours, and the gate is not compressed.