From the team
What a 0.78 faithfulness score never told us
Before quots, we built the eval stack for a WhatsApp booking agent. The metrics were green, the customers were not. This is the gap quots exists to close.
Before quots, we were the team on the other side of this product. We built the evaluation stack for a WhatsApp booking agent that talked to thousands of customers a week: six prompt versions, more than 150 hand-written test scenarios, and a faithfulness score that settled at 0.78.
By every convention of the trade, that was a good number. It went into the weekly report. It turned the release checklist green. And it told us almost nothing about what the agent was actually doing to the people talking to it.
The dashboard said ship it
Here is what 0.78 looked like from the inside. A returning customer opened with the fact that she had already reset her password twice; the agent cheerfully sent her the reset link a third time. A salon owner asked to move a booking and got a perfectly grounded answer about how to cancel one. Each of those conversations passed the checks that feed a single score, because the answers were faithful to the context the agent was given.
The score was not wrong. It was answering a different question. Faithfulness measures whether an answer stays grounded in its context. It does not measure whether a specific kind of customer, with a specific history and a specific patience budget, got what they came for.
What the average was hiding
When we finally read transcripts segment by segment, the 0.78 fell apart into pieces that did not resemble each other. First-time bookers: nearly flawless. Concise, tool-savvy regulars: strong, as long as the request stayed on the happy path. Returning customers with a problem, the ones who had already tried something and said so, failed so consistently it might as well have been policy.
The agent was not 78% good. It was excellent with some people and quietly burning others, and the number had no way to say so.
No dashboard named those people. No metric explained why their trust collapsed at turn four (they were being walked through steps they had already ruled out). We learned all of it by reading conversations after midnight, which works exactly once and does not scale.
"Which" and "why" are the whole job
The uncomfortable conclusion was that the midnight reading was not a workaround. It was the actual job: group conversations into the customer types that genuinely exist, test against each type, grade per segment, and keep the conversation behind every verdict so the finding can be audited.
So we built the tool we wanted back then. quots derives personas from your own chat history, has each persona interrogate your agent for thousands of turns, and reports which segments your agent wins, which it loses, and why, with the receipts attached.
Rehearse before anyone real
The same mechanism answers the question that kept us up before every release: what will the new version break? A candidate prompt, model or setup runs against the same personas before a single real customer meets it. You ship the version that wins more segments, not the one that demoed well.
If you run a customer-facing agent and your reporting is one number, the people quietly giving up on it are inside that number. We built quots to name them.