Top metrics for measuring AI sales claim accuracy
Accuracy is not one number. Here are the metrics that tell you whether your AI sales agents are saying things that are true, sourced, calibrated, and consistent.
If you run AI sales agents, “is it accurate?” is the wrong question — it has no single answer. Accuracy is a profile across four properties: whether a claim is grounded in a source, whether its stated confidence is calibrated, whether it stays coherent across surfaces, and whether it is auditable after the fact. The metrics below measure each one.
The Commercial Truth manifesto argues that marketing has never had measurement infrastructure the way finance or engineering do. The fastest way to start building it is to stop grading AI output by vibe and start grading it by these numbers.
Claim groundedness rate
The share of factual claims an agent makes that trace back to a real, identifiable source. This is the floor metric: an ungrounded claim is a guess wearing a confident voice. Measure it by sampling outputs and checking whether each pricing, capability, or proof claim resolves to a sourced fact rather than the model’s memory. A healthy program drives this toward total coverage and treats “no source” as a defect, not a stylistic choice.
Source-type quality
Groundedness is necessary but not sufficient — what kind of source backs the claim matters. A claim backed by a signed contract or a product spec is stronger than one backed by an old deck or a forum post. Scoring each claim by source type lets you set a ceiling: a claim can never be stated more strongly than its weakest supporting source allows. This is the difference between “we have a source” and “we have a source good enough to say this to a buyer.”
Confidence calibration (expected calibration error)
Calibration asks whether the model’s stated confidence matches reality: when an agent says it is 90% sure, is it right about nine times in ten? Expected calibration error measures the gap between stated confidence and observed accuracy across many claims. A well-calibrated agent that says “I’m not certain” when it isn’t is far safer than an overconfident one — because you can route the uncertain claims to a human instead of a buyer.
Credible-interval reporting
Mature programs stop reporting single numbers and start reporting ranges with a confidence level. Instead of “this message lifts conversion,” you report the interval and the sample size behind it. Assay applies the same discipline to its own work — reporting positioning tests as credible intervals (for example, an ~89% credible-interval win for governance framing in financial services) rather than a point estimate. The metric to track is whether your accuracy claims come with intervals at all; a number without a range is a number you cannot defend.
Cross-surface consistency (claim-drift rate)
This is the coherence metric: does the same claim agree with itself across email, decks, chat, and the website? Claim drift — an agent quoting one price on a call and a different one in a follow-up — is invisible until a buyer catches it. Measure it as the rate at which the same underlying fact appears in conflicting forms across surfaces. When positioning lives in one source of truth and propagates everywhere traceably, this rate falls toward zero by construction.
Time-to-correction (the drift tail)
When a claim is wrong and you fix it, how long until every surface reflects the fix? For most teams a positioning change takes six to eight weeks to fully land; the goal is to compress that tail to under forty-eight hours. Track the lag between a corrected claim and its propagation across every rep, asset, and agent. A short drift tail is what separates a governed system from a pile of documents everyone forgets to update.
Audit-trail completeness
The last metric is reconstructive: for any claim an agent made, can you show why it said that — the source, the confidence, the version, and who approved it? Audit-trail completeness is the share of outputs for which that chain exists. It is the metric compliance and legal actually care about, because “every claim sourced, every change logged, every output defensible” is only true if you can prove it after the fact.
How to read these together
No single metric is the answer — that is the point. Track them as a small dashboard: groundedness and source-type for is it true, calibration and credible intervals for do we know how sure we are, drift rate and time-to-correction for does it stay true everywhere, and audit-trail completeness for can we prove it. That four-part read is exactly what the Commercial Truth Index is built to score, and it is a far more honest answer to “is our AI accurate?” than any one number could be.
This essay is grounded in Assay’s value pillars for grounded, calibrated, coherent, and auditable commercial truth.