Resolution Audit · B2B SaaS · Python + React

Resolution Audit — Verify Your AI Support Bill

Your AI vendor grades its own homework. Then invoices you for the grade.

Start Portfolio · $7,500/mo Hands-on onboarding — we calibrate against your own labelled tickets

The whole support-agent market now bills by the outcome: Intercom's Fin charges $0.99 per resolution, Zendesk about $1.50–$2.00 per automated resolution — a term Zendesk itself defines as a conversation with no human escalation and 72 hours of customer silence — and Decagon runs a comparable model at a $432,750 median contract, per Vendr. In every case the vendor defines the billable event, counts it, and sends the invoice, and there is no independent party checking any of it. Salesforce just paid $3.6B for Fin; Gartner expects over 40% of agentic AI projects to be cancelled by end-2027 for unprovable value. Resolution Audit is the buyer's side of that table: it re-reads every closed conversation, classifies what actually happened, links the follow-up ticket that arrived two days later, and reconciles the month against the invoice in dollars. It publishes its own accuracy against a human-labelled set, because an auditor that won't show its error rate is just another opinion.

Built for: CX and support-ops leaders paying per-resolution for an AI agent, and the finance teams signing those invoices

4 screens · the running build

Inside the software

01 / 04

Resolution AuditCalibration
Resolution Audit — Calibration
our own accuracy, error bars and blind spots, published in the product

Founding-pilot plans

Brand

$2,500/mo

One helpdesk, one vendor.

  • One helpdesk
  • Up to 25k conversations a month
  • Monthly reconciliation report
  • Verdict explorer with transcripts
Start Brand

Founding-pilot plan — one company, cancel anytime.

Most popular

Portfolio

$7,500/mo

Multiple brands or vendors.

  • Everything in Brand
  • Multiple brands and vendors
  • Dispute packs + re-contact analysis
  • Priority support
Start Portfolio

One company, higher volume — cancel anytime.

Enterprise

$15,000/mo

API, custom rubrics, SSO.

  • Everything in Portfolio
  • Reconciliation API + SSO
  • Custom judging rubrics
  • Onboarding and a named contact
Start Enterprise

Multi-entity / enterprise terms — annual or monthly.

More from the store

MCP Server Production Kit

The production layer every MCP server needs — OAuth, rate limits, evals & Cloudflare deploy in one kit.

Agent Skill Production Kit

The review layer every publishable Claude Code skill needs — validate, scan, ship.

AI Usage Billing Kit

Credits that can't double-spend, caps that hold, and margin you can actually see.

Questions

Isn't this just adversarial to my AI vendor?

It's the same relationship you already have with any metered supplier — you meter the meter. Most engagements find the billing broadly right and a specific pattern wrong, like conversations closed by a deflection that the customer immediately re-opened under a new ticket. That's a conversation with your vendor about definitions, not an accusation, and having the transcripts to hand is what makes it a short conversation.

How do I know your judgement is any better than theirs?

Because we publish ours, and right now what we'd be publishing to you is an uncalibrated judge. The software ships a Calibration screen showing our agreement against a human-labelled sample with the confidence interval on it, per verdict class, including the classes we're bad at — and the product refuses to report a number as a judge accuracy when no model produced it. Being straight about where that stands today: the judge panel is built and its two-pass-plus-tiebreak logic is tested, but it has not been scored against a model on our side, so the screen currently reads "does not clear our 90% bar". Calibrating it against your labelled tickets is the first thing week one does, and if it doesn't clear the bar you shouldn't buy the second month.

What has actually been measured, then?

The parts that don't need a model. Given correct verdicts the pipeline recovers 86% of a known planted misbilling with under 2% false disputes; the re-contact tracker runs at a 0.3% false-dispute rate, having been rebuilt four times because the first version confidently linked 218 tickets of which 3 were real; re-running an audit costs nothing because every verdict is cached on the transcript and the rubric version. All of it on synthetic data — no real support estate has been audited and no dispute has ever been filed. The full ledger ships in the repo as docs/VALIDATION.md and we'll send it before you buy.

What does it need access to?

Read-only access to your helpdesk's closed conversations and a copy of the vendor invoice — nothing that can change a ticket, and nothing that touches your customers. A CSV export path exists for teams that would rather not connect an API at all, and it's the path we build against first on purpose: this product tells your vendor their invoice is wrong, and a design that only works through their API is a design they can switch off.

Was this scoped by AI?

Yes, and here's exactly how, because it matters. Six research agents worked one domain each against a strict evidence rule: every demand claim had to come from an analyst report, a vendor's own pricing page, a dated funding round, public earnings, a named survey or a regulatory text — no forum posts, no vibes. That produced 41 candidates scored on demand evidence, contract value, how verifiable the gap was, buildability and defensibility; this one ranked third, and it's the third of the ten we've actually built. What that does NOT mean: that anyone has paid for it, or that we've audited a real invoice. The research says the problem is real. Only a pilot says the product works.

01 / 05

Drag to read · 2880px capture