Catalog-building price — 25% off. Code tap to copy

Eval Kit · Runner-agnostic eval kit

Agent Eval Harness Kit

The eval layer free runners leave as an exercise — task sets, rubrics & a CI gate that says no.

⇩ Instant download — yours right after checkout

Production telemetry across 6,259 deployed agents measures 56.6% task success — a roughly 37% drop from benchmark to production — and only 7% of enterprises have scaled agentic AI past pilots. The eval runners are free and excellent; that's not the gap. The gap is what every team still hand-rolls on top: the golden task sets worth defending, judge rubrics calibrated against human labels so a score means something, thresholds someone will stand behind, and the CI wiring that turns a score into a merge gate. This kit is that layer — runner-agnostic, with adapters for the free open-source runners you already have, so last week's prompt change gets caught in CI instead of production.

Built for: Teams with a live LLM agent, traces wired, and no way to say whether last week's prompt change made it worse

Pick your license

Solo

$59

One engineer, one agent to defend.

  • The full kit
  • Task sets + rubrics + CI gate
  • Runner adapters
  • 12 months of updates

One developer's own projects.

Most popular

Team

$119

Your whole team gates with it.

  • Everything in Solo
  • Team-wide license
  • Priority email support

One team, internal and client agents.

Agency

$299

Unlimited client agents.

  • Everything in Team
  • Unlimited client agents
  • Source design files where provided

Unlimited client projects.

14-day money-back guarantee — if the kit doesn't fit your stack, reply to your receipt within 14 days and we refund it in full. Keep nothing, owe nothing.

More from the store

Signal Foundry — AI Product Studio Kit

The 40 hours before your first prompt — research, architecture, and an app that refuses to make things up.

Relay — Autonomous SDR

The AI SDR that can't make things up.

ReconBridge — CRM↔ERP Reconciliation Agents

Your integrations move data. Nothing proves it arrived right.

Questions

What exactly do I get?

An instant download: 36 authored golden tasks across two packs (18 agent tasks — tool use, refusal, ambiguity, planning, error recovery, instruction fidelity — and 18 RAG tasks with embedded fictional corpora), five judge rubrics with concrete 1–5 anchors plus the judge prompt template, the calibration procedure with a script that computes agreement stats against human labels, threshold guidance with per-number rationale, the CI merge gate and GitHub Actions workflow, a deterministic regression report, a promptfoo adapter plus a generic results contract for any runner, real docs, and 12 months of rubric and task-pack updates. npm test runs 24 tests — green at packaging on a clean install.

The eval runners are free — why would I pay for this?

Keep the free runners; the kit runs on them rather than replacing them. What they leave as an exercise is the content: which tasks, judged how, against what threshold, wired into CI so a regression actually blocks a merge. That authoring and calibration work is what you'd otherwise spend engineer-weeks on — and it's the part our own engine uses in production to gate its agent workflows, including the gate that scored this very product.

How does it plug into my runner?

The promptfoo adapter emits a ready config from any task pack, and the generic results contract ingests any runner's output into the gate — both tested. For DeepEval, Ragas and Phoenix the docs wire the concepts and point you at their current APIs rather than hardcoding fields that move; if a shape drifts, the ingester fails loudly with instructions instead of guessing.

Was this scoped by AI?

Yes — our commercial engine researched and scored it, adverse evidence included and every stat source-checked. Full honesty: the engine scored this concept below our own demand gate — no proof yet that teams pay for the content layer above free runners — and the owner shipped it anyway as a bet. What's not a bet: 24 tests green on a clean-room install, a calibration script verified against hand-computed statistics, and the 14-day refund if it's not what you needed.

01 / 01

Drag to read · 2880px capture