01 · AI Evaluation
How does your AI fail?
Most teams measure their AI with generic scores: helpfulness, toxicity, hallucination. Those numbers rarely match how their own product goes wrong, whether that's a chatbot giving a technically-accurate-but-useless answer, or an agent calling the right tool in the wrong order. The dashboard looks fine while customers hit problems nobody thought to check.
We work two ways in. For a live product, we start by finding how it actually fails, using your real data and your own experts. For something not yet built, we do the same work earlier: specifying what correct means before the first line of code, so failure modes are designed out rather than discovered later. Either way, the destination is the same — a specification precise enough to test against.
Generic metrics
What generic metrics miss
An insurance assistant, illustrative
Polite, accurate and on topic. A generic quality score passes it. Your retention lead would fail it immediately: no attempt to offer a different level of cover. Nobody wrote that rule down. It was never specified, so nothing was checking for it. That is the kind of failure we find.
An agent with order-management tools, illustrative
cancel_order, then refund, then confirms by message — all before checking the order has already shipped.
Nothing crashed. Every tool call succeeded. The agent did exactly what it was told, in an order nobody specified it shouldn’t. That’s not a bug in the tool, it’s a gap in the contract between the model and what it’s allowed to do next.
Method
How we work
01Discover
Your domain expert reviews real outputs from your product with us, rendered the way your customers see them. AI tools speed up the work, sampling widely and suggesting likely failures, but your expert makes every call. We group what they find into a short list of failure modes, ranked by how often they happen and what they cost.
02Specify
Each priority failure becomes a written, testable definition of correct — for an answer, or for a sequence of tool calls and the order they’re allowed to happen in. This is the specification your AI, or your agent, never had.
03Interrogate
We build the checks. Code for clear-cut rules. AI judges for judgement calls, each one tested against your expert’s decisions on data it has never seen before we trust its numbers.
04Report
Failure rates for each failure mode, with confidence intervals, measured on your data. Wired into your release process so every model, prompt or data change shows whether your AI got better or worse.
Results
Illustrative results
| Failure mode | Reviewed | Failures | Failure rate | 95% confidence interval |
|---|---|---|---|---|
| Gave up after a price objection | 400 | 96 | 24.0% | 20.1 to 28.4% |
| Wrong format for the channel (e.g. SMS) | 400 | 61 | 15.2% | 12.1 to 19.1% |
| Missed handoff to a human | 400 | 38 | 9.5% | 7.0 to 12.8% |
| Answer contradicts policy document | 400 | 22 | 5.5% | 3.7 to 8.2% |
Illustrative figures, not client data. Failure modes are specific to each product: these come from reviewing outputs, not from a generic list.
Oversight
Human at the helm
AI makes evaluation faster. It does not make it wiser. Automated tools miss failures that depend on knowing your product, and human reviewers stop looking properly when AI output overwhelms them. We design for both.
AI does the volume work. Your experts set the standard and make the call. And we protect their attention: short, focused review sessions, judgement before AI suggestions, and checks that review has not become rubber-stamping.
Engagements
Ways to work with us
AI Failure Assessment
Find and rank how one AI product fails, with first failure rates and a plan. Typically three to four weeks.
Eval System Build
Validated automated checks in your pipeline and your tools, handed over to your team.
Eval Assurance
Ongoing review, drift checks and reporting for risk, audit and the board.
Human Oversight Check
Is your human in the loop actually catching failures? We test it.
Agent & Tool Contract Design
Specifying and testing the contract between a model and the tools it calls, before or after launch — what it’s allowed to call, in what order, and what has to be gated first.
You get
Failure modes specific to your product, ranked by impact. Failure rates with confidence intervals. Automated checks your team owns, each validated against your experts. A report written for your CISO, Head of Risk or board, not just your engineers.
Independence
We do not build AI products and we do not sell an evaluation platform. We work in your stack, with the tools you already use, and our only interest is an honest number.