The problem
Systems are being built faster than they can be checked, using tools which are themselves unexamined. While a typical old failure was an outage, a new one is a system working exactly as built, just not as specified, doing damage nobody expected.
Who decides whether your AI's answers are good enough? Most teams measure generic scores like helpfulness or hallucination. Those rarely match how their own product goes wrong, or how an agent misuses the tools it's been given. Do you know how your AI fails, or only what you assumed to measure?
The API layers of digital applications are often a huge attack vector, but these load-bearing structures can be the least scrutinised area. Business logic flaws don't show up in automated vulnerability scans. An API must be deeply understood before we can creatively discover where the genuine vulnerabilities are.
It's the post-mortem. Someone senior wants to know why.
"Who defined what the behaviour should have been here?"
Nobody.
That is the hardest problem in software. And it is getting harder.
Our answer
We verify strictly against what the system was designed to do, not only what it's doing now. The gap between those two is where the damage often lives. For AI, that means finding how it fails before deciding what to measure — or specifying it up front, before anything's built, when we're brought in early enough to.
What we do
How does your AI fail, and how often? We find the failures that matter using your data and your experts, then measure them with confidence intervals. Not the vendor's benchmark. Yours.
AI Evaluation → 02 · AI & LLM SECURITYCan your AI, or the agents it runs, be made to misbehave or break their own rules? Adversarial testing of business logic and tool-call contracts, not just a published checklist.
AI & LLM Security → 03 · BUSINESS LOGIC & API SECURITYBusiness logic and API security, the flaws scanners don't find because the code does exactly what it was built to do. What matters is what it then lets someone do.
Business Logic & API Security → 04 · SPECIFICATION AS METHODTwenty-five years of asking the same question, before it was APIs or AI. The discipline behind everything above.
Read the philosophy →Who
Testclub was founded in 2017 by Omar El Dali, now 25 years a practitioner at the boundary between specification and implementation, across trading platforms, fintech, insurance, health, energy and government.
A skilled high-agency group who move fast and adapt to what your project actually needs. Every finding arrives with its evidence and the method that produced it, so an engineer can follow every step.
We are independent. We do not build AI products and we do not sell a platform, so our only interest is discovery.
Clients