testclub_

02 · AI & LLM Security

Can your AI be made to misbehave?

Security asks a different question to evaluation: not whether the system performs well, but whether it can be made to misbehave, or talked into breaking its own business rules. We test by adversarial example — crafted inputs, hidden instructions in retrieved content, requests that push an agent beyond its intended scope — and we test the contract between a model and the tools it calls: does it call the right tool, with the right arguments, gated correctly before it acts. That contract is the attack surface most teams haven’t specified yet, let alone tested. The OWASP Top 10 for LLM Applications is where this starts, not where it stops; what matters is whether a category applies to what your system actually does.

Method

We test by adversarial example: crafted inputs and hidden instructions in retrieved content, attempts to extract system prompts or training data, and requests that push an agent beyond its intended scope. Findings are reproducible and mapped to the OWASP category they sit under.

You get: reproduction steps, business impact, and a fix, in priority order, for every finding. Walked through with your engineers, not left in a PDF.

Typical initial engagement: two to four weeks.

Get in touch →

Related

Security asks whether it can be made to misbehave. Evaluation asks how often it gets things wrong on its own. Most teams need both.

See AI Evaluation →

AI agents act through your APIs. If an agent can be talked into a request, your API decides what happens next.

See Business Logic & API Security →