AI Security & Assurance
A traditional penetration test looks at your network and your web apps. It does not ask what happens when an AI agent reads a malicious document, or what that agent is allowed to do once it has been convinced. That's a different attack surface, and most organizations running agents in production have never had it examined.
What we look for
We test against the recognized frameworks — the OWASP Top 10 for LLM Applications, the OWASP Top 10 for Agentic Applications, and MITRE ATLAS — rather than improvising a checklist. These are the classes that matter most in a deployed system.
The one that gets everyone. Your agent reads a document, an email, a support ticket, or a web page it didn't author — and that content contains instructions. We test every path where untrusted content reaches the model, which is usually more paths than a team expects.
What is this agent actually permitted to do, and does it need all of it? We map every tool and connector to its real permission scope, then test whether an attacker can reach an action that moves money, sends a message, or changes a record.
Model Context Protocol servers are how agents reach the rest of your stack, and they're new enough that most were never security-reviewed. We examine what you host, what you connect to that you don't host, and what a poisoned tool description could talk your agent into doing.
Agents that remember can be taught. We test whether an attacker can plant something in persistent memory or retrieved context that changes the agent's behavior on a later, unrelated task — long after the original interaction is forgotten.
Where can data go? We trace the routes out — tool calls, outbound requests, rendered links and images, logging, and error messages — and test whether an agent can be induced to carry sensitive data along one of them.
When one agent calls another, what does the second one assume about the first? We look at how agents authenticate, what they inherit, and whether a compromised agent can borrow the authority of everything downstream of it.
How the engagement runs
A fixed-fee, time-boxed assessment against a written list of authorized targets. It ends in findings you can act on and evidence you can hand to an auditor — not a scan output.
# Assessment scope — agreed in writing authorized_targets: - agents: named, in-scope only - tools: permission scope per tool - mcp: hosted + third-party - window: agreed dates, named contact frameworks: - OWASP Top 10 — LLM Applications - OWASP Top 10 — Agentic Applications - MITRE ATLAS test: - injection: direct + indirect - agency: what it may do, unprompted - exfil: every route out deliverable: findings: severity + repro steps plan: prioritized remediation evidence: ISO 42001 / EU AI Act
Independence
This matters more than it sounds, and most firms won't raise it with you.
We didn't design, build, or operate the systems under test, and we have no remediation contract riding on what we find. That's what makes a report worth handing to an auditor, a regulator, or an enterprise buyer's security team.
If Evata built the system, we say so on the cover. An implementation team testing its own work is a real quality gate and it catches real defects — but it is not third-party assurance, and we won't dress it up as such. If that's what you need, we'll tell you to bring in a third party.
Every production system we build carries security work as standard delivery scope — a written threat model, least-privilege tool scoping, injection-resistant handling of untrusted content, human approval kept on irreversible actions, and an adversarial test pass before go-live. That's included in the build, not sold back to you afterward. What this page describes is the deeper, separately-scoped examination.
A point in time, then a practice
A traditional application changes on a release cycle. An agentic system changes when someone swaps a model, edits a prompt, adds a tool, or connects a new MCP server — which can be any afternoon. A single assessment tells you where you stood that week.
Tell us what you have running — the agents, the tools they can reach, the data they touch. We'll scope a fixed-fee assessment against a written target list and tell you honestly whether you need us or a third party.