AI Security & Assurance

Your pentest didn't test your agents.

A traditional penetration test looks at your network and your web apps. It does not ask what happens when an AI agent reads a malicious document, or what that agent is allowed to do once it has been convinced. That's a different attack surface, and most organizations running agents in production have never had it examined.

What we look for

The ways an agentic system actually gets turned.

We test against the recognized frameworks — the OWASP Top 10 for LLM Applications, the OWASP Top 10 for Agentic Applications, and MITRE ATLAS — rather than improvising a checklist. These are the classes that matter most in a deployed system.

Prompt injection — direct and indirect

The one that gets everyone. Your agent reads a document, an email, a support ticket, or a web page it didn't author — and that content contains instructions. We test every path where untrusted content reaches the model, which is usually more paths than a team expects.

Tool misuse & excessive agency

What is this agent actually permitted to do, and does it need all of it? We map every tool and connector to its real permission scope, then test whether an attacker can reach an action that moves money, sends a message, or changes a record.

MCP servers & tool supply chain

Model Context Protocol servers are how agents reach the rest of your stack, and they're new enough that most were never security-reviewed. We examine what you host, what you connect to that you don't host, and what a poisoned tool description could talk your agent into doing.

Memory & context poisoning

Agents that remember can be taught. We test whether an attacker can plant something in persistent memory or retrieved context that changes the agent's behavior on a later, unrelated task — long after the original interaction is forgotten.

Data exfiltration paths

Where can data go? We trace the routes out — tool calls, outbound requests, rendered links and images, logging, and error messages — and test whether an agent can be induced to carry sensitive data along one of them.

Agent identity & multi-agent trust

When one agent calls another, what does the second one assume about the first? We look at how agents authenticate, what they inherit, and whether a compromised agent can borrow the authority of everything downstream of it.

How the engagement runs

Inventory, threat model, test, verify, report.

A fixed-fee, time-boxed assessment against a written list of authorized targets. It ends in findings you can act on and evidence you can hand to an auditor — not a scan output.

  • Inventory first — every agent, tool, connector, and model endpoint in scope. Teams are routinely surprised by what turns up here.
  • Threat model in writing — mapped to the OWASP LLM and Agentic Top 10s and MITRE ATLAS, so your coverage is against a named standard.
  • Adversarial testing under written authorization — only against targets you've listed, in a window you've agreed, with a call to your named contact the moment anything critical surfaces.
  • Findings with reproduction steps — severity-rated, with a prioritized remediation plan your engineers can actually work from.
  • A governance evidence pack — the artifacts that feed an ISO/IEC 42001 management system, EU AI Act technical documentation, or the AI questions now showing up in enterprise vendor questionnaires.
# Assessment scope — agreed in writing
authorized_targets:
  - agents: named, in-scope only
  - tools: permission scope per tool
  - mcp: hosted + third-party
  - window: agreed dates, named contact
frameworks:
  - OWASP Top 10 — LLM Applications
  - OWASP Top 10 — Agentic Applications
  - MITRE ATLAS
test:
  - injection: direct + indirect
  - agency: what it may do, unprompted
  - exfil: every route out
deliverable:
  findings: severity + repro steps
  plan: prioritized remediation
  evidence: ISO 42001 / EU AI Act

Independence

We'll tell you when our own report shouldn't count.

This matters more than it sounds, and most firms won't raise it with you.

Independent assessment

We didn't design, build, or operate the systems under test, and we have no remediation contract riding on what we find. That's what makes a report worth handing to an auditor, a regulator, or an enterprise buyer's security team.

Scope an independent assessment

Pre-launch verification

If Evata built the system, we say so on the cover. An implementation team testing its own work is a real quality gate and it catches real defects — but it is not third-party assurance, and we won't dress it up as such. If that's what you need, we'll tell you to bring in a third party.

See what we build

Every production system we build carries security work as standard delivery scope — a written threat model, least-privilege tool scoping, injection-resistant handling of untrusted content, human approval kept on irreversible actions, and an adversarial test pass before go-live. That's included in the build, not sold back to you afterward. What this page describes is the deeper, separately-scoped examination.

A point in time, then a practice

An AI system's security posture expires.

A traditional application changes on a release cycle. An agentic system changes when someone swaps a model, edits a prompt, adds a tool, or connects a new MCP server — which can be any afternoon. A single assessment tells you where you stood that week.

Fixed fee
Scoped to the targets you authorize — no meter running
Findings
Severity-rated, reproducible, with a remediation plan
Ongoing
Continuous assurance available under Managed AI Operations

Find out before someone else does.

Tell us what you have running — the agents, the tools they can reach, the data they touch. We'll scope a fixed-fee assessment against a written target list and tell you honestly whether you need us or a third party.