Offensive security testing

Your estate has changed.
Your testing hasn't.

Web applications, infrastructure and full red team engagements plus the two surfaces almost nobody is testing yet, the AI systems you've deployed, and the machine identities running underneath them.

Web application testing Infrastructure Red team AI system security Non-human identity

Capability

We test like a real attacker

Findings you can act on, with reproduction steps, evidence and a fix

The gap

You shipped AI faster than you governed it.

Every organisation deployed agents, copilots and retrieval pipelines in the last eighteen months. Almost none of them extended their testing or their identity governance to cover it. That gap is measurable.

109:1
Machine identities per human identity — up from 82:1 a year earlier. 79 of the 109 are AI agents.
Palo Alto Networks, 2026
78%
Of organisations have no documented policy for creating or removing AI agent identities.
CSA / Oasis Security, 2026
92%
Are not confident legacy IAM can manage AI and non-human identity risk.
CSA / Oasis Security, 2026
68%
Could not tell what an agent had done as distinct from what a person had done.
CSA / Oasis Security, 2026

None of this is a new discipline. It is appsec and trust boundaries applied to a component you added last year.

AI system security

What we actually test.

An agent with a tool, a credential and a text input is a remote code execution surface with better manners. We test it as one.

01

Prompt injection

Direct and indirect. Content that reaches the model from documents, retrieved chunks, tickets, emails or web pages or anywhere the system treats untrusted input as instruction.

02

Tool and function abuse

What the agent can be persuaded to call, with what arguments, on whose authority. Where the tool boundary is enforced, and where it is merely described in a system prompt.

03

RAG and retrieval leakage

Whether the retrieval layer respects the access model of the source data, or whether embedding everything into one index quietly flattened your permissions.

04

Agent containment

Blast radius. Standing credentials, autonomy limits, chained agent calls, and what happens when one step in the chain returns something hostile.

05

Model supply chain

Model and dependency provenance, third-party integrations, MCP servers, plugin surfaces, and the packages that arrived with the framework.

06

Guardrail testing

Whether the controls hold under adversarial pressure rather than in the demo, and whether failures are detected, logged and attributable afterwards.

Non-human identity

Boring plumbing. Highest yield.

Nobody wants to own this work, which is precisely why it is where the standing risk accumulates. There is no CVE for a five-year-old service account with domain admin and no owner.

Method

How an engagement runs.

Authorisation first, evidence throughout, and a report your engineers can act on without a translation layer.

01

Scope and authorisation

Targets, rules of engagement, testing windows, escalation contacts and written authorisation. Nothing starts without it, including for systems you own but a third party hosts.

02

Passive discovery

Mapping the real surface before touching it — exposure, technology, identities and integrations.

03

Active testing

Automated coverage where it earns its place, manual testing everywhere it matters. Business logic, access control and trust boundaries are hand-worked.

04

Verification

Every finding manually confirmed and evidenced. No unvalidated scanner output reaches your report.

05

Reporting

Risk-rated findings with reproduction steps, evidence, business impact and a specific fix. Executive summary that a board can read, technical detail an engineer can act on.

06

Retest

Remediation verified and the report updated, so you have something that stands up to an auditor or a customer security review.

Why Smartsec

Twenty years of adversarial thinking, pointed at new components.

01

Senior testers only

The person scoping your engagement is the person testing it. No pyramid, no junior delivery behind a senior sales conversation.

02

Manual where it counts

Tooling handles coverage. Judgement handles logic, chained flaws and trust boundaries — the findings that actually matter.

03

Reports engineers respect

Reproducible, evidenced, prioritised and specific. Written to be fixed from, not filed.

04

Ahead on the new surface

AI systems and machine identity are where the exposure is growing fastest and testing coverage is thinnest. We already work there.

Contact

Tell us what you've built.

Scoping conversation, fixed quote, agreed dates. No discovery call theatre.

Emailandy@smartsec.co.uk Phone+44 (0) 7516 143769
CompanySmartsec Information Security Ltd
BasedWakefield, West Yorkshire