Prompt injection
Direct and indirect. Content that reaches the model from documents, retrieved chunks, tickets, emails or web pages or anywhere the system treats untrusted input as instruction.
Offensive security testing
Web applications, infrastructure and full red team engagements plus the two surfaces almost nobody is testing yet, the AI systems you've deployed, and the machine identities running underneath them.
Capability
Findings you can act on, with reproduction steps, evidence and a fix
Manual, authenticated testing across roles and trust boundaries. Business logic and access control flaws that scanners structurally cannot find, plus full API and integration coverage.
External and internal network testing, cloud configuration, segmentation validation, privilege escalation paths and build reviews. What is actually exposed, not what the asset register claims.
Objective-led adversary emulation against people, process and technology. Run stealth to test detection and response, or purple to build it with the tradecraft mapped so your defenders can replay every step.
Adversarial testing of LLM deployments, agentic systems, RAG pipelines and MCP integrations. Prompt injection, tool and function abuse, retrieval data leakage, model supply chain, and whether your guardrails hold under pressure.
Discovery and assessment of service accounts, tokens, keys and agent identities. Who created them, what they can reach, whether anyone owns them, and how they get retired. Unglamorous, and the most over-permissioned surface you have.
The gap
Every organisation deployed agents, copilots and retrieval pipelines in the last eighteen months. Almost none of them extended their testing or their identity governance to cover it. That gap is measurable.
None of this is a new discipline. It is appsec and trust boundaries applied to a component you added last year.
AI system security
An agent with a tool, a credential and a text input is a remote code execution surface with better manners. We test it as one.
Direct and indirect. Content that reaches the model from documents, retrieved chunks, tickets, emails or web pages or anywhere the system treats untrusted input as instruction.
What the agent can be persuaded to call, with what arguments, on whose authority. Where the tool boundary is enforced, and where it is merely described in a system prompt.
Whether the retrieval layer respects the access model of the source data, or whether embedding everything into one index quietly flattened your permissions.
Blast radius. Standing credentials, autonomy limits, chained agent calls, and what happens when one step in the chain returns something hostile.
Model and dependency provenance, third-party integrations, MCP servers, plugin surfaces, and the packages that arrived with the framework.
Whether the controls hold under adversarial pressure rather than in the demo, and whether failures are detected, logged and attributable afterwards.
Non-human identity
Nobody wants to own this work, which is precisely why it is where the standing risk accumulates. There is no CVE for a five-year-old service account with domain admin and no owner.
Method
Authorisation first, evidence throughout, and a report your engineers can act on without a translation layer.
Targets, rules of engagement, testing windows, escalation contacts and written authorisation. Nothing starts without it, including for systems you own but a third party hosts.
Mapping the real surface before touching it — exposure, technology, identities and integrations.
Automated coverage where it earns its place, manual testing everywhere it matters. Business logic, access control and trust boundaries are hand-worked.
Every finding manually confirmed and evidenced. No unvalidated scanner output reaches your report.
Risk-rated findings with reproduction steps, evidence, business impact and a specific fix. Executive summary that a board can read, technical detail an engineer can act on.
Remediation verified and the report updated, so you have something that stands up to an auditor or a customer security review.
Why Smartsec
The person scoping your engagement is the person testing it. No pyramid, no junior delivery behind a senior sales conversation.
Tooling handles coverage. Judgement handles logic, chained flaws and trust boundaries — the findings that actually matter.
Reproducible, evidenced, prioritised and specific. Written to be fixed from, not filed.
AI systems and machine identity are where the exposure is growing fastest and testing coverage is thinnest. We already work there.
Contact
Scoping conversation, fixed quote, agreed dates. No discovery call theatre.