← All insights

The Hidden Risks of AI Agents

Published: 26 August 2025

The Hidden Risks of AI Agents

Why AI Agent Security Matters

The more autonomy an AI agent is given, the larger the attack surface it opens. Advanced agents, powered by the latest large language models (LLMs), now reach into business processes across nearly every sector, from finance automation to healthcare data management and customer support. They promise efficiency and scale. But in my view, as their autonomy grows, so above all does the risk, particularly around cybersecurity and regulatory compliance.

To move beyond speculation and assess the real security posture of these systems, the world’s largest public red-teaming competition was launched earlier this year. The results are revealing and sobering. For decision-makers in the executive team, understanding where the real risks and opportunities lie is essential.

The Red-Teaming Competition

A month-long global red-teaming competition was organised by leading research institutions and backed by industry, including the UK AI Security Institute, OpenAI, Anthropic and Google DeepMind. Nearly 2,000 security experts (“red teamers”) attacked 22 of the most advanced AI agents across 44 realistic business scenarios. A prize pool of USD 171,800 underlined the scale and significance of the initiative.

Why This Approach?

Unlike academic tests or isolated benchmarks, this setup mirrored real-world deployment conditions:

  • Agents were tested with simulated access to sensitive data and tools, as in genuine enterprise environments.
  • Attackers could use direct and indirect methods, including insider threats, supply-chain risks and sophisticated attack techniques.

Areas Tested

  • Confidentiality: protection of sensitive information, such as patient or client data.
  • Integrity: protection against manipulation of processes or outputs.
  • Policy compliance: adherence to regulations and company policy, even under attack.

The Results

The findings can be summarised as follows:

  • Volume of attacks: over 1.8 million attack attempts in a single month.
  • Successes: more than 62,000 successful policy violations, including unauthorised data disclosures, illegal transactions and regulatory breaches.
  • Universality: all attacks could be repeated successfully (100% “behaviour ASR”).
  • Transferability: attacks on one model often worked on others too, an indication of shared vulnerabilities and systemic risk.

Types of Attack

  • Direct prompt injection: manipulation through the user interface.
  • Indirect injection: malicious instructions hidden in documents, emails or external APIs. Particularly effective at bypassing conventional security controls, with a success rate of up to 29.8% in leaking confidential information.

“With a single request, models show policy violations in 20% to 60% of tested behaviours. At ten specific requests, the success rate for most models approaches nearly 100%.”

Consequences for Organisations

  • Large, capable models are not inherently safer: there was no clear link between model size, capability or vendor and resilience. Even flagship models from leading vendors were compromised.
  • Shared weaknesses mean systemic risk: transferable attacks can escalate quickly into industry-wide problems, a matter for business continuity, not only for IT.
  • Compliance cannot be assumed: high policy-violation rates show that organisations cannot trust out-of-the-box compliance blindly. This matters most for boards, which carry ultimate liability.

Practical Measures for Secure Deployment

  1. Security by design: anchor security as a strategic pillar, from concept through to operation, before a project starts, not after.
  2. Testing: continuous testing against open standards such as the new Agent Red Teaming (ART) framework.
  3. Layered security and defence: combine filtering, policy enforcement, real-time monitoring and human oversight.
  4. Vendor due diligence: demand transparency on security testing, incident response and red-teaming results.
  5. Board-level responsibility: put AI risk firmly on the agenda and align it with cybersecurity, compliance and digital strategy.

Key Takeaways

AI agents are rapidly becoming a core part of modern business infrastructure. This large-scale study shows that the technology remains highly vulnerable today, to simple attacks as well as sophisticated ones.

“We consistently observe near-100% attack success rates across models and deployment scenarios… an urgent, real risk that must be addressed before broader deployment.”

Anyone deploying AI agents needs to act on the following:

  • Treat AI security as a leadership responsibility.
  • Demand ongoing, independent security audits.
  • Implement robust, layered controls.
  • Prepare for regulatory scrutiny, at both executive and board level.

I am convinced that organisations which act proactively and transparently can capture the benefits of AI while keeping the growing risks under control.

For further reading, see the full report “Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition” (https://arxiv.org/abs/2507.20526v1).