AI TechnologyCiscoJul 26, 2026 09:20 UTC

88% of AI Models Breached by Multi-Turn Attacks

Amy Chan, Head of AI Threat Intelligence at Cisco, presented research at the 2026 VB Transform Conference showing that "multi-turn attacks"—which manipulate AI models through multiple conversation exchanges—can breach AI model safeguards with a probability of up to 88.3%. Based on a large-scale evaluation of 15 major AI models, the research also revealed that single-turn testing alone cannot capture this risk. In a corporate survey, 54% of enterprises have experienced or prevented agent-related security incidents, making security enhancement of AI agents an urgent priority across the industry.

88% of AI Models Breached by Multi-Turn Attacks

Multi-turn attacks have emerged as a critical threat to AI models. According to research presented by Amy Chan, Head of AI Threat Intelligence and Security Research at Cisco, at the 2026 VB Transform Conference (an agent-type AI security panel), multi-turn attacks—which manipulate conversations through multiple exchanges—can breach AI model safeguards with a probability of up to 88.3%. This figure was obtained from 6,986 attack attempts targeting 15 major AI models. While breakthrough rates ranged from 7.89% to 88.3% depending on the model, all models exhibited non-negligible risks.

Multi-turn attacks differ from single malicious queries (single-turn attacks) by guiding AI through multiple conversation exchanges to elicit harmful information or inappropriate actions that would normally not be output. Chan explains that this approach "more closely resembles how we actually interact with AI models, agents, and applications," making it a more accurate reflection of real-world attack scenarios. In contrast, single-turn testing alone cannot detect vulnerabilities arising from accumulated conversation exchanges and risks overestimating model safety. In fact, the research revealed that the two testing methods did not even produce consistent model risk rankings.

This research, co-authored by Chan and Nicholas Conley, is based on large-scale evaluation combining 30,090 single-turn prompts and 6,986 multi-turn attacks. Cisco currently publishes evaluation results for 105 AI models as an "LLM Security Leaderboard," continuously measuring attack resilience for each model. Additionally, Chan noted that the company is developing a framework in which AI agents autonomously generate and evaluate attack scenarios.

The current state of enterprise security measures has also been revealed. According to a survey conducted by VentureBeat in June 2026 of 107 corporate representatives, 54% had experienced or prevented agent-related security incidents. However, only 32% of enterprises assign dedicated scoped credentials to each agent, and merely 30% isolate high-risk agents in sandboxed (isolated secure execution) environments. 82% of enterprises rely on cloud service providers' standard features as their primary security means, revealing inadequate implementation of additional defensive layers.

In response to this situation, major security companies have pursued successive acquisitions. Palo Alto Networks completed its acquisition of CyberArk for $25 billion in February 2026, and Crowdstrike agreed to acquire SGML for $740 million in January of the same year. According to reports, Cisco announced the acquisition of Astrix Security at a scale of $400 million. All these acquisitions aim to strengthen identity and isolation mechanisms that control "who and what" can access AI systems, representing efforts by many enterprises to fill security gaps still under construction.

Chan brings approximately 20 years of experience spanning cybersecurity operations, government agencies, and the military. She led the cyber threat intelligence team at JP Morgan Chase and served as senior staff to the U.S. House Committee on Foreign Affairs and as a naval reserve officer. She currently teaches cybersecurity and emerging threats at Middlebury Institute of International Studies.

Today's announcement underscores that organizations relying solely on single-turn testing for AI agent security evaluation face unseen risks. As AI evolves from a tool answering one-off questions to an "agent" autonomously performing multiple tasks, attackers are similarly shifting toward methods exploiting conversational flow. Going forward, multi-turn attack response and enhanced privilege management are expected to become priority issues for all organizations operating agent-type AI.

#AISecuritySecurity#AIAgent#LLM#CyberSecurity#Vulnerability#GenerativeAI#Cisco
AI issue Staff

This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.

Comments

Log in to comment