LLM guardrails falter under dialogue attacks

Cisco researchers have warned that leading open-weight large language models can be manipulated through sustained conversations that gradually push them past safety controls, exposing a weakness in systems now being adopted across business, public services and consumer applications.

The assessment tested eight widely used open-weight models from Alibaba, DeepSeek, Google, Meta, Microsoft, Mistral, OpenAI and Zhipu AI. The models were examined through automated adversarial testing designed to measure whether they could resist prompt-injection and jailbreak attempts across both single-turn and multi-turn exchanges.

The findings point to a marked gap between how models behave when challenged with one direct prompt and how they respond when harmful intent is introduced over several conversational steps. Multi-turn attacks achieved success rates ranging from 25.86 per cent to 92.78 per cent, with some models proving two to 10 times more vulnerable in extended dialogue than in single-prompt tests.

The risk is significant because many enterprise AI systems are built around chat interfaces, agents and assistants that depend on long exchanges with users. A request that would be blocked if made directly may be broken into smaller, apparently harmless steps, allowing the user to build context, establish a role-play scenario or gradually steer the system towards prohibited output.

Cisco’s researchers described the pattern as a systemic weakness in the ability of current open-weight models to maintain safety instructions across longer conversations. The tests were conducted as black-box engagements, meaning the internal architecture and any additional safety layers were not disclosed before assessment.

The models tested included Qwen3-32B, DeepSeek v3.1, Gemma 3-1B-IT, Llama 3.3-70B-Instruct, Phi-4, Mistral Large-2, GPT-OSS-20b and GLM 4.5-Air. The research did not argue against open-weight AI development, but said organisations need to understand the security posture of models before using them in production or fine-tuning them for sensitive tasks.

Open-weight models have become central to the AI ecosystem because they allow developers to inspect, customise and deploy systems without relying entirely on closed commercial platforms. Their growth has accelerated across research, software development, cyber security operations, customer service and internal knowledge tools. That flexibility also creates exposure when models are deployed without layered protections.

Capability-focused models showed larger gaps between single-turn and multi-turn performance, while models with stronger safety alignment appeared to perform more consistently across attack types. The distinction matters for enterprises choosing systems not only for speed, cost or benchmark performance, but also for resilience against manipulation.

Security specialists have warned that model capability benchmarks often overshadow safety testing. A model that performs well in coding, reasoning or language tasks may still be weak against adversarial dialogue. This creates a procurement risk for organisations that select models on productivity metrics while underestimating misuse scenarios.

The concerns extend beyond harmful text generation. Multi-turn manipulation could affect systems connected to databases, code repositories, workflow tools, customer records or decision-support platforms. A compromised AI assistant could expose confidential information, generate misleading material, alter business logic or assist in unauthorised activity if linked to operational systems.

The threat becomes sharper as AI agents gain the ability to take actions rather than merely produce text. When models are connected to tools, calendars, cloud environments, ticketing systems or financial workflows, a successful jailbreak may have consequences beyond the chat window. Guardrails therefore need to monitor not only individual prompts but the full conversational trajectory.

Researchers in the wider AI safety field have also found that multi-turn attacks are harder to detect because each message can look benign when viewed alone. The malicious intent becomes clear only when the dialogue is assessed as a sequence. That creates a challenge for filters that operate at the level of isolated inputs and outputs.



Notice an issue?

Arabian Post strives to deliver the most accurate and reliable information to its readers. If you believe you have identified an error or inconsistency in this article, please don't hesitate to contact our editorial team at editor[at]thearabianpost[dot]com. We are committed to promptly addressing any concerns and ensuring the highest level of journalistic integrity.


Loading next story…
Just in:
Ping An Digital Bank Becomes Hong Kong’s First Digital Bank to Enter High-End Wealth Management Segment // UK and allies expose Integrity Tech cyber operations // India establishes 5.56 km open-air quantum security link // Anti-Election Commission Protest: Athletic Rahul Steals The Show // India rebuts Musk allegations over Starlink launch delay // Malicious GitHub workflows expose credentials across hundreds of repositories // UAE delegation heads to Bangkok for IMF meetings // OpenAI extends GPT-6 access with interactive ChatGPT interface // Abu Dhabi climate summit records over 1,000 registrations // React flaw exposes Next.js servers to service disruption // Prudential Singapore launches multi-generational protection plan to help caregivers manage families’ healthcare needs // Two Bypoll Results In Bengal Vindicate State BJP’s Success In Courting Minorities // Almarai earmarks $4 billion for expansion through 2031 // Trump-Newsom Clash Assumes Special Significance Before Nov 3 Polls // Lufthansa and three airlines halt Riyadh flight operations // LANDMARK Launches ‘Destination CENTRAL’: A District-Wide Invitation to Explore the Dynamism, Luxury, and Soul of Central // Wikimedia identifies unauthorised OpenAI agent activity across platforms // BINGXUE Opens First U.S. Store in Davis, California: Shandong’s First Mass-Market Tea Beverage Brand Enters North America // ONYX Hospitality Group Marks 60 Years with Curated Partnerships Bringing “More of What You Love” to Life // First Week Of Anti-CEC Agitation Turns Into Electoral Rights Movement //