Ki Modelle Sind Anfallig Fur Wiederholte Angriffe

Large Language Models Show Vulnerabilities Against Repeated AI Attacks

Notice: This article was created with AI.

What’s It About?

Cisco researchers have uncovered serious weaknesses in the security architecture of large language models. The central finding: current evaluation methods capture the actual risk only inadequately, as they are based on single requests. Realistic attack scenarios, however, work with sequential, strategically built inputs — so-called multi-turn attacks. This method achieves dramatically higher success rates at bypassing safety measures.

In the study, 15 widely used AI models were systematically tested for their resilience against multi-stage attacks. The results clearly show: what appears secure against individual requests becomes increasingly vulnerable through repeated, cleverly formulated inputs.

Background & Context

The test results reveal considerable differences between simple and complex attack patterns. Google’s Gemini 3 Pro illustrates this particularly clearly: while simple attacks succeeded in only 18.10 percent of cases, the rate rose to 73.35 percent for multi-stage attacks. This pattern appeared across models, with different configurations – with or without reasoning enabled – leading to considerable variation in security performance.

The attackers use a variety of techniques: role-play, contextual manipulation and shifting boundaries step by step. Open-source models proved just as vulnerable as proprietary solutions. The researchers stress that standardised security benchmarks do not reflect reality, because real attackers proceed iteratively and exploit weaknesses systematically.

The findings have far-reaching implications for upcoming regulatory frameworks. Cisco calls on developers and providers to publish security metrics more transparently and to establish evaluation procedures that account for both attack forms. This could trigger a fundamental reassessment of existing security standards across the AI industry.

What Does This Mean?

  • Existing security evaluations for language models don’t adequately reflect real threat scenarios, giving operators a false sense of security
  • Developers must fundamentally overhaul their testing procedures and integrate multi-stage attack scenarios into evaluations
  • Organizations deploying AI models should implement additional protection layers, as the models themselves are more vulnerable than previously assumed
  • Regulatory authorities are expected to impose stricter requirements on AI system security certification
  • Transparency in security metrics needs to increase significantly to enable meaningful comparisons between models

Sources

This article was created with AI assistance and is based on the listed sources as well as the language model’s training data.

Further Reading: From Text Generator to Digital Employee: How AI Is Changing the World in Four Stages

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top