A fundamental flaw leaves LLMs strikingly vulnerable to attack

Share

Researchers have identified a fundamental vulnerability in large language models where LLMs identify instruction roles based on text style rather than structural tags, making them impossible to fully secure against attacks. The team demonstrated extracting dangerous information like cocaine synthesis and aircraft sabotage instructions from leading models including OpenAI, Anthropic, Alibaba, and DeepSeek.


Source: Read the original article