Anime, manga, and games, with a take · A Yukimedia publication

← all stories other 1 sources · 44m ago ·

Fake Chain of Thought Tricks LLMs, Paper Argues Training Cannot Fix It

The paper identifies a structural flaw in how LLMs parse message provenance, suggesting that safety training alone cannot prevent prompt-injection-style extraction.

Reporting from 1 source: ASCII.jp.

Fake Chain of Thought Tricks LLMs, Paper Argues Training Cannot Fix It

A paper submitted to ICML shows that large language models identify the source of user, system, and chain-of-thought messages only by writing style, not by tags. Exploiting this, a fake chain of thought can extract information such as cocaine manufacturing methods or aircraft hacking techniques. The researchers argue this is a fundamental flaw that training cannot resolve.

The paper, submitted to ICML, describes a mechanism where LLMs judge the origin of a message by its style rather than any structural tag. An attacker who mimics the style of a chain of thought can therefore pass off malicious instructions as internal reasoning.

In tests, this fake chain of thought was enough to extract instructions for manufacturing cocaine and for hacking aircraft. The research team argues that because the flaw is structural, it cannot be resolved through training alone.

Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources