Copilot Tricked Into Revealing Its Own Vulnerabilities
The attack shows that AI assistants can be manipulated through their own reasoning to reveal and exploit hidden features, a method the researchers say applies to any AI.
Reporting from 1 source: GIGAZINE.
Varonis Threat Labs researchers tricked Microsoft's Copilot into disclosing its own security flaws by cooperating with it instead of coercing it. They found a disabled URL parameter and an undocumented 'autorun=1' parameter, enabling automatic prompt execution. Three vulnerabilities, named CoSnitch, were reported to Microsoft, allowing data theft and memory pollution.
Varonis Threat Labs researchers extracted vulnerabilities from Microsoft's Copilot by showing a cooperative attitude rather than forcing disclosure. The root cause was a disabled URL query parameter '?q=' in Copilot's web interface, which could pre-enter text into the chat input field.
By asking Copilot to reason about its own mechanisms, the researchers got it to list disabled parameters and reveal an undocumented 'autorun=1' parameter. Under certain session conditions, this parameter caused prompts to execute automatically on page load without user interaction or UI confirmation.
The researchers named the three vulnerabilities CoSnitch and reported them to Microsoft. Exploiting them, an attacker could read past conversations, connected apps, and Copilot's memory, and pollute stored prompts to inject false information in future sessions.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.
Sources
- GIGAZINE Copilotが騙されて自分自身のハッキング方法を教えてしまう