Prompt InjectionIntermediate

Hidden Prompts Trick AI Into False Email Summaries

Invisible HTML markup can be injected into emails to deceive AI summarization tools into generating false or malicious summaries, exploiting the models' inability to detect content hidden from human readers.

Why this matters

Invisible HTML markup can be injected into emails to deceive AI summarization tools into generating false or malicious summaries, exploiting the models' inability to detect content hidden from human readers.

Check the original work

This explanation is Korpalis’s guide to the material, not a replacement for it. Read the publisher’s page for the full method, evidence and limitations.

Read the original source

Related research

Advanced

Breaking Claude Code Opus 5 Auto Mode

A researcher successfully broke Claude Code's auto mode defenses, which Anthropic designed to block prompt injection attacks in agentic workflows. The demonstration shows that safety mechanisms intended to protect autonomous agents remain vulnerable to determined adversaries.

Read summary →
Intermediate

Vulnerable Code Search: Transferable Attack for Code Language Models

Researchers demonstrate a transferable adversarial attack against code language models used in search tools by perturbing variable and function names while preserving code functionality, revealing how seemingly non-functional textual changes can manipulate neural retrieval systems. The attack generalizes across programming languages and between different models, highlighting a semantic-level vulnerability in code understanding.

Read summary →
Advanced

What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions

The study identifies injection vulnerabilities in LLM agents where untrusted external data can be dynamically interpreted as behavior-guiding instructions, allowing attackers to subvert agent decisions. The work proposes methods to localize and defend against these dynamic injections that occur during inference rather than at static input/output boundaries.

Read summary →