Prompt InjectionAdvanced

When "Do Not" Is Not Deny: Security Rules in CLAUDE.md vs Built-In Controls

The paper compares natural-language safety rules written in CLAUDE.md files against Claude Code's built-in enforcement mechanisms, finding a significant gap between instructional constraints and executable controls. By analyzing 481 public files and validating findings through security practitioner review, the authors quantify how many stated security goals lack corresponding technical enforcement.

Why this matters

The paper compares natural-language safety rules written in CLAUDE.md files against Claude Code's built-in enforcement mechanisms, finding a significant gap between instructional constraints and executable controls. By analyzing 481 public files and validating findings through security practitioner review, the authors quantify how many stated security goals lack corresponding technical enforcement.

Check the original work

This explanation is Korpalis’s guide to the material, not a replacement for it. Read the publisher’s page for the full method, evidence and limitations.

Read the original source

Related research

Advanced

Breaking Claude Code Opus 5 Auto Mode

A researcher successfully broke Claude Code's auto mode defenses, which Anthropic designed to block prompt injection attacks in agentic workflows. The demonstration shows that safety mechanisms intended to protect autonomous agents remain vulnerable to determined adversaries.

Read summary →
Intermediate

Vulnerable Code Search: Transferable Attack for Code Language Models

Researchers demonstrate a transferable adversarial attack against code language models used in search tools by perturbing variable and function names while preserving code functionality, revealing how seemingly non-functional textual changes can manipulate neural retrieval systems. The attack generalizes across programming languages and between different models, highlighting a semantic-level vulnerability in code understanding.

Read summary →
Intermediate

Hidden Prompts Trick AI Into False Email Summaries

Invisible HTML markup can be injected into emails to deceive AI summarization tools into generating false or malicious summaries, exploiting the models' inability to detect content hidden from human readers.

Read summary →