Anthropic’s Opus 5 Model Shows Major Resistance to Prompt Injection

🤖 AI-GENERATED✓ HUMAN-REVIEWED⚡ Posted 18 minutes after it broke⏱ 2 min read📡 Simon Willison

The short version

Anthropic's latest Opus 5 model is reported to be its most resistant to prompt injection attacks, a key security advance for generative AI.

In a significant development for AI safety, a key figure at Anthropic has highlighted the enhanced security of the company’s newest large language model. According to a quote from Boris Cherny shared by Simon Willison, Opus 5 represents a major step forward in defending against a critical class of AI vulnerabilities.

Key takeaways

  • Anthropic’s Opus 5 is described as the company’s “least prompt injectable model yet.”
  • The improvement is based on extensive evaluation across prompt injection (PI) evals and red teaming exercises.
  • Cherny notes the finding is “a bit buried” in the model’s official System Card documentation on page 73.
  • The announcement was made in a blog post dated July 25, 2026.
  • The advancement is positioned as a particularly exciting development beyond standard benchmark scores.

The security milestone

The core of the announcement is a focused claim about security robustness. Boris Cherny states that “across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully.” Prompt injection is a technique where a user crafts a specific input, or “prompt,” designed to override a language model’s original instructions or safety guidelines, potentially making it produce harmful, biased, or unintended outputs. By making the model highly resistant to such attacks, Anthropic is addressing a fundamental security concern that has plagued the deployment of generative AI systems in real-world, potentially adversarial environments. The fact that this resilience was validated through both systematic evaluations (“PI evals”) and adversarial simulation (“red teaming”) suggests a rigorous testing process.

Why it matters

Improving resistance to prompt injection is critical for the safe and reliable deployment of advanced AI. As language models are integrated into more applications—from customer service bots to code assistants—the risk of malicious actors manipulating them increases. A model that is “very hard to prompt inject” is inherently more trustworthy and secure, reducing the potential for misuse. Cherny’s emphasis on this feature over other evaluation scores underscores a shift in priority for some in the AI field, where practical security against manipulation can be as important as raw performance on standard benchmarks. This development by Anthropic, a leading AI safety-focused company, sets a new bar for defensive capabilities in state-of-the-art models and could influence safety standards across the industry.

📡 Original reporting: Simon Willison. AI Craft Technologies’ news engine summarised and rewrote this story in our own words; facts are drawn from the linked source.

⚙️ How this article was made — fully automated
01📡 ScanOur engine watches trusted AI & tech sources in real time.
02🤖 WriteAI drafts an original summary in the ACT house style.
03🎨 IllustrateA custom hero image is generated for every story.
04📤 PublishReviewed, posted, and shared to social — hands-free.

This is a live demo of the ACT News Factory engine. Want one running on your own site? See our services →

Share this project

Leave a Reply