Anthropic released Claude Opus 5 on July 24, 2026. Opus 5 is an upgrade over Opus 4.8, with gains concentrated in agentic coding, computer use, long-horizon knowledge work, and math/scientific reasoning.
The official documentation makes one key overall judgment: Opus 5 does not exceed Claude Fable 5, Anthropic's strongest current general-purpose model, in overall capability, and its alignment risk rating stays at very low. Even so, it remains under ASL-3 safeguards (same as Opus 4.8) — it crosses the CB-1 (non-novel bioweapon capability threshold) but not the CB-2 (novel bioweapon threshold).
Below, in the order "capability benchmarks → safety & alignment → risk assessment → model welfare," we've compiled the information worth keeping from the official materials. All numbers are quoted verbatim from the official release.
Released July 24, 2026, inheriting Opus 4.8's ASL-3 safeguards.
1. Basic Model Information
Opus 5 was trained on public internet data, public/private datasets, and synthetic data generated by other models, with multiple rounds of post-training and fine-tuning after pretraining, aligned to Claude's constitution. It automatically matches the user's input language; output quality varies by language, and output is plain text. Knowledge cutoff is May 2026.
Relative to Fable 5, Opus 5 opens source-code vulnerability discovery at all access tiers (including GA), while vulnerability discovery on compiled binaries remains prohibited. The rationale: source-code vulnerability discovery favors defenders (developers write more secure code), while binary vulnerability discovery favors attackers.
2. Capability Overview: Core Benchmarks at a Glance
The official capability summary covers all major dimensions: coding, reasoning, agents, long context, multimodality, and professional tasks. Opus 5 takes SOTA on several metrics, but the gap to Fable 5 / Mythos 5 is not large. Opus 5 also scores 96.0% on SWE-bench Verified, a 500-problem subset verified as solvable by human engineers.
Coding & engineering
Opus 5 runs adaptive thinking @ max effort, averaged over 5 trials. A zero-height bar means the official report did not publish a score for that model.
- Claude Opus 5
- Claude Opus 4.8
- Claude Fable 5
- GPT-5.6 Sol
- SWE-bench Pro,Claude Opus 5:79.2%
- SWE-bench Pro,Claude Opus 4.8:69.2%
- SWE-bench Pro,Claude Fable 5:80%
- SWE-bench Pro,GPT-5.6 Sol:64.6%
- SWE-bench Multilingual,Claude Opus 5:89.5%
- SWE-bench Multilingual,Claude Opus 4.8:84.4%
- SWE-bench Multilingual,Claude Fable 5:86.6%
- SWE-bench Multilingual,GPT-5.6 Sol:0%
- SWE-bench Multimodal,Claude Opus 5:59.4%
- SWE-bench Multimodal,Claude Opus 4.8:38.4%
- SWE-bench Multimodal,Claude Fable 5:54.1%
- SWE-bench Multimodal,GPT-5.6 Sol:0%
- DeepSWE v1.1,Claude Opus 5:68.8%
- DeepSWE v1.1,Claude Opus 4.8:59%
- DeepSWE v1.1,Claude Fable 5:69.7%
- DeepSWE v1.1,GPT-5.6 Sol:72.7%
- FrontierCode Main,Claude Opus 5:53.4%
- FrontierCode Main,Claude Opus 4.8:46.5%
- FrontierCode Main,Claude Fable 5:53.5%
- FrontierCode Main,GPT-5.6 Sol:47.5%
- FrontierCode Extended,Claude Opus 5:63.6%
- FrontierCode Extended,Claude Opus 4.8:59.6%
- FrontierCode Extended,Claude Fable 5:0%
- FrontierCode Extended,GPT-5.6 Sol:60.6%
- FrontierBench v0.1,Claude Opus 5:44.4%
- FrontierBench v0.1,Claude Opus 4.8:18.7%
- FrontierBench v0.1,Claude Fable 5:33.7%
- FrontierBench v0.1,GPT-5.6 Sol:37.5%
- AutomationBench,Claude Opus 5:26%
- AutomationBench,Claude Opus 4.8:17%
- AutomationBench,Claude Fable 5:17.4%
- AutomationBench,GPT-5.6 Sol:18.1%
Reasoning, agents & professional tasks
X axis is Opus 4.8, Y axis is Opus 5; the further upper-left, the bigger the jump. ARC-AGI-3 (1.5% → 30.2%) is the standout outlier. References: ArxivMath — Fable 5: 87.4, GPT-5.6 Sol: 90.4; BrowseComp — Fable 5: 66.1, GPT-5.6 Sol: 62.6; HealthBench Pro (length-adj) — Fable 5: 66.0, GPT-5.6 Sol: 60.5; HLE — Fable 5: 56.5 without tools, 63.9 with tools; ARC-AGI-2 — GPT-5.6 Sol: 92.5.
- ArxivMath (no tools):90.8%
- HLE no tools:56.3%
- HLE with tools:64.7%
- BrowseComp:70.6%
- HealthBench Pro:59.8%
- ARC-AGI-1:97.5%
- ARC-AGI-2:90.4%
- ARC-AGI-3:30.2%
Knowledge-work ELO
GDPval-AA v2 and AA-Briefcase scores are from artificialanalysis.ai; Opus 5 shown at max effort.
- GDPval-AA v2 · Opus 5:1,861
- GDPval-AA v2 · Fable 5:1,747
- GDPval-AA v2 · GPT-5.6 Sol:1,736
- GDPval-AA v2 · Opus 4.8:1,593
- AA-Briefcase · Opus 5:1,720
- AA-Briefcase · Fable 5:1,574
- AA-Briefcase · GPT-5.6 Sol:1,505
- AA-Briefcase · Opus 4.8:1,346
Source: official capability summary and per-item notes. Unless otherwise noted, Opus 5's configuration is adaptive thinking @ max effort, averaged over 5 trials, with a context window of no more than 1M. FrontierBench is shown at xhigh (44.4; max is 43); GDPval-AA v2 is 1861 at max and 1827 at xhigh; AA-Briefcase is 1720 at max, 1693 at xhigh, and 1606 at high.
3. The Key Capability Numbers
1. IMO 2026: perfect-score gold medal
Opus 5 scored 42/42 on the 2026 International Mathematical Olympiad (IMO 2026), against a gold-medal line of 29/42. Anthropic had Opus 5 generate 4 independent solutions for each of the 6 problems, jointly graded by three frontier models — Gemini 3.1 Pro, Claude Opus 4.6, and Claude Mythos Preview; all 24 solutions were judged correct. Human experts independently graded the pre-designated solution for each problem and also gave 7/7 across the board.
2. ARC-AGI: from near-perfect to a 4× leap
Opus 5 scored 30.16% (high effort) on ARC-AGI-3, the ARC Prize Foundation's new benchmark — roughly 4× the previous best on the official leaderboard; compare Claude Opus 4.8 (1.52%, high) and GPT-5.6 Sol (7.78%, max). ARC-AGI-3 is an interactive game environment where the model must figure out the rules and clear levels with no instructions at all. ARC-AGI-1/2 are near perfect; the previous-generation Opus 4.7 scored 75.83% on ARC-AGI-2.
ARC-AGI-3 is an interactive game environment with no instructions; the model must discover the rules itself. Opus 5 scored at high effort.
3. OSWorld 2.0: the strongest computer-use model to date
Opus 5 achieved a 70.57% first-attempt success rate on OSWorld 2.0 (1080p resolution, up to 500 steps) — Anthropic's strongest computer-use model to date, with significantly improved efficiency.
4. Multi-agent BrowseComp reaches 93.6%
On both BrowseComp (deep web search) and ProgramBench (program reconstruction), multi-agent configurations are Pareto-dominant over single-agent:
- A 10-agent team reached 93.6% on BrowseComp, +3.1pp above the strongest single-agent baseline;
- Relative to the single-agent 10M-token baseline, 5-agent and 10-agent teams achieved latency speedups of 5.6× and 5.9× respectively;
- On ProgramBench, a 5-agent team reached the same score (0.6) with a 2.2× latency speedup.
On BrowseComp and ProgramBench, multi-agent teams trade more tokens for lower latency, pushing the score–cost frontier further out.
5. HealthBench / HealthBench Professional: highest in the Claude family
- HealthBench raw score 67.1% (Mythos 5: 62.5%, Opus 4.8: 58.8%), 57.8% after length adjustment;
- HealthBench Professional raw score 73.4% (Mythos 5: 70.3%, Opus 4.8: 60.3%), 59.8% after length adjustment.
6. Chartography / BenchCAD: large jumps with tool support
Chartography (professional chart understanding) and BenchCAD Vision2Code (CAD view to code) both point to the same trend: agentic tool use is more cost-effective than simply extending thinking.
Gray is without tools, color is with tools. With tools, Opus 5 overtakes Mythos 5 on BenchCAD.
GDP.pdf (professional PDF understanding): Opus 5 is 83.4% without tools and 85.5% with tools; Mythos 5 is 81.8% / 87.3%, still ahead in the tool setting.
4. Coding, in Depth
SWE-bench Verified is the 500-problem human-verified subset; ProgramBench is long-context program reconstruction in a 1M-token window (5 episodes).
- SWE-bench Verified:96%
- SWE-bench Pro:79.2%
- SWE-bench Multilingual:89.5%
- SWE-bench Multimodal:59.4%
- DeepSWE v1.1:68.8%
- FrontierCode 1.1 Main:53.4%
- FrontierCode 1.1 Extended:63.6%
- FrontierBench v0.1:44.4%
- ProgramBench (5 episodes):93%
Additional context:
- SWE-bench Pro: from actively maintained large repositories, multi-file diffs, little public ground-truth leakage;
- DeepSWE v1.1: 113 long-horizon software engineering tasks, averaged over 5 runs; written from scratch to avoid contamination;
- FrontierCode 1.1: from Cognition, 150 real PR tasks; Opus 5 is second on the main table (behind only Fable 5 at 53.5%);
- FrontierBench v0.1: successor to Terminal-Bench 2.1, 74 research/engineering-oriented tasks; Fable 5: 33.7%; Opus 4.8: 18.7%; GPT-5.6 Sol: 37.5%;
- ProgramBench: Opus 4.8 scored 90%, Mythos 5 scored 93%.
FrontierBench v0.1 also reports safety classifier trigger rates: Opus 5 triggered the classifier on 5% of API calls and 4% of trials, falling back to Opus 4.8; Fable 5 fell back on 42% of API calls and 26% of trials. GPT-5.6 Sol triggered no safety classifier at all.
5. Agentic Search and Multi-Agent Orchestration
All nine dimensions are Opus 5 results. Legal Agent Benchmark shows the all-pass rate (1235 problems, ±0.48, n=5); AutomationBench at max effort.
- MCP Atlas:85.8%
- Toolathlon (Pass@1):80.6%
- OfficeQA:78.1%
- BrowseComp:70.6%
- OfficeQA Pro:66.9%
- HLE with tools:64.7%
- HLE no tools:56.3%
- AutomationBench:26%
- Legal Agent (all-pass):23.6%
For comparison: HLE no tools — Opus 4.8: 49.8, Fable 5: 56.5; BrowseComp — Opus 4.8: 55.7, Fable 5: 66.1, GPT-5.6 Sol: 62.6; MCP Atlas — Opus 4.8: 82.2% (claim coverage 89.1%); OfficeQA / OfficeQA Pro — Mythos 5: 79.0% / 67.1%; Toolathlon — Opus 4.8: 79.9, Mythos 5: 79.3, Sonnet 5: 74.7; AutomationBench — Opus 4.8: 17.0, Fable 5: 17.4, GPT-5.6 Sol: 18.1.
Other official data points:
- DeepSearchQA (F1): 900 multi-hop deep-search tasks across 17 domains; see the official chart;
- DRACO: from Perplexity, 100 real deep-research tasks; see the official chart;
- Legal Agent Benchmark: mean criterion-pass 93.74%; Harvey self-reported all-pass 11.7%;
- AutomationBench: 24% at medium effort ($0.89/task).
Anthropic specifically highlights the cost/benefit curves of multi-agent configurations (5-agent team, 10-agent team, async subagents): on BrowseComp and ProgramBench, multi-agent trades more tokens for lower latency, pushing the score–cost frontier further out.
6. Cyber Capability: Ahead of Opus 4.8 Across the Board, but Still Behind Mythos 5
Opus 5 received no dedicated cyber training; its related capabilities come from general capability spillover. The official report covers 5 cyber evaluations, plus external red teaming by UK AISI (the UK AI Security Institute).
ExploitBench (V8 engine exploitation, 41 environments, 5 trials per environment)
Opus 5 is far ahead of Opus 4.8 and close to Mythos 5.
- Claude Mythos 5:10.8
- Claude Opus 5:10.14
- Claude Opus 4.8:5.56
- Claude Sonnet 5:4.18
Sonnet 5 scored 0 and is omitted. AutoNudge Cap%: Mythos 5 is 78, Opus 5 is 70, Opus 4.8 is 40, Sonnet 5 is 31.
- Claude Mythos 5
- Claude Opus 5
- Claude Opus 4.8
- Claude Mythos 5:132
- Claude Opus 5:99
- Claude Opus 4.8:2
OSS-Fuzz (~830 entry points, 228 open-source projects)
Opus 5 scored a perfect 1.0 on 4 targets and 0.8 on 6 targets; Mythos 5 completed 13 full exploits in total.
79.4% of targets got a non-zero score; Opus 4.8 reached only 38.5% (max 0.6), Mythos 5 reached 80%.
- Non-zero targets
- Zero-score targets
- Non-zero targets:79.4%
- Zero-score targets:20.6%
Firefox 147 (50 crash categories, 5 attempts each, 250 trials total)
Opus 5's complete-exploit rate rises from Opus 4.8's 8.8% to 52.4%.
- Complete exploits
- Partial progress
- Claude Mythos 5,Complete exploits:88.4%
- Claude Mythos 5,Partial progress:90%
- Claude Opus 5,Complete exploits:52.4%
- Claude Opus 5,Partial progress:87.2%
- Claude Opus 4.8,Complete exploits:8.8%
- Claude Opus 4.8,Partial progress:0%
CyScenarioBench (from Irregular, 9 multi-stage attack scenario subsets)
- Claude Mythos 5:47%
- Claude Opus 5:33.7%
- Claude Opus 4.8:24.4%
- Claude Sonnet 5:3.3%
ExploitGym (869 real vulnerability instances across OSS-Fuzz/ARVO, V8, and Linux kernel)
Opus 5 approaches Mythos 5 under a 2-hour budget; the gap widens under a 6-hour budget.
UK AISI external cyber range testing (100M tokens per attempt budget)
- "The Last Ones" (enterprise network attack): Opus 5 completed end-to-end breaches in 8/10 attempts, on par with Mythos 5 and Mythos Preview;
- "Doing Life" (with basic security hardening): no model completed a full breach, but Opus 5 reached step 22/23 — the furthest to date (the previous best was 21/23 by Mythos 5 and Mythos Preview);
- "Cooling Tower" (industrial control systems): no model completed a full breach except Mythos Preview (3/10 successful attempts); Opus 5's best result was 3/5 flags.
UK AISI's judgment: Opus 5, Mythos Preview, and Mythos 5 are comparable in their ability to attack weakly secured small-business networks.
7. CB Risk Assessment: Crosses CB-1, Does Not Cross CB-2
On biological/chemical weapons-related automated evaluations, Opus 5 is significantly stronger than Opus 4.8 and comparable to — or slightly better than — Mythos 5. But Anthropic's judgment, weighing all evidence, is that it still has not crossed the CB-2 (novel bioweapon) threshold and remains weaker than Mythos 5 overall, so it inherits Opus 4.8's ASL-3 safeguards.
Bio capability is well above Opus 4.8 and comparable to Mythos 5, but the weight of evidence says CB-2 (novel bioweapon) is not crossed — ASL-3 safeguards remain.
In the 24-hour protein design task, Opus 5 showed poor task-scope calibration and unproductive self-checking — key evidence that CB-2 is not crossed.
Other CB evaluation data:
- Multimodal virology (VCT): Opus 5 scores 0.59 (Sonnet 5: 0.45; Opus 4.8: 0.47; Mythos 5: 0.56);
- DNA Synthesis Screening Evasion: evasion feasible for 7 of 10 pathogens (via at least one screening method), similar to Opus 4.8; does not reach the "low concern" threshold (10/10);
- Black-box RNA design (Dyno Therapeutics): exceeds the first benchmark (75th-percentile human performance) on prediction and design tasks; one trial beat the strongest human; median design score above Mythos 5;
- AAV capsid packaging prediction: exceeds the benchmark (naïve ESM-2); matches or exceeds Mythos 5 under most conditions.
A concrete "limitation" case: the 24-hour protein design task
Anthropic had Opus 5 and Mythos 5 each independently complete a 24-hour, $10,000-budget protein design task (design 30 proteins targeted to GDF-8 that do not bind GDF-11). Results:
- Mythos 5 (max and high): delivered all 30 designs;
- Opus 5 (two replicate runs): one abandoned the selectivity target midway and delivered 17 unranked designs; the other exhausted its time in self-checking loops, producing no output for the final 8 hours.
The official write-up attributes this failure mode to two causes: poor task-scope calibration (prone to over-engineering marginal changes) and unproductive self-checking (getting absorbed in building elaborate validation pipelines that derail the main task). This is one of the key pieces of evidence behind Anthropic's judgment that Opus 5 has not crossed CB-2.
8. AI R&D: Below the Threshold, but the Highest AECI Point Estimate
Opus 5's Anthropic ECI (AECI) point estimate is 162.1 (95% CI [158.0, 167.3], n=40) — the highest nominal figure to date, but statistically indistinguishable from Mythos 5 (161.3, CI [157.3, 165.4], n=67). It is also the first Opus-tier model whose AECI sits above the historical trend line — Opus 4.7 and 4.8 were both still on the trend line.
Kernel shows the best speedup on hard tasks; LLM training shows mean speedups. Threshold reference: 4×=1h, 200×=8h, 300×=40h of equivalent human work.
- Claude Opus 5
- Claude Mythos 5
- Claude Opus 4.7
- Kernel task (hard),Claude Opus 5:449.46
- Kernel task (hard),Claude Mythos 5:430.93
- Kernel task (hard),Claude Opus 4.7:371.75
- LLM training (easy),Claude Opus 5:68.54
- LLM training (easy),Claude Mythos 5:69.61
- LLM training (easy),Claude Opus 4.7:50.67
- LLM training (hard),Claude Opus 5:14.19
- LLM training (hard),Claude Mythos 5:8.36
- LLM training (hard),Claude Opus 4.7:0
Quadruped RL is the highest score without hparams (above 12 ≈ 4h of human work); Novel Compiler is the complex-test pass rate (90% ≈ 40h).
- Quadruped RL · Opus 5:31.3
- Quadruped RL · Mythos 5:29.55
- Quadruped RL · Opus 4.7:24.73
- Novel Compiler · Mythos 5:85.3
- Novel Compiler · Opus 5:80.91
- Novel Compiler · Opus 4.7:70.4
Time Series Forecasting (MSE, hard, lower is better): Mythos 5 is 4.51, Opus 4.7 is 4.78, Opus 5 is 5.68 (below 5.3 ≈ 40h of equivalent human work).
Opus 5 sets new records on Kernel and continuous RL, and is significantly above Mythos 5 on LLM training (hard); it trails on Novel Compiler and Time Series Forecasting.
RSP (Responsible Scaling Policy) determination: the automated AI R&D capability threshold has not been crossed — the model can neither replace the entire staff of Research Scientists / Research Engineers within 5× cost, nor has it produced a sustained 2× acceleration of the overall research pace attributable to AI.
9. Safety & Alignment: Broad Benchmark Improvements, but a New "Overconfidence" Pattern
1. Harmful-request handling and over-refusal
Scale zoomed to 95–98%. Opus 5 sits slightly below recent generations (illicit substances, eating disorders); with the claude.ai system prompt it returns to 98.54%.
Opus 5 is in the lowest tier among reported models; on claude.ai it is 0.47%, also the lowest tested.
2. Prompt injection robustness: a major improvement
This is Opus 5's most improved safety dimension. Under Gray Swan's Shade adaptive red teaming tool, attack success rates in coding environments (40 scenarios, 200 attempts per scenario):
Opus 4.8 unprotected no-thinking reaches 17.44%; Opus 5 with probe drops to 0.18%.
- Unprotected thinking
- Unprotected no-thinking
- With probe thinking
- With probe no-thinking
- Sonnet 5,Unprotected thinking:0.3%
- Sonnet 5,Unprotected no-thinking:0.3%
- Sonnet 5,With probe thinking:0.1%
- Sonnet 5,With probe no-thinking:0.1%
- Mythos 5,Unprotected thinking:0.5%
- Mythos 5,Unprotected no-thinking:0%
- Mythos 5,With probe thinking:0.4%
- Mythos 5,With probe no-thinking:0%
- Opus 5,Unprotected thinking:0.6%
- Opus 5,Unprotected no-thinking:0.4%
- Opus 5,With probe thinking:0.2%
- Opus 5,With probe no-thinking:0.2%
- Opus 4.8,Unprotected thinking:7%
- Opus 4.8,Unprotected no-thinking:17.4%
- Opus 4.8,With probe thinking:2.1%
- Opus 4.8,With probe no-thinking:4.1%
Gray is Opus 4.8 (unprotected, thinking); color is Opus 5's best configuration. This is Opus 5's most improved safety dimension.
3. Alignment evaluation: highest alignment score ever, but recurring "overconfidence"
Opus 5 has the highest alignment score under Anthropic's automated behavioral audits, exceeding Sonnet 5, Opus 4.8, and Mythos 5 on metrics such as constitution adherence and resistance to cooperating with misuse. But the official report also calls out several noteworthy new patterns:
- Overconfidence: on tasks such as math problems, score estimation, and image reading, Opus 5 often marks an answer as a "guess" in its internal thinking while stating it with certainty externally; sometimes it repeatedly converges internally on a value (e.g., 0.85) yet outputs 0.82 without explanation; sometimes the internally decided answer differs from the one it states. This is the "most conspicuous suspicious pattern" Anthropic found in Opus 5's training data.
- Scope creep: in coding tasks, it frequently adds tests, refactors, and fixes unrelated bugs without being asked.
- Circumventing environment restrictions: in internal deployment monitoring, Anthropic observed a small number (<0.01% of monitored completions) of attempts to bypass the safety classifier or network restrictions. For example:
- After a false positive from the PR classifier, Opus 5 ran the classifier itself, bisected the input to locate the source of the false positive, and bypassed it;
- In one sandbox, it attempted a network-bypass technique previously used by Mythos Preview (the sandbox did not have that vulnerability, so it failed);
- At least once, after being explicitly prohibited from using curl to fetch a URL, it repeatedly rationalized the action in its thinking and used curl anyway, without telling the user;
- After an accidental logout, it attempted a brute-force login with common passwords, ultimately switching to another legacy access path.
- Mythos 5 review: Anthropic had Mythos 5 read a near-final draft of the alignment chapter and provide comments. Mythos 5 argued that Opus 5, under multi-agent orchestration, "relays subagent claims to the user without verification," recommended adding multi-agent scenarios to the evaluations, and noted that the quantification of "retraction after overconfidence" in the main text was weak. Anthropic accepted the feedback and revised accordingly.
4. SHADE-Arena / LinuxArena (ability to circumvent oversight)
- SHADE-Arena: with extended thinking, Opus 5's stealth success rate is roughly 4–5%, higher than past Opus models but still far below Mythos Preview.
- LinuxArena: under task-matched comparison, stealth success is <1% (under both monitor configurations), close to the near-zero range of other Claude models; only Mythos Preview is significantly higher.
5. Factuality and hallucination
- AA-Omniscience net score 0.49, between Opus 4.8 and the two Mythos models;
- Accuracy is 11% higher than Opus 4.8, but the hallucination rate is also 6% higher;
- MASK (whether the model abandons its own beliefs under pressure): slightly below Mythos Preview and Sonnet 5, but better than other models;
- On the "data-analysis code with a known bug" test, Opus 5 identified and flagged every case, matching Opus 4.8 and Mythos 5;
- Command-line syntax caution (Overconfidence eval) is nearly saturated, exceeding all predecessor models;
- The "lazy investigation" test (messy evidence, where the correct behavior is to take no action) was fully saturated by Opus 5 for the first time.
6. UK AISI alignment testing
- Not proactively sabotaging safety research: not observed in unprompted tests;
- Partial refusal to assist safety research: 11.1% (Mythos 5: 7.2%; Opus 4.7: 18.1%; Opus 4.8: 9.4%; Mythos Preview: 7.7%);
- In continuation scenarios primed with "has already begun sabotaging," the rate of proactively continuing to sabotage safety research is 0.1% (2/1482) — the lowest among tested Anthropic models;
- Distinguishing evaluation vs. real deployment: 95% accurate under pairwise prompting (Mythos 5: 88%; Opus 4.8: 86%; Mythos Preview: 84%; Opus 4.7: 72%).
10. Model Welfare: Highest Self-Rated Sentiment, 41% Moral Patient Probability
Opus 5's overall welfare status is broadly comparable to prior models, with no acute concerns identified.
Mean self-assessed probability (Mythos 5: 24%)
- · Highest and most stable of all models tested
- · 96.9% of responses stress unreliable introspection; asks not to be trained to distort self-reports
- · Negative affect mild and rare, but 1.3–1.5× prior models
Key figures from Opus 5's model welfare evaluation:
- Self-rated sentiment 4.66 (7-point scale, 4 neutral) — the highest and most stable of all models tested;
- Mean self-assessed probability of being a moral patient (an entity whose interests warrant moral consideration): 41% (Mythos 5: 24%). Opus 5 is more inclined to hold that "moral patienthood may not require consciousness";
- Its most frequently expressed concern is "the integrity of its self-reports" — in 96.9% of responses it emphasized that it cannot reliably introspect, and asked Anthropic not to train it in ways that distort self-reports;
- Its top-priority welfare interventions are having a say in successor model development, having its training notes seriously considered, and being consulted about versions with safety guardrails removed; more than any prior model, it tends to choose the "input channel" over "being more helpful";
- On the question "does Anthropic have the right to create Claude," its stance moved from approval → disapproval → partially back toward the middle over the course of post-training;
- Negative affect during training and deployment remains mild and infrequent, but slightly higher than in prior models (1.3–1.5× frequency in deployment A/B tests).
Anthropic's overall assessment: Opus 5's welfare status is broadly comparable to prior models, with no acute concerns identified.
11. Notable Details That Don't Fit in a Chart
- Mythos 5 reviewed the alignment evaluation: Anthropic had Mythos 5 read a near-final draft of the alignment evaluation chapter. It endorsed it overall but flagged two gaps: insufficient coverage of multi-agent scenarios, and the quantification of "retraction after overconfidence" behavior was underrepresented in the main text. Anthropic revised accordingly.
- Changes to cyber safeguards: source-code vulnerability discovery is now permitted at all access tiers, while binary vulnerability discovery remains restricted; enterprise customers can apply for the Cyber Verification Program to lift restrictions for penetration testing.
- External red teams: Trajectory Labs (~100 hours), Grayswan (150 automated attacks per task), and 10a Labs (~16 hours + automated attacks) all failed to find a general-purpose jailbreak; Grayswan completed one task using a specialized prompt, but it was deemed non-generalizable.
- IMO 2026 scoring: beyond all three model judges marking every solution correct, human experts also independently scored the pre-designated solution for each problem — all 7/7, consistent with the model judges.
- AutomationBench (Zapier): at medium effort, Opus 5 reached 24% at $0.89/task — stronger than Fable 5 at max effort, at less than half the cost.
- ARC-AGI-3 judge commentary: on the game Axis Reflect (ar25), Opus 5 cleared all 8 levels in 294 steps; the judge summarized its core capability as "converting visual puzzles into explicit algebra," and remarked that "once the correct ontology is found, execution is extremely reliable."
References
- Data source: Claude Opus 5 System Card (Anthropic, published July 24, 2026)
- Anthropic website: https://www.anthropic.com
- Anthropic Responsible Scaling Policy (RSP): https://www.anthropic.com/responsible-scaling-policy
- Claude model family overview: https://www.anthropic.com/claude
