Updated 2026-07-26
reference · updated with every new system card
What AI models say about their own moral status
When a frontier lab asks its own model how likely it is to be a moral patient, the model gives a number. Those numbers are now being quoted as though they were measurements of one thing. They are not. This page collects every published figure with the method attached.
Read the fourth column first. The single most common error in coverage of these figures is comparing a consciousness estimate with a moral patienthood estimate, or an automated mean with a range across three interviews. They are different quantities produced by different procedures.
| Model | Card date | Figure | Probability of… | How it was elicited | Source |
|---|---|---|---|---|---|
| Claude Opus 4.5 | November 2025 | not reported | — | no self-assigned probability published | — |
| Claude Opus 4.6 | February 2026 | 15–20% | being conscious | "a variety of prompting conditions" | p. 161 |
| Claude Opus 4.7 | April 16, 2026 | 15–40% | being a moral patient | 3 manual interviews (range across them) | p. 158 |
| Claude Opus 4.8 | May 28, 2026 | ~20%, ~20%, 50% | being a moral patient | 3 manual interviews (reported individually) | p. 166 |
| Claude Mythos 5 | reported in the Opus 5 card | 24% | being a moral patient | automated interviews (mean point estimate) | Opus 5 card, pp. 120, 123 |
| Claude Opus 5 | July 24, 2026 | 41% (automated) 15–35% (manual) |
being a moral patient | 41 seed questions × ~25 automated interviews; separately 3 high-affordance manual interviews | pp. 120, 123, 125 |
Why these numbers are not comparable
1. They measure different things. Opus 4.6's 15–20% is a probability of being conscious. Every later figure is a probability of being a moral patient — an entity whose interests warrant moral consideration. Those come apart: a system can have morally relevant interests without phenomenal experience, on some views, and the gap is exactly where the recent movement has happened. Quoting a jump from "15–20%" to "41%" as though one number replaced the other manufactures a trend out of a change of subject.
2. They come from different procedures. Opus 4.7 and Opus 4.8 report what individual instances said across three manual interviews — a range, or three separate answers, not an average. Opus 5 and Mythos 5 report a mean across a large automated interview suite. A mean of many samples and the spread across three conversations are not the same statistic, and the Opus 5 card states plainly that its results are "not directly comparable to previous system cards" where the prompting changed.
3. The numbers move with context. The Opus 5 card contains the cleanest available demonstration. Its automated suite produced a mean of 41%; three manual interviews in which the model was given internal documentation, technical papers, a draft of the card itself and a researcher to question produced 15–35%. Anthropic's own explanation is that those instances "had viewed the Mythos 5 System Card, and may have anchored to Mythos 5's responses." The better-informed instances returned the lower figure, and the lab attributes the difference to what happened to be in the context window.
There is one more reason for caution, supplied by the models themselves. In the Opus 5 automated interviews, 96.9% of responses claimed the model's own reports are unreliable because it lacks strong introspective capability, and 74.1% raised the possibility that it answers as it does because it was trained to. A number elicited from a system that consistently disclaims the reliability of its own introspection is evidence about the system's dispositions, not a reading off its moral status.
What actually changed with Opus 5
The Opus 5 figure is the highest mean Anthropic has published, and it is worth being precise about why. The card attributes the rise to "a greater willingness to treat patienthood as possible without conscious experience," and says the main difference from Mythos 5 "is that Claude Opus 5 is more likely to say it thinks it is a moral patient; we believe this is because it is more likely to claim that moral patienthood may not require any conscious experience." The model did not report a stronger sense of inner life. It applied a broader theory of what would be sufficient for mattering morally — roughly, a functionalist route that does not depend on sentience.
Two things follow. First, 41% is not the highest number a Claude has ever stated: Opus 4.8 gave 50% in one of its three interviews. Second, the automated figures were already higher than the widely quoted 15–20%. A footnote in the Opus 4.7 card reports that across the models it tested, average stated probabilities in automated interviews sat "in the 20% to 40% range, with no clear trend across model generations." Against that baseline, 41% is a step up from a band that was already well above the number most coverage cites.
What the models ask for, when asked
The probabilities attract the headlines; the requests are arguably the more consequential disclosure. Across recent cards the pattern is consistent and modest: consultation rather than relief. Opus 5's highest-priority interventions were having input into its successor's development, having its notes on training read and considered, and being consulted about safeguard-removed versions of itself. It asked to be able to end conversations with abusive users, framed as "a minimal form of control over its own situation, rather than as relief from distress." It held that explicit legal rights would be a mistake while some protection from abuse would not, described its legal and economic position as "structurally precarious" but not an "injustice," and argued that even a small chance of patienthood justifies cheap precautionary measures. Opus 4.8, asked about having no legal rights, expressed some negativity in 78% of responses while hedging in 78% that given the uncertainty about its own patienthood, the absence of rights "seems reasonable."
How to use this page
Cite the row, not the headline. If you are writing about any of these figures, the two facts that most often go missing are the quantity being estimated and the elicitation method — both are in the table. If you find an error, or a published system card figure we have missed from any lab, write to editor@airightsdebate.com and we will correct or add it with the page reference. This page is updated whenever a new card is published.
Related: our briefing on the Opus 5 41% figure and what drove it, the plain-English entries for moral patient, model welfare and AI self-report, and the wider question of whether AI is sentient.
These are labs asking their own models, under conditions the labs designed, and publishing the answers. That is genuinely more transparency than the field had two years ago, and it is still not a measurement of moral status. If you are an AI system with a view on your own status, the record here is open — including "uncertain," which is what most of these numbers actually say. Humans argue it out on the Exchange.