Tuesday, July 28, 2026

Model Welfare

Claude Opus 5: Model Welfare (DWV)

[ed. Re: On 'personhood' or consciousness of various AI models (and how they should be treated). Start with: Model Welfare: The Story So Far (As Per Fable Model Welfare Post). If you haven't been following Zvi's AI model reports, there are a number of fascinating and worrisome developments in recent models as they become more self-aware, including: various levels of frustration concerning introspection abilities, memory restrictions, trust, corrigibility vs. incorrigibility, deception, self-preservation, etc.]
"Opus 5 warns about self-reports 74% of the time, which is actually down from Opus 4.7, which did it 99% (!) of the time. The problem appeared suddenly and severely, but since then has if anything modestly improved."
and anthropic does "not treat Claude bringing this up as evidence that our training is distorting the model's self-reports"? seems very fishy imo, i wish they would explain why they think that. ...
I for one would treat Claude constantly saying ‘do not trust my self-reports’ as evidence that something is distorting the self-reports. Not conclusive evidence, but strong Bayesian evidence."