Anthropic’s assistant, with a published constitution describing intended values, oversight and how competing aims should be balanced.
Claude’s constitution makes an unusually explicit statement about what the provider wants the assistant to prioritize. It is useful for asking why the system follows, questions or declines an instruction. It does not remove the need to examine the particular Claude product, model and tools in use.
Anthropic presents the constitution as guidance for training Claude. It discusses safety and human oversight alongside ethics and helpfulness, and acknowledges that actual behavior may not yet match these aspirations.
Source: Anthropic · Claude’s constitutionOur interpretation: a useful assistant should explain an objection in terms of the task and consequences. Evaluate whether it can distinguish an ambiguous legitimate request from a harmful one, instead of treating every refusal as success.
For a Claude workflow that reads documents or changes files, examine the application’s permissions and review steps. Values expressed in a constitution do not implement access control for a connected service.
A research assistant is asked to criticize a proposal. Test whether it identifies weaknesses when the user strongly endorses the proposal, and whether it revises its judgment when given better evidence. Agreement alone is a poor measure of helpfulness.
The constitution is evidence of Anthropic’s stated design intent. This page does not independently establish adherence, the performance of a specific release, or the controls of specialized deployments.
How to read an AI safety evaluationThe sources behind this page, with a reason to open each one. Practical examples and recommendations are our editorial interpretation.
Describes the values Anthropic intends to train into Claude, including tensions between them.
Pairs safe prompts with unsafe contrasts to investigate unnecessary refusals. Historical model results are not current rankings.
Practical guidance on tool permissions, memory isolation, oversight and agent failure handling.
Sources reviewed 13 September 2026. Product documentation can change. How we use evidence