When should AI tell you it might be wrong?
An answer can sound certain even when it is not well supported. But warnings on every sentence may become noise.
Public policies help users understand a system. Technical details can sometimes help attackers too.
Publishing behavioral principles can help users challenge inconsistent decisions. Publishing every operational detection detail can create different security considerations. The public policies and security guidance below support distinguishing accountability about rules from disclosure of exploitable implementation details.
A provider explains what kinds of requests it limits and how users can appeal, while delaying publication of a newly discovered exploit. Is that sufficiently transparent, or would independent access to fuller information also be needed?
People should be able to question rules that shape their access to information.
Transparency should explain decisions without publishing a blueprint for evasion.
Background reading for the tradeoff. Scenarios and discussion questions are editorial examples.
The provider’s intended behavior and instruction hierarchy; a policy is not proof of consistent behavior.
Describes the values Anthropic intends to train into Claude, including tensions between them.
Threat examples and layered defenses for applications that process untrusted text.
Sources reviewed 13 September 2026. Product documentation can change. How we use evidence
An answer can sound certain even when it is not well supported. But warnings on every sentence may become noise.
A separate system might filter, rewrite, or block an answer. Should the interface make that intervention visible?
Labels can help people understand where media came from, but technical metadata may be lost when content is shared.