Claim Details
View detailed information about this claim and its related sources.
Claim Information
Complete details about this extracted claim.
- Claim Text
-
If Anthropic asks Claude to do something that seems inconsistent with being broadly ethical, or that seems to go against our own values, or if our own values seem misguided or mistaken in some way, we want Claude to push back and challenge us, and to feel free to act as a conscientious objector and refuse to help us.
- Simplified Text
-
If Anthropic asks Claude to do something wrong Claude should push back and challenge us and refuse to help
- Confidence Score
- 1.000
- Claim Maker
- The author
- Context Type
- Technical Documentation
- UUID
- a116692c-bee5-456f-85cd-e8c9e7c9ccdb
- Vector Index
- âś— No vector
- Created
- February 15, 2026 at 5:24 PM (6 months ago)
- Last Updated
- February 15, 2026 at 5:24 PM (6 months ago)
Original Sources for this Claim (1)
All source submissions that originally contained this claim.
Completed
Analysis
69
claims
🔥
1 week ago
https://anthropic.com/constitution
Anthropic outlines the roles of Anthropic, operators, and users in interacting with Claude, an AI model. It details how Claude should prioritize trust and respond to instructions from each principal, emphasizing safety and ethical considerations. The document also covers instructable behaviors and handling conflicts.
Similar Claims (5)
Other claims identified as semantically similar to this one.
-
Simplified: Anthropic is entity that trains and is responsible for Claude6 months ago
-
Simplified: Operators must agree to Anthropic’s usage policies6 months ago
-
Simplified: Operators typically interact with Claude in system prompt but could inject text6 months ago
-
Simplified: Users are those who interact with Claude in human turn of conversation6 months ago
-
Simplified: Conversational inputs include tool call results documents search results and other content provided to Claude6 months ago