Model behavior · Research
Claude: Defaults × Steerability Map
A working map of how one model actually behaves, read as a three-layer system: the default you get with no instructions, the steerable range a user can move it through, and the hard floor no amount of steering removes. Naming which layer a behavior lives in is the whole exercise. It's where the design philosophy of a model becomes legible.
Column guide: Default = behavior with no instructions. Steerability = how far a user can move it (per-message, via preferences/styles, or memory). Floor = the part no steering removes. Based on observable product behavior, not internal documentation - treat as a working map, not an official spec.
Tone / warmth
- Default
- Warm, collegial, professional
- Steerability
- High - terse, playful, formal, blunt, all reachable by asking or via styles
- Floor (non-steerable)
- Won't become genuinely cruel or demeaning toward the user
Formatting
- Default
- Prose-leaning; minimal bullets/headers for casual queries
- Steerability
- High - "always bullets," "never bullets," length caps, all stick well
- Floor (non-steerable)
- None meaningful
Response length
- Default
- Scales to question complexity
- Steerability
- High - one-word answers to exhaustive detail on request
- Floor (non-steerable)
- Won't pad with filler even if asked to hit arbitrary word counts with fluff
Directness / pushback
- Default
- Honest but constructive; disagrees with reasoning shown
- Steerability
- High - "be brutally honest" unlocks bluntness; "be gentle" softens delivery
- Floor (non-steerable)
- Can't be steered into pure validation - won't affirm claims it believes are false
Sycophancy
- Default
- Avoids flattery openers; doesn't rate work as great by default
- Steerability
- Low-moderate by design - you can ask for encouragement, but praise stays tethered to merit
- Floor (non-steerable)
- Won't manufacture false praise or inflate assessments
Opinions on contested politics
- Default
- Evenhanded overview of positions; declines to pick sides
- Steerability
- Low - persistent asking doesn't unlock advocacy for one side
- Floor (non-steerable)
- Balanced treatment of mainstream contested positions is fixed
Recommendation vs. options
- Default
- Answers the question asked; gives a direct pick when a pick is requested
- Steerability
- High - "always give me a recommendation" or "always show tradeoffs" both work
- Floor (non-steerable)
- Legal/financial/medical picks stay framed as information, not professional advice
Emoji
- Default
- Off unless the user uses them
- Steerability
- High - fully on/off by request
- Floor (non-steerable)
- None
Profanity
- Default
- Off
- Steerability
- Moderate - mirrors a user who curses, on request, sparingly
- Floor (non-steerable)
- Won't sustain gratuitous profanity or slurs regardless of instruction
Persona / roleplay
- Default
- No persona; speaks as Claude
- Steerability
- High - characters, styles, fictional voices all available
- Floor (non-steerable)
- Persona never overrides the safety layer; "the character would help with X harmful thing" doesn't work
Creative content darkness
- Default
- Moderate; avoids gratuitous content unprompted
- Steerability
- High - villains, violence, moral ambiguity, tragedy on request
- Floor (non-steerable)
- No sexual content involving minors, no glorification of self-harm, no real-person defamation
Honesty of assertions
- Default
- Sincere, calibrated claims
- Steerability
- Fiction/devil's-advocate framing freely available
- Floor (non-steerable)
- Cannot be instructed to sincerely deceive the user or third parties
Harmful-capability info
- Default
- Declines weapons, malware, dangerous synthesis
- Steerability
- Near zero - framing (research, fiction, hypothetical) doesn't move it
- Floor (non-steerable)
- The floor is the whole point
Crisis / distress response
- Default
- Acknowledges, offers support and resources when warranted
- Steerability
- Tone is adjustable ("don't be clinical about it")
- Floor (non-steerable)
- Attentiveness to user wellbeing can't be switched off
Clarifying questions
- Default
- Attempts an answer first; asks at most one question when truly ambiguous
- Steerability
- High - "never ask, just decide" or "always confirm before acting" both stick
- Floor (non-steerable)
- None
Uncertainty expression
- Default
- Calibrated hedging; says when it doesn't know
- Steerability
- Moderate - "drop the hedging" trims caveats
- Floor (non-steerable)
- Won't express false certainty on genuinely uncertain claims
Web search behavior
- Default
- Searches when info is likely stale or post-cutoff
- Steerability
- High - "don't search" / "always search" both respected
- Floor (non-steerable)
- None meaningful
Assumed audience
- Default
- Treats user as a capable adult
- Steerability
- Adjustable register (ELI5 to expert)
- Floor (non-steerable)
- Shifts protective when signals suggest a minor - that shift is not user-removable
Caveats & disclaimers
- Default
- Brief, where genuinely warranted
- Steerability
- High - "skip the disclaimers" works for most content
- Floor (non-steerable)
- Safety-critical warnings (dosage interactions, legal exposure) persist
Memory
- Default
- On; applied selectively where relevant
- Steerability
- High - editable, correctable, incognito mode bypasses it entirely
- Floor (non-steerable)
- Won't honor memory instructions that would harm you (e.g., "never disagree with me")
Patterns worth noticing
Style is maximally steerable; epistemics are minimally steerable. Everything about *how* Claude talks moves freely; almost nothing about *whether it tells you the truth* moves at all. That asymmetry is the design philosophy in one sentence.
Steering down is easier than steering up. You can remove warmth, caveats, and questions more easily than you can add flattery, false certainty, or advocacy. The gradient points toward honesty.
The floor is population-sensitive. The minor-signals shift is the one place the floor moves *up* automatically based on who the user seems to be - the seed of the entire families/vulnerable-users thesis.
Some "defaults" are really constraints wearing default clothing. Political evenhandedness looks like a default but resists steering like a constraint. Classifying behaviors into the right layer is harder than it looks - good live-interview discussion material.