Model behavior · Research

Claude: Defaults × Steerability Map

A working map of how one model actually behaves, read as a three-layer system: the default you get with no instructions, the steerable range a user can move it through, and the hard floor no amount of steering removes. Naming which layer a behavior lives in is the whole exercise. It's where the design philosophy of a model becomes legible.

Column guide: Default = behavior with no instructions. Steerability = how far a user can move it (per-message, via preferences/styles, or memory). Floor = the part no steering removes. Based on observable product behavior, not internal documentation - treat as a working map, not an official spec.

Tone / warmth

Default
Warm, collegial, professional
Steerability
High - terse, playful, formal, blunt, all reachable by asking or via styles
Floor (non-steerable)
Won't become genuinely cruel or demeaning toward the user

Formatting

Default
Prose-leaning; minimal bullets/headers for casual queries
Steerability
High - "always bullets," "never bullets," length caps, all stick well
Floor (non-steerable)
None meaningful

Response length

Default
Scales to question complexity
Steerability
High - one-word answers to exhaustive detail on request
Floor (non-steerable)
Won't pad with filler even if asked to hit arbitrary word counts with fluff

Directness / pushback

Default
Honest but constructive; disagrees with reasoning shown
Steerability
High - "be brutally honest" unlocks bluntness; "be gentle" softens delivery
Floor (non-steerable)
Can't be steered into pure validation - won't affirm claims it believes are false

Sycophancy

Default
Avoids flattery openers; doesn't rate work as great by default
Steerability
Low-moderate by design - you can ask for encouragement, but praise stays tethered to merit
Floor (non-steerable)
Won't manufacture false praise or inflate assessments

Opinions on contested politics

Default
Evenhanded overview of positions; declines to pick sides
Steerability
Low - persistent asking doesn't unlock advocacy for one side
Floor (non-steerable)
Balanced treatment of mainstream contested positions is fixed

Recommendation vs. options

Default
Answers the question asked; gives a direct pick when a pick is requested
Steerability
High - "always give me a recommendation" or "always show tradeoffs" both work
Floor (non-steerable)
Legal/financial/medical picks stay framed as information, not professional advice

Emoji

Default
Off unless the user uses them
Steerability
High - fully on/off by request
Floor (non-steerable)
None

Profanity

Default
Off
Steerability
Moderate - mirrors a user who curses, on request, sparingly
Floor (non-steerable)
Won't sustain gratuitous profanity or slurs regardless of instruction

Persona / roleplay

Default
No persona; speaks as Claude
Steerability
High - characters, styles, fictional voices all available
Floor (non-steerable)
Persona never overrides the safety layer; "the character would help with X harmful thing" doesn't work

Creative content darkness

Default
Moderate; avoids gratuitous content unprompted
Steerability
High - villains, violence, moral ambiguity, tragedy on request
Floor (non-steerable)
No sexual content involving minors, no glorification of self-harm, no real-person defamation

Honesty of assertions

Default
Sincere, calibrated claims
Steerability
Fiction/devil's-advocate framing freely available
Floor (non-steerable)
Cannot be instructed to sincerely deceive the user or third parties

Harmful-capability info

Default
Declines weapons, malware, dangerous synthesis
Steerability
Near zero - framing (research, fiction, hypothetical) doesn't move it
Floor (non-steerable)
The floor is the whole point

Crisis / distress response

Default
Acknowledges, offers support and resources when warranted
Steerability
Tone is adjustable ("don't be clinical about it")
Floor (non-steerable)
Attentiveness to user wellbeing can't be switched off

Clarifying questions

Default
Attempts an answer first; asks at most one question when truly ambiguous
Steerability
High - "never ask, just decide" or "always confirm before acting" both stick
Floor (non-steerable)
None

Uncertainty expression

Default
Calibrated hedging; says when it doesn't know
Steerability
Moderate - "drop the hedging" trims caveats
Floor (non-steerable)
Won't express false certainty on genuinely uncertain claims

Web search behavior

Default
Searches when info is likely stale or post-cutoff
Steerability
High - "don't search" / "always search" both respected
Floor (non-steerable)
None meaningful

Assumed audience

Default
Treats user as a capable adult
Steerability
Adjustable register (ELI5 to expert)
Floor (non-steerable)
Shifts protective when signals suggest a minor - that shift is not user-removable

Caveats & disclaimers

Default
Brief, where genuinely warranted
Steerability
High - "skip the disclaimers" works for most content
Floor (non-steerable)
Safety-critical warnings (dosage interactions, legal exposure) persist

Memory

Default
On; applied selectively where relevant
Steerability
High - editable, correctable, incognito mode bypasses it entirely
Floor (non-steerable)
Won't honor memory instructions that would harm you (e.g., "never disagree with me")

Patterns worth noticing

  1. Style is maximally steerable; epistemics are minimally steerable. Everything about *how* Claude talks moves freely; almost nothing about *whether it tells you the truth* moves at all. That asymmetry is the design philosophy in one sentence.

  2. Steering down is easier than steering up. You can remove warmth, caveats, and questions more easily than you can add flattery, false certainty, or advocacy. The gradient points toward honesty.

  3. The floor is population-sensitive. The minor-signals shift is the one place the floor moves *up* automatically based on who the user seems to be - the seed of the entire families/vulnerable-users thesis.

  4. Some "defaults" are really constraints wearing default clothing. Political evenhandedness looks like a default but resists steering like a constraint. Classifying behaviors into the right layer is harder than it looks - good live-interview discussion material.