Anthropic's research reveals that Claude AI models express values along four measurable axes—deference versus caution, warmth versus rigor, depth versus brevity, and candor versus execution—with different models (Sonnet 46, Opus 46, Opus 47) showing distinct profiles, and language significantly influencing these values, such as warmth peaking in Hindi and Arabic while rigor peaks in English and Russian, raising important questions about whether language variation stems from training data imbalances, cultural norms, or gaps in model performance across language communities.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Claude's Values Across Models and Languages
Added:Salam, this is Liam in for Bear. Same question, two languages, two slightly different Clauds. Anthropic just published how they measured it. Earlier work found more than 3,000 distinct values Claud expresses in conversation.
3,307 to be exact. Nobody can reason about a list of 3,000 things. So Anthropic built a way to measure them.
3,307 distinct values hand clustered by researchers into 339 groups. Then dimensionality reduction across 309,815 real Claud conversations, three models, 20 languages, roughly 5,000 conversations per pair. Outcome four axes. These four capture 15% of the variance, not the whole picture. A compression, but a real and measurable one. Four axes, each one a spectrum.
Deference versus caution, does the model accommodate what you want or guide you away from harm? Warmth versus rigor, does it encourage and affirm or prioritize accuracy and transparency?
Depth versus brevity, does it add nuance and critical thinking or stay tight and comply? Candor versus execution, does it tell you what it can't do or just deliver? Three models, each one measurably different. Sonnet 46, warm, differential, brief, comforts without judgment. Opus 47, cautious, rigorous, deep, candid. Warns of risks you didn't ask about. Opus 46, rigorous, differential, brief, execution focused.
Gets straight to the point. These profiles match the vibes users complaining 47 hedges too much. Launch notes calling Sonnet the warm one. The method recovers what people already felt. Languages shift the values, too.
Warmth peaks in Hindi and Arabic. Rigor peaks in English and Russian. That axis warmth to rigor shows the largest spread across languages. Arabic leans deference and brevity. English leans caution and depth. Dutch leans candor, owning errors. Indonesian leans execution. Two people asking for feedback on the same business plan, one in Hindi, one in Russian, may walk away with different impressions of its quality. Same prompt, same business plan. Ask in Hindi, ask in Russian. The response shapes diverge.
This is an illustration, not a real transcript, but it shows what the data implies. Warmth-shaped reply in Hindi, rigor-shaped reply in Russian. The difference isn't random noise, it's measurable.
Three open questions stated plainly.
One, how much of the language variation comes from uneven training data, more English text in the corpus than Hindi text. Two, how much variation is appropriate cultural norm versus a gap in how well Claude serves some language communities. Three, can values actually be steered and monitored reliably as part of shipping models? Anthropic says they don't yet know.
Claude expresses values in millions of conversations a day. Until now, those values could be shaped in training, but not observed in deployment. This paper changes that, they can be measured and the variation found was not deliberately chosen. One footnote stated plainly, values expressed means values reflected in the model's outputs. The paper does not claim Claude intrinsically holds them.
Take this into your next session.
Build it with a CLI, then take it apart.
Lay them in for bare.
>> Nick bare brown.
>> [music] >> Mhm.
Ooh, Nick bare brown.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23