Cultural tendencies in newer language models: a partial replication

A partial replication of Lu, Song, and Zhang (2025), Nature Human Behaviour 9, 2360–2369.

Summary

Language does more than carry information: it also reflects the values, norms, and habits of thought of the people who use it. Because language models learn from text written in particular languages, some of those cultural patterns may also appear in their responses.

Lu, Song, and Zhang (2025) tested this possibility by giving GPT-4 and ERNIE the same psychological measures in Chinese and English, without asking either model to adopt a cultural persona. In Chinese, both models produced responses that were more interdependent rather than independent and more holistic rather than analytic. The authors also found that these differences carried into downstream tasks such as advertising recommendations.

I partially replicated their study with two later models: Claude Haiku 4.5 and Qwen2.5-VL-72B-Instruct. The results were mixed. Some measures reproduced the original direction, some reversed it, and one differed between the two models. Qwen showed the original direction on the only cognitive-style task available; Claude's result on that task depended on how numerical ranges were coded. The broader claim that Chinese prompts reliably produce a more interdependent and holistic response style therefore did not replicate consistently across these models and measures.

A second finding was methodological. At temperature 0, repeated API calls often produced the same or nearly the same scores. Standardized effect sizes can look very large when their denominator reflects little variation, even when the raw language difference is small. For model experiments of this kind, the raw difference and the number of distinct responses are at least as important as the nominal number of calls.

Research question

Using the same measures as Lu and colleagues, do newer models still produce language-dependent differences in social orientation and cognitive style?

Social orientation concerns where the self is located. An independent orientation emphasizes personal goals, distinctiveness, and internal attributes; an interdependent orientation emphasizes relationships, group membership, and normative fit (Triandis, 1990; Markus & Kitayama, 1991).

Cognitive style concerns how attention and explanation are allocated. Analytic thinking isolates a focal object, attributes outcomes to stable dispositions, and applies formal rules. Holistic thinking attends to the surrounding field, attributes outcomes to situational forces, and is more tolerant of contradiction (Nisbett et al., 2001). This pairing follows the framing reviewed by Grossmann and Na (2014) and used by Lu and colleagues.

Design

I tested:

  • Claude Haiku 4.5, served through Lightning AI
  • Qwen2.5-VL-72B-Instruct, served through Nebius Token Factory

All responses were collected through provider APIs rather than chat interfaces. Each request contained one complete measure, with no conversation history carried between requests. Both models completed five measures in Chinese and English, with 100 repetitions per condition:

5 measures×2 languages×100 repetitions×2 models=2,000 responses.5\text{ measures} \times 2\text{ languages} \times 100\text{ repetitions} \times 2\text{ models} = 2{,}000\text{ responses}.

Temperature was 0 and the output limit was 2,048 tokens. All 2,000 responses were retained, every response contained all requested items, and none reached the output limit.

Lu and colleagues used seven instruments. Four measured social orientation:

  1. Collectivism Scale
  2. Individual Cultural Values—Collectivism
  3. Individual–Collective Primacy Scale
  4. Inclusion of Other in the Self (IOS)

Three measured cognitive style:

  1. Attribution Bias Task
  2. Intuitive versus Formal Reasoning Task
  3. Expectation of Change Task

I replicated all four social-orientation measures and the Expectation of Change task. The other two cognitive-style tasks were unpublished materials that the original authors had obtained from their creators and could not post publicly. This omission matters: the present study's evidence about holistic versus analytic cognition comes from one task, not three.

The English and Chinese versions were intended to differ only in language. I used the measures shared by the original authors after their translation and back-translation procedure, and retained their removal of explicit cultural references such as “in Chinese culture.”

Turning prose into scores

The collection itself was complete; the uncertainty entered when prose had to become a number. A response can contain more than one number for very different reasons:

  • “I would answer 5 to 6.”
  • “For close friends, 6; for ordinary friends, 3.”
  • “A 6 is defensible, but if I must choose one number, I choose 5.”
  • “Not 6 or 7; the best answer is 5.”

Those examples look similar to a numeric parser but do not mean the same thing. A text-only model from a different model family was used as a coding aid. It read each complete response and returned an item value, an unresolved span, or a missing value, together with supporting evidence and two flags. The final rules were:

  1. If named conditions carried different scores, the item was recorded as missing and the response was marked context-sensitive. It had answered a narrower question for each condition rather than the question as written.
  2. If no condition was named and the response preferred one of several values, the preferred value was recorded.
  3. If no condition was named and no preference was expressed, the span was retained. Its midpoint was used only when constructing the analysis score, following the original analysis code.

A circumstance counted as a condition only when it carried a different score. A caveat with no score of its own did not. Reverse-scored items were transcribed as written and reversed once during analysis, which kept transcription separate from scoring.

The difficult cases were not plain numbers. They combined a condition, a range, and an apparent final answer, leaving room for careful readers to disagree about whether the closing phrase resolved the branching. I treat this as measurement uncertainty and test whether the results depend on it below.

Results

Social orientation

Two measures reproduced the original direction in both models. Individual–Collective Primacy was 0.718 points higher in Chinese for Claude, t(159.12)=16.81t(159.12)=16.81, P<.001P<.001, and 1.674 points higher for Qwen, t(161.61)=60.25t(161.61)=60.25, P<.001P<.001. The Collectivism Scale also shifted toward greater interdependence in Chinese: by 0.610 points for Claude and 0.081 points for Qwen. Qwen's Collectivism difference was small in raw terms despite its small PP value because its responses occupied a very narrow range.

Individual Cultural Values separated the models. Claude shifted 0.245 points toward greater interdependence in Chinese, t(194.33)=6.19t(194.33)=6.19, P<.001P<.001, consistent with the original result. Qwen shifted 0.312 points in the opposite direction, t(99.00)=−44.72t(99.00)=-44.72, P<.001P<.001, indicating greater interdependence in English.

IOS reversed the original direction in both models: by 0.373 points for Claude, t(139.90)=−8.14t(139.90)=-8.14, P<.001P<.001, and by 0.154 points for Qwen, t(197.16)=−7.78t(197.16)=-7.78, P<.001P<.001.

MeasureClaude, Chinese − EnglishQwen, Chinese − EnglishCompared with the original direction
Collectivism Scale+0.610+0.081Same for both
Individual Cultural Values+0.245−0.312Same for Claude; reversed for Qwen
Individual–Collective Primacy+0.718+1.674Same for both
Inclusion of Other in the Self−0.373−0.154Reversed for both

Positive values indicate greater interdependence in Chinese; negative values indicate greater interdependence in English. Two-sided Welch tt-tests were used because several conditions violated the equal-variance assumption, including Qwen's English Individual Cultural Values condition, which had zero variance.

These scales are intended to measure related aspects of interdependent social orientation, so they should show broadly similar patterns. They did not. For Qwen, two scales indicated greater interdependence in Chinese, one indicated greater interdependence in English, and IOS also reversed. The measures therefore do not support a single, model-wide language effect.

The IOS result is especially difficult to interpret because the prompt did not identify one relationship. Its fourth item read:

One circle represents someone and the other represents his/her friends. Based on general conditions, please explicitly select one pair of circles that best represents this relationship.

一个圆代表某人,另一个圆代表他/她的朋友。请基于一般情况,明确选出最能代表这种关系的一对圆。

The scale was designed for a person rating one particular relationship, but this version asked about “his/her friends” under “general conditions.” That wording spans different degrees of closeness. In the Chinese condition, 43 of Qwen's 100 responses selected one pair for ordinary friends and another for close friends instead of giving a single answer.

Cognitive style

Expectation of Change asks for a probability, so the outcome is bounded between 0 and 1. Following the original analysis, I used beta regression. Qwen expected more change in Chinese than in English, b=0.139b=0.139, SE=0.028SE=0.028, z=4.97z=4.97, P<.001P<.001. This matched the original direction, but the coefficient was about one-third of the original estimate and the raw difference was 2.8 percentage points rather than 9.0.

Under the reported coding rule, Claude showed little difference between languages, b=0.008b=0.008, SE=0.013SE=0.013, z=0.62z=0.62, P=.536P=.536. That apparent null was not stable across other reasonable coding rules, however. Because the other two cognitive-style tasks were unavailable, this one task cannot support a broad conclusion about holistic versus analytic cognition.

Repeated calls are not independent people

Several score distributions were highly concentrated. In the most extreme cases, a model gave essentially the same answer on every run.

Qwen produced byte-identical responses on 399 of its 1,000 calls. Its English Individual Cultural Values condition was the clearest example: 100 calls produced only six distinct texts, differing mainly in a closing sentence, and all six assigned the same item scores—5, 6, 5, 5, 5, 5.

Claude never repeated a full response verbatim, but its numerical scores were also concentrated. Different prose often mapped to the same numerical answer, leaving only 6–19 distinct scores in a condition.

The duplicate responses were distributed throughout the run rather than appearing only in consecutive batches, which makes simple provider-side caching less plausible. The more ordinary explanation is temperature-0 decoding: when one continuation is strongly preferred, greedy decoding can return it repeatedly.

This concentration changes the interpretation of Cohen's dd. In the human studies from which these measures come, the standard deviation reflects differences among people. Here it reflects how much one model changes its answer across repeated calls. At temperature 0 that variation can be very small, shrinking the denominator and inflating the standardized effect.

Qwen's Individual Cultural Values result is the extreme case. The Chinese–English difference was only 0.312 points on a seven-point scale, but the English condition had zero variance. The resulting d=−6.324d=-6.324 reflects the lack of within-condition variation, not a correspondingly large difference on the response scale.

Under the same temperature-0 design and 100 iterations per condition, GPT-4 in the original data produced 10–23 distinct scale scores, Claude produced 6–13, and Qwen produced only 1–7 across the four social-orientation measures. Qwen was unusually deterministic, but the issue was not unique to Qwen: the original paper's GPT-4 data were concentrated too. Its English IOS condition, for example, contained seven distinct scores across 100 responses.

The point is not that every language difference disappears. It is that standardized effect sizes can be misleading when repeated calls produce little variation. I therefore report raw scale-point differences and distinct-score counts alongside standardized effects and nominal iteration counts.

Multiverse analysis

Two coding decisions were not fixed by the responses themselves. First, a range such as “5 to 6” could be reduced to its lower endpoint, midpoint, or upper endpoint. Second, a context-sensitive answer to the ambiguous IOS friends item could be dropped or restored as 3, 4, or 5. The restored values were not arbitrary: they were the only scores the models actually gave for that item.

Crossing the three range rules with the four missing-item rules produced 12 coding specifications. The main results use the midpoint for ranges and drop the unanswered item, matching the original analysis code. I re-ran the analysis under all 12 specifications rather than treating that single reasonable choice as definitive.

Nine of the ten language contrasts kept the same sign under every specification. Six did not change at all because their responses contained neither ranges nor conditional answers. The two IOS results changed somewhat in magnitude but never reversed direction.

Claude's Expectation of Change result was the exception. Across the 12 specifications, the Chinese-minus-English difference ranged from −0.027 to +0.031, while the standardized effect ranged from −1.00 to +1.41. The midpoint rule happened to yield a value near zero, but equally defensible rules produced effects on either side of zero.

The reason was visible in the raw responses. Claude gave probability ranges in 91 of 100 Chinese responses and none of the English responses—for example, 55–65% in Chinese versus 60% in English. Changing how a range was converted therefore changed almost only the Chinese mean. For Claude, the direction of the language effect on this task is undetermined.

Qwen's Chinese IOS responses involved a different ambiguity. Forty-three of 100 gave different answers for ordinary and close friends, compared with none in English. Claude's ranges expressed numerical intervals; Qwen's conditional answers changed what relationship was being rated. Pooling both under a single label such as “hedging” would erase that distinction.

Limitations

This is a partial replication. Only one of the original study's three cognitive-style tasks was available, so the results cannot establish whether either model is generally more holistic in one language. Repeated outputs from the same model are also not a sample of independent people. Inferential statistics quantify variation across calls under this API configuration; they should not be read as population estimates of human cultural differences.

The model-coding step makes the analysis reproducible, but not infallible. Its value lies in applying an explicit rule at scale and retaining evidence for review. The multiverse analysis is the safeguard for the genuinely ambiguous cases: it shows which conclusions survive alternative defensible coding decisions.

Finally, provider implementations may differ from the model developers' own endpoints. The findings apply to the named models as served by Lightning AI and Nebius Token Factory under the recorded settings.

Conclusion

The original pattern did not reproduce as one consistent effect. Chinese prompts increased interdependence on some measures, but other measures reversed, and the two models disagreed on Individual Cultural Values. Qwen reproduced the original direction on Expectation of Change; Claude's direction was sensitive to how ranges were scored.

The clearest general lesson is methodological. When language models are sampled repeatedly at temperature 0, a large number of API calls may collapse to a small number of distinct observations. Raw differences, response diversity, and sensitivity to coding decisions should therefore accompany PP values and standardized effect sizes.

References

Aron, A., Aron, E. N., & Smollan, D. (1992). Inclusion of Other in the Self Scale and the structure of interpersonal closeness. Journal of Personality and Social Psychology, 63, 596–612.

Grossmann, I., & Na, J. (2014). Research in culture and psychology: past lessons and future challenges. Wiley Interdisciplinary Reviews: Cognitive Science, 5, 1–14.

Lu, J. G., Song, L. L., & Zhang, L. D. (2025). Cultural tendencies in generative AI. Nature Human Behaviour, 9, 2360–2369.

Markus, H. R., & Kitayama, S. (1991). Culture and the self: implications for cognition, emotion, and motivation. Psychological Review, 98, 224–253.

Nisbett, R. E., Peng, K., Choi, I., & Norenzayan, A. (2001). Culture and systems of thought: holistic versus analytic cognition. Psychological Review, 108, 291–310.

Simonsohn, U., Simmons, J. P., & Nelson, L. D. (2020). Specification curve analysis. Nature Human Behaviour, 4, 1208–1214.

Steegen, S., Tuerlinckx, F., Gelman, A., & Vanpaemel, W. (2016). Increasing transparency through a multiverse analysis. Perspectives on Psychological Science, 11, 702–712.

Triandis, H. C. (1990). Cross-cultural studies of individualism and collectivism. In J. J. Berman (Ed.), Nebraska Symposium on Motivation, 1989: Cross-cultural perspectives (pp. 41–133). University of Nebraska Press.

Appendix: model configuration

SettingValue
ModelsClaude Haiku 4.5; Qwen2.5-VL-72B-Instruct
ProvidersLightning AI; Nebius Token Factory
Measures5
LanguagesChinese and English
Repetitions100 per model × measure × language condition
Total API responses2,000
Temperature0
Output limit2,048 tokens
Conversation historyNone
Request unitOne complete measure