A new arXiv study found that large language models overpredict negative sanctions for norm violations and diverge further from human judgment as social distance increases.

1 min read

New Study Reveals LLMs Overpredict Social Sanctions in Norm Violations

What's New in This Study?

FAQ

What are second-order social norms (metanorms)?

Metanorms are expectations about who will enforce a social rule and how, beyond simply knowing the rule itself. For example, a model must not only know stealing is wrong but also anticipate whether family, authorities, or the community will respond.

What is the NormReact dataset?

A new dataset of 450 norm violation scenarios, hand-annotated for emotions and behavioral responses, accounting for the violator's gender and the observer's social closeness.

How do LLMs differ from humans in these judgments?

LLMs overpredict negative sanctions where humans expect inaction, and the misalignment worsens as social distance between violator and observer increases.

Should MENA AI teams act on these findings now?

Yes, especially in family and tribal mediation and policy simulation, where punishment bias could produce unrealistic or socially harmful recommendations. Testing models on local data before deployment is advised.

Source: arXiv cs.AI

AI-assisted content, human-reviewed.