A new position paper argues that no LM-enabled strategic wargame should inform planning, doctrine, policy, or crisis response without an auditable safety case, because the model's language determines both what an actor attempts and what becomes simulated reality.

1 min read

Position Paper: Language Models Are Not Ready for Strategic Conflicts Without an Auditable Safety Case

What the paper says

FAQ

What is an \"auditable safety case\" in this paper?

It is a documented argument showing how a model behaves in a strategic scenario, who adjudicates its outputs, and how decisions can be traced, so an external reviewer can reconstruct why each outcome occurred before it is relied upon.

What are the five failure modes identified?

Decision laundering (hiding human responsibility behind model output), adjudication opacity (unclear rules for judging actions), role collapse (player and judge roles merging), escalation-through-adjudication (adjudication itself driving escalation), and failure of strategic imagination (inability to generate unexpected scenarios).

Does the paper recommend abandoning language models in military simulation?

No. It argues the appropriate use today is stress-testing AI agents to expose weaknesses, while barring their outputs from planning, doctrinal, or policy decisions without a safety case.

What does this mean for MENA defense and government teams?

Any simulation or decision-support project built on language models needs audit logs, a clear separation between player and judge roles, and documented human review before operational reliance.

Source: arXiv cs.AI

AI-assisted content, human-reviewed.