A new study reveals that large language models change their answers up to 23% when the same question is rephrased differently, challenging their reliability in critical applications like healthcare and government services in the Middle East.

1 min read

Study Reveals LLM Accuracy Varies by Question Wording Despite Semantic Equivalence

Study Summary

FAQ

What is the new study about LLM reliability?

An arXiv study examines how LLM answers change when the same question is rephrased, finding that models may give different answers up to 23% of the time.

How does rephrasing affect model accuracy?

While overall accuracy changes modestly, instance-level behavior is unstable, with answers flipping between correct and incorrect.

Can model reliability be improved?

Yes, using a self-paraphrasing strategy can recover latent knowledge and improve inference-time performance.

Why is this important for MENA?

In critical applications like healthcare and government, such instability could lead to wrong decisions, requiring more rigorous evaluation before deployment.

Source: arXiv cs.AI

AI-assisted content, human-reviewed.