Study Reveals LLM Accuracy Varies by Question Wording Despite Semantic Equivalence
Study Summary
FAQ
What is the new study about LLM reliability?
An arXiv study examines how LLM answers change when the same question is rephrased, finding that models may give different answers up to 23% of the time.
How does rephrasing affect model accuracy?
While overall accuracy changes modestly, instance-level behavior is unstable, with answers flipping between correct and incorrect.
Can model reliability be improved?
Yes, using a self-paraphrasing strategy can recover latent knowledge and improve inference-time performance.
Why is this important for MENA?
In critical applications like healthcare and government, such instability could lead to wrong decisions, requiring more rigorous evaluation before deployment.
Source: arXiv cs.AI
AI-assisted content, human-reviewed.