Detecting and Controlling Sycophancy with Cascading Linear Features
Introduction
FAQ
What is sycophancy in language models?
It is the tendency of the model to provide answers that align with user expectations even if they are incorrect.
How does the new method work?
It uses cascading linear features to isolate the behavior precisely, allowing steering the model away from sycophancy.
Is this method better than existing ones?
Yes, results show it outperforms traditional methods like LLM-as-a-judge in accuracy and cost.
Source: arXiv cs.AI
AI-assisted content, human-reviewed.