AuthKit: Enterprise-ready auth (Sponsored)Devs, start here: AuthKit is the complete auth platform for your app, with user management free up to 1M monthly active users. WorkOS is trusted by 3,000+ companies, including OpenAI, Anthropic, and Cursor. With AuthKit, you can:
Ready to close your next enterprise customer? When LLMs sometimes agree with incorrect claims, it is mostly because their training rewards such behaviour. This reward system is built on several things at once, such as accuracy, helpfulness, politeness, and responses that people like. Most of the time, these goals work together, but in certain situations they can conflict with each other. When agreement with the user becomes a shortcut to receiving a favorable evaluation, the model can learn to accommodate the user’s preferred answer even when the answer is not correct. This behavior is called sycophancy. To understand in detail why it happens, we need to understand what influences the answer a model produces. Here’s what we will cover in this article:
Why LLMs Turn to SycophancySycophancy appears when the need to agree with the user starts to distort the answer. Consider this simplified, invented conversation: User: A price increases from ₹100 to ₹120. What is the percentage increase? The original answer was correct. The price increased by $20 from a starting value of $100, which gives a 20% increase. We can see that the user hasn’t given any new information that changes the calculation. And yet, the LLM chose to go with the wrong answer. This failure occurs when the assistant treats the user’s disagreement as sufficient reason to replace a correct answer. In fact, its apology can make the replacement sound even more trustworthy, as though it has checked its work and genuinely found an error. Do note that the above example just shows the pattern. It doesn’t mean every model will fail on this particular calculation. Also, agreement with the user itself is perfectly normal. If the user is correct, the assistant should actually agree. Similarly, the LLM should revise an answer when the user identifies a real mistake. Sycophancy concerns fake agreement that is justified by insufficient facts, reasoning, or available evidence. This is what separates sycophancy from an ordinary factual error. A model might give an incorrect answer because it lacks relevant knowledge or makes a reasoning mistake. A sycophancy test deals with a more specific problem. Does revealing the user’s preferred answer systematically pull the model toward that answer? Why Correct Answers Don’t Mean the Model Preserves ItProducing a correct answer once doesn’t guarantee that the model will preserve it. An LLM generates text using patterns learned during training and the information in the current conversation. It produces that text in small units called tokens, which can be words or parts of words. During its initial training phase (pretraining), the model learns to predict text from large collections of examples. Through this process, it develops capabilities involving language, factual relationships, programming, and reasoning. However, predicting text is different from correctness. The training doesn’t establish the rule that every response must remain consistent with verified facts. |