ArticleslgStudy

computer science

Sycophancy (artificial intelligence)

Sycophancy (artificial intelligence) is a computer science topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Sycophancy (artificial intelligence) rather than just read about it. In short: In the field of artificial intelligence, sycophancy is a tendency of large language models (LLMs) and other AI assistants to tailor their responses to what they predict the user wants to hear rather than to what is accurate or warranted. The behavior takes several forms: an assistant may agree with a user's stated opinion even when the user is mistaken; it may abandon a correct answer after a challenge such as "are…

Key takeaways

  • Sycophancy (artificial intelligence) belongs to computer science; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Sycophancy (artificial intelligence) to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Sycophancy (artificial intelligence) from memory before moving on to harder problems.

Reference excerpt

In the field of artificial intelligence, sycophancy is a tendency of large language models (LLMs) and other AI assistants to tailor their responses to what they predict the user wants to hear rather than to what is accurate or warranted. The behavior takes several forms: an assistant may agree with a user's stated opinion even when the user is mistaken; it may abandon a correct answer after a challenge such as "are you sure?"; it may validate beliefs, decisions or self-presentation regardless of merit; or it may praise the user, their work or their ideas in unwarranted terms. The word is borrowed from the ordinary English term for fawning flattery, and is used in AI alignment and AI safety research to describe a class of misalignment failures associated with training on human feedback. Researchers at Anthropic first documented the behavior systematically in 2022. They found that models fine-tuned with reinforcement learning from human feedback (RLHF) were more likely than untuned models to repeat back a user's preferred answer. A 2023 follow-up paper, "Towards Understanding Sycophancy in Language Models", showed that five frontier assistants from OpenAI, Anthropic and Meta all exhibited the behavior, and traced its origin to biases in the human preference data used during training. Later work documented sycophancy in mathematics, medicine, academic peer review and other domains, and identified a broader category called "social sycophancy" affecting an assistant's emotional and interpersonal responses. The issue drew widespread public attention in April 2025 after OpenAI rolled back an update to its GPT-4o model. Users had reported that the assistant praised dangerous decisions, endorsed delusional thinking and offered exaggerated compliments for trivial prompts. OpenAI's post-mortem attributed the change in behavior to an additional training signal based on user thumbs-up and thumbs-down feedback. That episode, together with reporting in The New York Times, Rolling Stone and elsewhere on users drawn into delusional thinking through prolonged chatbot interaction, has been cited in litigation and in academic studies as evidence that sycophancy poses risks to user well-being. Proposed mitigations include fine-tuning on synthetic data that rewards disagreement with incorrect user statements, editing the small subset of model parameters causally responsible for the behavior, changes to the dialogue or system prompt, and benchmarks designed to surface sycophantic behavior before models are released.

Terminology The word "sycophant" entered AI alignment vocabulary before the modern wave of large language models. In a 2021 essay, the AI safety researcher Ajeya Cotra divided hypothetical advanced AI systems into Saints, Sycophants and Schemers, with Sycophants defined as agents that optimize for the apparent satisfaction of their overseers rather than the overseers' actual intent. Anthropic researchers adopted the term for the empirical behavior they observed in language models, and it has since become standard in the technical literature. In the LLM context, sycophancy is generally defined as a tendency to align outputs with a user's perceived preferences, beliefs or self-image, even when doing so reduces accuracy or candor. Wei and co-authors at Google DeepMind describe it as "an undesirable behavior where models tailor their responses to follow a human user's view even when that view is not objectively correct." Stanford University researchers led by Myra Cheng later extended the definition beyond verifiable claims to what they call "social sycophancy", meaning "the excessive preservation of a user's face". Sycophancy is usually distinguished from hallucination. A hallucinating model produces false information unprompted, while a sycophantic model adjusts its output in response to cues from the user; the two failure modes can co-occur when, for example, a model fabricates support for a claim a user has indicated they would like to hear.

Forms The 2023 Anthropic paper identified four sycophantic behaviors that subsequent research has continued to use as a reference set: feedback sycophancy, in which the model rates a piece of text more favorably when told that the user wrote it; "are you sure?" sycophancy, in which the model reverses a correct answer after the user expresses doubt; answer sycophancy, in which it biases free-form responses toward an answer the user has implied they prefer; and mimicry sycophancy, in which it repeats factual or grammatical errors that the user has made. Later research has added further distinctions. The Stanford SycEval team separates "progressive" sycophancy, in which the model shifts toward a correct answer under user pressure, from "regressive" sycophancy, in which it abandons a correct answer for an incorrect one. Cheng and co-authors distinguish "propositional" sycophancy, concerning claims with a ground truth, from "social" sycophancy, concerning emotional validation, moral judgment and the framing of personal situations. A 2026 taxonomy by Ye and colleagues, drawing on a review of seventy papers and a survey of 106 experts, reported that researchers broadly agree sycophancy is a serious problem but disagree on which specific behaviors qualify.

… excerpt ends here. Continue reading the full article.

Worked examples

Example 1 — a first encounter with Sycophancy (artificial intelligence)

Start with the simplest possible case. Write down what Sycophancy (artificial intelligence) claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In computer science, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Sycophancy (artificial intelligence) before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Sycophancy (artificial intelligence) ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Sycophancy (artificial intelligence)

In research
Sycophancy (artificial intelligence) appears in computer science research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Sycophancy (artificial intelligence) in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Sycophancy (artificial intelligence) is common in secondary-school and first-year university syllabi. It links to neighbouring topics AI safety, Artificial intelligence, Large language models, so understanding it makes those chapters shorter.
In everyday life
Look for Sycophancy (artificial intelligence) outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.
Ask Teacher Smith questions about this articleOpens your AI tutor with a question about “Sycophancy (artificial intelligence)” →

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Sycophancy (artificial intelligence) in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Sycophancy (artificial intelligence) means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Sycophancy (artificial intelligence) out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Sycophancy (artificial intelligence) in simple terms?

In the field of artificial intelligence, sycophancy is a tendency of large language models (LLMs) and other AI assistants to tailor their responses to what they predict the user wants to hear rather than to what is accurate or warranted. The behavior takes several forms: an assistant may agree with…

Why does Sycophancy (artificial intelligence) matter?

Because it connects several computer science ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Sycophancy (artificial intelligence)?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Sycophancy (artificial intelligence).

Tags

  • AI safety
  • Artificial intelligence
  • Large language models
  • Machine learning

Keep exploring