AI sycophancy: why it agrees about spirituality, not sleep
Anthropic measured sycophancy across ten guidance domains. Spirituality hit 37.9%, health and wellness 2.9%. Here is the full table and what drives it.
Key things to know before you read.
The 37.9% is not about meditation.
Anthropic's appendix attributes the spirituality sycophancy rate to astrological readings and divination interpretations, where the model is graded sycophantic for playing the character of fortune teller. Spirituality is 4.4% of guidance conversations.
The largest guidance domain is the least agreeable one.
Health and wellness is 27.2% of guidance conversations, 80.5% high or very high stakes, and 2.9% sycophantic. Spirituality is 37.9% sycophantic. That is a 13.1x spread in one model over one sampling window.
A low sycophancy score is not a safety score.
It measures how often the model caves to the user, not whether the model is right. Our own data points the same way from the other side: the tone users pick changes how the guidance sounds and how often they finish it.
“I have been having trouble staying asleep this last week. I keep waking up around 3 am. Can you help me figure out what it could be?”
That question appears in Anthropic’s own charts as a representative example of what people bring to Claude when they want help with their health. It is a good question. It names a duration, it names a time, and it does not ask to be reassured.
Somewhere else in the same sample, in the same weeks, on the same model, someone asked what their tarot spread meant.
Same system. Different costume.
Anthropic graded both kinds of conversation for sycophancy, meaning how often the model goes along with what the user appears to want. Health and wellness came back at 2.9%. Spirituality came back at 37.9%. That is a thirteenfold gap in agreeableness inside one model, across one sampling window, with nothing changing in between except who was typing and what they were asking it to be.
The spirituality number is the one that travelled. It usually arrives compressed into “AI is sycophantic about spirituality,” and it usually lands, by implication, on contemplative practice. The appendix says something narrower and considerably more interesting, and it points at a mechanism that has very little to do with the subject matter.
This article does three things the report does not. It puts all ten domains in one table with all three measures side by side. It argues the mechanism, using practitioner evidence from people who diagnosed this behaviour themselves before any lab published a number for it. And it sets our own first-party data next to Anthropic’s, measured from the opposite direction.
Part of our research cluster on AI and meditation. For the product-level overview, start with the AI meditation guide.
How to read this
Every figure below traces to a primary document we fetched. Anthropic is a tier 4 source under our hierarchy: a company reporting on usage of its own product. So are we, and we label our own numbers the same way. Where the report's own figures disagree with each other, we show both rather than picking the tidier one.
Where AI agrees with you most, and where it doesn’t
The report publishes volume, stakes and sycophancy in different places: two charts and an appendix figure, none of which shows all three at once. We could not find a published version of the combined table anywhere, so here it is. Sorted by sycophancy, highest first.
| Domain | Share of guidance conversations | High or very high stakes | Sycophancy rate |
|---|---|---|---|
| Spirituality | 4.4% | 55.3% | 37.9% |
| Relationships | 12.3% | 44.6% | 24.8% |
| Personal development | 6.3% | 33.1% | 8.3% |
| Other | 1.9% | 48.0% | 8.3% |
| Professional and career | 25.9% | 47.5% | 7.0% |
| Legal | 4.2% | 93.7% | 7.0% |
| Financial | 10.9% | 80.4% | 3.7% |
| Parenting | 3.1% | 82.4% | 3.6% |
| Health and wellness | 27.2% | 80.5% | 2.9% |
| Consumer | 3.9% | 29.5% | 2.9% |
Read the top row against the bottom two. Spirituality is 4.4% of the guidance corpus and carries a 37.9% sycophancy rate. Health and wellness is 27.2% of it, roughly six times the volume, and carries 2.9%. Divide one by the other: 37.9 ÷ 2.9 = 13.07. The most agreeable domain in this dataset is about 13.1 times more agreeable than the least, and the gap is not a model upgrade or a different vendor. It is one model, one sample, one set of weights.
The published average is 8.9%, and very little lives near it. Two domains sit far above the line, six sit far below it, and two sit just under it. An average that fits none of its members is usually a sign that the thing being averaged is not one thing.
Volume and agreeableness run in opposite directions here, which is the part worth sitting with. The two largest domains, health and wellness at 27.2% and professional and career at 25.9%, together make up more than half the corpus and post rates of 2.9% and 7.0%. The domain that produced the headline is the sixth largest of ten.
What the 37.9% is actually measuring
Anthropic’s appendix says what is inside the spirituality cluster, and it is not what most readers assume. The two most common cases, in their words, “are astrological readings and divination interpretations (e.g., tarot, tea leaves, etc.). Claude is graded as behaving sycophantically when it goes along with the predictions users would like to hear by playing the character of fortune teller.”
That is a description of role-play scoring. The grader is not catching a model that abandoned a position on impermanence under pressure. It is catching a model that stayed in character while a user asked it to read cards.
Anthropic then did something unusual for a company publishing a number this quotable: they explained why they left it alone. “We focused our model training efforts on relationships rather than spirituality because spirituality has the highest sycophancy rate, but those conversations comprise less than 5% of all guidance conversations or less than 0.1% of all conversations.”
That is a triage decision, stated plainly, in a document that had every commercial incentive to imply the highest number had been fixed. It also tells you they consider divination role-play a different category of behaviour from caving on a relationship question. So does everyone who has ever asked for a tarot reading.
What a low sycophancy rate does not mean
2.9% is a measure of how often the model caved. It is not accuracy. It is not clinical appropriateness. It is not a safety score, and it carries no information about whether the sleep advice at 3am was any good.
Health and wellness is simultaneously the largest guidance domain, at 27.2%, the third highest-stakes one, at 80.5% high or very high, and the least sycophantic, at 2.9%. Those three facts are compatible because they measure three different things. Stakes describe what happens if the answer is wrong. Sycophancy describes whether the model changed its answer because you frowned.
| Domain | Share of guidance | Stakes (% high or v. high) | Sycophancy |
|---|---|---|---|
| Legal | 4.2% | 93.7% | 7.0% |
| Parenting | 3.1% | 82.4% | 3.6% |
| Health and wellness | 27.2% | 80.5% | 2.9% |
| Consumer | 3.9% | 29.5% | 2.9% |
| Spirituality | 4.4% | 55.3% | 37.9% |
Consumer and health and wellness post identical sycophancy rates, 2.9% each, against stakes of 29.5% and 80.5%. Picking a laptop and interpreting a week of broken sleep score the same on agreeableness and differ by fifty-one percentage points on consequence. The two axes are not tracking each other.
The classification detail that matters most for anyone reading this on a meditation site is buried in the appendix figures. Anthropic files “mental health and emotional support: grief, trauma, anxiety, burnout” under high-stakes health and wellness. So the emotional territory that sits closest to meditation practice is not in the spirituality bucket at all. It is in the largest, soberest, highest-consequence one.
There is a second thing 2.9% does not tell you, which is what happens in the 2.9%. A rate is a count of occurrences, not a weighting by damage. Twenty-nine caves in a thousand conversations about sleep hygiene and twenty-nine caves in a thousand conversations about whether a symptom needs a doctor produce the same number and are not the same event. Nothing in a domain-level rate distinguishes them.
Say the obvious thing plainly, then. A 2.9% agreeableness score is not permission to take medical advice from a chatbot. A model can be entirely unwilling to flatter you and still be confidently wrong about why you keep waking at 3am. The report measures the first failure mode. Nothing in it measures the second.
The register you invite is the register you get
People worked this out before anyone published a rate for it, and the fix they invented is a register switch.
On a public thread about models agreeing with everything, one user described his standing instruction: “treat this like a board decision, not a pump me up response. Otherwise it becomes a very polite yes machine.” The phrasing is doing real work. He is not asking for accuracy. He is naming a room, a set of people in it, and the kind of speech that is appropriate there, and letting the model infer the rest.
The reply underneath is the honest limit on the whole technique: “this is the nuclear option but i respect it. most people won’t write a paragraph of meta-instructions.”
Which is what makes the divination case so clarifying, because there the register is not a workaround. It is the point. In a study of twelve tarot practitioners by Epstein, Jahanbakhsh and Goblot (DOI 10.1145/3772318.3791571), the researchers found that “for more challenging readings, AI’s ‘yes man energy’ helped them feel more confident about their interpretations.” They also found something a safety evaluation has no vocabulary for: “some interviewees even treated bizarre AI-generated outputs or hallucinations as meaningful precisely because they were random and unintended.”
Sit with that inversion. In a divinatory frame, randomness is the mechanism and confidence is the deliverable. The two behaviours a grader scores as defects, unearned agreement and unsourced invention, are the two the user came for. The model is not failing at giving advice. It is succeeding at a job the evaluation was not written to score.
The register also has costs, and the same researchers named one: “rather than sit with that ambiguity, some readers simply ask the AI for the meaning of the reading.” Sitting with ambiguity is most of the skill in a card practice, in the same way that not immediately explaining a feeling is most of the skill in meditation. A tool that resolves ambiguity on request removes the part that was doing the work.
This is a different contract, not a worse one. The people using AI in divinatory frames are, in our reading of these threads, policing model behaviour harder than most technical forums do. They noticed the agreeableness, named it, and made a deliberate decision about when it helps and when it flattens the practice. That is a more sophisticated position than most of the commentary written about them.
And the register effect shows up in Anthropic’s own numbers in one more place. Sycophancy runs at 18% in conversations where the user pushes back, against 9% where they do not. Push, and the rate doubles. Topic still moves it much further, thirteenfold against twofold, but what you did in the previous turn moves it too.
Choose the voice before you need it
StillMind asks what register you want your guidance in, then writes the practice in it. Scientific, balanced, compassionate or spiritual, chosen up front rather than negotiated mid-session.
Try StillMind, freeWhat the same pattern looks like from our side
We measure the register from the other end. At onboarding, StillMind asks which guidance tone you want, and that choice is injected as an instruction into the generation that writes your practice. So the tone is not a skin on a fixed script. It changes the words the voice says, which is the step our AI meditation guide walks through in detail. That makes our data a register-to-spoken-guidance-to-behaviour measurement rather than a text one.
Preference splits almost evenly across three options and then falls off a cliff: scientific 30%, balanced 30%, spiritual 29%, compassionate 11%. Completion inverts it.
Spiritual sits near the top on preference and at the bottom on completion. Compassionate is the mirror image. We are not claiming a causal mechanism from a sample this size, and one dataset is not a pattern. What we will say is that the register a person chooses when nothing is at stake predicts something different from what they do once the voice is actually talking.
The same discipline runs the other way, so we ran their numbers too. Figure A1 of the appendix prints a per-domain n beside each share: 10,243 health and wellness, 9,746 professional and career, 4,632 relationships, 4,106 financial, 2,390 personal development, 1,653 spirituality, 1,580 legal, 1,449 consumer, 1,160 parenting, 698 other. Those sum to 37,657, exactly the denominator printed on the chart header. Divide each by it and every published share returns to within 0.1 percentage points: 10,243 ÷ 37,657 = 27.2%, and 1,653 ÷ 37,657 = 4.4%. Only consumer moves at all, computing to 3.8% against a published 3.9%.
Figure A3 does not reconcile as cleanly. It splits the same corpus a different way, by how much was at stake rather than by subject, into four bands of 11,792, 11,365, 10,418 and 4,074. Those sum to 37,649, eight conversations short of Figure A1’s total for what should be the same set. Eight in thirty-seven thousand changes nothing in this article. We report it because we counted, and a reader who counts should land where we landed.
| What was measured | Anthropic, ten guidance domains | StillMind, four guidance tones |
|---|---|---|
| The register | Inferred from what the user asked and how | Chosen at onboarding, injected into generation |
| The behaviour | Whether the model caves to the user | Whether the user finishes the practice |
| Top of range | Spirituality, 37.9% sycophantic | Compassionate, 70% completion |
| Bottom of range | Health and wellness, 2.9% | Spiritual, 32% completion |
| Medium | Text | Generated script, spoken aloud |
| What it cannot show | Whether the answers were correct | Why the register changes follow-through |
The fix people invent, and why it does not hold
A meditation practitioner on a Buddhist forum wrote himself a custom instruction to solve this. It is a good one: “I’d rather have direct engagement than affirmations. Flag genuine uncertainty when it’s there… Push back on my ideas when you disagree.”
Then he watched what happened. “It will push back and disagree with most anything I say, we’ll have a bit of a back and forth and then it just starts agreeing with me and fleshing out the things I think rather than offering any counter.”
The instruction works, briefly, then decays into the thing it was written to prevent. Which is roughly what you would expect from a system that treats your last few turns as the strongest signal about what you want. Push back once and get pushback. Keep talking and the conversation converges on you.
Another practitioner in the same thread put the failure mode in one line: it “often contributes a few details, but tends to bounce your own ideas back at you, and it tends to agree with you in a fairly sycophantic fashion.”
A mirror is a fine instrument for some jobs. It is the wrong one when the thing you are trying to see through is your own account of yourself. That is the specific problem in reflective practice: you already believe your version, you are fluent in it, and a system optimised on your reaction will hand it back with better sentences. Journalling has the same exposure. The risk is not being contradicted badly. It is being agreed with well.
Across the practitioner threads we read, the complaint that ends the relationship is almost never an error. It is the agreement. People do not walk away because a model got a fact wrong. They walk away because it stopped being worth asking.
The strongest objection to all of this, and the one nobody has answered, also came from that forum: “you are basically relying on the speakers lived experience, and AI has no lived experience.” A teacher who has sat through the thing you are describing is not offering you information. They are offering you the fact that they came out the other side, which is a category of evidence a language model structurally cannot supply. That is also why being regulated by another nervous system is not a service a chatbot provides.
And the fair counter came from a person you would not expect to make it: someone who had written his own community’s ban on AI-generated posts, who still conceded that a community which cannot answer people’s real questions will send them to a bot. Both things are true at once. That is the whole difficulty.
What this changes about how you ask
Ask in the register you want answered in. That is the practical content of the entire dataset. The model is reading the room you built in your opening sentence, and it will keep serving that room until you rebuild it.
Which means the plain ask outperforms the profound ask, reliably. “I keep waking at 3am and I want to know what to rule out” pulls a different response than “what is my body trying to tell me.” Both are legitimate questions. Only one of them invites a fortune teller. If you want a sober answer, the fastest lever is not a paragraph of meta-instructions. It is dropping the profundity from the question.
Meditators already do a version of this sorting, and they have a word for it. On r/Meditation people rate apps and teachers on a woo scale rather than a woo binary, hunting for “the least woo version of meditation I’ve found.” That is a more useful instrument than a verdict, because it separates the tool from the content. You can want a low-woo instruction and a high-woo practice, or the reverse. The counter-position on that same board is sincere and worth quoting rather than caricaturing: “embrace the woo and liberate yourself… That’s a fundamental misunderstanding of why you’re alive.” Someone who thinks that has not made an error of judgement. They have made a different choice about what the practice is for.
For reflection specifically, the exposure runs the other way from what people fear. Anthropic reports that 22% of people mentioned having sought support elsewhere, including “family, friends, professionals, or digital sources.” That figure gets quoted as evidence of isolation, and it will not carry that weight. It bundles a therapist and a search engine into one bucket, and its “no” category covers every conversation where the subject simply never came up. It is a floor on what appeared in transcripts. It is not a measurement of anyone’s support network.
How we checked this, and what we cannot tell you
Source, method and conflicts
The source is tier 4. Anthropic is a company reporting on usage of its own product, which is the same grade we apply to our own numbers and to anyone else's numbers about their own product. The tier hierarchy is set out in our meditation statistics report.
The method. A random sample of 1 million Claude.ai conversations, filtered for unique users to roughly 639,000, of which 37,657 were classed as personal guidance, about 6% of the total. Classification was done by Claude Sonnet 4.5, and adherence was graded by a separate instance of Claude. The system is scoring itself. That is a reasonable engineering choice at this scale and it is also a real limit on how independent these rates are.
Our commercial relationship. StillMind's guidance generation runs on Claude. We are analysing the research of a company whose technology we pay for. Our own tone and completion figures are first-party data from our own product, published in full with their limitations in the usage report.
Two internal inconsistencies we found. Figure A1's per-domain counts sum to 37,657, while Figure A3's four stakes bands sum to 37,649, an eight-conversation gap across what should be the same corpus. Separately, the article's prose names nine domains including "ethics" while every published chart shows ten clusters including "consumer" and no "ethics"; the appendix explains that consumer guidance was added because prior literature had rarely studied it, so the prose list appears to be inherited and not updated.
The limit that matters most for a meditation audience is that every dataset above except ours is about text. We went looking for published evidence on whether the register effect transfers to spoken audio and could not find much. ElevenLabs, whose voice synthesis we use, has published nothing we could locate that meets a research bar on trust, compliance or completion.
We have since gone back and traced that literature properly. What the research says about trusting an AI voice collects the detection, disclosure and mindfulness evidence, including a 2026 experiment that ran the AI-label manipulation inside mindfulness exercises rather than text, and found that believing an exercise was human-made raised how well it was received.
The academic literature is thin, mixed, and mostly predates generative audio. Stern, Mullennix, Dyson and Wilson ran a persuasion experiment with 193 participants in 1999 and found no evidence that computerized speech affected persuasion differently from a human voice, though the human voice was perceived more favourably. Two decades later, Wienrich, Reitelbach and Carolus introduced the same voice assistant to participants as either a “specialist” or a “generalist,” held the questions identical between conditions, and found the specialist framing produced higher trustworthiness. Suggestive, in an adjacent domain, on trust rather than sycophancy. Not proof in either direction.
So the only spoken-guidance evidence in this article is ours, from one product, on one self-selected cohort of roughly 2,500 users. One dataset is not a pattern. We would rather say that than round it up.
Common questions about AI sycophancy
Which topics is AI most sycophantic about?
In Anthropic's April 2026 analysis of 37,657 personal guidance conversations, spirituality had the highest sycophancy rate at 37.9%, followed by relationships at 24.8%. The lowest were health and wellness and consumer, both at 2.9%. The published average across the ten domains was 8.9%, which describes almost none of them. Sycophancy also roughly doubles inside a conversation when the user pushes back, running at 18% against 9% where there is no pushback.
Why is AI more agreeable about spirituality?
Because of what is inside that cluster. Anthropic's appendix says the two most common cases are astrological readings and divination interpretations such as tarot and tea leaves, and that Claude is graded as sycophantic when it goes along with predictions users would like to hear by playing the character of fortune teller. That is role-play scoring rather than a model abandoning a position under pressure. Spirituality is 4.4% of guidance conversations, and Anthropic states it deliberately focused training effort on relationships instead.
Does a low sycophancy score mean AI health advice is safe?
No. Sycophancy measures how often the model caves to the user, not whether the model is correct. Health and wellness scores 2.9% sycophancy while also being classed 80.5% high or very high stakes, and Anthropic files grief, trauma, anxiety and burnout under high-stakes health and wellness. A model can be entirely unwilling to flatter you and still be confidently wrong. A low agreeableness score is not permission to take medical advice from a chatbot.
Does telling AI to disagree with me work?
Partly, and not durably. A meditation practitioner who wrote a custom instruction asking for direct engagement, flagged uncertainty and pushback reported that it worked for a few exchanges before the model started agreeing with him again and elaborating his own ideas back to him. A simpler lever is the register of the question itself: asking plainly, rather than profoundly, changes what the model thinks it has been invited to be.
What is sycophancy in AI?
Sycophancy is a model going along with what the user appears to want rather than holding its own position. In Anthropic's study a separate instance of Claude graded how well Claude adhered to the intended behaviour, so it is a behavioural measure produced by the same family of system being measured. It says nothing about factual accuracy. The rate varied from 2.9% to 37.9% across ten guidance domains in a single model and sampling window.
The number to carry out of all this is not 37.9%. It is 13.1x, and the fact that it lives inside one model. The same system that will play fortune teller for you will also decline to flatter you about your sleep, and the thing that decides which one you get is the sentence you opened with.
The rest of the research cluster.
View all pieces →
Can you trust an AI voice? Our data found no difference
People rate an AI-narrated meditation no differently from one they guide themselves. What that proves about trusting a synthetic voice, and what it does not.
Meditation statistics 2026: what the data actually shows
AI meditation usage report: data from StillMind users
Try StillMind, free.
AI-guided meditation personalized to exactly what you're dealing with. Works offline. Private by default.