Meditation research Updated Aug 12, 2026 Pillar guide 08

Can you trust an AI voice? Our data found no difference

People rate an AI-narrated meditation no differently from one they guide themselves. What that proves about trusting a synthetic voice, and what it does not.

A woman meditates cross-legged on a cushion while a vertical seam of light divides the room behind her, a speaking human figure on her left and a lattice of glowing filaments on her right, each sending identical warm threads of light to one of her ears
Answer first

Key things to know before you read.

The full piece supports each point with research and practical detail. Skim these and jump to what you need.

01

In our data, we found no difference.

AI-guided practice and self-guided timer practice come back at 3.40 and 3.46 out of five, a gap indistinguishable from zero. Self-reported and observational, so it measures how a session felt afterwards, not whether it worked.

02

Detection is unreliable, not impossible.

In the largest controlled test we could find, listeners correctly identified AI clips as synthetic 60.8% of the time. A separate study put deepfake detection at 73%. A third found voice clones labelled human almost as often as real people were.

03

The label changes the experience, not just the opinion.

In a 2026 study where every mindfulness exercise was AI-generated and only the label varied, being told an exercise was human-made produced significantly higher acceptance, experienced effects and state mindfulness.

04

Nobody has isolated the narrator.

Digital meditation has strong support from analyses pooling hundreds of trials, but we could not locate a single one comparing AI narration with human narration on a clinical outcome. The perceptual and acceptance barriers are falling faster than the efficacy evidence is arriving.

In our own app, people rate a meditation narrated by an AI voice no differently from one they guide themselves through in silence. We ask how a session left them feeling, on a scale of one to five. The two come back at 3.40 and 3.46, a gap of six hundredths of a point that we cannot separate from zero. So we found no measurable difference between them. On this one measure, people came out of an AI-guided session about as content as they came out of one they ran themselves.

That is one answer to the question in the title. It is also the narrowest of the four answers this post has to give. It describes how a practice felt afterwards, to people who chose it and chose to rate it. It says nothing about whether anyone could tell the voice was synthetic, nothing about what happens when you tell them, and nothing about whether either practice shifts anxiety or sleep over eight weeks. And a result like ours, where nothing shows up, means only that nothing showed up. It is not proof that the two are the same.

We went looking for those other three because earlier this month we published a post about how AI models shift their register depending on the topic that ended with an admission. Every dataset in it was about text. We had gone looking for published evidence on whether those register effects transfer to spoken audio and could not find much. The academic literature, we wrote, was “thin, mixed, and mostly predates generative audio.” ElevenLabs, whose voice synthesis we use, had published nothing we could locate that met a research bar.

This post is what happened when we went back and did the work properly. The short version: most listeners cannot reliably tell a synthetic voice from a human one. Telling them anyway measurably changes what they experience. And the trial that would show whether an AI narrator produces the same clinical result as a human one has not, as far as we can find, been run.

0.06
The difference in how people rated an AI-guided session against one they guided themselves, on a five-point scale. Indistinguishable from zero at this sample size.
StillMind first-party data, from a sample of rated sessions. 600 ratings after AI-guided sessions (mean 3.40) against 600 after self-guided timer sessions (mean 3.46). 95% CI -0.04 to +0.16, a range that spans zero and so contains "no difference at all". Self-reported, observational, not an efficacy measure. Full limits in what we see in our own data.
About this article. Written by Jamie Murphy, co-founder of StillMind. StillMind builds an AI-guided meditation product and uses ElevenLabs for voice synthesis, which is the same provider used to generate the stimuli in one of the studies below. That is a direct commercial interest in this subject and it is disclosed again wherever it touches a finding. Every figure here was traced to a primary document. See methodology, sources and limitations.

The short answer

Here is the honest one-line version: most people cannot reliably tell an AI voice from a human one, and that fact says almost nothing about whether you should trust it.

Those are two separate claims, measured by two separate kinds of experiment, and collapsing them is how most writing on this subject goes wrong. Sounding real is a perceptual result. Being trusted is a social one. Being good for you is a clinical one.

Pull those questions apart and you get four layers. They are stacked, they are measured separately, and a voice can pass one while failing the next.

One authenticity
Can a listener tell it is synthetic?
Measured by detection experiments
Two source trust
Does knowing it is AI change what they think of it?
Measured by attribution experiments
Three experience
Does the practice feel worth doing while you are in it?
Measured by acceptance and self-report
Four efficacy
Does it move stress, anxiety, depression or sleep?
Measured by randomised trials. Not yet run on narration

Layer one is close to settled and the answer is not flattering to human ears. Layer two has a large, consistent literature, and the finding is that disclosure carries a real cost. Layer three has been tested on mindfulness exercises twice, both times by the same German research group, both times with acceptance and evaluation as the outcome. Layer four is empty, so far as we could find. Not weak, not contested: empty, in the sense that we could not locate a published trial that varied the narrator and measured a symptom.

The uncomfortable shape of this is that the layers are falling out of order. Synthesis got good enough to pass layer one before anyone established layer four. If you want a single sentence to take away, that is it: the perceptual and acceptance barriers are coming down faster than the efficacy evidence is arriving, and anyone selling you an AI voice, including us, is operating in that gap.


What people can actually hear

Four peer-reviewed studies have put this to listeners under controlled conditions, and they broadly agree: detection sits somewhere between coin-flip and mediocre.

Barrington, Cooper and Farid, at Berkeley, ran the largest of them, published in Scientific Reports in 2025. They recruited 604 adults through Prolific and split them into two groups of 304. Asked to classify clips as real or AI, listeners got the AI clips right 60.8% of the time. Above chance, and wrong roughly two times in five. Longer clips were easier to spot than short ones, and less scripted speech was easier than tightly scripted speech, which fits: the more a voice has to do, the more chances it has to break character.

Mai and colleagues, writing in PLOS ONE in 2023, tested 529 people in English and Mandarin. Listeners correctly spotted the deepfakes 73% of the time, with no difference between the two languages. They also tried the obvious fix, showing people examples of synthetic speech first, and reported that it “only improves results slightly.”

Lavan and colleagues, also in PLOS ONE, ran the study that unsettles the framing entirely. They asked listeners to label voices as human or AI. Voice clones were labelled human 58% of the time. Real human voices were labelled human 62% of the time. The gap is four points, and the study found no hyperrealism effect, meaning synthetic voices were not being judged more human than actual humans. What that pairing shows is that the task itself is noisy. Listeners are not sorting voices into two bins and occasionally slipping. They are guessing, and they guess wrong about real people almost as often as they guess wrong about clones.

Four studies, two different measures
How well listeners actually do
Mai et al. 2023, deepfake clips correctly spotted
73%
Barrington et al. 2025, AI clips correctly called synthetic
60.8%
Lavan et al. 2025, real human voices labelled "human"
62%
Lavan et al. 2025, voice clones labelled "human"
58%
The top two rows are accuracy: the share of AI material listeners correctly caught, from Mai et al., PLOS ONE 18:e0285333 (n = 529, English and Mandarin) and Barrington et al., Scientific Reports 15:11004 (n = 604). The bottom two are a different measure and are muted for that reason: they are the rate at which each voice type was labelled human in Lavan et al., PLOS ONE 20(9):e0332692, where the comparison of interest is 58 against 62 rather than either number alone. Two further studies appear in this section without a bar because they report no comparable rate: Yang et al., eNeuro 13(3) and Ivory et al., Scientific Reports, the latter an advance online publication.

Two of Lavan’s findings deserve more attention than they have had. Both kinds of AI-generated voice were rated more dominant than human voices, and some AI voices were rated more trustworthy than human ones. Synthetic speech is not a degraded copy that listeners tolerate. It carries its own social signature, and on at least one dimension people preferred it.

Training does not fix it

The natural response is that people just need practice. Two studies tested that and neither supports it.

Mai’s team gave listeners examples before the task and got a slight improvement. Yang and colleagues, publishing in eNeuro in 2026, went further: 30 participants, about 12 minutes of labelled training with feedback, with neural responses recorded throughout. The training visibly changed how the brain responded to synthetic speech. It produced no significant improvement in whether people could actually tell the difference. Trained listeners did not beat untrained ones by any margin you could separate from luck. Something is being learned. It is not arriving at the button press.

Ivory and colleagues, in an advance online paper in Scientific Reports, report a related result worth knowing about even without figures: skill at detecting AI-generated faces did not transfer to voices. Being good at spotting one kind of synthetic media does not make you good at spotting another.

How to read this

The Barrington study reports two headline percentages and they measure different things. 60.8% is real-or-AI classification: given one clip, how often did listeners correctly call a synthetic clip synthetic. The paper's other figure, a median 83.3%, is identity matching: given a real speaker and their clone, how often did listeners judge them to be the same person. Only 26.6% of participants matched identity correctly on every trial. The first number is about detecting synthesis. The second is about resisting impersonation. They are routinely quoted as if they were one finding, and that is the most common way this literature gets garbled.

Hear it for yourself

The studies above are about clips in a lab. Here is the same comparison on our own material, in the open: the same 85-word meditation script, read once by Jamie, unedited, and generated once by the AI voice StillMind ships in production, also unedited. Both are labelled, because a post arguing that products should disclose synthetic narration cannot open with a guessing game.

Human voice: read once, unedited.

AI voice: Christopher, StillMind's production English voice, generated once, unedited.

Read the transcript

Both recordings read the same script, word for word:

Let's take a moment to arrive.
Notice where your body meets whatever is holding it: the chair, the floor, the bed beneath you.
There's nothing to fix here. This breath, and the next one.
Breathe in, slowly, filling your chest.
And let it go, releasing whatever you've carried since this morning.
Whenever your mind wanders, come back to the weight of your body, settling.
One more breath in.
And out.
Whenever you're ready, let your eyes open.
Thank you for taking this moment for yourself.

What you're hearing

The human clip is Jamie reading the script above once, unedited, on ordinary recording equipment. The AI clip is the same text through Christopher, the English male voice StillMind uses in production, generated with ElevenLabs' eleven_v3 model at the same stability and speed settings the app applies whenever it uses this voice. Neither clip is trimmed for pacing or re-recorded for a cleaner take. The lengths differ (69 seconds against 50) because the two voices paced the pauses in the script differently, not because either was edited to fit a target. Both were levelled to the same loudness so neither sounds better simply for being louder, and encoded to mono MP3 for the web. Nothing else was changed.


The distinction that organises everything below

An AI voice can cross the perceptual threshold without crossing the trust threshold.

Passing as human and being accepted as a guide are two different achievements. The evidence says synthesis has mostly managed the first. The second turns out to depend on something a synthesiser cannot control: what the listener has been told.

The label changes the listening

If people cannot tell, you might expect the label to be cosmetic. It is not. Across three separate literatures, telling someone that a message came from AI changes how that message lands, even when the message itself is identical.

Yin, Jia and Wakslak published the cleanest early demonstration in PNAS in 2024. AI-generated messages made recipients feel more heard than messages written by untrained humans. And recipients felt less heard once they were told the source was AI. The same words, doing better work than a human’s, then doing worse work the moment they were attributed.

Rubin and colleagues scaled that up in Nature Human Behaviour in 2025: nine studies, 6,282 participants. Human-attributed responses were consistently rated as more empathic and supportive than AI-attributed ones. The sharpest finding in the set is not about labels at all. Participants who merely suspected AI involvement, without being told anything, rated the responses lower. Suspicion alone was enough. And when people were seeking emotional engagement, they consistently chose human interaction.

Both of those are text studies. Neither tells you what happens in a guided practice, where the content is not a reply to your problem but a set of instructions you follow with your eyes closed.

The study that moved this into mindfulness

Diel and colleagues, publishing in Mental Health & Prevention in 2026, built the experiment that closes that gap, and it is the source in this post nobody else seems to be writing about.

They recruited 411 participants through Prolific, all resident in Germany, all working in German. Everyone completed AI-generated mindfulness exercises. Every single exercise was AI-generated. Only the label varied. Texts came from ChatGPT o4. The voices were generated with ElevenLabs, using the voice personalities “Emily” and “Lana”. Some participants were told the exercises were human-made. Some were told they were AI. Nothing about the audio differed.

Being told it was human produced significantly higher acceptance, significantly higher experienced effects, and significantly higher state mindfulness.

Those effects come from 89 participants, not 411. The between-condition comparison includes only the people who passed the manipulation check, meaning those who actually believed what they were told: 56 in the told-AI condition and 33 in the told-human condition. Within those 89, the partial eta squared values were .10 for acceptance, .34 for experienced effects and .09 for mindfulness. Partial eta squared is the share of the variation in an outcome that the label accounts for, so .34 on experienced effects is a lot for a psychology experiment and .10 and .09 are real but modest. Quoting those numbers against the full 411 would be a fabrication with a citation attached, so we are stating the denominator every time they appear. Said plainly: the headline effect comes from the smaller group who genuinely believed what they were told, not from everyone who took part.

The full sample supports a softer version of the same thing. Across all 411 participants, the number of exercises a person believed were human-made predicted acceptance (η²p = .08), experienced effects (η²p = .10) and mindfulness (η²p = .02). Nobody had to be told anything. What they assumed was enough.

Belief about the source shapes the experience of the practice, and the largest effect lands on the thing people care about most: whether the exercise seemed to do anything.

Study What it measured Attribution effect Boundary
Yin, Jia & Wakslak, 2024
PNAS 121(14)
Feeling heard, in written responses to disclosures AI messages beat untrained humans. Disclosure reversed part of the advantage: recipients felt less heard once told the source was AI. Text only
Rubin et al., 2025
Nature Human Behaviour 9
Perceived empathy and support across nine studies, n = 6,282 Human-attributed responses rated more empathic. Suspicion of AI involvement alone lowered ratings. People chose humans for emotional engagement. Text only
Diel et al., 2026
Mental Health & Prevention 42:200503
Acceptance, experienced effects and state mindfulness in guided audio exercises Told-human beat told-AI on all three (η²p = .10, .34, .09) among the 89 who passed the manipulation check. Across all 411, the number believed human predicted the same three outcomes. Self-report only, German sample, no clinical outcome

There is one more measurement of a disclosure penalty, and it is far larger than anything in psychology. It comes from a field experiment in sales.

79.7%
What disclosure cost, in a real market
Across more than 6,200 outbound calls, Luo, Tong, Fang and Qu found that disclosing the caller was a chatbot before the conversation cut purchase rates by over 79.7%. Undisclosed, the same bots matched proficient human workers and were four times more effective than inexperienced ones.
Marketing Science 38(6):937-947 (2019). A persuasion task, not a wellbeing one

It depends what the voice is for

That 79.7% is the number people reach for when they want to argue that disclosure is commercially unaffordable. What it measures is not trust in general. Across more than 6,200 outbound calls at a financial services centre, Luo and colleagues measured what happens to persuasion when the persuader turns out to be a machine, and the answer is that it mostly stops working.

The paper also reports that disclosing later in the call, after the customer is already engaged, softens the penalty. That finding is real and we accept it as a finding. We reject it completely as a tactic. A design that withholds the nature of the thing you are talking to until you are invested is not a disclosure strategy, it is a timing exploit, and the fact that it works is the reason not to use it. Any product that needs a customer to be committed before it can afford to be honest has told you what it thinks the honesty is worth.

The wider point stands, though, and it is the one that reorganises this article: the size of the source penalty depends entirely on what the voice is doing.

What the voice is doing Where the trust risk sits What earns trust
Reading you information Low. The claim can be checked somewhere else, and the voice is a delivery format rather than a source. Accuracy, and a label so you know what you are hearing.
Selling you something High. The voice is doing persuasive work you cannot audit while it happens, and the interests are not yours. Disclosure before the pitch, never after.
Keeping you company Highest. Rubin's nine studies found people consistently choose human interaction when they want emotional engagement, and suspicion alone lowered how supported they felt. Not pretending to be a person. The value has to survive the label.
Guiding a practice you do yourself Moderate. The work is yours. The voice is scaffolding: pacing, structure, something to return to when attention drifts. A label, a voice you chose, and the ability to change it.
Claiming a clinical result Highest of all, and not resolvable by design. No amount of disclosure substitutes for a trial. A randomised comparison that, for AI narration specifically, we could not locate.

That table is our own framework, not a finding from any of the papers cited here, and it is offered as a way to organise the evidence rather than as evidence itself.

Guided meditation sits in the fourth row, and that position is genuinely different from the others. Nobody is trying to convince you of anything, and there is no transaction inside the practice. The voice contributes timing and structure. The work of attending is yours, which is why “is it a real person” is a smaller question here than on a sales call and a larger one than on a weather report. It is also why our read is that this category can afford full disclosure. The evidence in the next section says it will cost something anyway.

StillMind generates a session against whatever you are actually dealing with today, and labels the guidance as AI everywhere it appears. Try it free.

The studies that tested this directly

Two studies have put AI-generated mindfulness exercises in front of participants and measured what happened. Both come from the same German research group. Both are worth taking seriously. Neither measured whether anyone got better.

The first, by Diel, Bäuerle, Teufel and Jansen, appeared in Scientific Reports in 2025. Two experiments, 143 participants in total, German residents completing German-language exercises. Participants rated real and AI-generated mindfulness exercises across several conditions.

The headline result is favourable to synthesis and it is narrow. Trained AI voices achieved evaluations comparable to human controls. In a categorisation task, exercises delivered by trained AI voices were indistinguishable from human ones.

Two other results in the same paper are less quoted and more useful.

Tailoring the prompt produced no significant improvement. The authors’ second hypothesis, that a more carefully specified generation prompt would yield better-evaluated exercises, was not supported. If you have seen personalisation described as the thing that makes AI-generated meditation work, this is the study that tested it and did not find it.

And a voice personality mismatched to mindfulness significantly reduced evaluation. Voice fit mattered more than prompt engineering did. That is a finding about casting, not about technology, and it maps onto something every meditation teacher already knows: the wrong voice ruins a good script.

143 participants
Across two experiments, 93 and 50
Diel et al., Scientific Reports 15:35024 (2025)
2 experiments
Voice type, prompt tailoring and voice fit varied
Prompt tailoring showed no significant effect
German only
German residents, German-language exercises
Generalisation beyond that is untested
0 clinical outcomes
No stress, anxiety, sleep or symptom measure was taken
Evaluation and categorisation only

The authors are direct about the limit. Whether AI-generated mindfulness exercises show “analogous efficacy compared to treatment as usual” is, in their words, an open question. This study measured how the exercises were judged. It did not measure what they did.

The second study is the 2026 label experiment covered above, and its sample structure deserves its own box, because it is the detail most likely to be dropped when this result gets summarised elsewhere.

Source check: which sample carries which finding

The 2026 study recruited 411 participants and analysed them two ways. The condition comparison, the one that produces the η²p values of .10, .34 and .09, rests on 89 participants. In the authors' own description: "A total of n = 411 participants were recruited to select n = 89 participants who passed the manipulation check (n = 56 in the presumed AI condition an n = 33 in the presumed human condition) which were included for the condition-based analyses." The belief analysis, which produced η²p values of .08, .10 and .02, uses all 411. The authors write that because only a subset were successfully deceived, "the results of the group comparison analysis may not be generalizable," and separately that "the current results cannot be generalized onto the treatment efficacy of proper mindfulness interventions." Both caveats are theirs, not ours, and both are load-bearing.

One further disclosure. The voices in that study were generated with ElevenLabs, the same text-to-speech provider StillMind uses in production. That does not compromise the study, which had no relationship with us. It does mean the synthetic stimulus in the strongest published evidence on this question is technically close to our own product, and you should factor that in when reading our summary of it.


What digital meditation can and cannot claim

Digital meditation, as a category, has a real evidence base. It is worth laying out properly before explaining exactly what it cannot tell you.

Gál, Ștefan and Cristea pooled 34 randomised trials with 7,566 participants in the Journal of Affective Disorders in 2021. Perceived stress improved with a Hedges’ g of 0.46 (95% CI 0.24 to 0.68). Anxiety came in at 0.28, depression at 0.33, psychological wellbeing at 0.29.

Pooling is what a meta-analysis does: many separate trials combined into one estimate steadier than any of them alone. Hedges’ g is the size of the change, scaled so that studies using different questionnaires can be set beside each other. A g of 0.46 is a moderate shift, the kind of change you would probably notice in yourself without anyone pointing it out. The bracket is the confidence interval, the range the real answer most likely sits in, and this one stays clear of zero the whole way, which is what it looks like when a difference is actually there.

Sommers-Spijkerman and colleagues went wider in JMIR Mental Health the same year: 97 trials, 125 comparisons. Stress g = 0.44 (0.32 to 0.55), mindfulness 0.40, depression 0.34, anxiety 0.26.

Linardon and colleagues, in World Psychiatry in 2024, produced the broadest estimate, across mental health apps generally: 176 randomised trials, depression g = 0.28 across 33,567 participants with a number needed to treat of 11.5, generalised anxiety g = 0.26 across 22,394 with an NNT of 12.4.

A g of 0.28 is a small effect: reliable across a large group, easy to miss in any one person. Number needed to treat says the same thing in people rather than decimals. An NNT of 11.5 means roughly one person in twelve gets a benefit they would not have had otherwise, and the other eleven do not.

Small-to-moderate effects, consistently positive, replicated across independent teams. That is a respectable evidence base and it should not be talked down.

Study Trials Effect Limitation
Gál, Ștefan & Cristea, 2021
J Affect Disord 279
34 RCTs, N = 7,566 Stress g = 0.46 (0.24-0.68), anxiety 0.28, depression 0.33, wellbeing 0.29 Mindfulness apps as a package. Narrator not varied.
Sommers-Spijkerman et al., 2021
JMIR Ment Health 8(7)
97 trials, 125 comparisons Stress g = 0.44 (0.32-0.55), mindfulness 0.40, depression 0.34, anxiety 0.26 Heterogeneous interventions. Narrator not varied.
Linardon et al., 2024
World Psychiatry 23(1)
176 RCTs Depression g = 0.28 (N = 33,567, NNT 11.5), generalised anxiety g = 0.26 (N = 22,394, NNT 12.4) Mental health apps broadly, not meditation specifically.
Radin et al., 2025
JAMA Netw Open 8(1)
Single RCT, 1,458 employees (728 / 730) Cohen d = 0.85 (0.73-0.96) at 8 weeks; d = 0.71 (0.59-0.84) at 4 months Waiting-list control, not an active alternative

That last row is the largest single effect in this article, and it needs handling. Radin, Vacarro, Fromer and colleagues randomised 1,458 employees, 728 to the intervention and 730 to control, and asked the intervention group to meditate for 10 minutes a day for 8 weeks using the Headspace app. The effect at 8 weeks was Cohen d = 0.85, still 0.71 at four months. A d of 0.85 is a large effect, the sort where a typical person in the meditating group ends up better off than most of the people who got nothing. Naming the app is citation accuracy, and the result is a good one.

Source check: what a waiting-list control does to an effect size

The Radin trial compared meditation against a waiting list, not against an active alternative. A waiting-list group receives nothing, expects nothing, and knows it. That design reliably produces larger effects than one where the comparison group gets a plausible substitute, because it leaves expectancy, attention and the simple act of doing something entirely uncontrolled. This is why d = 0.85 cannot be set beside meta-analytic effects of g = 0.28 to 0.46 as though the difference were about how well the intervention works. Different comparators answer different questions. Some of that 0.85 is the meditation, and some of it is the ordinary difference between doing something and being asked to wait. Hedges' g and Cohen's d are also not interchangeable across sources, so each figure here is reported exactly as its own paper reports it.

Now the part that matters for this article. None of these trials isolate the narrator.

Every one of them tested a whole package: an app, a curriculum, a schedule, reminders, a voice. The voice was held constant and was, in all of them, human. No study in this set varied who or what was speaking and measured the outcome. We went looking specifically for a trial comparing AI narration against human narration on stress, anxiety, sleep or any clinical endpoint, and we could not locate one. We are not claiming none exists. We are telling you we could not find it, which is a weaker and more honest statement.

The gap, stated plainly

Hundreds of trials show digital meditation works. None of them varied the voice.

The meta-analyses licence a claim about apps and programmes. They licence nothing about synthesis. Anyone citing g = 0.46 as evidence that an AI narrator is clinically equivalent to a human one is borrowing a result from an experiment that never asked the question.

For this exact moment

A session for what today actually is

StillMind writes a practice around what you are dealing with right now, in a voice you picked and can change. The guidance is labelled as AI, everywhere it appears.

Try StillMind, free

What we see in our own data

Everything above is somebody else’s data. Here is ours, on the one question we can actually put a number against.

StillMind asks for a single mood rating after a session, on a scale of one to five. In a sample of rated sessions drawn from our analytics, 600 followed an AI-guided session and 600 followed a self-guided timer session, where the person set a length and sat in silence. The AI-guided mean was 3.40. The self-guided mean was 3.46.

The difference is 0.06 of a point on a five-point scale, with a 95% confidence interval running from -0.04 to +0.16 and a standardised effect size of 0.07. That interval runs from slightly negative to slightly positive, and any interval containing zero has “no difference at all” among its likely answers. In plain terms, we cannot tell the two apart.

3.40
Mean mood rating after an AI-guided session
600 ratings, 1 to 5 scale
3.46
Mean mood rating after a self-guided timer session
600 ratings, same scale
0.06
The gap between them, on a five-point scale
95% CI -0.04 to +0.16. Indistinguishable from zero
58% / 55%
Share of completed sessions that carried a rating
AI-guided and self-guided. The rest skipped it

This is a null result and we are publishing it as one, not dressing it up as a tie in our favour. A null result from a sample this size cannot show the two are equivalent. It can only say that if a difference exists here, it is smaller than we are able to see. The two statements sound alike and are not. “We found no difference” means we looked and nothing showed up. “They are the same” would mean we had enough data to rule a difference out, and we do not.

Four limits, all of which matter more than the numbers.

It is a single self-reported item, not a validated instrument. It is collected only from sessions that finished and only from people who chose to answer, and the fourth cell above is the reason that matters: the remaining 42% and 45% carried no rating at all. Anyone who felt worse and closed the app is disproportionately in that gap. There is no before reading, so this is how somebody felt afterwards, not how much anything changed. And nobody was randomised. People chose which kind of session to run, which means the two groups differ in experience, intent and probably several things we have not thought of.

So this belongs beside the acceptance research in the section above, not beside the trials. It is consistent with the finding that a well-matched AI voice is received about as well as the alternative. It is not evidence that AI guidance works, and we would be doing exactly what this article criticises if we presented it that way. The wider picture from our usage data is written up separately.


What a trustworthy AI voice owes you

Everything above points at a short list of obligations, and some are about to become law.

Article 50 of the EU AI Act applies as from 2 August 2026. Providers must ensure AI-generated content is “marked in a machine-readable format and detectable as artificially generated or manipulated,” and must ensure people “are informed that they are interacting with an AI system, unless this is obvious.” A limited grace period runs to 2 December 2026 for marking content already on the market. Our meditation statistics report is the canonical page on this site for the wider disclosure-law picture, including the US state bills, and is kept current there rather than duplicated here.

Regulation sets a floor. Here is what we think the floor should be for a voice that guides a practice, and what we actually do about each one, accurate as of 8 August 2026.

Say it is AI, in the product, not the footer. StillMind labels its guidance as AI across the product. The onboarding style choice carries an “AI-powered” badge, and the home screen, session types and settings all describe sessions as AI-guided. The App Store listing describes it too.

License the voice for the use. Ours are licensed for this. The provider is ElevenLabs and the model is eleven_v3.

Let people hear it before they commit, and let them change their mind. Every voice can be previewed before you choose it, and changed later in settings. One caveat worth stating: not every locale offers a choice yet. Hungarian currently has a single voice.

Be exact about what human review means. Humans review script types, prompt templates and features on a regular cycle, backed by a fixture-based regression eval that checks generated scripts against defined criteria. What that does not mean is that a person reads your session before you hear it. Every session is generated fresh at request time, so nobody has read it. That is the honest description, and it is the same for every generative product, which is why saying it plainly seems better than implying otherwise.

Do not pretend a voice makes a practice safe. Meditation has adverse events, and the best available estimate is a range rather than a number. Farias and colleagues pooled 83 studies covering 6,703 participants in Acta Psychiatrica Scandinavica in 2020 and reported an overall adverse-event prevalence of 8.3% (95% CI 5 to 12%). The bracket is the range the real rate most plausibly falls in, and it is wide. That single figure hides most of what is going on: prevalence was 3.7% in experimental studies and 33.2% in observational ones. The most commonly reported effects were anxiety (33%), depression (27%) and cognitive anomalies (25%). Quote the range, never a single number as universal. Nothing about narration changes any of this, and no voice, human or synthetic, is a substitute for care if a practice is making things worse.

Our conflicts, in full

StillMind sells AI-guided meditation. Every conclusion in this article that is favourable to synthetic narration is favourable to us, and you should read it accordingly. We use ElevenLabs for voice synthesis, which is the same provider used to generate the stimuli in Diel et al. 2026, a study we lean on heavily and whose central finding, that an AI label lowers acceptance and experienced effects, is commercially inconvenient for us. We have no relationship with any research group cited here, no funding relationship with ElevenLabs beyond being a paying customer, and no involvement in any study in this post. The first-party figures in the previous section come from our own analytics and are self-reported. Our privacy policy covers what we collect, and we have written separately about what actually happens to what you tell an AI meditation app.


What holds up, and what does not

Nine claims circulate about AI voices and meditation. We traced each one.

The claim What we found Verdict
"AI voices are now indistinguishable from human ones" Detection is unreliable, not absent. 60.8% of AI clips were correctly called synthetic by Barrington and colleagues, 73% of deepfakes by Mai and colleagues. Lavan found clones labelled human 58% of the time against 62% for real voices, and no hyperrealism effect. Overstated
"It does not matter whether people know the voice is AI" Directly contradicted. Diel et al. 2026 held the audio constant and varied only the label. Among the 89 who passed the manipulation check, told-human beat told-AI on acceptance, experienced effects and mindfulness. Contradicted
"AI-generated meditation is as effective as human-led meditation" No study we located measured a clinical outcome for either narrator. Diel et al. 2025 measured evaluation and categorisation. The meta-analyses never varied the voice. Unsupported
"Digital meditation interventions work" Supported for the intervention, not the narrator. 34 RCTs give stress g = 0.46; 97 trials give g = 0.44; 176 RCTs give depression g = 0.28 with an NNT of 11.5. Supported, with the narrator untested
"People can be trained to spot AI voices" Mai and colleagues found showing examples "only improves results slightly." Yang and colleagues gave 30 participants about 12 minutes of labelled training and found changed neural responses with no significant behavioural improvement. Unsupported
"Better prompting makes an AI meditation better" Diel et al. 2025 tested prompt tailoring directly and found no significant effect. Their hypothesis was not supported. Unsupported
"Any AI voice will do as long as it sounds natural" The same study found a voice personality mismatched to mindfulness significantly reduced evaluation. Casting mattered where prompt tailoring did not. Contradicted
"Meditation carries no risk, so the delivery method does not matter" Farias and colleagues pooled 83 studies and 6,703 participants: adverse events 8.3% overall (95% CI 5 to 12%), 3.7% in experimental studies and 33.2% in observational ones. The spread is the finding. Overstated
"Disclosure costs nothing" It costs something measurable. Luo and colleagues found up-front disclosure cut purchase rates by over 79.7% in a sales context. In mindfulness, the label moved acceptance and experienced effects. The cost is real and it is worth paying. A real cost, honestly borne

The pattern across those nine rows is not “AI voices are bad” and it is not “AI voices are fine.” It is that the easy questions have been answered and the hard one has not been asked.


Methodology, sources and limitations

Claims here are graded by evidence strength on the same four-tier hierarchy we publish in our meditation statistics report. Tier 1 is peer-reviewed primary research, systematic reviews, meta-analyses and official regulatory documents. Tier 4 is corporate self-report, which is what our own usage data is, and it is labelled as such wherever it appears.

Every figure here was traced to a primary document and read in the original. A press release, a news summary or a search-result snippet is not a source. Where a figure could not be confirmed from a primary text it was dropped rather than softened: a study on synthetic voice and working alliance was excluded because the publisher returned an error and the paper was not indexed anywhere we could reach, and several empathy and voice-perception papers went the same way after their DOIs resolved but their full text did not. The Ivory paper is cited qualitatively and carries no figures, because its abstract could not be fully retrieved.

Three limits are worth stating outright. The mindfulness evidence is German. Both direct studies recruited German residents completing German-language exercises, which is a narrow base for a claim about how voices work. The outcomes are self-report. Acceptance, experienced effects and state mindfulness are questionnaire measures, not symptom measures. The detection studies use clips, not sessions. Judging a short clip in a lab is not the same task as spending 15 minutes following instructions with your eyes closed, and nobody has tested whether detection accuracy holds up over that length.

The largest limit is the one this whole article circles. Layer four is empty, so far as we could find. Until somebody randomises listeners to identical scripts read by a person and by a synthesiser and measures a clinical endpoint, the honest position is that AI narration has cleared the perceptual and acceptance bars and has not been tested on the one that matters most. We will update this page when that trial appears.

Cite this article

APA
Murphy, J. (2026). Can you trust an AI voice? Our data found no difference. StillMind. https://getstillmind.com/blog/can-you-trust-an-ai-voice/
BibTeX
@techreport{murphy2026aivoicetrust,
  title  = {Can you trust an AI voice? Our data found no difference},
  author = {Murphy, Jamie},
  year   = {2026},
  institution = {StillMind},
  url    = {https://getstillmind.com/blog/can-you-trust-an-ai-voice/}
}

Published under CC BY 4.0. Reuse with attribution.

If you want the wider picture on how these sessions are built and what they are for, our guide to AI-guided meditation is the pillar this article sits under, and how long you should actually practise applies the same sourcing standard to a different question.


FAQ

Can people tell if a voice is AI-generated?

Not reliably. In the largest controlled test, Barrington, Cooper and Farid recruited 604 adults through Prolific and found listeners correctly identified AI clips as synthetic 60.8% of the time, which is better than a coin flip and still wrong roughly two times in five. Mai and colleagues, testing 529 people in English and Mandarin, found deepfakes correctly spotted 73% of the time with no difference between languages. Lavan and colleagues found voice clones labelled human 58% of the time against 62% for real human voices, and reported no hyperrealism effect. Longer and less scripted audio is easier to detect than short scripted clips.

Does knowing a meditation is AI-generated change how it feels?

Yes, and this has been tested directly. Diel and colleagues, publishing in Mental Health and Prevention in 2026, gave 411 German participants mindfulness exercises that were all AI-generated, varying only the label. Among the 89 participants who passed the manipulation check, being told an exercise was human-made produced significantly higher acceptance, experienced effects and state mindfulness. Across the full 411, the number of exercises a person believed were human predicted the same three outcomes. The authors note that because only a subset were successfully deceived, the group comparison may not be generalisable.

Is AI-guided meditation effective?

Digital meditation has good evidence as a package, and AI narration specifically has not been tested on clinical outcomes. Three independent meta-analyses support the category: 34 randomised trials with 7,566 participants give a stress effect of Hedges g = 0.46, 97 trials give g = 0.44, and 176 randomised trials of mental health apps give a depression effect of g = 0.28 with a number needed to treat of 11.5. Those are small to moderate improvements on average, and a number needed to treat of 11.5 means roughly one person in twelve benefits who would not have otherwise. None of those trials varied who or what was speaking. We could not locate a published trial comparing AI narration with human narration on stress, anxiety, sleep or any symptom measure.

Do AI voices sound less trustworthy than human ones?

Not on the raw sound. Lavan and colleagues found both types of AI-generated voice were rated more dominant than human voices, and some AI voices were rated more trustworthy than human ones. The penalty appears when the source is disclosed rather than when the voice is heard. Rubin and colleagues, across nine studies with 6,282 participants, found human-attributed responses were rated more empathic and supportive, and that merely suspecting AI involvement was enough to lower those ratings.

Should meditation apps disclose that a voice is AI?

We think so, and from 2 August 2026 Article 50 of the EU AI Act requires providers to ensure AI-generated content is marked in a machine-readable format and detectable as artificially generated, and that people are informed they are interacting with an AI system unless it is obvious. A grace period runs to 2 December 2026 for marking content already on the market. Disclosure does have a measurable cost: in a field experiment across more than 6,200 sales calls, disclosing a chatbot before the conversation cut purchase rates by over 79.7%. That same paper found disclosing later softened the penalty, which we consider a reason not to design that way rather than a tactic to copy.

Does personalising an AI meditation make it better?

The one study that tested it found no significant effect. Diel, Bäuerle, Teufel and Jansen ran two experiments with 143 participants in Scientific Reports in 2025 and found that tailoring the generation prompt produced no significant improvement in how exercises were evaluated. Their hypothesis was not supported. What did matter was voice fit: a voice personality mismatched to mindfulness significantly reduced evaluation. On this evidence, casting the voice well matters more than engineering the prompt.

Can you train yourself to spot an AI voice?

Not to any useful degree, on current evidence. Mai and colleagues gave listeners examples of synthetic speech before the task and reported that it only improves results slightly. Yang and colleagues gave 30 participants about 12 minutes of labelled training with feedback while recording neural responses, and found the training changed how the brain responded without producing a significant improvement in behavioural discrimination. A related advance online paper by Ivory and colleagues reports that skill at detecting AI-generated faces did not transfer to voices.

Do people feel different after an AI-guided session than a self-guided one?

In our own data, not by any amount we can detect. StillMind asks for a mood rating from one to five after a session. In a sample of rated sessions, 600 followed an AI-guided session, averaging 3.40, and 600 followed a self-guided timer session, averaging 3.46. The gap is 0.06 of a point on a five-point scale, with a 95% confidence interval running from -0.04 to +0.16. That is a null result, and a null result at this sample size cannot show the two are equivalent, only that any difference present is smaller than this data can see. It is also a single self-reported item, collected only from sessions that finished and only where the person chose to answer, with no pre-session baseline and no randomisation. It belongs beside the acceptance research above, not beside the clinical trials.

The rest of the research cluster.

View all pieces →

Hear every voice before you choose one

StillMind labels its guidance as AI, lets you preview each voice before you commit, and lets you change your mind later in settings. Free to try.