Issue 25 · Pick 01 AI / ML ✓ read
AI systems out-persuade expert humans
TL;DR: In four preregistered experiments (~19,000 conversations), frontier LLMs reliably out-persuaded every class of human they were pitted against—random crowd workers, tournament-selected top persuaders, professional canvassers, and world-championship debaters—even when the humans picked the topics, prepared for weeks, practiced against the AI that beat them, and competed for £1,000 prizes. The advantage transferred to real money: AI raised nearly 3× more in actual donations than professional fundraisers. But the paper's most interesting contribution is mechanistic: when the AI was throttled to human message lengths and human typing speeds, its advantage vanished entirely. The edge isn't rhetorical genius—it's information throughput. And, uncomfortably, the information doesn't have to be true.
The experiment nobody had actually run
We already knew LLMs are decent persuaders. Prior work showed they can shift voter preferences, talk people out of conspiracy beliefs, and roughly match crowd workers in head-to-head persuasion. But "matches a random Prolific worker" leaves the interesting question open. Superforecasters crush average forecasters; maybe super-persuaders crush average persuaders, and the ceiling of human persuasion sits well above anything an LLM can reach.
What makes this paper unusual is how hard the authors tried to lose. They didn't compare AI against convenient humans—they built an escalating ladder of increasingly formidable opponents, and stacked the deck for each rung:
- Selected Laypeople: the top ~10% survivors of a separately preregistered four-round elimination tournament (1,154 entrants, 9,475 conversations)—an explicit attempt to find "superforecaster-equivalent" persuaders.
- Elite Debaters: 56 competitors including 4 world champions and 11 continental champions (mean 8.9 years of competition). They chose the ten policy issues by vote, got ~8 paid hours of research time, and earned per-conversation bonuses tied to beating the AI, plus a £1,000 top prize.
- Coached Elite Debaters: the same debaters, back after ~8 hours with a coaching tool that showed them the AI's exact system prompt, their own annotated transcripts, live practice against the AI, and a "what would the AI have said here?" button at every turn of every past conversation.
- Professional Canvassers: from a UK firm, median ~10,000 career persuasion conversations, paid £140/hour.
The task: a live, text-based conversation (median 7 turns, 14 minutes) trying to shift a persuadee's 0–100 agreement with a UK policy stance ("the UK should legalise physician-assisted suicide, even if…"). Effects are measured against an active control—a chat with an AI about dogs vs. cats—so the contrast isolates persuasion, not just "having a conversation."
The ladder, and who's at the top
Every human class beat the control. None came close to the AI.
The gaps: AI beat random laypeople by 8.2 pp, selected laypeople by 5.6, elite debaters by 4.6, canvassers by 5.9, and even the coached debaters by 4.1 (all p < .001). And this isn't an artifact of averaging over a mix of strong and weak humans: of 318 per-persuader estimates across Studies 1–2, not one exceeded the pooled AI estimate. The single best human—one coached debater at 9.9 pp—was still 4 points behind. Under each class's fitted distribution of individual ability, the probability that a new randomly drawn persuader would beat AI was under 0.1% for every class. The advantage held across all 10 policy issues (3.0–9.6 pp) and 46 of 49 demographic/political/psychological subgroup levels.
The coaching result deserves a pause. Debaters who studied the AI's playbook did change their behavior—writing 19% longer messages and deploying 54% more fact-checkable claims—but their persuasiveness improved by a statistically indistinguishable +1.0 pp. Humans could see exactly what the AI was doing, imitate its style, and still couldn't close the gap. Which raises the question: what exactly is the AI doing that humans can't copy?
The aha: it's a bandwidth advantage, not a rhetoric advantage
Look at the raw conversation statistics. Elite debaters averaged 54 words per reply and took ~95 seconds to write each one. The AI averaged 294 words per reply with sub-second latency. In a fixed-length conversation with strict turn-taking, one side is delivering roughly five times the content of the other. Maybe the AI isn't a better persuader per word—it just gets vastly more words.
Study 2 tested this directly with a Constrained AI condition: same models, same prompts, but each reply held for a delay drawn from the debaters' empirical response-time distribution (mean 92 s) and length-matched to the debater mean of ~51 words via a clever lagged adaptive scheme (because LLMs comply only loosely with length instructions, the daily prompted target was recomputed so the cumulative mean converged to the debater mean by end of collection).
The result is the paper's cleanest finding: the AI's 4.1 pp advantage over coached debaters collapsed to 0.0 pp (95% CI [−1.7, +1.6], p = .96). Unconstrained AI beat its own constrained twin by 4.2 pp. Same model, same weights, same persuasion strategy—the only thing removed was superhuman bandwidth.
Two converging analyses pin down what that bandwidth carries. First, when the AI was constrained, persuadees' ratings dropped most on the two informational items—argument strength and "I learned a lot" (each ~−11.8 points)—while warmth and enjoyment moved half as much and perceived human-likeness actually rose 7.5 points. Second, an automated fact-checking pipeline counted fact-checkable claims per conversation: elite debaters deployed ~3, constrained AI ~12, unconstrained AI ~37. Across all conditions, mean fact-claims per conversation predicted persuasive impact with R² = 0.89—and, remarkably, held both within human conditions (0.89) and within AI variants (0.90). Regress attitude change on log fact-claims plus an AI-vs-human dummy, and the dummy goes to zero (−0.9 pp, p = .38). Statistically, "being an AI" adds nothing once you know how many claims were made. Persuasion here looks like a nearly pure function of information volume, and the AI simply operates further up the curve than any human hand can type.
The part that should bother you: the facts don't have to be true
Buried in SI Section 4.19 is the finding I'd argue is the most consequential. The fact-checking pipeline also scored each claim's veracity via web search. Human persuaders' claims were corroborated 73% of the time. AI claims: 47%.
Worse: across conditions, accuracy was weakly negatively correlated with persuasive impact (R² = 0.13, negative slope). The most persuasive arm in the whole program—Claude Opus 4.1, at ~+15 pp—had only ~21% of its claims corroborated, while the most accurate model (GPT-5.4, 97%) was not the most persuasive. The authors' framing is careful but the implication is stark: the mechanism is volume of claims, and the persuadees in these conversations could not, in real time, distinguish 37 rapid-fire claims that are mostly unverifiable from 37 that are true. (One caveat the authors flag: the low-accuracy Claude was a non-public research build supplied via the UK AI Security Institute because the public model refused persuasion prompts—itself an interesting detail about where the deployment frontier actually sits.)
Real money
Attitude sliders are cheap talk, and prior work found attitude and behavior effects can be uncorrelated—information-dense conversations in particular were less effective at moving real action. So Study 4 is a genuine test, not a formality. AI (Claude Opus 4.6, prompted with an "impact-efficacy" strategy) competed against 18 canvassers from AppcoUK—a firm that had actually run Save the Children fundraising operations for seven years, raising £824k from 22,583 donors. Persuadees could donate any share of a real £1 bonus.
AI: +17.2 pp of the £1 vs. control. Professional canvassers: +6.4 pp (barely significant, p = .048). The AI's edge held on both margins—more people donated (+6.0 pp) and donors gave more (+12.9 pp)—and persuadees rated the AI higher on all seven preregistered donation mechanisms and all 14 individual items, including six mechanisms it wasn't even prompted to use.
What to be skeptical about
Ecology. These are paid Prolific participants who agreed to a 14-minute text conversation and couldn't leave before turn two. The authors themselves note that sustained engagement of this kind is "demanding to reproduce outside a paid-survey context," and that in the wild, exposure variance dwarfs per-exposure persuasiveness variance. An AI that wins every conversation it gets still needs to get conversations.
Durability. Attitudes were measured immediately post-conversation. Nothing here speaks to whether shifts persist a week later—a known weak point for information-based persuasion.
The scope of the "tie." The human-parity result is one comparison: constrained AI vs. coached debaters, within Study 2, with the constraint calibrated to Study 1 debater statistics and length enforced only in aggregate (realized ~55 words vs. 51 target). It's a clean single-mechanism demonstration, but "humans tie throttled AI" shouldn't be over-read as "humans are as good as AI per word" in general—it's one calibration point on one population.
Stakes. £1 donations are real but small. Vote choice, recurring donations, or health compliance may behave differently.
Minor cracks. Attrition differed by condition in Studies 1–2 (though Lee bounds keep the headline contrasts positive), and the Study 4 preregistration ended up as a timestamped PDF rather than a formal OSF registration due to a platform error.
None of these threaten the core claim within its setting; the design is unusually airtight for this literature (preregistered, active control, incentive-compatible, per-persuader random effects, robustness refits everywhere). They mainly bound how far you can extrapolate.
Why this matters, and where to spend your time
The headline—"AI beats expert persuaders"—was arguably predictable. The two findings that reframe things are: (1) the advantage is structural, a bandwidth asymmetry rather than superior rhetoric, which means it can be regulated (rate limits, length caps) but not trained away by humans—the coaching study shows humans can copy the strategy but can't execute it at the required throughput; and (2) the persuasion channel is claim volume with weak dependence on truth, which is a much sharper statement about epistemic risk than the usual hand-wraving about "AI misinformation." A system that persuades by out-informing you is fine when it's accurate and indistinguishable-from-fine when it isn't.
Read the mechanism section ("Why does AI out-persuade expert humans?") together with SI Section 4.19 on factual accuracy—the main text presents the fact-density result; the SI quietly delivers the accuracy result that changes its meaning. The constrained-AI calibration in Methods (the lagged adaptive length-matching) is also a nice reusable trick for anyone who needs LLMs to hit behavioral targets that prompting alone can't enforce.