A study drops. A headline says 'Coffee linked to lower dementia risk.' You've seen this headline before — also for chocolate, red wine, and standing desks, all of which were going to save or kill you depending on the week. You don't know if you should change anything. Neither does anyone, off one study, but most of us were never taught why a single study is so weak. So we treat the 50,000-person trial and the 200-person survey as if they carry the same weight, because the headlines do, and the headlines are written by people whose job is clicks, not calibration.
The good news is that the method section of almost any paper contains nearly everything you need to grade it, and you can extract that grade in about four minutes. The trick is knowing which three questions to ask, and then making something do the boring work of finding the answers buried in a paragraph that starts with 'Participants were recruited via.' I don't want to read that paragraph. I want its verdict.
The three questions that do most of the work
How many people were studied? Was it randomized and controlled, or just observed? Who paid for it? A 30,000-person randomized trial funded by a public agency and a 180-person observational study funded by the industry it flatters are not the same kind of evidence, even when their headlines rhyme. When I run the Plain-Language Paper Summary with Honest Caveats prompt on Claude Opus 4.8, the method section of the output answers all three in language I can actually repeat to a colleague, and the caveats section surfaces the limitations the paper admits in its own discussion — the ones the press release quietly forgot to mention.
That last part matters more than people realize. Papers are usually honest in their limitations section and oversold in their abstract. The abstract is marketing; the limitations are the fine print. Most readers never reach the fine print because it's on page eleven, after the part that's hard to read, in the section that doesn't make headlines. A good summary reads the fine print for you and drags it to the top where you can't pretend you didn't see it. That single move — limitations promoted to the front — has changed more of my opinions than any finding ever has.
- Sample size: under 500 for a behavioral or health claim, stay skeptical; under 100, treat it as a pilot, not proof
- Design: observational shows things move together; only experiments hint at cause, and even then carefully
- Funding: a disclosed conflict doesn't void a finding, but it earns extra scrutiny on every claim that flatters the funder
- Replication: one unreplicated result is a hypothesis wearing a finding's clothes — wait for the second study
The mistake the funding question catches
I used to think the funding question was paranoid — surely a conflicted study isn't automatically wrong? And it isn't. The point of the funding flag isn't to dismiss the study; it's to tell you where to aim your skepticism. A study funded by a company will rarely fabricate data, but it will often design the question, choose the comparison, and frame the result in the most flattering way available. The summary's caveat section names the funder in one clause, and that clause is enough for me to re-read the 'finding' looking for the favorable framing. About a third of the time, it's there: a real result, dressed up one size too generous.
Why the one-sentence takeaway is the discipline
The feature I lean on hardest is the forced one-sentence takeaway. Compressing a study into a single honest sentence is the same skill as not overstating it, because every weasel word you'd need to add — 'in a small sample,' 'over four weeks,' 'compared to nothing in particular' — is a word that won't fit in a clean sentence. If your one sentence sounds too tidy and triumphant, the study was almost certainly messier than your sentence, and you're about to repeat the tidy version to someone who'll quote you back to themselves later.
I don't use any of this to win arguments. I use it to lose them more gracefully — to be the person who says 'the evidence is thinner than that headline suggests, here's the asterisk' instead of picking a side I can't actually defend when pushed. It's a less satisfying way to talk about studies and a much more honest one, and over time people start trusting the boring, hedged version more than the confident wrong one. Grab the Plain-Language Paper Summary with Honest Caveats prompt on Prompt Dock and run it on the next study a friend sends you with the caption 'see??' You'll either confirm it or, far more often, find the asterisk the headline politely left off.