Bad DetoxWellness claims, checked

Reading the Evidence

Statistically Significant Does Not Mean Big Enough to Notice

The phrase sounds like an endorsement and is closer to a technical footnote. It concerns whether an effect exists, not whether it matters.

Detailed close-up of a blue bar graph showing data analysis on printed paper.
Photograph by RDNE Stock project via Pexels
General information. This article is journalism, not medical advice, and it cannot know your circumstances. Speak to a qualified professional about anything that concerns you. How we work.

These are listed in the order worth acting on, which with statistical significance and effect size is not the order they are usually presented in.

What matters most

  • Significance addresses whether a result is likely to be chance, not how large it is.
  • Very large studies can make trivially small effects significant.
  • Confidence intervals communicate more than a yes or no verdict.

What the test is actually asking

A significance test asks how likely data like this would be if the treatment had no effect at all. A small probability leads researchers to conclude that chance alone is an unsatisfying explanation for what they saw. It does not tell you the probability that the treatment works, which is a different question and a common misreading.

It also says nothing about how large the effect is, which is the question a reader almost always cares about. The conventional threshold used is a convention rather than a law of nature, and it was chosen for convenience.

Significance and size are independent

A study can find a statistically significant difference that is far too small for anyone to notice in daily life. It can also find a large, potentially important difference that misses significance because the study was small.

Follow the mechanism and reporting only the verdict discards the information that would let a reader judge whether the finding matters. This is why journals increasingly require effect sizes and intervals alongside any significance statement. A supplement that shifts a measurement by an amount no person could perceive has still cleared the significance bar.

Why large studies complicate this

The larger the study, the smaller the difference it can detect reliably, which is usually a good thing. It also means that in very large datasets almost any difference will eventually reach significance. Nutritional epidemiology often involves enormous datasets, so significance there is a weak signal on its own.

On the label, the important question becomes how large the association is and whether confounding could account for it. Headlines rarely distinguish a large effect in a small study from a tiny effect in an enormous one.

What a confidence interval tells you

An interval gives a range of effect sizes compatible with the data rather than a single point estimate. A narrow interval indicates a well-estimated effect, while a wide one indicates considerable remaining uncertainty. An interval spanning both meaningful benefit and meaningful harm is a genuinely uninformative result.

That situation is common in small studies and is frequently reported as showing a trend towards benefit. Reading the interval rather than the verdict is the single biggest upgrade available to a non-specialist reader.

Multiple comparisons and the significance that appears

A study measuring twenty outcomes will typically produce one significant result through chance alone. If the significant one is then presented as the finding, the study has tested nothing and reported a coincidence. Statistical corrections exist for this and are frequently omitted in smaller nutritional and supplement research.

Against the trial data, declaring the primary outcome before collecting data is the structural protection against this problem. Checking a trial registry for that declaration is quick and often changes how a paper reads entirely.

Supplements interact with prescribed medicines, so tell whoever prescribes for you what else you are taking.

Reading a headline with this in mind

Ask how large the difference was in units you understand rather than in percentages of a percentage. Ask what the range of plausible effects was, and whether that range includes no effect at all.

Ask how many outcomes were measured and whether the reported one was the declared primary. Ask whether the difference is large enough to change anything you would actually do. Most health headlines fail at least two of those questions, and knowing which two is genuinely useful.

Everything above, in order of what to do first

  1. What the test is actually asking. A significance test asks how likely data like this would be if the treatment had no effect at all.
  2. Significance and size are independent. A study can find a statistically significant difference that is far too small for anyone to notice in daily life.
  3. Why large studies complicate this. The larger the study, the smaller the difference it can detect reliably, which is usually a good thing.
  4. What a confidence interval tells you. An interval gives a range of effect sizes compatible with the data rather than a single point estimate.
  5. Multiple comparisons and the significance that appears. A study measuring twenty outcomes will typically produce one significant result through chance alone.
  6. Reading a headline with this in mind. Ask how large the difference was in units you understand rather than in percentages of a percentage.

The takeaway

Significance answers whether, not how much. Ask for the effect size and the interval before deciding a result matters.

An extraordinary mechanism needs better evidence than testimonials, and usually has less.

Questions readers ask

Is a p-value the chance the result is wrong?

No, and that is the most common misreading. It describes how surprising the data would be if there were no real effect, not the probability that a conclusion is correct.

Why do studies disagree so often?

Small studies produce unstable estimates, and different populations, doses and outcome measures produce genuinely different answers. Systematic reviews exist to pull those results together rather than to pick a winner.

Reading the Evidencestatisticssignificancereading
More in Reading the Evidence
Sarika Bhandarkar
Editor, Bad Detox

Sarika edits Bad Detox and cuts anything that overstates what a study found.

Also by Sarika Bhandarkar