April 4, 2026
August 26, 2026 · 8 min read
A p-value of 0.04 does not mean there is a 96% chance the supplement works. It does not mean a 4% chance the result was luck, either. Both readings are wrong, and both sell a great deal of product.
Here is what the number is. A p-value is the probability of seeing data at least this extreme if the supplement did nothing at all. It starts by assuming no effect. So it cannot hand you back the odds that there is no effect.
The American Statistical Association made this its second published principle in 2016, in a statement drafted by twenty statisticians over several months. Their sixth principle is blunter: a p-value near 0.05 offers only weak evidence against the null hypothesis.
David Colquhoun ran the arithmetic in 2014. Observe p close to 0.05 in a realistic experiment, call it a discovery, and you will be wrong at least 30% of the time. Not 5%. Thirty.
Treat 0.05 as a hint. It is not a verdict.
Erick Turner's team pulled the FDA files on 74 registered antidepressant trials covering 12,564 participants. The FDA judged 51% of them positive. In the published journal literature, 94% looked positive. Negative trials were either never published or written up to read as wins. Across the drugs, the apparent effect size was inflated by 32%.
Supplements are not exempt from this. Robert Kaplan and Veronica Irvin reviewed 55 large NHLBI trials of drugs and dietary supplements for cardiovascular outcomes, run between 1970 and 2012. Of the 30 published before 2000, 17 reported a benefit on the primary outcome — 57%. Of the 25 published afterwards, 2 did. That is 8%.
What changed after 2000 was pre-registration. Researchers now have to state their primary outcome publicly before they collect the data. Look for a registration number on any trial you are asked to take seriously. Its absence tells you a lot.
The ASA's fifth principle is the one supplement marketing ignores most often. A p-value says nothing about how big an effect is. Any effect, however trivial, produces a small p-value if the sample is large enough.
Run that in reverse and the trap becomes clear. A large trial can make a meaningless difference statistically significant. A small trial can miss a real one entirely.
VITAL randomised 25,871 US adults to 2,000 IU of vitamin D3 daily or placebo for a median of 5.3 years. Neither primary endpoint moved. Invasive cancer came in at a hazard ratio of 0.96, with a confidence interval of 0.88 to 1.06.
Source: Manson et al., New England Journal of Medicine, 2019 (n=25,871).
That interval is the useful part. With nearly 26,000 people and five years, the trial had the power to detect a modest benefit. It found none, and it ruled out anything large. A null result from a trial that size carries real information.
Ask for the effect size in units you understand. Millimetres of mercury. Percentage points of HbA1c. Events per thousand people per year. A relative percentage with no baseline attached is not a number, it is a decoration.
Test one outcome at the 0.05 threshold and you have a 5% chance of a false positive. Test five and it rises to 23%. Test twenty independent outcomes and the chance that at least one clears the bar by luck alone is 64%.
The ISIS-2 investigators made this point in 1988, and made it memorably. Their trial of aspirin in 17,187 heart attack patients found a large, real benefit. They then reported the effect split by astrological birth sign. Patients born under Gemini or Libra showed no benefit at all.
They published that subgroup deliberately, so readers would see how easily slicing data manufactures a finding. Nearly forty years later, supplement press releases still run on exactly this fuel.
VITAL's omega-3 arm shows the softer version. The pre-specified primary endpoint, major cardiovascular events, returned a hazard ratio of 0.92 with an interval of 0.80 to 1.06 — not significant. Total myocardial infarction, a secondary endpoint, came in at 0.72 (0.59 to 0.90). Interesting. Worth a follow-up trial. Not established.
Find the pre-specified primary outcome. Everything after it is a hypothesis wearing a result's clothes.
Michael Walsh and colleagues took 399 randomised trials from the New England Journal of Medicine, the Lancet, JAMA, Annals of Internal Medicine and the BMJ. All had reported a significant result in the abstract. Median sample size: 682 patients.
Then they asked a simple question. How many patients would have to switch from non-event to event before the p-value crossed back above 0.05?
Across 399 trials, the median fragility index was 8. A quarter of the results hinged on three patients or fewer. Ten percent had an index of zero — significance disappeared the moment Fisher's exact test was applied. In 53% of trials, more participants were lost to follow-up than the number needed to erase the finding.
Source: Walsh et al., Journal of Clinical Epidemiology, 2014 (399 trials).
The LIMIT-2 trial is the cautionary case. It randomised 2,316 patients, reported a 24% relative reduction in mortality, and cleared the bar at p=0.04. Its fragility index was 1. One patient. Three years later a trial of 58,050 patients found no benefit whatsoever.
If a supplement study rests on 40 participants and a p-value of 0.045, assume it rests on two or three people.
The reverse mistake is just as common. "No significant difference" gets read as "no effect", and sometimes that is right.
ASCEND randomised 15,480 adults with diabetes to 840 mg of omega-3 fatty acids daily or placebo, and followed them for a mean of 7.4 years. Serious vascular events occurred in 8.9% of the omega-3 group and 9.2% of placebo, p=0.55. That is a genuinely null result: large, long, pre-registered, and precise enough to exclude a meaningful benefit.
Now picture a 30-person crossover study returning p=0.4. It excludes nothing. It was never capable of finding anything. Reporting it as evidence of no effect is as misleading as reporting it as evidence of one.
The confidence interval separates these two cases. The p-value does not. Read the interval and ask what it rules out.
None of this requires statistical training. It requires six questions, in order.
That last chart explains why 72 statisticians and scientists proposed in 2018 that new discoveries should clear p<0.005 rather than 0.05. They were not being fussy. They were looking at replication rates.
We sell NMN, so apply the six questions to it. NMN raises blood NAD+ — that finding replicates and is not in dispute. What has not held up is the metabolic benefit built on top of it.
Chen and colleagues pooled eight randomised trials of NMN covering 342 middle-aged and older adults, at 250 to 2,000 mg a day for 14 days to 12 weeks. The meta-analysis found no significant benefit on fasting glucose, fasting insulin, HbA1c, insulin resistance or lipids. That is our own category, reported as it stands.
The most cited positive trial, from Yoshino's group at Washington University, gave 250 mg of NMN to postmenopausal women with prediabetes for ten weeks. Muscle insulin sensitivity improved. It involved 25 women — 13 on NMN, 12 on placebo — at a single site, and it has not been replicated.
That is a thin base for a molecule sold as confidently as this one is. We think it is worth taking, and we say why in our honest comparison of UK NMN products and our look at what the human safety data reports. We would rather you knew the size of the evidence before you spent anything on our NMN + Resveratrol.
The same six questions work on the marker panels people track alongside supplements. We used them on hs-CRP, and the genetics turned out to undercut a story that had looked settled for years. Use them everywhere, including here.
April 4, 2026
August 24, 2026