Skip to content

What a p-Value Actually Tells You About a Supplement Study

August 26, 2026 · 8 min read

Supplement capsules and research papers — reading the statistics behind a trial result

The p-value answers a question you did not ask

A p-value of 0.04 does not mean there is a 96% chance the supplement works. It does not mean a 4% chance the result was luck, either. Both readings are wrong, and both sell a great deal of product.

Here is what the number is. A p-value is the probability of seeing data at least this extreme if the supplement did nothing at all. It starts by assuming no effect. So it cannot hand you back the odds that there is no effect.

The American Statistical Association made this its second published principle in 2016, in a statement drafted by twenty statisticians over several months. Their sixth principle is blunter: a p-value near 0.05 offers only weak evidence against the null hypothesis.

David Colquhoun ran the arithmetic in 2014. Observe p close to 0.05 in a realistic experiment, call it a discovery, and you will be wrong at least 30% of the time. Not 5%. Thirty.

Treat 0.05 as a hint. It is not a verdict.

Why the published literature looks better than the evidence

Erick Turner's team pulled the FDA files on 74 registered antidepressant trials covering 12,564 participants. The FDA judged 51% of them positive. In the published journal literature, 94% looked positive. Negative trials were either never published or written up to read as wins. Across the drugs, the apparent effect size was inflated by 32%.

Supplements are not exempt from this. Robert Kaplan and Veronica Irvin reviewed 55 large NHLBI trials of drugs and dietary supplements for cardiovascular outcomes, run between 1970 and 2012. Of the 30 published before 2000, 17 reported a benefit on the primary outcome — 57%. Of the 25 published afterwards, 2 did. That is 8%.

Share of trials reporting a positive result Same drugs, different vantage point Antidepressant trials, as published 94% Same 74 trials, FDA's own review 51% NHLBI trials before 2000 57% NHLBI trials after registration 8% Registration and full disclosure both cut the positive rate sharply.
The interventions did not change between these bars. The reporting rules did. Source: Turner et al., NEJM, 2008 (74 trials, 12,564 participants); Kaplan & Irvin, PLoS ONE, 2015 (55 trials).

What changed after 2000 was pre-registration. Researchers now have to state their primary outcome publicly before they collect the data. Look for a registration number on any trial you are asked to take seriously. Its absence tells you a lot.

Statistical significance is not effect size

The ASA's fifth principle is the one supplement marketing ignores most often. A p-value says nothing about how big an effect is. Any effect, however trivial, produces a small p-value if the sample is large enough.

Run that in reverse and the trap becomes clear. A large trial can make a meaningless difference statistically significant. A small trial can miss a real one entirely.

What 25,871 Adults Told Us About Vitamin D

VITAL randomised 25,871 US adults to 2,000 IU of vitamin D3 daily or placebo for a median of 5.3 years. Neither primary endpoint moved. Invasive cancer came in at a hazard ratio of 0.96, with a confidence interval of 0.88 to 1.06.

Source: Manson et al., New England Journal of Medicine, 2019 (n=25,871).

That interval is the useful part. With nearly 26,000 people and five years, the trial had the power to detect a modest benefit. It found none, and it ruled out anything large. A null result from a trial that size carries real information.

Ask for the effect size in units you understand. Millimetres of mercury. Percentage points of HbA1c. Events per thousand people per year. A relative percentage with no baseline attached is not a number, it is a decoration.

Every extra outcome buys another chance at 0.05

Test one outcome at the 0.05 threshold and you have a 5% chance of a false positive. Test five and it rises to 23%. Test twenty independent outcomes and the chance that at least one clears the bar by luck alone is 64%.

The ISIS-2 investigators made this point in 1988, and made it memorably. Their trial of aspirin in 17,187 heart attack patients found a large, real benefit. They then reported the effect split by astrological birth sign. Patients born under Gemini or Libra showed no benefit at all.

They published that subgroup deliberately, so readers would see how easily slicing data manufactures a finding. Nearly forty years later, supplement press releases still run on exactly this fuel.

VITAL's omega-3 arm shows the softer version. The pre-specified primary endpoint, major cardiovascular events, returned a hazard ratio of 0.92 with an interval of 0.80 to 1.06 — not significant. Total myocardial infarction, a secondary endpoint, came in at 0.72 (0.59 to 0.90). Interesting. Worth a follow-up trial. Not established.

Find the pre-specified primary outcome. Everything after it is a hypothesis wearing a result's clothes.

How fragile a significant result usually is

Michael Walsh and colleagues took 399 randomised trials from the New England Journal of Medicine, the Lancet, JAMA, Annals of Internal Medicine and the BMJ. All had reported a significant result in the abstract. Median sample size: 682 patients.

Then they asked a simple question. How many patients would have to switch from non-event to event before the p-value crossed back above 0.05?

The Median Answer Was Eight Patients

Across 399 trials, the median fragility index was 8. A quarter of the results hinged on three patients or fewer. Ten percent had an index of zero — significance disappeared the moment Fisher's exact test was applied. In 53% of trials, more participants were lost to follow-up than the number needed to erase the finding.

Source: Walsh et al., Journal of Clinical Epidemiology, 2014 (399 trials).

The LIMIT-2 trial is the cautionary case. It randomised 2,316 patients, reported a 24% relative reduction in mortality, and cleared the bar at p=0.04. Its fragility index was 1. One patient. Three years later a trial of 58,050 patients found no benefit whatsoever.

If a supplement study rests on 40 participants and a p-value of 0.045, assume it rests on two or three people.

A non-significant result is not always a null result

The reverse mistake is just as common. "No significant difference" gets read as "no effect", and sometimes that is right.

ASCEND randomised 15,480 adults with diabetes to 840 mg of omega-3 fatty acids daily or placebo, and followed them for a mean of 7.4 years. Serious vascular events occurred in 8.9% of the omega-3 group and 9.2% of placebo, p=0.55. That is a genuinely null result: large, long, pre-registered, and precise enough to exclude a meaningful benefit.

Now picture a 30-person crossover study returning p=0.4. It excludes nothing. It was never capable of finding anything. Reporting it as evidence of no effect is as misleading as reporting it as evidence of one.

The confidence interval separates these two cases. The p-value does not. Read the interval and ask what it rules out.

Six questions worth asking of any supplement study

None of this requires statistical training. It requires six questions, in order.

  1. Was it registered before it started? No registration number, no primary outcome you can trust.
  2. What was the pre-specified primary outcome, and did it move? Secondary findings are leads, not conclusions.
  3. How large was the effect, in units that mean something to you?
  4. What does the 95% confidence interval exclude? That is the real result.
  5. How many people, for how long? Fourteen days in twelve volunteers answers almost nothing.
  6. Has anyone replicated it? One trial is a finding. Two independent trials are evidence.
How many patients flip a significant result Median fragility index, 399 trials in five major journals p between 0.05 and 0.01 3 patients p between 0.01 and 0.001 11 patients p below 0.001 26 patients Fewer events needed means a more fragile finding.
Where the p-value sits inside the significant range changes how much it is worth. Below 0.001 is a different animal from 0.049. Source: Walsh et al., Journal of Clinical Epidemiology, 2014 (399 trials).

That last chart explains why 72 statisticians and scientists proposed in 2018 that new discoveries should clear p<0.005 rather than 0.05. They were not being fussy. They were looking at replication rates.

What that means for what we sell

We sell NMN, so apply the six questions to it. NMN raises blood NAD+ — that finding replicates and is not in dispute. What has not held up is the metabolic benefit built on top of it.

Chen and colleagues pooled eight randomised trials of NMN covering 342 middle-aged and older adults, at 250 to 2,000 mg a day for 14 days to 12 weeks. The meta-analysis found no significant benefit on fasting glucose, fasting insulin, HbA1c, insulin resistance or lipids. That is our own category, reported as it stands.

The most cited positive trial, from Yoshino's group at Washington University, gave 250 mg of NMN to postmenopausal women with prediabetes for ten weeks. Muscle insulin sensitivity improved. It involved 25 women — 13 on NMN, 12 on placebo — at a single site, and it has not been replicated.

That is a thin base for a molecule sold as confidently as this one is. We think it is worth taking, and we say why in our honest comparison of UK NMN products and our look at what the human safety data reports. We would rather you knew the size of the evidence before you spent anything on our NMN + Resveratrol.

The same six questions work on the marker panels people track alongside supplements. We used them on hs-CRP, and the genetics turned out to undercut a story that had looked settled for years. Use them everywhere, including here.

References

  1. Wasserstein RL, Lazar NA (2016). "The ASA's Statement on p-Values: Context, Process, and Purpose." The American Statistician, 70(2), 129-133. doi:10.1080/00031305.2016.1154108
  2. Colquhoun D (2014). "An investigation of the false discovery rate and the misinterpretation of p-values." Royal Society Open Science, 1(3), 140216. PMID 26064558. doi:10.1098/rsos.140216
  3. Turner EH, Matthews AM, Linardatos E, Tell RA, Rosenthal R (2008). "Selective Publication of Antidepressant Trials and Its Influence on Apparent Efficacy." New England Journal of Medicine, 358(3), 252-260. PMID 18199864. doi:10.1056/NEJMsa065779
  4. Kaplan RM, Irvin VL (2015). "Likelihood of Null Effects of Large NHLBI Clinical Trials Has Increased over Time." PLoS ONE, 10(8), e0132382. PMID 26244868. doi:10.1371/journal.pone.0132382
  5. Walsh M, Srinathan SK, McAuley DF, et al. (2014). "The statistical significance of randomized controlled trial results is frequently fragile: a case for a Fragility Index." Journal of Clinical Epidemiology, 67(6), 622-628. PMID 24508144. doi:10.1016/j.jclinepi.2013.10.019
  6. Manson JE, Cook NR, Lee IM, et al. (2019). "Vitamin D Supplements and Prevention of Cancer and Cardiovascular Disease." New England Journal of Medicine, 380(1), 33-44. PMID 30415629. doi:10.1056/NEJMoa1809944
  7. Manson JE, Cook NR, Lee IM, et al. (2019). "Marine n-3 Fatty Acids and Prevention of Cardiovascular Disease and Cancer." New England Journal of Medicine, 380(1), 23-32. PMID 30415637. doi:10.1056/NEJMoa1811403
  8. Bowman L, Mafham M, Wallendszus K, et al. (ASCEND Study Collaborative Group) (2018). "Effects of n-3 Fatty Acid Supplements in Diabetes Mellitus." New England Journal of Medicine, 379(16), 1540-1550. PMID 30146932. doi:10.1056/NEJMoa1804989
  9. ISIS-2 (Second International Study of Infarct Survival) Collaborative Group (1988). "Randomised trial of intravenous streptokinase, oral aspirin, both, or neither among 17,187 cases of suspected acute myocardial infarction: ISIS-2." The Lancet, 332(8607), 349-360. PMID 2899772.
  10. Chen F, Zhou D, Kong APS, et al. (2025). "Effects of Nicotinamide Mononucleotide on Glucose and Lipid Metabolism in Adults: A Systematic Review and Meta-analysis of Randomised Controlled Trials." Current Diabetes Reports, 25, article 4. PMID 39531138. doi:10.1007/s11892-024-01557-z
  11. Yoshino M, Yoshino J, Kayser BD, et al. (2021). "Nicotinamide mononucleotide increases muscle insulin sensitivity in prediabetic women." Science, 372(6547), 1224-1229. PMID 33888596. doi:10.1126/science.abe9985
  12. Benjamin DJ, Berger JO, Johannesson M, et al. (2018). "Redefine statistical significance." Nature Human Behaviour, 2, 6-10. doi:10.1038/s41562-017-0189-z

Read Our Evidence the Way You Would Read Anyone Else's

We publish the trials behind NMN, including the meta-analysis of eight studies in 342 adults that found nothing on glucose or lipids. Our NMN + Resveratrol lists 500 mg of NMN and 600 mg of trans-resveratrol per serving, with a batch certificate of analysis you can read first.

See the Label and the CoA →