๐Ÿ“Š FormStat
FormStat Guides ยท By Abdul Hannan

What Is a P-Value? A Simple Explanation for Medical Students

Almost every medical research paper reports p-values, and almost every medical student finds them confusing. You are not alone โ€” surveys show even experienced researchers misinterpret them. This article explains the p-value in plain language, with one worked example and the three most common mistakes to avoid.

The One-Sentence Definition

A p-value is the probability of seeing your result (or something more extreme) if there were actually no real difference or relationship.

Read that twice. The p-value does NOT tell you the probability that your hypothesis is true. It answers a narrower question: "If nothing were really going on, how surprising would my data be?"

An Everyday Analogy: The Coin Toss

Imagine a friend claims a coin is rigged to land on heads. You toss it 10 times and get 9 heads. Is the coin rigged, or was that just luck?

A fair coin gives 9 or 10 heads out of 10 only about 1% of the time by pure chance. So the p-value here is about 0.01. In plain words: if the coin were fair, you would see a result this extreme only 1% of the time. That is surprising enough that you reasonably conclude the coin is probably rigged.

Medical statistics works the same way. "The coin is fair" is your null hypothesis (nothing going on). Your data is the 9 heads. The p-value tells you how surprising your data would be in a world where nothing is going on.

Why 0.05?

By convention, researchers call a result "statistically significant" when p < 0.05. This cutoff was suggested by the statistician Ronald Fisher nearly a century ago, and it stuck. It means: we tolerate up to a 5% chance of being fooled by luck.

It is just a convention, not a law of nature. A result with p = 0.049 is not magically real while p = 0.051 is magically false โ€” they are nearly identical pieces of evidence. But journals and examiners expect the 0.05 line, so you need to know it.

What "Statistically Significant" Actually Means

When p < 0.05, all you can honestly claim is: "This result is unlikely to be due to chance alone." That is it. It does NOT mean:

Worked Mini-Example

Recall the smoking and cough survey from our chi-square article: 20 people, and the test gave p โ‰ˆ 0.0073.

In plain words: if smoking truly had nothing to do with cough, there is only a 0.73% chance of seeing smokers cough this much more than non-smokers just by luck. Because 0.0073 is well below 0.05, we call the association statistically significant โ€” the link between smoking and cough in this data is very unlikely to be a fluke.

Notice what we did NOT say: we did not say "there is a 99.27% chance smoking causes cough." The p-value says nothing about causation, and nothing about how likely our explanation is. It only quantifies surprise under the "nothing going on" assumption.

Three Misinterpretations to Avoid

  1. "p = 0.03 means there is a 97% chance my hypothesis is true." Wrong. The p-value is calculated assuming there is no real effect. It cannot tell you the probability your hypothesis is true โ€” that is a different (Bayesian) question entirely.
  2. "p = 0.06 is almost significant / trending toward significance." Wrong. 0.06 is above 0.05, so by the convention it is not statistically significant. Describing it as "almost" significant is a way of smuggling a positive result past the cutoff. Report it honestly: not significant at the 5% level.
  3. "A smaller p-value means a bigger effect." Wrong. The p-value mixes together the size of the effect AND the size of your sample. A tiny, meaningless difference can give p = 0.0001 in a study of 10,000 people. Always report the actual effect (percentages, means, differences) alongside the p-value.

What If p > 0.05? (Not Significant โ‰  No Effect)

Students often write "p = 0.20, therefore there is no difference between the groups." That is wrong. A non-significant p-value means "we did not find enough evidence" โ€” not "we proved there is nothing there." Absence of evidence is not evidence of absence.

The usual culprit is a small sample. Remember the 10-person smoking table from our chi-square article: the pattern was striking (80% vs 20%), yet Fisher's test gave p โ‰ˆ 0.21 โ€” not significant โ€” simply because 10 people cannot rule out luck. The effect might be completely real; the study was just too small to show it.

Correct language: write "no statistically significant difference was found" โ€” never "there is no difference." And if you genuinely expected an effect, the honest next step is a larger sample or a more precise measurement, not creative re-analysis until a p-value drops below 0.05.

What to Report Instead of Just "p < 0.05"

Good papers report three things together: the effect size (e.g. "80% of smokers vs 20% of non-smokers had a cough"), the test used (e.g. "chi-square test"), and the exact p-value (e.g. "p = 0.007"), not just "p < 0.05". This lets readers judge for themselves whether the finding matters.

Calculate P-Values Free on Your Phone

Understanding p-values is step one; getting them for your own data is step two โ€” and you do not need SPSS for it. Try FormStat free: run chi-square, t-tests, ANOVA, and correlations on your survey data and get exact p-values with plain-English interpretations, all offline on your phone.

Try it yourself โ€” free.
Run this analysis in seconds with the FormStat app. No signup, no laptop, works offline.

โ† All guides