Understanding Health Statistics
Reviewed by Dr C. J. Odike, MRCGP
A number can be accurate and still mislead when its outcome, denominator, time frame or comparison is missing. This lesson explains risk, effect size, statistical significance and uncertainty without turning group evidence into personal certainty.
Begin with what was counted A health statistic summarises observed data. It may describe how common an outcome is, compare groups or estimate an effect. Before interpreting a number, ask five questions: What outcome was measured? Which population was studied? Over what time frame? What were the numerator and denominator? What was the comparison? A statement such as "30% improved" is incomplete. It must define improvement, the time period, who was studied and what happened in the comparison group. Risk means the chance of a defined outcome in a defined group during a stated period. A rate also includes time, such as events per 1,000 person years. Risk and rate are not interchangeable. Group evidence can inform, but not determine, an individual outcome A group statistic does not decide what will happen to one person. It can still inform an individual estimate when the evidence is relevant. Applicability depends on how closely the person, setting and treatment resemble those in the study. Age, health conditions, baseline risk and follow up time can all matter. An average result can also hide variation. Some people may benefit more, some less and some not at all. Separate absolute and relative effects Baseline risk is the risk in the comparison group before considering the new exposure or treatment effect. It gives the starting point needed to understand a relative change. Suppose 1 in 1,000 people has an outcome over five years without an exposure. With the exposure, 2 in 1,000 have it. The absolute risks are 1 in 1,000 and 2 in 1,000. The absolute risk difference is one extra outcome in every 1,000 people over five years. The relative risk is 2.0 because 2 divided by 1 equals 2. This can be described as a doubling of risk. Both descriptions are correct, but they answer different questions. Relative risk shows the proportional comparison. Absolute risk difference shows how many more or fewer outcomes occurred. Relative figures alone can make a small absolute change sound dramatic. Absolute figures alone can also mislead if the outcome's seriousness or time frame is omitted. For clearer communication, use natural frequencies and the same denominator. Compare 1 in 1,000 with 2 in 1,000, not 0.1% with 1 in 500. It can also help to show both positive and negative framing. Saying 90 in 100 avoided an outcome gives different emphasis from saying 10 in 100 experienced it. Sample size affects precision, not every form of quality Sample size is the number of people or observations included. An adequate sample usually reduces random uncertainty and produces a more precise estimate. A large sample does not automatically make a study trustworthy. It cannot repair biased selection, poor measurements, missing data, confounding or an unsuitable comparison. Bias is a systematic distortion that can push a result away from the truth. Increasing the sample size can make a biased estimate more precise without making it correct. Who was included also matters. A result can be precise but have limited applicability to people who differ from the study population. Statistical significance has a narrow meaning Many studies calculate a p value and compare it with a chosen threshold, often 0.05. Results below the threshold are commonly called statistically significant. A p value asks how compatible the observed data are with a specified statistical model. It is calculated while assuming a stated no effect hypothesis and other model assumptions. More precisely, it considers the probability of obtaining the observed result, or a more extreme result, under those assumptions. A p value does not tell you the probability that the hypothesis is true. It does not measure the probability that chance caused the result. It also does not show the effect size, clinical importance, study quality or likelihood that another study will reproduce the finding. Crossing a threshold does not create a sharp boundary between truth and falsehood. A p value of 0.049 and one of 0.051 usually provide very similar information. Confidence intervals show statistical uncertainty An effect estimate is the study's best single estimate of a difference or association. A confidence interval gives a range of values compatible with the data and statistical model. A narrow confidence interval usually indicates greater precision. A wide interval indicates more random uncertainty. The range may include effects representing benefit, no important difference or harm. This helps show which conclusions remain reasonably possible. A confidence interval is not a guarantee that the true value lies inside it. It also does not include every uncertainty caused by bias, missing evidence or poor applicability. Statistical importance and clinical importance are different A very large study can find a tiny difference that is statistically significant. The difference may still have little clinical importance. A smaller study may estimate an important effect but remain too imprecise for a firm conclusion. A non significant result is not proof that no effect exists. Clinical importance depends on the effect size, outcome severity, treatment burden, possible harms and what matters to the person. Study design and the wider evidence also matter. One precise number should not outweigh serious bias or conflict with a stronger body of evidence. A practical reading check When you see a health statistic, check several things. Is the outcome clearly defined, and are the population and time frame stated? Is there an appropriate comparison, and are absolute and relative effects both available? Are the denominators consistent, and what does the confidence interval show? Finally, could bias or poor applicability affect the result, and is the effect clinically important as well as statistically noticeable? Statistics support decisions by describing patterns and uncertainty. They do not remove the need for judgement, context or discussion of benefits and harms.
Never interpret a percentage alone. Identify the outcome, population, time frame, comparison, baseline risk, absolute effect, uncertainty, study quality and applicability.
Medical words made simple
- Outcome
- The event or change measured in a study, such as pain improvement, hospital admission or a side effect.
- Risk
- The chance of a defined outcome occurring in a defined group during a stated period.
- Rate
- How often an event occurs while accounting for time, such as cases per 1,000 person-years.
- Baseline risk
- The starting risk in a comparison group before considering the effect of a treatment or exposure.
- Absolute risk
- The chance of an outcome within one group over a stated period, such as 2 in 1,000 over five years.
- Absolute risk difference
- The difference between two absolute risks, expressed as how many more or fewer outcomes occur.
- Relative risk
- A ratio comparing the risk in one group with the risk in another group.
- Numerator
- The number of outcomes or events counted in a statistic.
- Denominator
- The group or amount of observation time against which the numerator is compared.
- Sample size
- The number of people or observations included in a study. A larger sample often improves precision but does not remove bias.
- Bias
- A systematic problem in design, conduct, measurement or reporting that can distort a result.
- P-value
- A measure of how compatible the observed data are with a specified statistical model and no-effect hypothesis, under stated assumptions.
- Statistical significance
- A label often used when a p-value passes a chosen threshold. It does not show effect size, truth or clinical importance.
- Confidence interval
- A range of effect values compatible with the data and statistical model. Its width helps show statistical precision.
- Precision
- How much random uncertainty surrounds an estimate. Narrower confidence intervals usually indicate greater precision.
- Clinical importance
- Whether an effect is large and meaningful enough to matter in practice, considering benefits, harms and treatment burden.
Quick recap
- A statistic needs a defined outcome, population, time frame, numerator, denominator and comparison.
- Relative risk shows a proportional comparison, while absolute risk difference shows how many more or fewer outcomes occurred.
- Larger samples usually improve precision, but they do not remove bias or guarantee applicability.
- A p value does not give the probability that a hypothesis is true or that chance caused the result.
- A confidence interval shows a range compatible with the data, but it does not capture every source of uncertainty.
- Clinical decisions also depend on effect size, outcome severity, harms, study quality and individual circumstances.