Statistics and Research
Understanding p-values
A p-value is the probability, under a specified null hypothesis and statistical model, of obtaining a test statistic at least as incompatible with that null as the one observed. It is not the probability that the null hypothesis is true, not an effect size, and not the chance that the result occurred by random variation alone. Its interpretation depends on the study design, test assumptions, analysis plan, and the evidence considered alongside it.
Start with the null hypothesis and test statistic
A significance test begins with a null hypothesis, such as no difference in population means under a specified model, no association, or no departure from an expected distribution. Data are summarized by a test statistic designed to measure discrepancy from that null relative to expected sampling variation. The p-value locates the observed statistic within its null distribution.
At least as extreme must be defined by the alternative hypothesis and test. A two-sided test counts discrepancies in both directions; a one-sided test counts a prespecified direction. Choosing one-sided after observing the sign invalidates the intended error control. Different tests applied to the same data can produce different p-values because they encode different nulls, statistics, assumptions, and conditioning.
The p-value is conditional: it assumes the null model and analysis procedure. It does not directly assign probabilities to hypotheses in the frequentist framework. To discuss probability of a hypothesis, one needs a framework that includes prior information and an explicit probabilistic model for hypotheses, such as a Bayesian analysis. Mixing these interpretations produces confident but incorrect statements.
What p below 0.05 means and does not mean
A conventional significance level alpha may be set before analysis, often 0.05 in some fields. If p is less than or equal to alpha under a correctly specified procedure, the result is called statistically significant and the null is rejected according to that decision rule. The threshold controls a long-run Type I error rate under repeated use and stated assumptions; it does not create a natural boundary between true and false findings.
Values 0.049 and 0.051 provide nearly the same continuous evidence, yet a rigid label can make them sound fundamentally different. Report the p-value with the effect estimate, confidence interval, sample size, assumptions, and context. Do not describe a nonsignificant result as proof of no effect. The study may be imprecise, underpowered, biased, or compatible with both meaningful benefit and harm.
A significant result does not mean there is a 95 percent probability the alternative is true, a 5 percent probability the finding is due to chance, or a 95 percent chance of replication. It also does not measure the size or practical importance of the effect. These statements require other quantities and assumptions.
Statistical significance versus practical significance
An effect estimate answers how large a difference or association appears to be. A p-value addresses compatibility with a null model. With a very large sample, a tiny and practically unimportant effect can produce a small p-value. With a small sample, an important effect can produce a large p-value because uncertainty is wide. Decisions should be driven by effect magnitude, uncertainty, costs, benefits, and domain thresholds rather than significance alone.
Confidence intervals help by showing a range of parameter values compatible with the data and method at a stated confidence level. An interval can reveal whether clinically, scientifically, or engineering-relevant values remain plausible. It should not be interpreted as the probability that the fixed parameter lies in this particular computed interval under the ordinary frequentist definition.
Standardized effect sizes can support comparison across scales but may hide direct practical units. A five-millimetre deflection difference, two-degree temperature change, or three-point score change may be easier to evaluate in context. Report both natural-unit and standardized quantities when useful, and predefine the smallest effect that would matter when planning the study.
How sample size affects p-values
Many test statistics compare an estimate with its standard error. As sample size grows under suitable sampling, standard error often decreases, so the same estimated effect can yield a larger test statistic and smaller p-value. This is not manipulation by itself; larger samples can distinguish smaller effects from noise. The interpretation still requires asking whether the effect matters.
Small samples can produce unstable estimates and wide uncertainty. A large p-value may reflect low information rather than evidence supporting the null. Power analysis uses an assumed effect, variability, significance level, and design to estimate the chance of detecting that effect. Post hoc power calculated from the observed effect adds little beyond the confidence interval and can be misleading.
Very large datasets do not remove bias. Convenience sampling, confounding, measurement error, missingness, batch effects, and dependence can make standard errors artificially small or estimates systematically wrong. Effective sample size may be much lower than row count when observations are clustered or repeated. Model the design rather than treating every record as independent.
Worked interpretation example
Suppose a two-group study estimates a mean difference of 4.2 units with a 95 percent confidence interval from 0.5 to 7.9 and a two-sided p-value of 0.027 for the null of zero difference. Under the test assumptions, data at least this incompatible with a zero-difference model would occur with probability 0.027. The result crosses a prespecified 0.05 threshold, but that is only the start of interpretation.
Ask whether 4.2 units is practically meaningful and whether the interval includes values too small or too large to change a decision. Review allocation, blinding, missing data, outcome definition, exclusions, and whether the test was planned. If ten outcomes were tested and only this one was highlighted, multiplicity changes the evidential context. If the estimate arose from an observational comparison, confounding remains possible.
A correct report might state the estimated difference, interval, p-value, sample sizes, test, and key limitations without claiming proof. A result with p = 0.08 and interval -0.7 to 9.1 would not prove no difference; it would show greater uncertainty that includes both a small negative value and a potentially important positive value.
Multiple testing and selective reporting
If many true null hypotheses are tested at alpha = 0.05, some significant results are expected by chance. Testing twenty independent nulls gives a substantial chance of at least one p below 0.05. Family-wise error and false-discovery-rate procedures address different multiplicity goals. The choice should be planned for the scientific question rather than applied selectively after seeing results.
Researcher degrees of freedom include trying multiple outcomes, subgroups, transformations, covariate sets, exclusions, stopping points, and tests. Reporting only the smallest p-value understates uncertainty. Preregistration, analysis plans, transparent reporting of all tested outcomes, and validation data reduce this problem. Exploratory analyses are valuable when labelled as exploratory and followed by confirmation.
Repeated looks at accumulating data also change Type I error unless a sequential design accounts for them. In engineering monitoring, thousands of correlated sensor comparisons can produce alerts under naive thresholds. Use procedures designed for the monitoring process and assess practical signal magnitude, persistence, and measurement quality.
Assumptions belong in the interpretation
A p-value is calibrated only under the test's assumptions. A t-test may require independent observations and an appropriate model for means and standard errors. ANOVA adds structure for group comparisons. Chi-square tests need suitable expected counts and sampling. Regression tests depend on model specification and standard-error assumptions. Violations can make p-values too small, too large, or simply target the wrong question.
Inspect distributions, residuals, group sizes, missingness, outliers, dependence, and study design. Robust or permutation methods can help in some cases but have their own conditions. Do not choose a test solely by a preliminary normality p-value; graphical assessment, design, sample size, and the target estimand matter. Testing assumptions with the same small dataset can have low power.
Data cleaning decisions affect inference. Removing outliers, recoding categories, and handling missing values must be justified independently of whether they improve significance. Preserve an audit trail and conduct sensitivity analyses. Statistical software cannot infer whether an observation is a data error or a valid extreme case from magnitude alone.
Common p-value mistakes
The most common mistake is saying the p-value is the probability that the null is true. Another is treating it as effect size or practical importance. A third is turning 0.05 into an absolute truth boundary. P-values are continuous summaries under a model, and their evidential meaning depends on design, prior plausibility, multiplicity, and data quality.
Nonsignificant is not equivalent to no effect, and significant is not equivalent to replication. A confidence interval that spans meaningful values should prevent a confident null conclusion. Rounding p = 0.0496 to 0.05 and then changing the decision inconsistently is avoidable; specify reporting precision and use the unrounded value for a prespecified rule.
Do not write p = 0.000. Report a bound such as p < 0.001 when software rounds below its display precision. Avoid excessive digits that imply numerical certainty unsupported by the test approximation. State whether the test was one- or two-sided, identify the statistic and degrees of freedom where relevant, and report the estimate and interval.
Using ScholarTool inference calculators
Choose a test from the study question and design before entering data. The T-Test Calculator supports specified mean-comparison modes; the One-Way ANOVA Calculator addresses group mean comparisons under its assumptions; the Chi-Square Test Calculator handles appropriate count-table questions; and the Confidence Interval Calculator focuses on estimation. Review input format, independence, scale, and missing-data handling for each tool.
The Statistical Test Selector can organize candidate methods, but it cannot replace protocol-specific statistical judgement. After calculation, report the effect, uncertainty, test statistic, p-value, sample size, and assumptions. Inspect raw data and diagnostics. Do not use the calculator output to certify a research claim, clinical conclusion, or regulatory decision.
Related in this workflow: T-Test Calculator, One-Way ANOVA Calculator, Chi-Square Test Calculator, Confidence Interval Calculator, Statistical Test Selector.
Limitations and cautions
This guide gives general frequentist interpretation and cannot determine whether a particular study design, test, stopping rule, or model is valid. Specialized analyses may require survey weights, mixed models, survival methods, causal inference, equivalence or noninferiority testing, Bayesian models, or multiplicity procedures. A single p-value cannot summarize all evidence.
For consequential health, policy, safety, regulatory, or research decisions, use a documented analysis plan, appropriate software validation, complete reporting, and qualified statistical and subject-matter review. Preserve uncertainty and limitations even when a threshold is crossed.
Related ScholarTool tools
- T-Test Calculator
- One-Way ANOVA Calculator
- Chi-Square Test Calculator
- Confidence Interval Calculator
- Statistical Test Selector
Related categories
References and recommended sources
- ASA Statement on p-Values: American Statistical Association, Statement on Statistical Significance and P-Values.
- ASA Statistical Inference Supplement: American Statistical Association, Moving to a World Beyond p less than 0.05, The American Statistician.
- Statistical Tests: D. J. Sheskin, Handbook of Parametric and Nonparametric Statistical Procedures, Chapman and Hall/CRC.
Continue with the working tools
Use the related calculators to apply the concept, then verify inputs, assumptions, method limits, and references before using an output in consequential work.
Explore Statistics Tools