Statistics cannot repair a weak design. They help you describe the evidence, quantify uncertainty, and decide how far a conclusion can travel.
Start with the data you actually have
Plot the observations before choosing a test. Look for shape, clusters, outliers, missing values, ceiling/floor effects, and changes in variability.
- Use the mean when the arithmetic average matches the question and extreme values are not distorting the summary.
- Use the median when the middle observation is more representative of a skewed distribution or when extremes are influential.
- Always pair center with variability: range, interquartile range, standard deviation, or another measure appropriate to the data.
- Report the sample size and what counts as an independent sample.
Three repeated readings from one sensor may show repeatability. They do not establish how the sensor performs across different devices, days, environments, or populations.
Show uncertainty honestly
Error bars are not self-explanatory. State whether they show standard deviation, standard error, a confidence interval, or something else.
A confidence interval estimates a range produced by a model and sampling procedure. It is not a guarantee that every future result will fall inside it, and it does not erase bias in the sample or measurement.
Effect size answers “how much?” Statistical testing addresses how compatible the observed result is with a specified model or null hypothesis. Practical importance asks whether the size matters in the real system.
Visibly large
The plotted groups look far apart.
Statistically supported
The design and uncertainty support a difference beyond the chosen model’s noise assumptions.
Practically meaningful
The effect is large enough to matter in the real decision or system.
A project may satisfy one, two, all three—or none. Report them separately.
Understand what a p-value is not
A p-value is calculated under assumptions, including a null model. It is not:
- the probability that your hypothesis is true;
- the probability that the result happened “by chance” in ordinary language;
- the size or importance of an effect;
- proof that the study is unbiased or reproducible;
- a substitute for showing the data and design.
Do not turn p < 0.05 into “proved.” Report the exact value when appropriate, the effect estimate, uncertainty, sample, and limitations.
Choose analysis from the design
Ask these questions before naming a test:
- What type of outcome is measured: continuous, count, binary, category, time-to-event, rank, or something else?
- Are groups independent, paired, or repeatedly measured?
- How was the sample selected or assigned?
- What distributional and independence assumptions are plausible?
- Is the question a difference, association, prediction, estimation, or equivalence question?
Consult a knowledgeable teacher or mentor when the analysis affects a major claim. An advanced-looking test is not better if its assumptions do not fit.
Correlation is not causation
An association may reflect reverse causation, selection, measurement error, or a third variable. Causal language requires a design and assumptions that justify it—not merely a small p-value or a machine-learning model with high accuracy.
Avoid multiple-comparison fishing
If you test many outcomes, subgroups, time points, or model variants, some may look unusual by chance. Define a primary question, distinguish planned analyses from exploration, and disclose the full set of relevant comparisons. Do not hide trials or keep changing tests until one crosses a threshold.
Plan sample size before collection
Required sample size depends on the effect worth detecting, expected variability, design, acceptable uncertainty, and analysis—not a universal minimum. A small careful pilot can estimate procedure reliability, but broad claims from tiny samples are rarely justified.
Primary sources
Go deeper
Use these first-party references to check rules, study the method further, or adapt this guide to your field.