← Back to Blog
Statistical Data Analysis

Which Statistical Test Should I Use? A Step-by-Step Guide

9 min readBy TimelyScholar Research Team
Statistical analysis from test choice to an APA 7 results sentence, with a dot plot of two groups
Quick answer

To choose a statistical test, answer three questions: (1) What is your aim? Comparing groups, or looking at a relationship or prediction? (2) What type is your outcome variable? Continuous, ordinal, or categorical? (3) How many groups, and are they independent or paired? Then match: two independent groups with a continuous outcome use an independent t-test (or Mann-Whitney U if assumptions fail); paired data use a paired t-test (or Wilcoxon signed-rank); three or more groups use one-way ANOVA (or Kruskal-Wallis); two categorical variables use a chi-square test (or Fisher's exact test for small samples); relationships use Pearson or Spearman correlation; prediction uses linear or logistic regression.

Unsure about your analysis? Our statisticians can plan, run and explain it with you. Get a free quote →

The three questions that decide your statistical test

Choosing a statistical test feels complicated because there are so many of them, but for most student projects, dissertations and journal articles, the choice comes down to three questions. Answer them in order and the right test usually becomes obvious.

  1. What is your research question trying to do? Compare groups ("Is there a difference?") or examine a relationship ("Is X associated with Y?", "Does X predict Y?").
  2. What type of data is your outcome (dependent) variable? Continuous, ordinal or categorical.
  3. How many groups are you comparing, and are the observations independent or paired? Different people in each group, or the same people measured twice?

Step 1: Identify your variable types

Types of variables
TypeDefinitionExamples
Continuous (interval / ratio)Numeric measurements on a scale with meaningful distancesBlood pressure, age, weight, exam score, reaction time
OrdinalOrdered categories with unequal or unknown gapsPain rated mild/moderate/severe, Likert items, cancer stage
Nominal (categorical)Categories with no natural orderSex, blood group, treatment group, yes/no outcomes
Binary (dichotomous)A nominal variable with two categoriesReadmitted yes/no, passed/failed, alive/dead

Identify both your outcome (dependent) variable and your predictor or grouping (independent) variable. For example, in "Does a new teaching method improve exam scores?", the outcome is exam score (continuous) and the predictor is teaching method (nominal, two groups).

Likert-type data cause frequent debate. A single Likert item is ordinal. A total score summed from many items is often treated as continuous, especially if it has a wide range and a reasonable distribution. State and justify your choice.

Step 2: Parametric or non-parametric?

Parametric tests (t-tests, ANOVA, Pearson correlation, linear regression) make assumptions about the data, typically that the outcome (or the residuals) is approximately normally distributed, that variances are similar between groups, and that observations are independent. When those assumptions hold, parametric tests are more powerful.

Non-parametric tests (Mann-Whitney U, Wilcoxon, Kruskal-Wallis, Spearman) make fewer assumptions. They are usually based on ranks and are appropriate for ordinal outcomes, for continuous outcomes that are clearly skewed with small samples, or when there are extreme outliers.

Step 3: Tests for comparing groups

Choosing a test to compare groups
OutcomeGroupsIndependent groupsPaired / repeated
Continuous (normal)2Independent-samples t-test (Welch)Paired t-test
Continuous (not normal) or ordinal2Mann-Whitney U testWilcoxon signed-rank test
Continuous (normal)3 or moreOne-way ANOVARepeated-measures ANOVA
Continuous (not normal) or ordinal3 or moreKruskal-Wallis testFriedman test
Categorical2 or moreChi-square test of independence (Fisher's exact for small expected counts)McNemar test (binary, 2 time points)

Independent-samples t-test

Compares the means of two independent groups, for example exam scores for students taught with method A versus method B (NIST/SEMATECH, n.d.-c). Calculate it with our t-test calculator, and see our step-by-step guide to the independent-samples t-test in SPSS.

Paired t-test

Compares two related measurements, such as blood pressure before and after an intervention in the same patients. The analysis is based on the differences within each pair.

One-way ANOVA

Compares means across three or more independent groups, such as pain scores for three analgesics (NIST/SEMATECH, n.d.-b); see our worked one-way ANOVA in SPSS example, or the 2 × 2 factorial ANOVA guide when you have two independent variables. A significant ANOVA tells you that at least one group differs; follow it with post hoc tests (for example Tukey's HSD) to find which. Try the ANOVA calculator.

Mann-Whitney U and Kruskal-Wallis

Rank-based alternatives to the independent t-test and one-way ANOVA, suitable for ordinal or skewed outcomes. See the Mann-Whitney U calculator and Kruskal-Wallis calculator.

Chi-square and Fisher's exact test

Test whether two categorical variables are associated, for example whether infection rates differ between two wound dressings (NIST/SEMATECH, n.d.-a). Our chi-square test of independence guide walks through the SPSS output. The chi-square test relies on expected counts that are not too small; a common rule of thumb is that most expected counts should be at least 5. When they are not, especially in 2×2 tables, use Fisher's exact test. Use our chi-square calculator and Fisher's exact test calculator.

McNemar test

For paired binary data, such as the same patients classified as anxious or not before and after an intervention. Try the McNemar test calculator.

Step 4: Tests for relationships and prediction

Choosing a test for relationships
QuestionVariablesTest
Are two variables associated?Two continuous, roughly linear and normalPearson correlation (r)
Are two variables associated?Ordinal, skewed or with outliersSpearman rank correlation (rho)
Does X predict a continuous outcome?One or more predictors, continuous outcomeLinear regression
Does X predict a binary outcome?One or more predictors, yes/no outcomeLogistic regression
Does X predict time to an event?Predictors, survival time with censoringCox proportional hazards regression; Kaplan-Meier and log-rank test to compare groups
Do two raters agree?Categorical ratingsCohen's kappa
Do two raters or measurements agree?Continuous ratingsIntraclass correlation coefficient (ICC)

Correlation measures association, not causation, and Pearson's r only captures linear relationships. Always plot the data first. Our correlation calculator computes Pearson and Spearman coefficients with confidence intervals, and the linear regression and logistic regression calculators handle prediction.

Regression is often the most flexible choice, because it lets you adjust for confounders. A t-test is actually a special case of linear regression with one binary predictor, and ANOVA is linear regression with a categorical predictor.

Worked examples: matching questions to tests

Research questions and the right test
Research questionOutcomeGroups / designTest
Do nursing students who use simulation score higher on a skills exam than those who don't?Exam score (continuous)2 independent groupsIndependent t-test (Mann-Whitney if skewed)
Does anxiety fall after a mindfulness course?Anxiety score (continuous)Same people, before and afterPaired t-test (Wilcoxon if skewed)
Do satisfaction ratings differ across three clinics?Rating (ordinal, 1–5)3 independent groupsKruskal-Wallis test
Is smoking status associated with post-operative infection?Infection yes/noSmoker vs non-smokerChi-square (Fisher's if small counts)
Is hours of sleep related to reaction time?Reaction time (continuous)Two continuous variablesPearson correlation, or linear regression
Which factors predict 30-day readmission?Readmitted yes/noSeveral predictorsLogistic regression
Does blood pressure change across four time points?BP (continuous)Same people, 4 timesRepeated-measures ANOVA or a mixed model

Beyond the p value: effect sizes, confidence intervals and power

A statistical test gives a p value, but a p value alone does not tell you how big or important an effect is. Report an effect size and its 95% confidence interval alongside every test.

Cohen's (1992) conventional benchmarks for effect sizes
Effect sizeSmallMediumLarge
Cohen's d (difference between two means)0.20.50.8
r (correlation)0.10.30.5

Cohen (1992) offered these benchmarks as rough guides when nothing better is available. Where possible, interpret effect sizes in the context of your field and what matters practically.

Before collecting data, calculate the sample size needed to detect a meaningful effect with adequate power (conventionally 80% or 90%) at your chosen alpha (usually 0.05). Underpowered studies often produce non-significant results that are simply inconclusive.

Checking assumptions before you run the test

Every test rests on assumptions. Check them before interpreting results, and report what you checked:

Key assumptions for common tests
TestMain assumptionsIf violated
Independent t-testIndependent observations; outcome roughly normal in each group; similar variancesWelch's t-test for unequal variances; Mann-Whitney U for skewed small samples
Paired t-testDifferences between pairs roughly normalWilcoxon signed-rank test
One-way ANOVAIndependence; normal residuals; equal variancesWelch's ANOVA; Kruskal-Wallis test
Chi-squareIndependent observations; adequate expected countsFisher's exact test; combine sparse categories if meaningful
Pearson correlationLinear relationship; no extreme outliers; roughly normal variablesSpearman correlation; transform variables
Linear regressionLinearity; independent, normally distributed residuals with constant variance; no severe multicollinearityTransformations, robust standard errors, or a different model
Logistic regressionIndependent observations; linear relationship between continuous predictors and the log odds; enough events per predictorFewer predictors, penalised methods, or combining categories

One-tailed or two-tailed? And what about multiple tests?

Use two-tailed tests by default. A one-tailed test is only justified when an effect in the opposite direction would be impossible or of no interest, and the direction was specified before seeing the data. Reviewers are rightly suspicious of one-tailed tests chosen to cross the significance threshold.

If you run many tests, the chance of at least one false positive rises quickly. Pre-specify a small number of primary analyses, and adjust for multiple comparisons where appropriate, for example with post hoc procedures after ANOVA, a Bonferroni correction for a few planned comparisons, or false discovery rate control for many tests.

Running the tests: SPSS, R, Stata, jamovi and Excel

All the tests above are available in standard statistical software. SPSS is common in health and social sciences; R is free and powerful; Stata is popular in epidemiology and economics; jamovi and JASP are free, point-and-click programs built on R; Excel can run basic t-tests and correlations but is limited and error-prone for anything more complex. Whatever you use, record the exact procedure and software version so your analysis can be reproduced.

Common mistakes when choosing a statistical test

Getting help with your statistics

If you are still unsure which test fits your design, or you need help running and interpreting the analysis, our statistical data analysis service works in SPSS, R, Stata and more, and explains each result in plain language. Once your analysis is done, see how to report statistics in APA 7.

Frequently asked questions

How do I know which statistical test to use?

Identify your aim (comparing groups or examining relationships), the type of your outcome variable (continuous, ordinal or categorical), and the number of groups and whether they are independent or paired. These three answers point to the right test.

When should I use a t-test vs ANOVA?

Use a t-test to compare the means of two groups. Use ANOVA for three or more groups, followed by post hoc tests to identify which groups differ.

When should I use a non-parametric test?

Use a non-parametric test when your outcome is ordinal, or when a continuous outcome is clearly non-normal with a small sample or has extreme outliers.

What test do I use for two categorical variables?

Use a chi-square test of independence, or Fisher's exact test when expected counts are small, particularly in 2×2 tables. For paired binary data use McNemar's test.

What is the difference between correlation and regression?

Correlation measures the strength and direction of an association between two variables. Regression models how one or more predictors relate to an outcome, allowing prediction and adjustment for confounders.

What test should I use for Likert scale data?

For a single Likert item, use non-parametric tests such as Mann-Whitney U or Kruskal-Wallis. Summed scale scores with many items are often analysed with parametric tests if the distribution is reasonable.

Sources

  1. Cohen J. A power primer. Psychol Bull 1992;112:155-159
  2. NIST/SEMATECH. e-Handbook of Statistical Methods: Two-sample t-test for equal means. n.d.-c
  3. NIST/SEMATECH. e-Handbook of Statistical Methods: One-way ANOVA. n.d.-b
  4. NIST/SEMATECH. e-Handbook of Statistical Methods: Chi-square test. n.d.-a
TimelyScholar Research Team
Written by TimelyScholar's PhD-led team of research methodologists and statisticians, who support systematic reviews, meta-analyses and data analysis. Facts and methods are checked against the sources listed above.
Need help with Statistical Data Analysis? See how our specialists can support you.View Service Page →

Need help with your statistical analysis?

Tell us what you need and a specialist will reply with a free, itemized quote within 2-4 business hours.