← Back to Blog
Healthcare Evidence Synthesis

Levels of Evidence: The Evidence Pyramid Explained

9 min readBy MSc. Michael Patel
Quick answer

The levels of evidence pyramid ranks research designs by how well they protect against bias when answering a question about treatment effects. From top to bottom, a typical pyramid shows: systematic reviews and meta-analyses; randomised controlled trials; cohort studies; case-control studies; cross-sectional studies and case series; case reports; and expert opinion or mechanism-based reasoning at the base. Higher levels are generally more trustworthy, but the right design depends on the question, and a well-conducted study lower down can outrank a flawed one higher up. GRADE (Guyatt et al., 2008) refines this by rating the certainty of a body of evidence as high, moderate, low or very low.

Building an evidence-based practice project? A specialist can guide you through it. Get a free quote →

What is the levels of evidence pyramid?

The levels of evidence pyramid, also called the evidence hierarchy or hierarchy of evidence, is a visual shortcut used in evidence-based medicine, nursing and allied health. It arranges study designs in layers, with designs that are less prone to bias at the top and those more prone to bias at the bottom. The pyramid shape reflects two things: stronger designs sit higher, and there are fewer of them.

Students meet the pyramid in almost every research methods course, and clinicians use it to decide how much to trust a finding. It is also the backbone of EBP frameworks in nursing, where appraising and "levelling" each article is a standard part of a literature review, capstone project or evidence table.

The levels of evidence, from top to bottom

A typical evidence pyramid for questions about treatment effects
LevelStudy designWhy it sits here
1 (top)Systematic reviews and meta-analyses of randomised trialsCombine all eligible trials with transparent methods, reducing the play of chance and selective citation
2Randomised controlled trials (RCTs)Random allocation balances known and unknown confounders between groups
3Cohort studiesFollow exposed and unexposed groups forward in time, but groups may differ in ways that affect outcomes
4Case-control studiesCompare people with and without an outcome, looking back at exposures; vulnerable to recall and selection bias
5Cross-sectional studies and case seriesMeasure exposure and outcome at one time point, or describe a group without comparison
6Case reportsDescribe individual patients; useful for spotting new problems, not for estimating effects
7 (base)Expert opinion, editorials and mechanism-based reasoningBased on experience or theory rather than systematically collected data

Different textbooks and institutions number the levels differently, split some layers and merge others, and some place clinical practice guidelines or "pre-appraised" evidence above systematic reviews. The order of the main designs is broadly consistent, though.

Each study design explained in plain language

Systematic reviews and meta-analyses

A systematic review finds, appraises and synthesises all the studies that answer a focused question, following a protocol. A meta-analysis statistically combines their results. They sit at the top because they summarise all the relevant evidence rather than one study, but only if they are well conducted and include good-quality studies. Read more in our comparison of systematic reviews, meta-analyses and scoping reviews.

Randomised controlled trials

In an RCT, participants are randomly allocated to the intervention or a comparison. Randomisation is powerful because it balances both measured and unmeasured characteristics between groups, so differences in outcome can be attributed to the intervention. Blinding, allocation concealment and complete follow-up protect the result further.

Cohort studies

Cohort studies follow groups of people who differ in an exposure (for example smokers and non-smokers) and compare what happens to them. They are the best practical design for many questions about harm and prognosis, where randomisation would be unethical or impossible, but they are open to confounding.

Case-control studies

Case-control studies start with people who have an outcome (cases) and people who do not (controls), then look back at past exposures. They are efficient for rare diseases, but depend on accurate records or recall.

Cross-sectional studies, case series and case reports

Cross-sectional studies, such as surveys, measure exposure and outcome at the same time, so they show associations but not which came first. Case series and case reports describe patients without a comparison group; they are often the first signal of a new adverse effect or disease.

Expert opinion and mechanistic reasoning

Opinion based on clinical experience or on how a treatment should work biologically sits at the base. It is valuable when nothing else exists, but history has many examples of treatments that made mechanistic sense and later proved harmful in trials.

The best evidence depends on the question

The classic pyramid is built for therapy questions: does this intervention work? For other types of question, the best design is different. The Oxford Centre for Evidence-Based Medicine (OCEBM) Levels of Evidence (OCEBM Levels of Evidence Working Group, 2011) make this explicit, with separate rows for different clinical questions, such as how common a problem is, whether a diagnostic test is accurate, what will happen without treatment, and whether a treatment helps or harms.

Best primary study design by type of question (examples)
Question typeExampleStrongest primary design
Therapy / interventionDoes early mobilisation shorten ICU stay?Randomised controlled trial
DiagnosisHow accurate is point-of-care ultrasound for pneumonia?Cross-sectional diagnostic accuracy study against a reference standard
PrognosisWhat is the five-year survival after diagnosis?Inception cohort study
Harm / etiologyDoes night-shift work increase breast cancer risk?Cohort study (RCTs are usually unethical)
PrevalenceHow common is burnout among ICU nurses?Cross-sectional survey with a random or representative sample
Meaning / experienceHow do patients experience dialysis?Qualitative study; qualitative evidence synthesis

At the top of every row sits a systematic review of the best primary design for that question. So a systematic review of diagnostic accuracy studies is top-level evidence for a diagnosis question, even though it contains no RCTs.

Levels of evidence in nursing: Melnyk and Johns Hopkins

Nursing programmes often use a numbered hierarchy from an EBP textbook or model instead of a generic pyramid. Two appear most often in capstone and EBP assignments:

Melnyk and Fineout-Overholt's seven levels

Melnyk and Fineout-Overholt (2023) rank evidence on seven levels:

  1. Level I: systematic review or meta-analysis of randomised controlled trials (or evidence-based guidelines based on them).
  2. Level II: a well-designed randomised controlled trial.
  3. Level III: a controlled trial without randomisation (quasi-experimental).
  4. Level IV: case-control or cohort studies.
  5. Level V: systematic review of descriptive and qualitative studies.
  6. Level VI: a single descriptive or qualitative study.
  7. Level VII: opinion of authorities and/or reports of expert committees.

The Johns Hopkins EBP model

The Johns Hopkins Nursing Evidence-Based Practice model (Dang et al., 2021) uses five levels, from experimental studies such as RCTs (Level I) and quasi-experimental studies (Level II), through non-experimental and qualitative studies (Level III), to clinical practice guidelines and consensus statements (Level IV) and literature reviews, quality improvement reports, case reports and expert opinion (Level V). It then adds a quality rating (A, B or C), so every article gets both a level and a grade.

Always use the hierarchy your programme or organisation specifies, and name it in your evidence table so markers know which scale you applied.

Limitations of the evidence pyramid

The pyramid is a useful heuristic, but it has well-known weaknesses:

Murad et al. (2016) proposed a new evidence pyramid in response. They replaced the straight lines between levels with wavy lines, to show that a study can move up or down depending on its quality, as in GRADE. They also removed systematic reviews from the top and presented them as a lens through which the other evidence is viewed.

Beyond the pyramid: GRADE certainty of evidence

Most guideline developers and systematic reviewers now use GRADE (Grading of Recommendations, Assessment, Development and Evaluation; Guyatt et al., 2008) to judge a whole body of evidence for each outcome, rather than levelling single studies. GRADE rates certainty as high, moderate, low or very low.

GRADE ratings appear in the summary of findings tables of Cochrane reviews and in many clinical guidelines. If you are writing a systematic review, rating certainty with GRADE is expected by most journals.

How to use the evidence levels in your own work

  1. Identify the question type (therapy, diagnosis, prognosis, harm, meaning).
  2. Identify each study's design from its methods section, not its title; authors sometimes call a before-and-after study a "trial".
  3. Assign a level using the hierarchy your programme requires.
  4. Appraise quality with a design-specific tool: RoB 2 for randomised trials (Sterne et al., 2019), ROBINS-I (Sterne et al., 2016) for non-randomised studies of interventions, the Newcastle-Ottawa Scale for cohort and case-control studies, JBI checklists for many designs, and AMSTAR 2 for systematic reviews.
  5. Summarise in an evidence table with columns for citation, design, level, sample, findings, quality and relevance. A synthesis matrix helps you compare studies by theme.
  6. Judge the body of evidence, not just individual studies, ideally with GRADE.

Worked example: levelling five studies on one question

Suppose your PICOT question asks whether chlorhexidine bathing reduces bloodstream infections in ICU patients, and your search finds five articles. Using the Melnyk and Fineout-Overholt (2023) hierarchy:

Assigning levels to a mixed set of studies (hypothetical)
ArticleDesign (from the methods)Level
AMeta-analysis of cluster-randomised trialsI
BMulticentre cluster-randomised crossover trialII
CBefore-and-after study on one unit, no control groupIII (quasi-experimental)
DRetrospective cohort comparing units that did and did not adopt bathingIV
EInterviews with nurses about barriers to daily bathingVI

Article E answers a different question (implementation barriers) and is still useful in a capstone, for example in your implementation plan. Articles A and B carry most weight for effectiveness, but only after appraisal: if article B had high loss to follow-up, its level stays the same but its quality rating drops.

Need help levelling and appraising evidence?

Levelling and appraising dozens of studies is time-consuming, especially when designs are poorly reported. Our healthcare evidence synthesis service can build your evidence tables and GRADE assessments, and our systematic review service covers the full process from search to summary of findings.

Frequently asked questions

What is the highest level of evidence?

For questions about treatment effects, a well-conducted systematic review and meta-analysis of randomised controlled trials is usually considered the highest level. For other questions, the top level is a systematic review of the best primary design for that question.

What is level 1 evidence?

In most hierarchies, level 1 (or Level I) evidence is a systematic review or meta-analysis of randomised controlled trials.

Where do qualitative studies fit in the evidence pyramid?

The traditional pyramid is built for effectiveness questions and places qualitative studies low. For questions about meaning and experience, qualitative studies and qualitative evidence syntheses are the most appropriate evidence.

Is expert opinion evidence?

Yes, but it is the lowest level, because it is not based on systematically collected data. It is useful when no research exists.

What is the difference between the evidence pyramid and GRADE?

The pyramid ranks individual study designs. GRADE rates the certainty of a whole body of evidence for each outcome as high, moderate, low or very low, taking into account risk of bias, inconsistency, indirectness, imprecision and publication bias.

What are the 7 levels of evidence in nursing?

In Melnyk and Fineout-Overholt's (2023) hierarchy: I, systematic reviews or meta-analyses of RCTs; II, RCTs; III, controlled trials without randomisation; IV, case-control or cohort studies; V, systematic reviews of descriptive and qualitative studies; VI, single descriptive or qualitative studies; VII, expert opinion.

Sources

  1. Murad MH, Asi N, Alsawas M, Alahdab F. New evidence pyramid. Evid Based Med 2016;21:125-127
  2. OCEBM Levels of Evidence Working Group. The Oxford Levels of Evidence 2. Oxford Centre for Evidence-Based Medicine; 2011
  3. Guyatt GH, Oxman AD, Vist GE, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ 2008;336:924-926
  4. Melnyk BM, Fineout-Overholt E. Evidence-Based Practice in Nursing and Healthcare: A Guide to Best Practice. 5th ed. Philadelphia, PA: Wolters Kluwer; 2023
  5. Dang D, Dearholt SL, Bissett K, Ascenzi J, Whalen M. Johns Hopkins Evidence-Based Practice for Nurses and Healthcare Professionals: Model and Guidelines. 4th ed. Indianapolis, IN: Sigma Theta Tau International; 2021
  6. Sterne JAC, Savović J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ 2019;366:l4898
  7. Sterne JA, Hernán MA, Reeves BC, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ 2016;355:i4919
Evidence Synthesis & Editorial Lead · MSc in Health Research Methods; training in qualitative synthesis
Michael has 6 years of experience in health research, nursing, public health, the social sciences, and qualitative evidence synthesis. He leads or supports projects that organize and interpret qualitative findings while preserving the context of the original studies. Areas of expertise Qualitative evidence…
Need help with Healthcare Evidence Synthesis? See how our specialists can support you.View Service Page →

Need help with your evidence-based practice project?

Tell us what you need and a specialist will reply with a free, itemized quote within 2-4 business hours.