To do a systematic review, you (1) define a focused question with PICO, (2) write and register a protocol, (3) search several databases with a peer-reviewed search strategy, (4) screen studies in duplicate against set criteria, (5) extract data with a piloted form, (6) assess risk of bias, (7) synthesise the results (with a meta-analysis if studies are similar enough), (8) rate the certainty of evidence with GRADE, and (9) report everything using the PRISMA 2020 checklist and flow diagram.
Working on a systematic review? A PhD-led reviewer can help at any stage. Get a free quote →
What is a systematic review?
A systematic review is a piece of research that answers a clearly defined question by finding, selecting, appraising and summarising all the relevant studies, using methods that are decided in advance and reported in enough detail for someone else to repeat them. That last point is what separates it from a traditional literature review. In a narrative review, the author chooses which studies to discuss; in a systematic review, the protocol decides, and every decision is written down.
Because the methods are explicit, a well-conducted systematic review limits the two biggest problems in summarising research: missing relevant studies and cherry-picking the ones that support a conclusion. That is why systematic reviews underpin clinical guidelines, health policy and funding decisions, and why they are one of the most requested types of research support we provide.
A systematic review may or may not include a meta-analysis. The review is the whole process of finding and appraising studies; a meta-analysis is the optional statistical step that pools their results into a single estimate. If the studies are too different to combine, you write a narrative (or structured) synthesis instead, and the work is still a systematic review.
Every systematic review reports a PRISMA 2020 flow diagram (shown above) that records how many studies were found, screened, excluded with reasons, and included. You can draw your own with our free PRISMA flow diagram generator.
Step 1: Define a focused research question (PICO)
Everything in a systematic review flows from the question, so it has to be precise. For questions about the effects of an intervention, the standard tool is PICO, first described as the "well-built clinical question" by Richardson et al. (1995):
- Population: who are you interested in? (for example, adults with type 2 diabetes)
- Intervention or exposure: what are you evaluating? (supervised resistance training)
- Comparison: what is it compared with? (usual care)
- Outcome: what will you measure? (glycaemic control, measured as HbA1c)
Adding T for time (PICOT) or S for study design (PICOS) helps when the follow-up period or eligible designs matter. For qualitative evidence, frameworks such as SPIDER or PCC (Population, Concept, Context) fit better, and PCC is also the standard for scoping reviews.
Broad topic: exercise and diabetes.
Focused question: In adults with type 2 diabetes, does supervised resistance training, compared with usual care, improve HbA1c over at least 12 weeks?
A good question is narrow enough to answer and broad enough to find evidence. If an early scoping search returns thousands of studies, narrow the population, intervention or outcomes. If it returns almost nothing, broaden one element or consider a scoping review. Our free PICO question generator turns each element into a question and a matching search.
Step 2: Write and register a protocol
A protocol is your plan: the question, eligibility criteria, databases, search approach, screening and extraction methods, risk of bias tools and planned analyses. Writing it before you see the results protects the review from decisions that are influenced, consciously or not, by what the studies show.
For reviews with a health-related outcome, register the protocol in PROSPERO, the international prospective register of systematic reviews run by the Centre for Reviews and Dissemination at the University of York (Centre for Reviews and Dissemination, n.d.). Registration is free of charge, and PROSPERO passed 500,000 registered reviews in September 2026. Registering publicly also shows editors and reviewers that your methods were fixed in advance, and helps you spot whether someone is already doing the same review.
- State eligibility criteria for each PICO element, plus study designs, languages, dates and settings.
- List every database and other source you will search.
- Describe how many reviewers will screen and extract, and how disagreements will be resolved.
- Name the risk of bias tool and the effect measures you plan to use.
- Pre-specify subgroup and sensitivity analyses rather than adding them after seeing the data.
Step 3: Build and run a comprehensive search
The search is where many reviews are criticised by peer reviewers. The aim is sensitivity: finding as many relevant studies as possible, even if that means screening many irrelevant ones.
Which databases should you search?
For health topics, most reviews search at least MEDLINE (often through PubMed), Embase and the Cochrane Central Register of Controlled Trials (CENTRAL), plus subject databases such as CINAHL for nursing or PsycINFO for psychology. Add trial registers (ClinicalTrials.gov and the WHO International Clinical Trials Registry Platform) and grey literature to reduce the risk of missing unpublished studies. Google Scholar is useful as a supplement, but its searches cannot be reproduced exactly, so it should not be your main database.
How to structure the search
- Split the question into concepts, usually population and intervention (outcomes are often left out to avoid missing studies).
- For each concept, combine controlled vocabulary (MeSH in MEDLINE, Emtree in Embase) with free-text synonyms, spelling variants and truncation.
- Join synonyms within a concept with OR, and join concepts with AND.
- Adapt the syntax for each database rather than copying it unchanged.
- Record the date, database, platform and number of results for every search.
The Cochrane Handbook (Lefebvre et al., 2025) strongly recommends that search strategies are peer reviewed by an experienced information specialist before they are run. It notes that running duplicate searches in parallel is not necessary, unlike screening, but getting a second expert to check the strategy is. Report the full strategies using PRISMA-S, a 16-item extension of PRISMA for reporting literature searches (Rethlefsen et al., 2021).
Step 4: Screen studies in duplicate
Export all results into a reference manager or screening tool (such as EndNote, Zotero, Covidence or Rayyan) and remove duplicates first. Our reference deduplication tool matches records by DOI, PMID and title.
Screening then happens in two stages: titles and abstracts, then full texts of anything that might be eligible. The Cochrane Handbook (Lefebvre et al., 2025) recommends using at least two people working independently to decide whether each study meets the eligibility criteria, and ideally screening titles and abstracts in duplicate as well. Disagreements are resolved by discussion or by a third reviewer.
- Pilot the criteria on a sample of records so reviewers interpret them the same way.
- Record a reason for every full-text exclusion; PRISMA asks you to report them.
- Keep count at every stage so you can draw the PRISMA flow diagram later.
- Measure agreement between reviewers with Cohen's kappa if your journal or supervisor asks for it; our kappa calculator does this.
Step 5: Extract data with a piloted form
Data extraction means copying the information you need from each included study into a standard form: study design and setting, participant characteristics, intervention and comparison details, outcomes, time points, results and funding. Extract more than you think you need, because returning to 40 full texts later is slow.
Pilot the form on two or three studies first, then refine it. As with screening, extraction is best done by two people independently, or by one person with full checking by a second, because errors in extracted numbers carry straight into the analysis. Where results are missing or unclear, contact study authors and record whether they replied.
Step 6: Assess risk of bias in each study
Risk of bias assessment asks whether a study's design or conduct could have systematically distorted its results. It is not the same as asking whether a study is well reported. Use a tool designed for the study design you are assessing:
| Study design | Recommended tool | Judgements |
|---|---|---|
| Randomised trials | RoB 2 (Sterne et al., 2019) | Low risk, some concerns, high risk |
| Non-randomised studies of interventions | ROBINS-I (Sterne et al., 2016) | Low, moderate, serious, critical |
| Diagnostic accuracy studies | QUADAS-2 | Low, high, unclear |
| Observational (cohort, case-control) | Newcastle-Ottawa Scale or ROBINS-E | Stars or domain judgements |
Assess each study in duplicate, record the reason for every judgement, and present the results in a table or traffic-light plot. Risk of bias should then feed into your synthesis and your certainty rating; it is not just a box to tick.
Step 7: Synthesise the results
If the included studies are similar enough in population, intervention, comparison and outcome, you can combine them statistically in a meta-analysis. The result is a pooled effect size with a confidence interval, usually shown in a forest plot. Random-effects models are common when studies differ in ways that may change the true effect, and the between-study variability is measured with statistics such as I² and τ² (Deeks et al., 2024).
If studies are too different, pool only what makes sense and describe the rest with a structured narrative synthesis, grouping studies by population, intervention or outcome and presenting results in tables. Avoid "vote counting" (tallying significant and non-significant studies), because it ignores study size and effect magnitude.
- Pre-specified subgroup analyses explore why effects might differ.
- Sensitivity analyses test whether conclusions change when you exclude high risk of bias studies or change the model.
- Assess small-study effects and publication bias with a funnel plot; the Cochrane Handbook (Page et al., 2024) advises that funnel plot asymmetry tests are used only when there are at least 10 studies.
Step 8: Rate the certainty of the evidence (GRADE)
A pooled result is only as trustworthy as the studies behind it. The GRADE approach (Guyatt et al., 2008) rates the certainty of evidence for each outcome as high, moderate, low or very low. Evidence from randomised trials starts high and can be rated down for risk of bias, inconsistency, indirectness, imprecision and publication bias; evidence from observational studies starts lower and can occasionally be rated up.
Present the ratings in a summary of findings table so readers can see, outcome by outcome, what the evidence shows and how confident they can be. Our GRADE evidence rating tool builds this table for you.
Step 9: Report with PRISMA 2020
Report the review using the PRISMA 2020 statement (Page et al., 2021), the reporting guideline for systematic reviews. Its 27-item checklist covers everything from the title and abstract to registration, funding and data availability, and its flow diagram shows how studies moved through the review. Most journals ask you to submit the completed checklist with page numbers.
Draft each section against the checklist from the start, rather than retrofitting it at the end. Our systematic review report builder lays out all 27 items in manuscript order and exports both the draft and the completed checklist to Word.
Common mistakes that delay systematic reviews
- A question that is too broad. Thousands of hits and no realistic way to screen them. Narrow one PICO element.
- No registered protocol. Reviewers increasingly ask why methods were not fixed in advance.
- Searching only one or two databases or relying on Google Scholar, which cannot be reproduced.
- Single-reviewer screening without justification, which Cochrane methods do not support (Lefebvre et al., 2025).
- Pooling studies that are too different, producing a precise-looking but meaningless average.
- Ignoring risk of bias in the conclusions, or using a generic quality checklist instead of a design-specific tool.
- Reporting gaps, such as missing search strategies or no reasons for exclusions.
How long does a systematic review take, and when to get help
Timelines depend on the size of the evidence base and the team. The screening, extraction and risk of bias steps involve at least two people at each stage, so a single researcher working alone cannot follow recommended methods without help. Plan for the review to take many months rather than weeks, and build in time for peer review of the search, author queries and revisions.
Researchers most often bring in expert support for the search strategy, second-reviewer screening, risk of bias assessment and the meta-analysis. Our systematic review service covers every stage from protocol to publication-ready manuscript, with two reviewers screening every study independently, and you stay in control of the question and conclusions.
Frequently asked questions
What are the steps of a systematic review?
Define a PICO question, write and register a protocol, search several databases, screen studies in duplicate, extract data, assess risk of bias, synthesise the results (with a meta-analysis if appropriate), rate certainty with GRADE, and report with PRISMA 2020.
Can one person do a systematic review?
One person can lead a review, but recommended methods such as the Cochrane Handbook's (Lefebvre et al., 2025) call for at least two people to judge eligibility independently. Most single researchers work with a second reviewer for screening, extraction and risk of bias.
Do I need to register my systematic review?
Registration is strongly recommended and increasingly expected by journals. Reviews with a health-related outcome can be registered free of charge in PROSPERO before the review begins.
What is the difference between a systematic review and a meta-analysis?
A systematic review is the whole structured process of finding and appraising studies. A meta-analysis is an optional statistical step within it that combines study results into one estimate.
How many databases should a systematic review search?
There is no fixed number, but health reviews usually search MEDLINE, Embase and CENTRAL plus subject databases, trial registers and grey literature, so that relevant studies are not missed.
What is PRISMA 2020?
PRISMA 2020 is the reporting guideline for systematic reviews. It has a 27-item checklist and a flow diagram showing how studies were identified, screened and included.
Sources
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 2021;372:n71
- Rethlefsen ML, Kirtley S, Waffenschmidt S, et al. PRISMA-S: an extension to the PRISMA Statement for reporting literature searches in systematic reviews. Syst Rev 2021;10:39
- Lefebvre C, Glanville J, Briscoe S, et al. Chapter 4: Searching for and selecting studies. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions version 6.5.1. Cochrane; 2025
- Deeks JJ, Higgins JPT, Altman DG, McKenzie JE, Veroniki AA, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions version 6.5. Cochrane; 2024
- Page MJ, Higgins JPT, Sterne JAC. Chapter 13: Assessing risk of bias due to missing evidence in a meta-analysis. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions version 6.5. Cochrane; 2024
- Centre for Reviews and Dissemination. PROSPERO: International prospective register of systematic reviews. University of York; n.d.
- Richardson WS, Wilson MC, Nishikawa J, Hayward RS. The well-built clinical question: a key to evidence-based decisions. ACP J Club 1995;123:A12
- Sterne JAC, Savović J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ 2019;366:l4898
- Sterne JA, Hernán MA, Reeves BC, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ 2016;355:i4919
- Guyatt GH, Oxman AD, Vist GE, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ 2008;336:924-926
