Figwise
Back to blog
How to Do a Systematic Review: 10 Steps From Idea to Paper

How to Do a Systematic Review: 10 Steps From Idea to Paper

Work through PICO, PROSPERO registration, search strings, two-stage screening, risk of bias, and PRISMA reporting — with the output each step must produce.

FigwiseFigwise Team

A systematic review finds, judges, and combines every study that answers one set question. The methods are fixed up front and reported in full. That last part is what makes it "systematic": a reader could rerun your process and land on the same studies.

It is also slow. A study of 195 finished reviews in the PROSPERO registry found a mean of 67.3 weeks from registration to published paper, and a mean of 5 authors each (Borah et al., 2017; the sample was drawn in 2014). So this guide does more than list the steps. For each one, it names the output you must produce — and the mistakes that force teams to start over.

Here are the 10 steps. Each ends with something you can check off.

Before step 1: make sure a systematic review is the right tool

Skip the systematic review — or switch to a different review type — in three cases.

  • Your question is too broad. "What treatments work for depression?" would return tens of thousands of records. The Cochrane Handbook warns against questions that pull in more records than a team can handle. Narrow the population, the intervention, or both.
  • A recent review already answers it. Search PROSPERO for registered and ongoing reviews, and PubMed for published ones. If a solid review from the past few years covers your question, a duplicate adds little.
  • The evidence is too thin or too mixed to combine. If you expect a handful of studies with wildly different designs, a scoping review fits better — it maps what exists instead of pooling results. Note that PROSPERO says it does not currently accept scoping reviews, so those register elsewhere (such as Open Science Framework).

If your question survives all three checks, go ahead.

Step 1: Frame the question with PICO and build a team

Turn your topic into one clear question using PICO. The letters stand for Population, Intervention, Comparator and Outcome.

Here is a full worked example. Topic: "exercise and depression." PICO version:

  • P — adults (18+) with a diagnosis of major depressive disorder
  • I — supervised aerobic exercise programs
  • C — usual care or waitlist control
  • O — change in depressive symptom score on a validated scale

The question then reads:

In adults with major depressive disorder, does supervised aerobic exercise, compared with usual care, reduce depressive symptom scores?

Every later choice flows from these four parts — search terms, inclusion rules, extraction fields. For qualitative questions, the SPIDER framework plays the same role.

Build the team now, not later. You need at least two reviewers. Screening and bias checks must be done twice, by two people working apart. Add a librarian for the search, and a statistician if you may pool results.

  • You produce: a one-sentence PICO question and a named team with roles.
  • Tools: none needed; a shared document is enough.
  • Common rework cause: a vague outcome ("effectiveness") that later splits into several outcomes no one agreed on. Pin down the exact measures now.

Step 2: Write a protocol and register it on PROSPERO

The protocol fixes your methods before you see any results. It covers the question, the in-and-out rules, planned databases, the screening process, extraction fields, the bias tool, and the synthesis plan. This is your defense against bias. No one can accuse you of bending the rules to fit findings if the rules sit timestamped in a public registry.

Register on PROSPERO, the global register for reviews with health outcomes in humans. The timing rule is strict: the register does not accept reviews that have already started data extraction. Register after your protocol is drafted and before formal screening begins — in practice, before or alongside your searches.

The current PROSPERO launched in February 2025. Its help page says what happens once every team member confirms the record:

"a registration number is assigned and the registration record is published immediately on the PROSPERO site. There is no delay to publication."

The trade-off is on the same page: PROSPERO "does not check record content (beyond built-in automatic checks) and does not provide peer review." The quality of the record is on you.

PROSPERO help page explaining that registration records are published immediately after all authors confirm, with no delay to publication

PROSPERO's "How to register" page describing the current immediate-publication process. Source: crd.york.ac.uk, captured 23 August 2026.

  • You produce: a registered protocol with a PROSPERO ID.
  • Tools: PROSPERO is free. A full protocol PDF can be attached to the record.
  • Common rework cause: registering after screening has started. Journals check the registration date against your PRISMA flow dates, and a late entry undercuts the whole point.

Step 3: Build the search strategy

The goal is reach over precision: catch every relevant study, and accept that most of what you pull will be thrown out. In the Borah study, included studies made up a mean of 2.94% of records found. A search that feels wastefully broad is usually correct.

Search at least three databases. For a review of trials, Cochrane sets the minimum as CENTRAL, MEDLINE and Embase:

  • PubMed — free
  • CENTRAL — the Cochrane trials register, free to search
  • Embase — paid, usually via your library
  • one subject database — CINAHL for nursing, PsycINFO for psychology

In each one, mix controlled terms (MeSH in PubMed) with free-text words, because indexing lags and varies.

Here is a runnable PubMed search for the step 1 example:

("Exercise"[Mesh] OR exercis*[tiab] OR "physical activity"[tiab] OR "aerobic training"[tiab])
AND
("Depressive Disorder"[Mesh] OR "Depression"[Mesh] OR depress*[tiab])
AND
(randomized controlled trial[pt] OR randomized[tiab] OR randomised[tiab])

Each block pairs a MeSH term with [tiab] free-text variants. The star in exercis* catches every ending of the word. The [pt] filter limits to trial reports, while the free-text terms catch trials not yet indexed.

Then cover grey literature:

Skip unpublished trials and you tilt your review toward positive results.

Have a librarian check the strategy before you run it. Search syntax differs across databases, and one missed synonym can hide a cluster of studies. Errors also creep in when a search is carried from one database to the next, and they are hard to spot.

This is the step where an hour of expert review saves months.

  • You produce: the full search string for every database, with run dates and hit counts — you must report these verbatim under PRISMA.
  • Tools: PubMed is free; Embase, CINAHL, and PsycINFO usually need a library subscription.
  • Common rework cause: finding a missed key term after screening is done, which forces a re-search and a second screening pass.

Step 4: Export and deduplicate records

Export all hits into one reference manager and remove duplicates. The same study shows up in several databases. Count it twice and both your screening numbers and your results are off.

Log two numbers before you touch anything: total records per database, and duplicates removed. Both go into your PRISMA flow diagram later, and they cannot be rebuilt after the fact.

  • You produce: one deduplicated library, plus the counts.
  • Tools: Zotero (free) or EndNote (paid); screening tools like Rayyan and Covidence also deduplicate on import.
  • Common rework cause: deduplicating each database export on its own and losing track of the combined totals.

Step 5: Screen in two stages, with two reviewers

Screening happens twice. First pass: titles and abstracts only, judged against your protocol's rules — when in doubt, keep it in. Second pass: full text of everything that survived, with a logged reason for every study dropped.

Both stages need two reviewers working on their own, with conflicts settled by talk or a third reviewer. This is not optional politeness: one person screening alone misses studies. PRISMA item 8 makes you say how many people screened each record and whether they worked apart. Cochrane treats double screening as a standard to meet.

Two rules protect you from subtle bias:

  • Do not exclude a study because it failed to report your outcome. If a trial measured depression scores but never published them, dropping it rewards selective reporting. Keep it in and chase the data.

  • Track full-text exclusion reasons in categories ("wrong population," "wrong comparator," "not randomized"). PRISMA requires them, and rebuilding reasons from memory is miserable.

  • You produce: a final list of included studies, plus counts and exclusion reasons for the flow diagram.

  • Tools: Rayyan has a free tier and handles blind dual screening. Covidence is paid (from $339 per year for one review as of August 2026; many university libraries hold site licenses) and covers screening through extraction.

  • Common rework cause: criteria that turn out to be fuzzy mid-screen ("adults" — does 16+ count?). Pilot the criteria on 50–100 records and fix the wording before the full pass.

Step 6: Extract data with a piloted form

Build a structured form from your PICO parts: study design, setting, sample size, who took part, what was done in each arm, outcome data (means, SDs, event counts), and funding source. Then pilot it on three to five studies before doing the rest. The pilot almost always reveals missing fields.

Extract in duplicate, or have a second reviewer check a sample. And merge multiple reports of the same study. A trial published as a protocol paper, a results paper, and a follow-up is one study, not three. Count it three times and you give it triple weight.

  • You produce: one completed extraction table covering every included study.
  • Tools: a spreadsheet is fine and free; Covidence has built-in extraction forms.
  • Common rework cause: an unpiloted form missing a field (subgroup data, adverse events), which sends you back through every PDF.

Step 7: Assess risk of bias — the tool depends on study design

Risk of bias asks whether each study's design could have skewed its result. Was the draw truly random? Were raters blinded? Did dropouts differ between arms?

A precise number from a flawed trial is still a flawed number.

The tool is not a free choice — it follows from the study designs you included:

Study design Tool Where
Randomized controlled trials RoB 2 Cochrane RoB 2 page / riskofbias.info
Non-randomized studies of an intervention, including cohort and case-control designs ROBINS-I Cochrane ROBINS-I page
Observational studies of a non-intervention question (aetiology, prognosis) Newcastle-Ottawa Scale Ottawa Methods Centre

All three are free. If your question is about an intervention, use ROBINS-I even for cohort and case-control designs. That is what the Cochrane Handbook asks for, and what reviewers expect. Note also that the Newcastle-Ottawa Scale gives a star rating, so it cannot produce the per-domain judgments this step asks for.

Rate in duplicate, judge each domain rather than one overall score, and plan to use the ratings — say, a check on whether results hold once high-risk studies are dropped. A bias table that never touches the results is decoration.

  • You produce: a per-study, per-domain bias table with reasons for each judgment.
  • Tools: the official templates above; RevMan and Covidence embed RoB 2.
  • Common rework cause: using a checklist that does not match the study design — NOS on trials, RoB 2 on cohort studies. Peer reviewers flag this at once.

Step 8: Synthesize — meta-analysis only when the studies allow it

Every systematic review needs a structured synthesis. Not every one gets a meta-analysis.

Pool results only when studies are alike enough to combine:

  • the same core question
  • similar people and treatments
  • outcomes that convert to one shared scale

If your studies differ at the core in what they compared or how they measured it, a pooled average is a precise number that means nothing.

When pooling is not sound, write a structured narrative synthesis instead. Group studies by treatment type or population. Report each group's direction and size of effect, and say plainly why you did not pool.

"We did not pool because the studies were too different clinically" is a fair, publishable finding. A forced meta-analysis is not.

If you do pool:

  • choose the effect measure (risk ratio, mean difference)

  • quantify how much studies disagree (I²)

  • set subgroup and sensitivity analyses in your protocol up front

  • check for small-study effects if you have enough trials

  • You produce: forest plots and pooled estimates, or a structured narrative synthesis with the reason stated.

  • Tools: R's metafor or meta packages (free), or RevMan (licensed, free for Cochrane review authors). Involve a statistician before you pool, not after.

  • Common rework cause: promising a meta-analysis in the protocol, then forcing unlike studies together to keep the promise. Protocols can state pooling as conditional — write it that way.

Step 9: Report with PRISMA 2020

Health journals expect the PRISMA 2020 statement. It has two parts: a 27-item checklist covering everything from search strings to funding, and the flow diagram that traces records from database hits to included studies.

The flow diagram is where the numbers you logged in steps 3–5 come home:

  • records per database
  • duplicates removed
  • records screened and dropped
  • full texts assessed, with reasons for each one dropped
  • studies included

The official template lives on the PRISMA flow diagram page. Our free PRISMA flow diagram generator draws the standard databases-and-registers version for you. It takes the counts and outputs a figure that follows the official 2020 layout.

While preparing figures, check your target journal's rules on AI-made images. The rules differ sharply by publisher, and flow charts are not treated like data images. We compared six publishers' actual policies in our guide to AI figures and journal rules.

Fill in the checklist as you write, not after. Most items are one honest sentence if the earlier steps were done properly.

  • You produce: a completed 27-item checklist (submitted with the manuscript) and the flow diagram.
  • Tools: the PRISMA checklist and templates are free.
  • Common rework cause: missing counts for the flow diagram because no one logged them at steps 3–5. There is no fix except redoing the bookkeeping.

Step 10: Write, update your registration, and submit

Structure the manuscript around PRISMA's headings:

  • background
  • methods, mirroring your protocol
  • results — flow, study traits, risk of bias, synthesis
  • discussion and limits

Where your review strayed from the registered protocol, say so and explain why. Reviewers will compare the two documents.

Pick the journal before polishing. Check that it publishes reviews like yours, and read its rules — many demand the PRISMA checklist and registration ID at submission. Then prepare the extras it asks for. If the journal wants a graphical abstract, check which tool you may use — several publishers bar general-purpose AI image tools from that one job, and we wrote up their actual rules.

Finally, close the loop on PROSPERO:

  • set the record's status to complete
  • add the citation once the paper is out

The register emails reminders, and a record left hanging reflects on the team.

  • You produce: a submitted manuscript with checklist, flow diagram, registration ID, and an updated PROSPERO record.
  • Tools: none new; your reference manager formats the (long) bibliography.
  • Common rework cause: a methods section that drifts from the registered protocol without saying so — the easiest thing for a peer reviewer to catch. A close second is a stale search. Cochrane asks you to rerun it if the first search is more than 12 months old at publication, so budget one more screening pass.

FAQ

What is the difference between a systematic review and a meta-analysis?

A systematic review is the whole process: question, search, screening, appraisal, synthesis. A meta-analysis is one optional step inside it that pools numeric results across studies. Every meta-analysis should sit inside a systematic review. Many good reviews contain no meta-analysis at all.

How long does a systematic review take?

The best number we have comes from 195 finished reviews on PROSPERO. The mean was 67.3 weeks from registration to a published paper (Borah et al., 2017). Scope drives the range — searches in that sample found anywhere from 27 to over 92,000 records.

Can one person do a systematic review?

Not to the accepted standard. Screening, extraction, and bias checks are all expected to be done by two people working apart. Solo work misses studies and bakes in one person's errors. The same 2017 study found reviews averaged 5 authors.

What is the difference between a systematic review and a literature review?

A literature review sums up sources the author chose, with no duty to be complete or repeatable. A systematic review registers its rules up front, tries to find every study that fits, and reports the method in enough detail to be rerun. A systematic-looking search alone does not qualify — PROSPERO excludes literature reviews that borrow the search but skip the other methods.

Do I have to register on PROSPERO?

No law requires it, but many journals expect it for health reviews, and it protects you against claims of outcome-switching. It must happen before data extraction begins — the register will not accept a review that has already started extracting. It is free and, since the 2025 rebuild, records publish at once when all authors confirm.