A systematic literature review is a research study that uses an explicit, pre-specified and reproducible method to search for, select, critically appraise and synthesise all the evidence answering one focused question. Knowing how to do a systematic literature review step by step means following that fixed order: frame the question, register a protocol, search several databases, screen records twice, appraise quality, extract data, then synthesise and report. A careful single-topic review usually takes four to twelve months of part-time work, so plan the time before you start.
Table of Contents
- 1What You Need
- 2Step-by-Step
- 3Step 1: Define the Review Question
- 4Step 2: Write and Register a Protocol
- 5Step 3: Identify Relevant Databases and Search Terms
- 6Step 4: Run and Record the Search
- 7Step 5: Screen Titles and Abstracts
- 8Step 6: Retrieve and Screen Full Texts
- 9Step 7: Extract Data and Assess Study Quality
- 10Step 8: Synthesize and Present the Evidence
- 11Common Mistakes
- 12Frequently Asked Questions
- 13Can ChatGPT be used for systematic literature reviews?
- 14How many steps are in a systematic review?
- 15Do I need a second reviewer for a systematic review?
- 16How long does a systematic review take?
- 17What is the difference between a systematic review and a meta-analysis?
- 18Conclusion
What You Need

You need six things in place before the first database search, and all six take longer to prepare than most students expect.
- An answerable review question. One focused question, not a topic. “Social media and adolescent mental health” is a topic; a PICO question naming population, exposure, comparator and outcome is a question.
- A written protocol. One to several pages that fix your eligibility criteria, outcomes, risk-of-bias tool and planned synthesis before you see results. Search and screening in the same week as a finished draft is where most reviews go wrong.
- A search strategy template. A table with one row per database: the exact string, the date you ran it, and the limits you applied.
- Eligibility criteria. Explicit inclusion and exclusion rules written so another person could apply them to a citation and reach the same decision you did.
- A reference manager. Zotero or EndNote, so deduplication produces one record per study instead of five.
- Screening software. Something that logs decisions, supports two reviewers and exports a flow diagram. Rayyan, AbstrackR and the Systematic Review Applications (SRA) tool have free options; Covidence and DistillerSR are subscription products.
Tell your supervisor, and a librarian if you have access to one, what topic you are considering before you commit to it. Searching for systematic reviews already published or in progress on your topic takes an afternoon and can save an entire wasted year.
Step-by-Step

The workflow below is the master sequence. Shorter versions you see elsewhere — five steps, seven steps — are compressions of the same eight stages, and different counting conventions exist because guides group the stages differently.
Step 1: Define the Review Question
Turn your topic into a question that names four elements, most often with the PICO framework: Population, Intervention, Comparator and Outcome. PICOC adds Context and SPIDER covers experience, setting and research type, which suits qualitative and social science topics. Once you have the question, write one sentence describing what a good answer would look like, and check whether a review already answers it.
Example: in adults with type 2 diabetes (P), does structured self-management education (I), compared with standard clinic advice (C), improve HbA1c at six months (O)?
Step 2: Write and Register a Protocol
A protocol is a dated plan you commit to before results can influence you. It should include the review question, eligibility criteria, information sources, search strategy, screening procedure, risk-of-bias tool, data extraction fields and planned synthesis. Register health-related reviews on PROSPERO and reviews outside health on OSF, then note the registration number in your write-up.
Registration also forces you to declare any changes you make later, which is exactly what a reviewer or examiner wants to see.
Step 3: Identify Relevant Databases and Search Terms
Pick databases that match your discipline and cover your topic’s grey literature. Searching three to five is normal; searching only one is a limitation a reviewer will flag.
| Database | Strongest for |
|---|---|
| PubMed, CINAHL, Embase | Health, nursing and allied health |
| Scopus, Web of Science | Multidisciplinary coverage with citation tracking |
| PsycINFO | Psychology and behavioural research |
| ERIC | Education research, including dissertations |
| IEEE Xplore, ACM Digital Library | Engineering and computing |
| Cochrane CENTRAL | Trials and systematic reviews |
Build each search string from three blocks joined with AND: one concept block for your population, one for your intervention or main factor, one for your outcome. Inside a block you join synonyms with OR, and each database has its own subject headings you should use alongside the free-text words.
A typical PubMed string reads (type 2 diabet*[tiab] OR T2DM[tiab] OR diabetes mellitus[MeSH]) AND (self-management education[tiab] OR structured education[tiab]) AND (HbA1c[tiab] OR glycated hemoglobin[MeSH]). The asterisk truncates a word so it catches diabetes and diabetic, and the MeSH tags pull in the controlled vocabulary. Have a librarian check the string before you commit to it; a poor string in the first week quietly poisons the whole review.
Step 4: Run and Record the Search
Run every string, record the database, the exact search, the date and any filters, then export all records into your reference manager. Deduplicate, and keep a note of how many duplicates you removed, because those counts feed the flow diagram later.
Add two things a single database cannot give you: citation searching forward from the key papers you already have, and citation searching backward through their reference lists. Pull in theses and dissertations, conference abstracts, trial registry records and relevant institutional reports. Researchers on research forums repeatedly name grey literature as the thing they most often miss, and they are usually right.
Step 5: Screen Titles and Abstracts
Two people screen independently and never see each other’s decisions, which is what stops one person’s reading habits from quietly shaping the evidence base. Pilot the criteria on the first 50 to 100 records, talk through disagreements, and tighten the wording before the full screen starts. Expect to exclude the large majority here; a title and abstract screen that keeps a third of your hits almost always means the criteria are too loose.
Screeners disagree most often on borderline records. The fix is a written tie-breaker rule, decided in advance, such as including mixed adult samples when at least 70 percent meet the population criterion.
Step 6: Retrieve and Screen Full Texts
Retrieve every remaining report, then read each against your criteria and record one primary reason for exclusion per report. “Not relevant” is not a usable reason; “wrong population: children rather than adults” is. Contact authors for missing full texts rather than dropping those records, because unavailable evidence biases the synthesis toward what got published.
Keep a folder structure with included and excluded subfolders. Most databases also let you file search results directly into folders while you search, which saves an export and import cycle.
Step 7: Extract Data and Assess Study Quality
Build a structured extraction form before you start pulling data: study characteristics, setting, sample size, design, intervention details, outcome measures with results, funding source and stated limitations. Two reviewers extract independently, at least for the outcomes that drive the conclusion.
Judge quality with a tool matched to the study design. Appraisal is a distinct step from extraction, and one tool does not fit every design.
| Study design | Typical risk-of-bias tool |
|---|---|
| Randomised trial | RoB 2 |
| Non-randomised intervention | ROBINS-I |
| Diagnostic accuracy | QUADAS-2 |
| Qualitative research | CASP |
| Systematic reviews | AMSTAR 2, ROBIS |
Rate each study, not just the body of studies overall. Then grade the certainty of the evidence across studies with GRADE, which moves an outcome down when studies share a bias or an outcome when the evidence is inconsistent or indirect.
Step 8: Synthesize and Present the Evidence
Synthesis has two forms. Narrative synthesis groups studies by theme or effect, compares where they agree and disagree, and explains disagreements rather than hiding them. Meta-analysis pools results statistically and is the right choice only when studies are similar enough in design, outcome and measure to be meaningfully combined.
Report with PRISMA 2020. It gives you a 27-item checklist and a flow diagram whose numbers come directly from your logs: records identified, duplicates removed, records screened, reports sought and retrieved, reports excluded with reasons, and studies finally included. Use PRISMA-ScR for a scoping review, MOOSE for a meta-analysis of observational studies, and the Cochrane Handbook as the methodological reference for the whole process.
If you did the review alone, say so plainly in your limitations and describe your checks, such as re-screening a sample of excluded records. If you used AI tools, state exactly where and how, because journals increasingly require that disclosure.
Common Mistakes
Most reviews that fail review are not badly written. They fail on method, and the failures repeat.
- Starting the search before the protocol. Criteria that shift after you see results are the main source of selection bias. Fix: write eligibility criteria and outcomes first, register, then search.
- Searching one database. Fix: search at least three appropriate ones, add citation chasing, and report which ones you used and why.
- A search string that only covers your own wording. If your first search returns forty results, the string is too narrow. Fix: use subject headings, synonyms, spelling variants and truncation, and check the string with a librarian.
- Criteria that cannot be operationalised. “Relevant studies” and “good quality studies” are judgements two reviewers will make differently. Fix: define thresholds, dates and designs in numbers.
- A single screener. Single screening introduces error in both directions and weakens the review. Fix: pair with a second reviewer, and where that is impossible, document the limitation and re-screen a random sample.
- Not deduplicating before screening. You will screen the same trial three times and misreport your counts. Fix: deduplicate in the reference manager at export, and log the number removed.
- Using AI to make the decisions. Language models help draft strings, organise citations and format code. They should not make final screening, eligibility or risk-of-bias calls, and any use needs disclosure to the journal and registry. Where automation ranks records to reduce screening load, a human still verifies a sample and the final decisions.
- Writing the conclusion before appraising quality. Fix: complete risk-of-bias and GRADE ratings before drafting any summary statement.
Three habits pay off more than any other. Keep a dated search log, record an exclusion reason for every full text, and never change a criterion without writing down that you changed it and why.
Frequently Asked Questions
Can ChatGPT be used for systematic literature reviews?
Use it for drafting search strings, tidying reference lists, formatting analysis code and spotting missing synonyms. Do not use it to make final screening, eligibility or risk-of-bias decisions, because these need human judgement against your criteria. Most journals and registries now expect any AI use to be disclosed, and automation used to prioritise records still needs human verification.
How many steps are in a systematic review?
It depends on who is counting. This guide uses eight stages: question, protocol, database choice, running the search, title and abstract screening, full-text screening, data extraction with quality appraisal, and synthesis with reporting. Guides that say five steps fold appraisal and extraction together, and guides that say four describe the same work in conceptual phases. The order never changes.
Do I need a second reviewer for a systematic review?
Formally, yes. Dual independent screening reduces errors and lets you report agreement between reviewers. When you work alone, screen everything yourself, re-screen a random 10 percent sample a few weeks later, and state the single-reviewer process and its limitation in your methods section.
How long does a systematic review take?
Expect four to twelve months of part-time work for a focused question, and longer than a single academic term in most cases. Planning and the protocol take weeks, searching takes days, screening takes weeks to months depending on volume, and writing takes several weeks. Reviews that start in the final months of a degree rarely finish.
What is the difference between a systematic review and a meta-analysis?
A systematic review is the whole structured process of finding, selecting, appraising and synthesising evidence. A meta-analysis is one optional part of it: the statistical pooling of quantitative results across studies. A systematic review can be entirely narrative when the studies are too varied to pool, and pooling numbers alone is not a systematic review.
Conclusion
Start with the question, not the database. Write it in PICO form, check no equivalent review already exists, and put your eligibility criteria and planned analysis in a protocol you register before you search. Everything after that — the search log, the exclusion reasons, the PRISMA counts — is bookkeeping, and it is what makes your conclusion defensible. A systematic review is credible because someone else could repeat it and reach the same answer, so document each stage as you go rather than reconstructing it at the end.


