Protocol-search-selection-bias appraisal
Compare the protocol with the published review
Prospective registration makes plans visible before reviewers know the complete pattern of results. Record the identifier, registration date, eligibility rules, outcomes and intended analyses. Compare them with the finished review. Criteria sometimes change for defensible reasons, but the change and rationale should be reported. Registration after screening offers less protection against result-driven choices than a genuinely prospective plan.
No registration does not make a synthesis automatically useless. It does make selective changes harder to assess. Search for a separately published protocol, methods paper or dated supplementary file. Treat the absence as an uncertainty rather than inventing a planning history.
Test coverage and reproducibility
Ask whether the databases suit the disciplines and study types in the question. Supplementary routes might include trial registers, reference lists, citation searching and grey literature. A larger database count is not automatically superior, but a narrow disciplinary search can miss relevant terminology or unpublished outcomes. At least one complete reproducible strategy should normally be available, with adaptations described for other platforms.
The final search date—not merely the publication year—sets the evidence cut-off. Inspect language, date and publication-status restrictions. Check how conference papers and preprints were handled. Unexplained limits may produce an evidence set that differs systematically from the intended question.
Look beyond the flow diagram
A flow chart accounts for records, but it cannot show whether eligibility decisions were sound. Read who screened titles, abstracts and full texts, whether decisions were checked and how disagreements were resolved. Full-text exclusions should use specific reasons rather than “not relevant”. Criteria need enough detail that a close case could plausibly be treated consistently.
One study may generate several articles, abstracts or follow-up reports. If a review treats these as independent studies, one participant sample can be counted repeatedly. Look for linkage of companion reports and a rule for selecting the most complete data. Duplicate checking helps reduce mistakes, but it cannot replace explicit extraction rules.
Require risk-of-bias findings to affect the review
First establish whether the appraisal tool fits the included designs. Then inspect the domain-level reasons rather than accepting coloured summary icons. Most importantly, ask what reviewers did with the judgements. Sensitivity analyses, subgroup decisions, certainty ratings or cautious conclusions should reflect serious concerns. A review that labels evidence high risk and then makes an unqualified recommendation has not integrated its own appraisal.
Keep study-level bias separate from reporting bias across the evidence base. Funnel-plot symmetry does not prove that no study is missing, especially with few or heterogeneous results. Search strategies, registers, outcome discrepancies and grey literature may provide additional clues. The absence of evidence about missing results is not evidence of their absence.
Interrogate synthesis and interpretation
For meta-analysis, inspect the effect measure, model, heterogeneity, dependent outcomes and sensitivity choices. Pooling produces a precise-looking number even when populations, interventions or measurements should not be combined. A narrative synthesis also requires method: studies should be grouped transparently, differences explored and conflicting findings retained rather than selected away.
Compare the abstract and any plain-language conclusion with the detailed results. Qualifications often disappear in compressed summaries. Your citation should preserve the relevant population, setting, period, outcome and certainty. A broad review title does not make every statement about that topic supported.
Write a transparent overall judgement
Avoid reducing the appraisal to one unexplained score. Identify critical strengths, critical weaknesses and what they mean for the claim you intend to cite. For example: the search is broad and reproducible, but the lack of a prospective protocol and failure to incorporate bias judgements weaken a strong causal conclusion. Such a statement remains understandable even when readers use a different formal appraisal tool.
Revisit the review when a living update, correction or newer search appears. Preserve the version you assessed and do not transfer your judgement automatically to an update with changed eligibility or analyses. Quality is a property of the particular review process and report, not a permanent badge attached to its title.