Finding something worth knowing…

Science

Pooling many studies into one sharper effect size

Meta-analysis combines quantitative results from independent studies that ask the same question, extracting effect sizes and variances to form a pooled estimate. Extra statistical power can settle conflicts between single papers—and help guide grants, clinical guidelines, and the next wave of experiments.

The method synthesizes data across studies addressing a shared research question by computing a combined effect size. Pulling effect sizes and variance measures together raises statistical power and can clarify discrepancies among individual findings. Meta-analyses often sit inside systematic reviews; they support grant proposals, treatment guidelines, health policy, and research agendas, and they are a core tool of metascience. Gene V. Glass coined the term in 1976, calling it “the analysis of analyses.” Karl Pearson’s 1904 British Medical Journal paper aggregating typhoid inoculation studies is widely seen as an early clinical meta-analytic approach. Smith and Glass published a landmark 1978 psychotherapy effectiveness synthesis; Hans Eysenck dismissed meta-analysis as “mega-silliness” and later “statistical alchemy,” yet publication counts rose from 334 in 1991 to 9,135 by 2014.

Workflow starts with disciplined searching—keywords, Boolean limits, databases such as PubMed, Embase, or PsycINFO, and snowballing from reference lists—then PRISMA flow accounting of inclusions and exclusions. Standardized forms capture effect sizes (often Pearson’s r for correlational work), moderators such as mean participant age, and study-quality scores; more than eighty tools assess observational risk of bias. Grey literature (abstracts, dissertations, preprints) can cut publication bias but may be weaker methodologically; conference data later disagree with published figures in almost twenty percent of cases studied. Evidence may be aggregate summaries (odds ratios, relative risks) or individual participant data, with one-stage versus two-stage modeling choices that usually agree but can diverge.

Fixed-effect models weight studies—commonly by inverse variance—so large samples dominate and smaller ones nearly vanish, under the strong assumption that all studies share one true effect. Random-effects models add a heterogeneity variance component that reweights toward equality as between-study scatter grows. Critics note that common random-effects intervals can undercover error and that redistribution of weight from large to small studies need not track genuine reliability. Prediction intervals around the pooled effect are sometimes offered to show the range of effects one might see in a new setting, though exchangeability of populations and treatments is often unrealistic in practice.

Source: Meta-analysis

Related

More in Science · All topics