The strongest answer to how effective GLP-1 weight loss drugs are is not a vague “they work.” In randomized trial evidence, the average effect is large enough to matter: tirzepatide is around 16% weight loss at 12 to 18 months, semaglutide is roughly 11% to 15%, and liraglutide is closer to 4% to 5%.[1][2] That is not the profile of a marginal obesity intervention.
The committee problem starts one sentence later. Across 21 randomized controlled trials with 7,024 participants, GLP-1 receptor agonists produced weight loss in 78.5% of participants compared with 26.5% on placebo, with a pooled odds ratio of 11.37 and a 95% confidence interval of 8.10 to 15.98; the same analysis reported high heterogeneity, with I² at 82%.[1] Cochrane’s 2025 WHO-commissioned reviews also found clinically meaningful weight loss, but emphasized that nearly all pivotal studies were manufacturer-funded and that sponsors were substantially involved in trial design, conduct, and reporting.[2]

So the procurement-grade verdict is split. The efficacy signal is strong. The certainty around its exact size, durability, and transportability into routine care is more constrained than a single pooled estimate suggests.
The drugs should not be scored as interchangeable
A formulary discussion that treats “GLP-1s” as one uniform category loses clinically and financially important resolution. These agents share incretin biology, but the observed weight-loss magnitudes are not the same. Semaglutide and liraglutide are GLP-1 receptor agonists; tirzepatide is a dual GIP and GLP-1 receptor agonist. Mechanism does not settle the adoption question, but it helps explain why the efficacy rankings are not flat.
| Agent | Approximate weight-loss magnitude in the cited evidence | Committee-relevant interpretation |
|---|---|---|
| Tirzepatide | Around 16% at 12 to 18 months | Highest efficacy signal among the compared agents; strongest weight-loss case, but still subject to trial-quality and durability questions |
| Semaglutide | Roughly 11% to 15% | Substantial efficacy; also supported by cardiovascular outcomes evidence in SELECT, though population fit must be examined separately |
| Liraglutide | Around 4% to 5% | Meaningful but materially smaller effect; should not be used as a proxy for newer incretin efficacy |
The network meta-analysis ranking reinforces that hierarchy. Ahmad et al. reported a SUCRA ranking of 91.2% for tirzepatide and 85.4% for semaglutide, placing both well ahead of older comparators in the analyzed network.[1] SUCRA is not a procurement decision by itself; it is a probability ranking derived from the network. But it is useful because it prevents an imprecise category label from hiding a real separation in expected effect.

For a committee, that separation matters in more than one column. A larger expected weight-loss effect can change cost-effectiveness assumptions, eligibility criteria, monitoring thresholds, and the tolerance for prior authorization burden. It can also change the consequences of interruption. If a patient loses substantially more weight on one agent, the operational cost of shortages, discontinuation, or forced switching is not merely administrative.
Large effects, fragile precision
High heterogeneity is not a technical footnote here. The Ahmad analysis reported I² of 82% for the pooled weight-loss outcome, and the research base includes heterogeneity estimates reaching into the 82% to 98% range across relevant analyses.[1][2] That means the pooled average is summarizing trial results that varied substantially rather than estimating one clean, stable treatment effect.
This does not make the weight-loss finding disappear. It changes how much confidence should be attached to the exact number used in a budget model. A procurement model built around “16%” or “15%” as if it were a fixed property of the drug is less defensible than a model that stress-tests lower persistence, lower target-dose attainment, and less selected populations.
Attrition is part of the same problem. Cochrane highlighted dropout in the 22% to 26% range in some trials.[2] In obesity pharmacotherapy, discontinuation is not a side detail. It determines whether the treated population in the final efficacy estimate still resembles the population a health system must cover at initiation.
Sponsor involvement also belongs in the evidence-quality column, not in an insinuation column. Manufacturer funding does not automatically invalidate a trial result, especially when the effect is large and repeatedly observed. But near-universal manufacturer funding, with sponsor participation across design, conduct, and reporting, reduces the comfort one should have that the published program has already answered every implementation-relevant question a payer or health system would ask.[2]
SELECT changes the value conversation, not the whole appraisal
The SELECT cardiovascular outcomes trial is an important adjacent signal because it moves semaglutide beyond weight loss alone. In 17,604 participants with overweight or obesity and established cardiovascular disease but without diabetes, semaglutide reduced major adverse cardiovascular events by 20%, with a hazard ratio of 0.80 and a 95% confidence interval of 0.72 to 0.90.[3]
That finding can legitimately alter a value assessment for patients resembling the SELECT population. A drug that reduces weight and lowers cardiovascular events in a high-risk population is not being judged on cosmetic weight change. The harder question is how far that result travels across the broader obesity population a formulary policy may cover.
SELECT’s demographic composition matters for that judgment. The trial enrolled 27.7% women and 3.8% Black participants.[3] Those proportions do not negate the cardiovascular result, but they do limit confidence that the magnitude of benefit and harms has been characterized with equal precision across all groups who will request or receive treatment after coverage expands.
There is also a category distinction. SELECT is cardiovascular outcomes evidence for semaglutide in a defined high-risk population. It should not be used as a blanket substitute for long-term obesity-specific safety, durability, or access evidence across all GLP-1 and dual incretin agents.
Routine care will not reproduce the trial machinery
The randomized trials answer efficacy under structured conditions. Coverage decisions have to estimate effectiveness after the trial machinery is removed: dose titration is uneven, visits are delayed, adverse effects are managed in crowded clinics, shortages interrupt therapy, and patients face copayments or authorization renewals.
The real-world evidence review summarized a substantial efficacy-effectiveness gap, including 20% to 50% discontinuation at 12 months and only about 12% of patients reaching the target semaglutide dose in some routine-care data.[4] Those figures are not a better estimate of biological efficacy than the RCTs. They are a warning about using RCT weight-loss percentages as if they will automatically become population-level outcomes.
Discontinuation also affects the durability assumption. Trial withdrawal and follow-up data indicate weight regain after cessation, and the BMJ editorial discussion emphasizes that stopping therapy commonly reverses part of the achieved weight loss.[5] The exact trajectory may differ with tapering, behavioral support, alternative agents, or future maintenance strategies, but a coverage model that assumes a one-time course produces durable weight loss without continued treatment is not aligned with the evidence now available.
The subgroup question is more encouraging, though still bounded. Johns Hopkins researchers reported that GLP-1 weight-loss drugs appeared comparably effective across age, race, and starting-weight subgroups in the analyzed data.[6] That supports a broader efficacy expectation, but it does not erase the representation limits in major outcome trials or the need to monitor access and outcomes after implementation.
Safety follow-up is still thinner than adoption scale implies
Most committees will already expect gastrointestinal adverse effects and discontinuation to appear in the safety review. The more difficult issue is time. Long-term obesity-specific safety data beyond 2 to 3 years remain sparse, and much of the longer safety familiarity comes from diabetes populations, often at lower doses than those used for obesity treatment.[2]
That gap matters because broad obesity coverage can move treatment into a larger and more heterogeneous population than the pivotal trials studied. Rare adverse events, differential tolerability, pregnancy-related policy questions, perioperative management, and effects of repeated stopping and restarting are all more visible after scale-up than before it.
The NAION signal is a useful example of how to hold uncertainty without exaggeration. Observational studies have reported an association between semaglutide and nonarteritic anterior ischemic optic neuropathy, with hazard ratios around 2.2 to 2.8, but the absolute excess risk was reported as low, about 1.4 additional cases per 10,000 person-years.[7] That is not a decisive safety verdict. It is a reason for surveillance, labeling attention, and careful interpretation of observational confounding.
What belongs on the formulary scorecard
The evidence is strong enough to justify serious adoption discussions. It is not strong enough to let a committee collapse every question into a single “clinically meaningful weight loss” box. The decision record should keep separate lines for effect size, evidence quality, population fit, persistence assumptions, funding risk, and safety follow-up.
- Efficacy magnitude: tirzepatide and semaglutide should be scored apart from liraglutide, and tirzepatide should not be treated as merely another GLP-1 category entry.
- Certainty of estimate: pooled results should be interpreted in light of high heterogeneity and attrition, not copied directly into budget forecasts as fixed expected outcomes.
- Population fit: SELECT strengthens the semaglutide value case for a specific cardiovascular-risk population, while its sex and racial composition limit overgeneralization.
- Durability: discontinuation, target-dose attainment, and weight regain should be explicit assumptions rather than afterthoughts.
- Evidence independence: manufacturer funding and sponsor involvement should trigger post-adoption audit requirements, not automatic dismissal of the efficacy signal.
- Safety follow-up: limited long-term obesity-specific data and emerging observational signals warrant active monitoring as coverage expands.
On that scorecard, the weight-loss evidence for GLP-1 and dual incretin drugs is substantial and clinically meaningful. Tirzepatide and semaglutide clearly outperform liraglutide in the cited evidence, and semaglutide has an important cardiovascular outcomes result in SELECT. The weaker part of the file is not whether weight loss occurs; it is how confidently a health system can price, staff, monitor, and sustain the observed effect at broad scale.
References
- Ahmad et al. 2026 systematic review, Medicine, https://pmc.ncbi.nlm.nih.gov/articles/PMC12991648/
- GLP-1 drugs are effective for weight loss, but more independent studies are needed, Cochrane, https://www.cochrane.org/about-us/news/glp-1-drugs-effective-weight-loss-more-independent-studies-needed
- Semaglutide and Cardiovascular Outcomes in Obesity without Diabetes, New England Journal of Medicine, https://www.nejm.org/doi/full/10.1056/NEJMoa2307563
- Real-world evidence review, https://pmc.ncbi.nlm.nih.gov/articles/PMC12000858/
- BMJ editorial on weight regain, BMJ, https://www.bmj.com/content/392/bmj-2025-085304
- GLP-1 weight loss drugs comparably effective for patients across age, race and starting weight, Johns Hopkins Bloomberg School of Public Health, 2026, https://publichealth.jhu.edu/2026/glp-1-weight-loss-drugs-comparably-effective-for-patients-across-age-race-and-starting-weight
- Expanding benefits review, https://pmc.ncbi.nlm.nih.gov/articles/PMC12281309/