PeptidePedia The community reference

Trial data comparison (revision 54)

Old revision·06:34, 19 Nov 2025·ParentCatPansy

This is an old revision of this page, as it stood at 06:34, 19 Nov 2025, saved by ParentCatPansy with the summary fix category — parent was wrong. It may differ substantially from the current revision, and any error it contains may since have been corrected.
This article needs additional citations for verification. (May 2026) Discussion: Cross-trial tabulation and the estimand problem.
Trial data comparisonInteractive reference tool
STEP 1 · sema 2.4−14.9%SURMOUNT-1 · tirz 15−20.9%SURPASS-2 · tirz 15−11.2%TRIUMPH · reta 12−24.2%SCALE · lira 3.0−8.0%mean weight change at primary endpoint
Effect sizes from separate trials. Placing them side by side is not a head-to-head comparison.
RowsMajor randomised trials of incretin agonists
ColumnsPopulation, n, duration, comparator, endpoint, result
Caveats
Head-to-head?Only where the comparator column says so
EstimandStated per row where known
PopulationsDiffer substantially between trials
Reference tool infobox · conventions

The trial data comparison tabulates the principal published results of the major randomised trials of GLP-1 receptor agonists and related compounds, so that a reader arriving at one figure can see the others alongside it. Each row is attributed to its trial, arm, duration and comparator.

The table is easy to misread in one specific way, and the article exists partly to prevent it. Two figures from two trials are not a comparison of two drugs. Trial populations differ in baseline weight, glycaemic status, diabetes duration, background therapy and geography; durations differ; and the statistical estimand used to summarise a weight-change endpoint differs between publications and sometimes within a single one.[1] A larger number in this table means a larger result in that trial, and nothing more.

Where a genuine head-to-head randomised comparison exists — SURPASS-2 compared tirzepatide with semaglutide 1 mg directly — the comparator column says so, and only those rows support a statement that one agent outperformed another.[2]

The comparison

[edit]

Filter by compound, trial programme or endpoint type. The bar is drawn to the magnitude of the primary result and is scaled within its endpoint type only.

Why the estimand matters

[edit]

A weight-change endpoint can be summarised in at least two defensible ways, and the ICH E9(R1) addendum names them.[1]

The treatment-policy estimand asks what happened to everyone randomised, including those who stopped the drug and those who started another one. It answers "what does prescribing this achieve", and it is the more conservative figure.

The trial-product estimand asks what happened while participants were taking the drug as intended. It answers "what does the molecule do", and it is the larger figure.

The gap between them is not small. In the STEP programme the two estimands for the same trial differed by roughly one to two percentage points of body weight, which is a difference of the same order as the gap between some of the drugs being compared.[3] A table that mixes estimands between rows therefore manufactures apparent differences between compounds that are artefacts of the analysis choice.

The same trial under two estimands
EstimandQuestion answeredEffect on the figure
Treatment policyWhat prescribing achievesSmaller; includes discontinuations
Trial productWhat the molecule does on treatmentLarger; censors discontinuation
Not statedUnusable for comparison

The SURMOUNT programme's publications state their estimand explicitly, which is why the tirzepatide rows in the table above can be compared with each other with more confidence than the cross-programme rows can.[4]

What differs between these trials

[edit]

Five differences are large enough to dominate any cross-trial reading.

Glycaemic status. Trials in people with type 2 diabetes consistently report smaller weight reductions than trials in people without it, at the same dose of the same drug. SURPASS and SURMOUNT are not comparable on weight for this reason alone.

Baseline weight. A percentage reduction from a higher baseline is a larger absolute loss. Reporting one and not the other changes the apparent ordering.

Duration. Weight-loss curves in this class had not fully plateaued at 68 weeks in several trials, so a 40-week and a 72-week figure are not measuring the same thing.

Background therapy. Trials differ in whether metformin, insulin or a sulfonylurea was permitted, which affects both glycaemic and weight endpoints and the hypoglycaemia rate.

Comparator. Placebo-controlled and active-controlled trials answer different questions. A placebo-adjusted difference and an absolute change are frequently confused; see Placebo-adjusted effect.

STEP 1 · sema 2.4−14.9%SURMOUNT-1 · tirz 15−20.9%SURPASS-2 · tirz 15−11.2%TRIUMPH · reta 12−24.2%SCALE · lira 3.0−8.0%mean weight change at primary endpoint
Primary weight endpoints from five separate trials. The visual comparison is the one this article warns against making.

Outcome trials are a different kind of evidence

[edit]

Three of the rows report hazard ratios for clinical events rather than changes in a measurement. These are the strongest evidence in the table and the least comparable to the rest of it: an event-driven trial reports a relative risk over years in a population selected for risk, and a hazard ratio cannot be placed alongside a percentage weight change in any meaningful ordering.

For those rows the useful companion figure is the number needed to treat, which converts a relative effect into an absolute one over a stated horizon — and which is meaningless without that horizon.[5]

See also

References

  1. ^ a b International Council for Harmonisation, E9(R1): Addendum on Estimands and Sensitivity Analysis in Clinical Trials (2019).
  2. ^ Frías JP, Davies MJ, Rosenstock J, et al. "Tirzepatide versus Semaglutide Once Weekly in Patients with Type 2 Diabetes." New England Journal of Medicine 385(6):503–515 (2021). DOI:10.1056/NEJMoa2107519. PMID 34170647.
  3. ^ Wilding JPH, Batterham RL, Calanna S, et al. "Once-Weekly Semaglutide in Adults with Overweight or Obesity." New England Journal of Medicine 384(11):989–1002 (2021). DOI:10.1056/NEJMoa2032183. PMID 33567185.
  4. ^ Jastreboff AM, Aronne LJ, Ahmad NN, et al. "Tirzepatide Once Weekly for the Treatment of Obesity." New England Journal of Medicine 387(3):205–216 (2022). DOI:10.1056/NEJMoa2206038. PMID 35658024.
  5. ^ Lincoff AM, Brown-Frandsen K, Colhoun HM, et al. "Semaglutide and Cardiovascular Outcomes in Obesity without Diabetes." New England Journal of Medicine 389(24):2221–2232 (2023). DOI:10.1056/NEJMoa2307563. PMID 37952131.