Methodology
How every rating in the document set was arrived at, and the decisions the Institute has taken about how to present them.
§1The certainty framework
A certainty rating describes confidence in an effect estimate for a stated population, comparator and outcome. It is not a recommendation, not an endorsement, and does not transfer between populations, comparators or outcomes.
Randomised evidence starts at high certainty. Each serious concern in one of five domains reduces the rating by one level and each very serious concern by two. Observational evidence starts at low and may, rarely, be upgraded where the effect is large, a dose–response gradient is present, or a plausible confounder would have reduced the observed effect.
Table 1. The four levels.
| Level | Meaning |
|---|---|
| High | The Institute is confident that the true effect lies close to the estimate. Further research is very unlikely to change confidence in the estimate. |
| Moderate | The true effect is likely to be close to the estimate, but there is a possibility that it is substantially different. Further research is likely to have an important impact on confidence in the estimate. |
| Low | Confidence in the effect estimate is limited. The true effect may be substantially different from the estimate. |
| Very low | The Institute has very little confidence in the effect estimate. The true effect is likely to be substantially different from the estimate. In this series a very low rating most often reflects an absence of controlled human evidence rather than conflicting evidence. |
Table 2. The five downgrading domains.
| Domain | What it assesses |
|---|---|
| Risk of bias | Limitations in the design and conduct of the contributing studies. |
| Inconsistency | Unexplained heterogeneity of results across contributing studies. |
| Indirectness | Differences between the population, intervention, comparator or outcome studied and those of the assessment question. |
| Imprecision | Width of the confidence interval relative to the decision threshold, and the number of events. |
| Publication bias | Risk that results were selectively reported or that unpublished studies exist. |
§2Standing decisions
Table 3. Decisions the Institute has taken about how it reports, each settled through consultation and each applied across the whole document set.
| Decision | Reasoning | Settled at |
|---|---|---|
| A relative measure is never reported without an absolute one | A hazard ratio without a baseline risk cannot be acted on and systematically overstates benefit in lower-risk populations. | Synthesis consultation |
| A clinical rating never inherits from a surrogate rating | A rating attaches to the outcome measured. Any transfer to a different outcome is an inference and is marked as one. | MASH synthesis consultation |
| An absence of events is reported with the detectable effect size attached | An absence of a rare event in a trial of conventional size is uninformative, not reassuring. | Safety review consultation |
| No effect estimate is reproduced from a sponsor announcement | A figure the Institute has not seen in a report is not a figure the Institute can assess. The document records that results exist and have not been seen. | Orforglipron consultation |
| No identifier is ever constructed | A constructed identifier makes a citation unresolvable in a way that is harder to detect than an absent one. | Citation policy |
| No ranking probability is published for a network whose transitivity is unsatisfied | Publishing a ranking under that finding would present as a conclusion something the document has just shown to be unsupported. | Transitivity consultation |
| Discontinuation for adverse events is a summary-of-findings row | A trial reporting a large effect with high discontinuation reports a different result from one reporting the same effect with high persistence. | Synthesis consultation |
| Regional estimates are not pooled without an interaction test | Percentage weight changes obtained at different baseline body-mass index distributions are not on a common scale. | Adolescent synthesis consultation |
§3Sponsorship of the evidence base
Every pivotal trial in the incretin class was conducted by the manufacturer of the compound under assessment. The Institute does not downgrade certainty for sponsorship as a separate domain, because doing so would double-count risk of bias as conventionally assessed. It records sponsorship on the front matter of every trial abstract, and treats the uniformity of sponsorship across the class as a limitation of the evidence base rather than of any individual trial.
Clinical study reports are requested as a matter of routine for every contributing trial. The outcome of each request, including refusals, is recorded in the amendment log of the synthesis that made it. A review that does not record refusals presents the sponsor's disclosure decision as the Institute's own.
The methodological basis for this position is set out in the review of sponsor-generated evidence.
§4Institute-constructed programme records
The trial register contains two kinds of record. 152 are published trials, carrying figures reproduced from their reports with fields left empty where the Institute does not hold them. 548 are Institute-constructed programme records: instalments, sub-studies and regional extensions of real programmes, constructed to carry a methodological point the parent report does not address.
A constructed record is labelled above its abstract, on every one of its five pages, and in every reference list in which it appears. The reference-list label was added after a submission observing that a notice on a page protects the reader of the page and not the reader of a citation to it; the label now travels with the reference. The reasoning is published on the programme-record disposition page.
§5Determinism and reproducibility
The document set is generated from its source data by a deterministic pipeline. No random number is used anywhere in it: every count, date, identifier, effect estimate and figure coordinate is derived from a string seed, so that two builds of the same sources produce byte-identical output. A figure that changes between builds would be an unrecorded amendment, and the Institute treats unrecorded amendment as the failure mode this constraint exists to prevent.