← Drug Development Suite Biostatistics · Vitalis Therapeutics Inc.

Statistical Analysis Plan Summary

Download Word

Vitalis Therapeutics Inc. — The Statistical Analysis Plan for the VitaFlow (VTX-401) pivotal program, v2.0, approved 2028-06-14: analysis sets, the estimand under ICH E9(R1), the testing hierarchy and why position four determined what could be claimed, missing-data handling, and the futility-only interim.

v2.0
SAP version
6
Endpoints in hierarchy
4th
Differentiation endpoint
35 days
Approval to lock
Contents
  1. Why the SAP Is a Separate Document
  2. Analysis Sets
  3. The Estimand
  4. The Testing Hierarchy
  5. Missing Data
  6. The Interim Analysis
  7. The Limits of Pre-Specification

1. Why the SAP Is a Separate Document

The protocol says what will be measured. The Statistical Analysis Plan says exactly how it will be analyzed — which populations, which models, which comparisons in which order, and what happens to data that is missing.

Versionv2.0
Approved2028-06-14
Database lock2028-07-19
Governing guidanceICH E9 and ICH E9(R1) addendum on estimands
The rule that makes the whole document meaningful: Finalized, reviewed and signed by the lead biostatistician BEFORE database lock and before any unblinding.

Thirty-five days separate SAP approval from database lock here, and the order is the entire point. An analysis plan written after anyone has seen the results is not an analysis plan — it is a description of the analysis that produced the most favorable answer.

That is not a hypothetical concern. Given enough freedom over populations, models, covariates and comparisons, almost any dataset can be made to yield a significant result somewhere. Fixing the analysis before the data exists is what makes the p-value mean anything.

For a program manager the practical consequence is a hard sequencing dependency that shows up on no clinical schedule: SAP approval gates database lock, which gates topline, which gates the Gate 5 submission. A late SAP does not delay statistics. It delays the filing.

2. Analysis Sets

Analysis setDefinitionUse
Randomized SetAll participants randomizedDenominator for disposition and demographics.
Full Analysis Set (FAS)All randomized who received at least one doseThe primary efficacy population. Analyzed as randomized, not as treated — the intention-to-treat principle.
Per-Protocol SetFAS excluding major protocol deviationsSupportive only. Never the primary analysis, because excluding participants on the basis of post-randomization events breaks randomization.
Safety SetAll who received at least one dose, analyzed as treatedThe one population analyzed by what participants actually received rather than what they were assigned.
One principle above is worth understanding properly rather than accepting. The per-protocol set excludes participants with major protocol deviations, and it is intuitively appealing — why judge a drug on people who did not take it correctly?

Because deviations are not random. Participants who discontinue, miss doses or violate the protocol differ systematically from those who do not, and often because of how the drug affected them. Excluding them removes the randomization that made the comparison valid in the first place. Per-protocol analysis is supportive, never primary.

The safety set is the one population analyzed as treated rather than as randomized. If someone assigned to placebo received drug in error, their adverse events belong to the drug. The efficacy question and the safety question require different answers to “which group is this person in?”

3. The Estimand

Under ICH E9(R1) a sponsor must state precisely what question the analysis answers. Five attributes define it.

AttributeThis program
TreatmentVitaFlow at target dose versus placebo, both plus lifestyle intervention
PopulationAdults meeting the inclusion criteria, per the FAS
VariablePercent change in body weight from baseline to week 68
Intercurrent eventsTreatment-policy strategy — the value is used regardless of whether the participant discontinued study drug or started rescue therapy
Population-level summaryDifference in means between arms
The fourth attribute is where the real decision sits, and it is a choice with consequences.

An intercurrent event is something that happens after randomization and changes the meaning of the measurement — discontinuing the drug, starting a different weight-loss therapy, undergoing bariatric surgery.

A treatment-policy strategy uses the week-68 value regardless: if a participant stopped taking the drug at week 12 and regained the weight, that regained weight counts. A hypothetical strategy would instead estimate what their weight would have been had they continued — a more flattering question, and one that describes a world that did not happen.

This program uses treatment policy, which is the harder standard and the one regulators increasingly expect. It answers “what happens if you prescribe this drug”, not “what happens to the people who tolerate it”.

And that choice is where the tolerability problem enters the efficacy result. Under a treatment-policy estimand, every participant who discontinued because of gastrointestinal adverse events remains in the analysis at their week-68 weight — which, having stopped the drug, has largely rebounded.

Those discontinuations are concentrated in the treated arm. So the GI tolerability problem does not merely sit alongside the efficacy result; it directly reduces it. The two endpoints the program cared about most are statistically coupled, and a program that treated them as independent would have been surprised twice by the same phenomenon.

4. The Testing Hierarchy

When multiple endpoints are tested, the chance of a false positive somewhere rises with each test. The hierarchy controls that — endpoints are tested in a fixed order, and testing stops at the first failure. Everything below an unmet endpoint becomes descriptive, however good its own result.

OrderEndpointTypeNote
1Mean percent change in body weight, week 68co-primarySuperiority to placebo
2Proportion achieving ≥5% weight reductionco-primarySuperiority to placebo
3Proportion achieving ≥10% weight reductionkey secondary
4GI-attributed treatment discontinuationkey secondaryThe differentiation endpoint — and it sits fourth.
5Proportion achieving ≥15% weight reductionkey secondary
6Change in waist circumferencekey secondary
Position four is the single most consequential number in this program's commercial history.

GI-attributed discontinuation is the differentiation endpoint. It is the reason the asset was funded as a differentiated entrant rather than a late follower. And it sits fourth, behind three weight endpoints, because the agency's expectation at the End-of-Phase-2 meeting was that weight endpoints lead in a weight-management indication.

A hierarchy position determines what can be claimed. An endpoint high in the hierarchy with a significant result supports a labeling claim. An endpoint at position four, in a hierarchy that holds, supports a descriptive statement in Section 6 — which is exactly what the approved label carried.

The Protocol Summary records that the argument to promote GI discontinuation was had at protocol finalization and lost. This document is where that decision became irreversible: the hierarchy is fixed in the SAP before unblinding, and it cannot be reordered once anyone has seen a result.

Why it could not simply be moved up. Reordering to protect a commercial claim would have been visible to the agency as exactly that. More practically, promoting GI discontinuation above a weight endpoint would have meant that a failure on tolerability stopped the hierarchy and rendered the weight endpoints below it descriptive — risking the entire efficacy claim to protect a differentiation claim.

The program chose to protect the indication. ⚠ That choice has to be defensible now, on what is known before the blind is broken — because the one thing certain about the result is that it is not yet available to justify the decision with.

5. Missing Data

In a 68-week trial, some participants will not have a week-68 measurement. How those absences are handled changes the answer.

AnalysisMethodWhat it assumes
Primary analysisMixed model for repeated measures (MMRM) on observed dataAssumes data are missing at random, conditional on observed values.
Sensitivity 1Multiple imputation under a washout assumptionAssumes participants who discontinue return toward their baseline weight — a deliberately conservative assumption for a weight-loss product.
Sensitivity 2Retrieved-dropout imputationUses observed data from participants who discontinued drug but stayed in follow-up.
Sensitivity 3Tipping-point analysisProgressively penalizes the treated arm until the result loses significance, and reports how far that is. Answers the reviewer's question before it is asked.
The reason this matters more here than in most trials: data in a weight-loss study is not missing at random.

Participants who discontinue are disproportionately those who experienced adverse effects, which means disproportionately those on active drug. Analyze only the people who finished, and the result flatters the drug — and flatters it most in the arm where discontinuation was most common.

The tipping-point analysis at the bottom of the table exists to answer this directly. It progressively penalizes the treated arm's imputed values until the primary result loses significance, and reports how far that had to go. A reviewer asking “how sensitive is this to your missing-data assumptions?” gets a number rather than an argument.

6. The Interim Analysis

Number of interim analyses1
TimingWhen 50% of participants have completed week 68
PurposeFutility only
Efficacy stopping boundaryNone
Conducted byConducted by the independent statistician and reviewed only by the DMC
Alpha spendingNo alpha spent — a futility-only analysis does not inflate type I error
Futility only, and no efficacy boundary — a deliberate asymmetry.

A futility analysis can stop a trial that is clearly not going to work, which spares participants exposure to a drug that will not help them and spares the sponsor the remaining cost. An efficacy boundary would allow stopping early on good news.

This program has the first and not the second, for the reason set out in the DMC Charter: stopping early on a chronic therapy produces a smaller long-term safety database than the label requires, and a program that can stop early for good news has an incentive to look for it.

Because the analysis is futility-only, no alpha is spent. Stopping rules that could declare success consume type I error and require the final analysis to clear a higher bar; a rule that can only stop for failure does not.

7. The Limits of Pre-Specification

The pattern across all four is the same, and it is the pattern of this entire suite. The decisions that determine what a program can claim are taken years before the data exists, by people choosing between options that all look reasonable at the time.

An SAP is where several of those decisions become irreversible — and writing one is therefore less an act of statistics than an act of governance. It is the last moment at which the program can still choose what question it is asking, and the first at which it can no longer change the answer.

What the program manager owns here

Not the statistics. The sequencing: SAP approval before database lock, database lock before unblinding, and enough calendar between them that a late statistical review does not force a choice between a rushed plan and a slipped filing. Thirty-five days, on this program, deliberately protected.