Vitalis Therapeutics Inc. — The Statistical Analysis Plan for the VitaFlow (VTX-401) pivotal program, v2.0, approved 2028-06-14: analysis sets, the estimand under ICH E9(R1), the testing hierarchy and why position four determined what could be claimed, missing-data handling, and the futility-only interim.
1. Why the SAP Is a Separate Document
The protocol says what will be measured. The Statistical Analysis Plan says exactly how it will be analyzed — which populations, which models, which comparisons in which order, and what happens to data that is missing.
| Version | v2.0 |
| Approved | 2028-06-14 |
| Database lock | 2028-07-19 |
| Governing guidance | ICH E9 and ICH E9(R1) addendum on estimands |
Thirty-five days separate SAP approval from database lock here, and the order is the entire point. An analysis plan written after anyone has seen the results is not an analysis plan — it is a description of the analysis that produced the most favorable answer.
That is not a hypothetical concern. Given enough freedom over populations, models, covariates and comparisons, almost any dataset can be made to yield a significant result somewhere. Fixing the analysis before the data exists is what makes the p-value mean anything.
For a program manager the practical consequence is a hard sequencing dependency that shows up on no clinical schedule: SAP approval gates database lock, which gates topline, which gates the Gate 5 submission. A late SAP does not delay statistics. It delays the filing.
2. Analysis Sets
| Analysis set | Definition | Use |
|---|---|---|
| Randomized Set | All participants randomized | Denominator for disposition and demographics. |
| Full Analysis Set (FAS) | All randomized who received at least one dose | The primary efficacy population. Analyzed as randomized, not as treated — the intention-to-treat principle. |
| Per-Protocol Set | FAS excluding major protocol deviations | Supportive only. Never the primary analysis, because excluding participants on the basis of post-randomization events breaks randomization. |
| Safety Set | All who received at least one dose, analyzed as treated | The one population analyzed by what participants actually received rather than what they were assigned. |
Because deviations are not random. Participants who discontinue, miss doses or violate the protocol differ systematically from those who do not, and often because of how the drug affected them. Excluding them removes the randomization that made the comparison valid in the first place. Per-protocol analysis is supportive, never primary.
The safety set is the one population analyzed as treated rather than as randomized. If someone assigned to placebo received drug in error, their adverse events belong to the drug. The efficacy question and the safety question require different answers to “which group is this person in?”
3. The Estimand
Under ICH E9(R1) a sponsor must state precisely what question the analysis answers. Five attributes define it.
| Attribute | This program |
|---|---|
| Treatment | VitaFlow at target dose versus placebo, both plus lifestyle intervention |
| Population | Adults meeting the inclusion criteria, per the FAS |
| Variable | Percent change in body weight from baseline to week 68 |
| Intercurrent events | Treatment-policy strategy — the value is used regardless of whether the participant discontinued study drug or started rescue therapy |
| Population-level summary | Difference in means between arms |
An intercurrent event is something that happens after randomization and changes the meaning of the measurement — discontinuing the drug, starting a different weight-loss therapy, undergoing bariatric surgery.
A treatment-policy strategy uses the week-68 value regardless: if a participant stopped taking the drug at week 12 and regained the weight, that regained weight counts. A hypothetical strategy would instead estimate what their weight would have been had they continued — a more flattering question, and one that describes a world that did not happen.
This program uses treatment policy, which is the harder standard and the one regulators increasingly expect. It answers “what happens if you prescribe this drug”, not “what happens to the people who tolerate it”.
Those discontinuations are concentrated in the treated arm. So the GI tolerability problem does not merely sit alongside the efficacy result; it directly reduces it. The two endpoints the program cared about most are statistically coupled, and a program that treated them as independent would have been surprised twice by the same phenomenon.
4. The Testing Hierarchy
When multiple endpoints are tested, the chance of a false positive somewhere rises with each test. The hierarchy controls that — endpoints are tested in a fixed order, and testing stops at the first failure. Everything below an unmet endpoint becomes descriptive, however good its own result.
| Order | Endpoint | Type | Note |
|---|---|---|---|
| 1 | Mean percent change in body weight, week 68 | co-primary | Superiority to placebo |
| 2 | Proportion achieving ≥5% weight reduction | co-primary | Superiority to placebo |
| 3 | Proportion achieving ≥10% weight reduction | key secondary | |
| 4 | GI-attributed treatment discontinuation | key secondary | The differentiation endpoint — and it sits fourth. |
| 5 | Proportion achieving ≥15% weight reduction | key secondary | |
| 6 | Change in waist circumference | key secondary |
GI-attributed discontinuation is the differentiation endpoint. It is the reason the asset was funded as a differentiated entrant rather than a late follower. And it sits fourth, behind three weight endpoints, because the agency's expectation at the End-of-Phase-2 meeting was that weight endpoints lead in a weight-management indication.
A hierarchy position determines what can be claimed. An endpoint high in the hierarchy with a significant result supports a labeling claim. An endpoint at position four, in a hierarchy that holds, supports a descriptive statement in Section 6 — which is exactly what the approved label carried.
The Protocol Summary records that the argument to promote GI discontinuation was had at protocol finalization and lost. This document is where that decision became irreversible: the hierarchy is fixed in the SAP before unblinding, and it cannot be reordered once anyone has seen a result.
The program chose to protect the indication. ⚠ That choice has to be defensible now, on what is known before the blind is broken — because the one thing certain about the result is that it is not yet available to justify the decision with.
5. Missing Data
In a 68-week trial, some participants will not have a week-68 measurement. How those absences are handled changes the answer.
| Analysis | Method | What it assumes |
|---|---|---|
| Primary analysis | Mixed model for repeated measures (MMRM) on observed data | Assumes data are missing at random, conditional on observed values. |
| Sensitivity 1 | Multiple imputation under a washout assumption | Assumes participants who discontinue return toward their baseline weight — a deliberately conservative assumption for a weight-loss product. |
| Sensitivity 2 | Retrieved-dropout imputation | Uses observed data from participants who discontinued drug but stayed in follow-up. |
| Sensitivity 3 | Tipping-point analysis | Progressively penalizes the treated arm until the result loses significance, and reports how far that is. Answers the reviewer's question before it is asked. |
Participants who discontinue are disproportionately those who experienced adverse effects, which means disproportionately those on active drug. Analyze only the people who finished, and the result flatters the drug — and flatters it most in the arm where discontinuation was most common.
The tipping-point analysis at the bottom of the table exists to answer this directly. It progressively penalizes the treated arm's imputed values until the primary result loses significance, and reports how far that had to go. A reviewer asking “how sensitive is this to your missing-data assumptions?” gets a number rather than an argument.
6. The Interim Analysis
| Number of interim analyses | 1 |
| Timing | When 50% of participants have completed week 68 |
| Purpose | Futility only |
| Efficacy stopping boundary | None |
| Conducted by | Conducted by the independent statistician and reviewed only by the DMC |
| Alpha spending | No alpha spent — a futility-only analysis does not inflate type I error |
A futility analysis can stop a trial that is clearly not going to work, which spares participants exposure to a drug that will not help them and spares the sponsor the remaining cost. An efficacy boundary would allow stopping early on good news.
This program has the first and not the second, for the reason set out in the DMC Charter: stopping early on a chronic therapy produces a smaller long-term safety database than the label requires, and a program that can stop early for good news has an incentive to look for it.
Because the analysis is futility-only, no alpha is spent. Stopping rules that could declare success consume type I error and require the final analysis to clear a higher bar; a rule that can only stop for failure does not.
7. The Limits of Pre-Specification
- It cannot change the protocol. If the protocol specifies an endpoint, the SAP describes how to analyze it — it does not get to substitute a different one.
- It cannot be revised after unblinding. Any change after that point is a post-hoc analysis and must be labeled as one in the submission.
- It cannot rescue an underpowered study. Sample size was fixed at protocol design; no analytical choice recovers power that was never there.
- It cannot make a fourth-ranked endpoint into a claim. That was decided at protocol finalization.
An SAP is where several of those decisions become irreversible — and writing one is therefore less an act of statistics than an act of governance. It is the last moment at which the program can still choose what question it is asking, and the first at which it can no longer change the answer.
What the program manager owns here
Not the statistics. The sequencing: SAP approval before database lock, database lock before unblinding, and enough calendar between them that a late statistical review does not force a choice between a rushed plan and a slipped filing. Thirty-five days, on this program, deliberately protected.