← Stage-Gate NPD Suite Product & Regulatory · Written for the Gate, Not the Testers

Master Test & Validation Strategy

Download Word
Program timeline · status 16 Oct 2026Read the full story →
Harborline
Aug 2025
Cancelled
Gate 0
Feb 2026
Go
Stage 1
Business case
Gate 1
Apr 2026
Recycled
Gate 1
Jun 2026
Go w/ conditions
Stage 2
Development
You are here
Gate 2
Apr 2027
Gate 3
Oct 2027
Gate 4
Feb 2028
Launch
Mar 2028
Gate 5
Sep 2028

Lighthouse Financial Services Company — How Beacon Index Advantage will be validated across six levels running from Stage 3 into Stage 4. Fully specified and entirely unstarted: testing is Stage 3 work and Stage 3's tranche of $7,450,000 has not been released. Status as at 16 October 2026.

6
Validation levels
12
Requirements under test
$7,450,000
Unreleased — funds all of it
0
Tests executed
Contents
  1. Who This Document Is Actually For
  2. The Six Validation Levels
  3. Why Illustration Validation Is Different
  4. Entry and Exit Criteria
  5. Defect Severity and What Blocks a Gate
  6. Coverage Gaps
  7. Dependencies

1. Who This Document Is Actually For

A test strategy normally addresses the people who will execute it. This one cannot, because they have not been funded yet.

Testing happens in Stage 3, and Stage 3 is not funded. So this document's first reader is not a QA lead planning execution — it is the Gate Review Board deciding, at Gate 2, whether the approach is good enough to release $7,450,000 against. You write the test strategy for the gate, not for the testers.

That inverts the usual audience and changes what the document must contain. A Board does not need scripts, environments or tooling choices; it needs to be shown that the approach will produce the specific evidence the next two gates require. So every level in §2 names the gate it feeds and the requirement it proves, and §4 states entry and exit criteria in terms a gate can certify rather than a team can interpret.

The corollary is uncomfortable and worth stating. If this strategy is wrong, the wrongness is discovered in Stage 3 — after Gate 2 has released the money, and with the launch date fixed. The strategy is therefore reviewed at Gate 2 as a should-meet criterion in its own right, rather than being treated as an operational detail delegated below the gate.

2. The Six Validation Levels

LevelNameWhat it provesFeedsOwnerRequirements
L1Configuration verificationThat the admin platform behaves as configured for the three launch crediting strategies, the surrender-charge schedule and the rider election flag.Gate 2 (readiness) / Gate 3 (evidence)V. SandovalPR-04, PR-05, PR-06, PR-15
L2Illustration validationThat illustration output reconciles to the filed methodology — not to the specification. This is a compliance test wearing a functional test's clothing.Gate 3E. KowalczykPR-12, PR-05
L3Actuarial reconciliationThat values produced by the admin platform reconcile to the pricing model within tolerance, across issue ages, premium bands and rider election states.Gate 3T. BrennanPR-01, PR-03
L4End-to-end new business and servicingThat a policy can be quoted, issued, serviced and surrendered by trained staff using production processes — not by the build team using workarounds.Gate 4L. MarchandPR-17, PR-02
L5Hedging shadow bookThat the hedging program can price, execute and reconcile against a simulated book before it holds a real one.Gate 4M. DelacroixPR-14
L6Distribution readiness rehearsalThat advisors can complete training, be appointed, and submit a clean application through the channel path they will actually use.Gate 4R. CastellanosPR-13, PR-18

L1–L3 produce Gate 3 evidence and are about correctness. L4–L6 produce Gate 4 evidence and are about readiness — whether the organization, not the software, can operate the product. The split matters because they fail differently: a correctness failure is fixed by a defect cycle, a readiness failure is fixed by training, hiring or process redesign, and only one of those fits inside a launch window.

L1 — what failure looks like, and the false pass to fear

A real L1 failure is visible and cheap: a surrender charge computed off the wrong schedule year, a rider flag that does not persist. The false pass is subtler and specific to configured platforms: the system behaves correctly for the cases the configuration team used while configuring — which are, by construction, the cases most likely to pass. L1 therefore executes from the requirement set imported through the RTM, not from Cordelane's own configuration checklist, so the vendor is never the author of the exam it sits.

L2 — the level where “matches the spec” is the wrong answer

L2's characteristic failure is an illustration that agrees with the PRD and disagrees with the filing — and §3 explains why that counts as a failure even though two internal documents agree. Its false pass is reconciliation performed against the draft methodology rather than the as-filed version; the filing goes through state review and can come back altered, so L2's input document is version-controlled against the Filing Plan's record of what was actually submitted.

L3 — tolerance is an actuarial judgment, recorded before the run

L3's failure mode is drift: platform values and pricing-model values that reconcile at issue and diverge at duration. Its false pass is a tolerance set after seeing the differences — wide enough to absorb whatever appeared. The tolerance is therefore fixed and signed by the Chief Actuary before execution, per cell of the age × band × election grid, and a reconciliation that passes only because the tolerance moved is treated as a failed control rather than a passed test.

L4 — the organization is the unit under test

By Gate 4 the software has already proven itself at L1–L3; L4 exists because a policy is issued by people, and its characteristic failure is a step that works only when a build-team member is in the room. The protocol therefore bars the build team from the floor during execution — their absence is not a staffing detail, it is the test condition. The false pass is a rehearsed run: the same operations staff processing the same scripted case until it goes clean. Case selection is drawn fresh from the L1 requirement grid on execution day.

L5 — mechanics now, market later, and honest about the difference

The shadow book proves the hedging program can price, execute and reconcile — the mechanics. TS-03 records what it cannot prove: behavior under the market conditions that make hedging matter. The month-long reconciliation window exists because the failure mode being hunted is cumulative — small daily breaks that net large — and a single clean day would hide it. L5 is also the level with an external critical path: no executed second ISDA, no shadow book (DEP-06), which is how a test strategy line becomes a Gate 2 must-meet risk.

L6 — the channel path is the product

L6's failure mode is an application that is clean by the carrier's rules and unmakeable by an advisor's — an appointment step that dead-ends, training that certifies without enabling. Its false pass is testing through an internal shortcut instead of the appointment-training-submission path an advisor will actually walk. The rehearsal therefore runs end-to-end through the IMO channel under R. Castellanos's ownership, because the distribution commitment the business case depends on is only as real as this path.

3. Why Illustration Validation Is Different

An illustration is not correct because it matches the specification. It is correct because it matches the methodology filed with the states. L2 therefore reconciles output against the filed methodology, not against the PRD — and where the two disagree, the filing wins and the specification is wrong.

That has a consequence no other test level carries: a whole class of defect here cannot be fixed by the program alone. If validation shows the engine producing output the filed methodology does not support, the options are to change the engine or to re-file — and re-filing restarts a review clock the program cannot compress (see the Regulatory Filing Plan).

Which document wins, decided in advance

The precedence order is written here because the moment it is needed is the worst moment to negotiate it: filed methodology > pricing model > PRD > platform behavior. When L2 finds a disagreement, that order assigns the defect without a meeting — if the engine matches the PRD and not the filing, the PRD is wrong; if the pricing model and the filing diverge, the filing wins and the model is corrected and re-signed. An organization that decides precedence per-incident will, under launch pressure, decide it in favor of whatever ships fastest.

Why this level reports to compliance as well as the gate

L1 and L3 failures embarrass the program; an L2 failure that reaches production exposes the carrier — an illustration that misstates filed values is a market-conduct matter, not a bug. L2 results therefore carry a second addressee: the compliance function under A. Nkemelu receives the reconciliation evidence directly, not summarized through the program. A control whose output is filtered by the party it controls is not a control, and this is the one level where that principle has a regulator standing behind it.

The program has already met the shape of this problem once. Issue I-02 recorded that the illustration engine could not produce compliant hypothetical performance for two of the five originally proposed crediting strategies. Those strategies were withdrawn at Gate 1 rather than engineered around, which is why the launch set is three. L2's scope is smaller today because a gate acted on that finding early, and it is the clearest example in the program of a gate reducing downstream test risk rather than merely observing it.

4. Entry and Exit Criteria

GateCriteria
Entry to Stage 3 testingGate 2 releases the Stage 3 tranche · configuration complete under Cordelane WP-1 · contract forms cleared by outside counsel · pricing signed and supported by the GC-03 external review.
Exit from L1–L3 (Gate 3 evidence)Zero open severity-1 defects · illustration reconciliation within tolerance against the filed methodology · actuarial reconciliation signed by the Chief Actuary.
Exit from L4–L6 (Gate 4 evidence)End-to-end cycle completed by trained operations staff without build-team intervention · hedging shadow book reconciled for a full month · advisor training completion records produced.

Exit criteria are written so a Board can certify them without interpreting them. “Zero open severity-1 defects” is checkable; “testing substantially complete” is not, and a criterion a gate cannot check is a criterion that will be argued rather than assessed.

5. Defect Severity and What Blocks a Gate

SeverityDefinitionEffect on the gate
S1Prevents issue, servicing or a compliant illustrationBlocks the gate. No exceptions.
S2Materially wrong output with a manual workaroundBlocks the gate unless the workaround is documented, staffed and accepted by Operations.
S3Cosmetic or low-frequency edge caseRecorded, scheduled, does not block.
The S2 rule is the one that will be tested in practice. A materially wrong output with a manual workaround is exactly the defect class that gets waved through under launch pressure. The rule requires the workaround to be documented, staffed and accepted by Operations — three things that each take somebody's time, which is the point. A workaround nobody has staffed is not a workaround; it is a defect with better wording.

Why the pressure lands on S2, and what the rule buys

S1 defects defend themselves — nobody argues a policy that cannot be issued onto the market. S3 defects threaten nothing. S2 is where every hard conversation in the final weeks will live, because each one arrives with an advocate and a plausible manual fix, and because the cost of accepting one is invisible on launch day: it appears later, as operational load compounding in a servicing team that never agreed to carry it. The acceptance signature from Operations is the mechanism that surfaces that cost before the gate — the accepting owner is the one who will staff the workaround, so the incentive to underestimate it is removed from the person estimating. The same rule appears, verbatim, as a no-go condition in the Launch Readiness Criteria, and the two documents are checked against each other at build time so they cannot drift apart — a claim of agreement between artifacts is enforced, not asserted.

6. Coverage Gaps

IDGap
TS-01Requirements under test: 12 of 15 in scope. The untested remainder (PR-09, PR-10, PR-16) are regulatory and reinsurance requirements whose validation is an approval or an executed treaty rather than a test result. They are verified, but not by this strategy — and a reader who assumed test coverage equaled requirement coverage would be wrong.
TS-02No performance or volume testing is scoped. Year 1 forecasts $185,000,000 of premium across roughly nine hundred policies — a volume the incumbent platform handles daily for other products. That reasoning is sound and it is an assumption, not a test.
TS-03L5 validates the hedging program against a simulated book only. A shadow book proves the mechanics work; it does not prove they work under the market conditions that make hedging matter. Nothing before launch can.
TS-04The strategy assumes test environments are available when Stage 3 opens. The configuration environment is shared with an in-flight servicing release (I-04), and contention is being scheduled around rather than resolved. If it is not resolved, L1 starts late and every level downstream of it moves.

7. Dependencies

IDDependencyEffect if it slips
DEP-03Cordelane's quarterly configuration windowsL1 cannot start until configuration lands. A missed window costs a quarter, not a sprint.
DEP-01/02Regulatory approvalsL2 can validate against the filed methodology before approval; it cannot confirm the filing was accepted. Those are different assurances and the gate evidence must not conflate them.
DEP-06Second ISDA counterpartyL5 cannot begin without executed documentation. This is the item currently placing Gate 2's MM-4 at risk.
A-08Actuarial and IT availability in 2027L3 depends on actuarial capacity the program does not hold. Untestable until the 2027 planning round.

Owned by V. Sandoval, Platform Delivery Lead, with E. Kowalczyk (illustration validation), T. Brennan (actuarial reconciliation) and M. Delacroix (hedging shadow book). Tabled through the gate process by C. Tyrrell, NPD Program Manager. Related: Requirements Traceability Matrix · Regulatory Filing Plan · Gate 2 Review Deck · RAIDD Log (I-02, I-04, DEP-03, DEP-06).