Declared analysis policy
AnalysisPlan
Section titled “AnalysisPlan”AnalysisPlanThe pre-registered decision rule: how evidence is judged.
Bindings (ExperimentMetric / MetricSpec) say how estimates are
computed; this says how they are judged.
Every role entry is a PlanEntry: either a bare metric name, or an
ExperimentMetric binding carrying per-metric overrides
(decision_method, sensitivity_methods, prior, guardrail margin). A metric
may hold only one role.
Error control by role
Section titled “Error control by role”A role is a claim about which error rate a metric’s verdict is protected against, so each role gets its own budget and its own procedure.
- primary — confirmatory, tested at
alpha / n_primariesand split again across the metric’s own non-control arms. Bonferroni familywise control atalphaacross the confirmatory family. - guardrail — a one-sided non-inferiority test at the full
alpha, never divided across guardrails, on the adverse side implied by the metric’spreferred_directionand margin. A guardrail asks “is this materially worse”, not “is this better”: dividingalphaacross guardrails would raise the false-negative rate for real harm, which is the error that matters here. - secondary — the discovery family at
q, a level deliberately looser thanalphabecause these verdicts are not confirmatory. The procedure depends oninference: BH/e-BH false-discovery control for fixed-horizon andAlwaysValidinference, and fixed-roster Bonferroni familywise control atqforAsymptoticMean, whose look structure has no established step-up composition. Familywise control atqimplies false-discovery control atq, so the weaker-procedure case never promises less than the step-up ones; it is only less powerful. - unassigned — reported at
alpha, no multiplicity claim.
Parameters
Section titled “Parameters”alpha : float
Familywise level for the confirmatory roles. A primary’s share is
alpha / n_primaries, split again across that metric’s own
non-control arms at estimation, so the shares sum to no more than
alpha. Guardrails test at alpha unsplit.
q : float
Secondary-family error level: BH/e-BH targets false discoveries for
fixed-horizon/exact registered inference. AsymptoticMean instead
uses fixed-roster Bonferroni familywise control at q, allocating
q / n_secondaries per secondary. It is not a quantile. These
guarantees are distinct from confirmatory allocation at alpha;
see increment.estimation.family.
view_multiplicity : MultiplicitySpec | None
Multiplicity for segmented as-of and breakout views. None keeps
route defaults: randomized breakout BH at q (fixed-roster
Bonferroni for AsymptoticMean), encouragement breakout uncorrected,
and as-of uncorrected unless segmented.
alternative : {“two-sided”, “greater”, “less”}
Direction of the test. A one-sided level is displayed through the
standard alpha-doubling convention.
primary : PlanEntry | tuple[PlanEntry, …] | None
The confirmatory metric(s). Accepts a single entry or a list;
None declares no primary.
secondaries : tuple[PlanEntry, …] | None
Metrics judged as a discovery family at q rather than against
alpha. A prior-bound entry sits outside the family.
guardrails : tuple[PlanEntry, …]
Metrics that must not move adversely. A guardrail needs an
explicit non-neutral preferred_direction on the metric, since
a non-inferiority margin has no adverse side otherwise.
inference : InferenceSpec | None
Look policy. None means fixed-horizon: one analysis, no
peeking guarantee.
Examples
Section titled “Examples”AnalysisPlan( … alpha=0.05, … q=0.10, … primary=“revenue_per_user”, … secondaries=[“signups”, “sessions”], … guardrails=[ExperimentMetric(metric=“latency_p95”, margin=0.01)], … inference=None, … ) # doctest: +ELLIPSIS AnalysisPlan(alpha=0.05, q=0.1, …)
MultiplicitySpec
Section titled “MultiplicitySpec”MultiplicitySpecMultiplicity policy for a segmented readout family.
ExperimentMetric
Section titled “ExperimentMetric”ExperimentMetricOne metric’s method-role and prior declaration within an experiment.
A bare metric NAME in a plan role (primary/secondaries/
guardrails) is shorthand for a binding with no overrides. The
decision method is optional because omission means the default
unadjusted estimator; sensitivity methods are additional estimates
reported after that decision estimate.
InferenceSpec
Section titled “InferenceSpec”InferenceSpecSerializable exact likelihood or explicitly asymptotic scalar-mean policy.
adjustments maps a mean metric to the pre-period CUPED coefficient and
covariate centre its automatic scalar-mean registration binds; an explicit
registration declares the same on ScalarMeanModel.adjustment instead.
baseline_rate is the conversion rate you expect in the control arm,
read off the same metric over the weeks before the experiment (the number
the power analysis used); an automatic always_valid registration
centres its Beta prior on it, and without it the prior is flat. A wrong
value costs power, never validity. Both automatic kinds bind their
registration from source metadata before any outcome is read.
With no outcome metrics, baseline_rate instead describes control uptake
and tunes its prior; composed outcome/compliance plans retain a flat uptake prior.
segments predeclares the one segment dimension (a frame column) and the
levels an automatic registration retains one hypothesis for, per metric and
arm, before any outcome is read; a level never observed stays a monitored
empty cell, and a value outside the declared levels joins no cell. Declaring
it asserts that every unit’s level was fixed before assignment: an
asymptotic_mean registration records exactly that assertion as its
segment_membership="pre_assignment", and nothing about membership is
checked or inferred from outcomes. An explicit registration names each
segment on its SequentialCell instead.
Winsorization
Section titled “Winsorization”WinsorizationOutcome bounds applied to a mean metric before aggregation.