Skip to content
Development documentation. The PyPI package predates these APIs. Install from GitHub instead: pip install 'increment @ git+https://github.com/kylejcaron/increment.git' Keep any extras requested by the guide, such as increment[dashboard].

Multiplicity

Testing several metrics at once inflates the chance that at least one looks significant by pure noise. Increment’s AnalysisPlan separates metrics into roles with different multiplicity treatment: a primary takes a share alpha / n_primaries of the familywise alpha, split again across that metric’s own non-control arms at estimation; a guardrail tests unsplit, at the full alpha, since it must not move adversely; a secondaries family is judged as a false-discovery-rate problem at level q instead of a per-metric significance test.

Declaring families: primary, guardrails, secondaries

Section titled “Declaring families: primary, guardrails, secondaries”
import numpy as np
import polars as pl
from increment import Analysis, AnalysisPlan, MetricSpec
rng = np.random.default_rng(9)
n = 3000
variant = np.where(rng.random(n) < 0.5, "treatment", "control")
sessions = 4.0 + np.where(variant == "treatment", 0.6, 0.0) + rng.normal(0, 2.0, n)
signups = 0.10 + rng.normal(0, 0.02, n)
df = pl.DataFrame(
{
"user_id": [f"u{i}" for i in range(n)],
"variant": variant,
"enrolled_on": np.arange(n),
"sessions": sessions,
"signups": signups,
}
)
plan = AnalysisPlan(primary="sessions", secondaries=["signups"], q=0.10)
results = Analysis.from_unit_summary(
df,
unit="user_id",
group="variant",
control="control",
metrics=[MetricSpec(name="sessions", type="mean"), MetricSpec(name="signups", type="mean")],
plan=plan,
).run()
for r in results:
print(f"{r.metric:>10} role={r.role:<10} discovery={r.discovery} lift={r.lift.value:+.2%}")
sessions role=primary discovery=None lift=+12.94%
signups role=secondary discovery=False lift=-0.85%

q=0.10 is a false-discovery level, not a second alpha or a quantile. It bounds the expected fraction of discovered secondaries that are false positives, not each metric’s type-I rate. A primary row always has discovery=None: family selection applies only to secondaries and breakouts.

correction and MultiplicitySpec for segmented views

Section titled “correction and MultiplicitySpec for segmented views”

AnalysisPlan.view_multiplicity: MultiplicitySpec | None governs segmented breakout and as-of views specifically, independent of the top-level secondaries/q family above. MultiplicitySpec(correction="bh", q=0.10), MultiplicitySpec(correction="bonferroni"), and MultiplicitySpec(correction="none") are the three options; q is only accepted under correction="bh" (and defaults to 0.10 there if omitted) — setting q under "bonferroni" or "none" is refused. Leaving view_multiplicity=None keeps the per-route defaults: randomized breakout BH at q (fixed-roster Bonferroni for a sequential AsymptoticMean registration), encouragement breakout uncorrected, and as-of uncorrected unless segmented.

For randomized sequential dataframe breakouts, InferenceSpec.segments declares the complete level roster before outcomes; automatic registration keeps absent levels in the family. Exact e-BH uses the breakout view’s q, including a MultiplicitySpec override, not a second family inferred from observed rows. Asymptotic breakouts retain fixed-roster Bonferroni and refuse BH. Exact multi-arm whole-window monitoring counts every secondary metric/arm cell in the family and splits each primary’s allocation across its treatment arms. An encouragement plan cannot combine predeclared segments with SequentialCompliancePolicy(family=True): its breakout policy is uncorrected. Use separately allocated uptake cells (family=False) or an unsegmented joint family; existing encouragement breakout-readout restrictions remain.

Under AnalysisPlan(inference=InferenceSpec(kind="asymptotic_mean")) or InferenceSpec(kind="always_valid"), the secondary family uses e-BH (Wang & Ramdas 2022: ordinary BH applied to 1/e) instead of ordinary p-value BH. e-BH controls FDR at <= q under arbitrary dependence and stopping time, but only when every input is a valid e-process: an e-value that remains valid at every stopping time, not just at one fixed look. e_bh_select assumes that property; it does not establish it. The asymptotic-mean route builds its e-process from the same delta-method sampling distribution as the fixed-horizon interval, so its time-uniform guarantee is asymptotic and empirical, not finite-sample. The exact kind="always_valid" route, limited to conversion and retention metrics, is the finite-sample alternative.

from increment import AnalysisPlan, InferenceSpec, Randomized
design = Randomized(control_group="control", allocation={"control": 0.5, "treatment": 0.5})
av_plan = AnalysisPlan(
primary="sessions",
secondaries=["signups"],
q=0.10,
inference=InferenceSpec(kind="asymptotic_mean"),
)
av_results = Analysis.from_unit_summary(
df,
unit="user_id",
group="variant",
design=design,
metrics=[MetricSpec(name="sessions", type="mean"), MetricSpec(name="signups", type="mean")],
plan=av_plan,
exposure_date="enrolled_on",
).run()
for r in av_results:
print(f"{r.metric:>10} inference={r.inference:<13} discovery={r.discovery}")
sessions inference=asymptotic_mean discovery=None
signups inference=asymptotic_mean discovery=False

A sequential AnalysisPlan needs a design= with an explicit allocation (not the dataframe shortcut control=), since the registered asymptotic-mean runtime must be declared before any outcome is read. A unit summary also names each unit’s exposure_date (a date or day index); units are revealed in that order, never in row order. See Sequential inference for the registration, capture and error-control mechanics this route shares with an ordinary (non-multiplicity) sequential primary.

A secondary or breakout cell selected under BH (fixed-horizon p-values) or e-BH (sequential e-values) gets a reported interval re-estimated at fcr_alpha = min(q * R / m, nominal_alpha). Here, R is the number of selected cells and m is the family size. The interval is wider than the nominal interval when few cells are selected from a large family, and narrower as more cells are selected, up to the nominal-alpha cap.

Treat a discovered secondary’s interval as already carrying this correction; it is not the same width as a non-discovered row’s interval. family_guarantee reports "finite_sample" only when every family member is exact (always_valid) and "asymptotic_sequential" when any member is asymptotic.

Fixed-horizon BH assumes independence or positive dependence (PRDS) across the metric x arm x segment family, and this is not checked — read Fixed-horizon FDR control assumes a dependence condition that is not checked. An always-valid discovery set’s guarantee holds for the current date; the union across dates is not controlled — read Stopping-date guarantees require the registered reveal contract.