Skip to content
Development documentation. The PyPI package predates these APIs. Install from GitHub instead: pip install 'increment @ git+https://github.com/kylejcaron/increment.git' Keep any extras requested by the guide, such as increment[dashboard].

Statistical limitations

Every number this library reports rests on assumptions. This page puts them in one place, so you can see what a result assumes without reading the source.

Most entries mark deliberate limits: a method that is standard but approximate, an estimand that is narrower than it looks, or a guarantee that holds under a condition the library cannot check for you. A few identify gaps: a safeguard or diagnostic that one entry point provides and a neighboring one does not. Where the package offers a better route, this page names it.

Read the section that matches the decision you are making. Each entry says what the library does, what it assumes, and when the assumption matters.

A capability is a claim only when its path is named. Analysis has six entry points in three families: dataframe (from_unit_summary, one row per unit; from_unit_panel, one row per unit per day; from_switchback_panel); warehouse (from_definitions, compiling from raw event facts; from_unit_day_artifact, reading unit × day relations previously published into the warehouse); and portable (from_moments, a file of pre-reduced moments exported from any of the others).

The warehouse is the primary route. The table records measured behavior on every path, using the dataframe unit-summary path as the oracle against which the others are checked.

Reproduce this table with uv run --extra demo --extra tables --extra dashboard python scripts/probe_capability_table.py.

CapabilityUnit summaryUnit panelSwitchback panelDefinitionsUnit-day artifactPortable moments
Mean CUPED, fixed horizonyesyesrefusedyesyesyes
Ratio CUPED, fixed horizonyesyesrefusedyesyesyes
Ratio denominator-precision advisory, fixed horizonyesyesrefusedyesyesyes
Informative Normal/Student-t/mixture log-lift priors, fixed-horizon mean/ratio/conversion/mean CUPEDyesyesnot comparable (SOURCE — matched parallel-arm data has no switchback schedule)yesyesyes
Fixed-threshold winsorization under sequential inferenceyesyesrefusedyesyescheckpoint replay
Ratio metric under sequential inferenceyesyesrefusedyesyescheckpoint replay
CUPED under sequential inferenceasymptotic routerefusedrefusedasymptotic routeasymptotic routecheckpoint replay
Ratio CUPED under sequential inferenceasymptotic routerefusedrefusedasymptotic routeasymptotic routecheckpoint replay
Conversion/retention metric under asymptotic sequential inferenceyesyesrefusedyesyescheckpoint replay
Conversion/retention metric under automatic exact Bernoulli monitoring (always_valid without a registration)yesyesrefusedyesyescheckpoint replay
Automatic exact multi-arm conversion/retention monitoringyesyesrefused (SOURCE — no sequential construction)yesyescheckpoint replay
Randomized registered segmented sequential family (automatic predeclared levels or explicit roster)yesyesrefused (SOURCE — no sequential construction)refused (SOURCE, sequential.route.unsupported — relational capture has no immutable segment-property contract)refused (SOURCE, sequential.route.unsupported — artifact capture has no immutable segment-property contract)retained-state replay only; breakout readout refused (SOURCE, readout.source.dimension — no breakout catalog)
Boolean and null segment labels in a registered segmented sequential family (canonical true/false/__null__, matching the string-labelled oracle)yes (previously missing_arm on summary; measured on pandas, Polars and Arrow Boolean columns, automatic and explicit rosters)yes (same three column kinds)refused (SOURCE — no sequential construction)refused (SOURCE, sequential.route.unsupported)refused (SOURCE, sequential.route.unsupported)retained-state replay only; breakout readout refused (SOURCE, readout.source.dimension)
Continuing a checkpoint recorded under earlier raw Boolean/null spellings (True, None) with canonically labelled unitsrefused (COMBINATION — earlier prefix by canonical source, sequential.continuation.rewrite); the same string spellings still continue, and the checkpoint replays unchangedrefused (COMBINATION, sequential.continuation.rewrite)refused (SOURCE — no sequential construction)refused (SOURCE, sequential.route.unsupported)refused (SOURCE, sequential.route.unsupported)replay only (SOURCE — no per-unit source to continue from)
Artifact context format 2 (hash-only source recipes; no source SQL)not comparable (SOURCE — no warehouse artifact)not comparable (SOURCE — no warehouse artifact)not comparable (SOURCE — no warehouse artifact)yes, publishes format 2yes, reads format 2 only; a format-1 context is refused (COMBINATION — stored older artifact by current reader, artifact.format.unsupported) and must be republished from trusted definitionsnot comparable (SOURCE — moments cubes carry no artifact context)
Experiment window days (start_day, end_day, observation_horizon_day) follow day_boundary: naive value = wall clock at the boundary, aware value convertednot comparable (SOURCE — from_unit_summary accepts no Experiment declaration, so there is no window to place)not comparable (SOURCE — same; a frame has no window lever)not comparable (SOURCE — same)yes; a native sequential registration made before this change is refused (COMBINATION — earlier registration by current recipe, sequential.source.invalid) and must be re-registeredyes, the context binds window_days; a context without it is refused (COMBINATION — stored older artifact by current reader, artifact.context.mismatch) and must be republishednot comparable (SOURCE — moments cubes carry no experiment window)
Publication abort after the manifest insert began invalidates the generation (tombstone first, then relation erase; reads refused with artifact.generation.dropped; if the tombstone cannot be written the relations are kept and named, and abandon_generation from a fresh process hides the generation before they are dropped)not comparable (SOURCE — no artifact store)not comparable (SOURCE — no artifact store)not comparable (SOURCE — no artifact store)yes, WarehouseArtifactStore publicationyes, reads of a dropped generation refusenot comparable (SOURCE — no artifact store)
increment.impute.pooled_mean with declared roles=; MetricSpec(missing="impute")roles honored: an outcome role refuses (CONSTRUCTION — filling an outcome with its pooled mean shrinks variance, impute.pooled_mean_outcome); an undeclared role warns (impute.pooled_mean_role_undeclared); MetricSpec(missing="impute") refuses (CONSTRUCTION, frame.metric.missing_impute)samesamenot comparable (SOURCE — a warehouse Metric cannot express missing="impute"; the helper prepares dataframes only)not comparable (SOURCE — same)not comparable (SOURCE — no per-unit rows)
Encouragement ITT, compliance and LATE, fixed horizon, declared via design:yesyesrefused (SOURCE — from_switchback_panel takes no design=)yesyesyes
Explicit ITT/compliance without a LATE exclusion declaration (fixed horizon or supported sequential monitoring)yesyesrefused (SOURCE — from_switchback_panel takes no design=)yesyesyes; sequential checkpoint replay
Fixed-horizon design-level compliance with an empty outcome catalogyes, metrics=[]yes, metrics=[]refused (SOURCE — no design=)yesyesyes, format-8 design_summary and metrics=[]
Encouragement guardrail with a declared non-inferiority marginyes, ITT row carries the margin verdict; refused (CONSTRUCTION, readout.encouragement.margin) if a request excludes ittyes, ITT row carries the margin verdict; refused (CONSTRUCTION) if itt is excludedrefused (SOURCE — no design=)yes, ITT row carries the margin verdict; refused (CONSTRUCTION) if itt is excludedyes, ITT row carries the margin verdict; refused (CONSTRUCTION) if itt is excludedyes, ITT row carries the margin verdict; refused (CONSTRUCTION) if itt is excluded
Encouragement BH-selected secondary’s LATE row re-estimated at the corrected levelyesyesrefused (SOURCE — no design=)yesyesyes
Encouragement ITT on a conversion/retention metric keeps the exact binomial routeyesyesrefused (SOURCE — no design=)yesyesyes
Encouragement conversion ITT breakouts retain exact confidence sets at zero control countsrefused (SOURCE — no declared breakout dimension)yesnot comparable (SOURCE — no switchback schedule)yesyesrefused (SOURCE — no per-unit breakout dimension)
Bounded-retention breakouts with mature exposure cohorts (audit-retention-breakout-cohorts)refused (SOURCE — no exposure-date cohort stream, source.frame.constructor)yes, run_breakoutnot comparable (SOURCE — no switchback schedule)yes, run_breakout and cohort-aware breakout_summariesyes, run_breakout and cohort-aware breakout_summariesrefused (SOURCE — no breakout operation, facade.analysis.operation)
Encouragement continuous ITT + Bernoulli uptake compliance, jointly monitored (composed sequential family, e-BH selected)yesyesrefused (SOURCE — no design=, no sequential construction)yesyescheckpoint replay
Uptake-only sequential compliance with an empty outcome catalog (always_valid and a compliance policy)yes, metrics=[]yes, metrics=[]refused (SOURCE — no design=, no sequential construction)yesyescheckpoint replay with metrics=[]
Observational IPTW covariate adjustmentyesyesrefused (SOURCE — no design=)yesyesrefused (SOURCE — a moments cube carries no per-unit rows to attach a covariate to)
Observational AIPW/DML covariate adjustmentyesyesrefused (SOURCE — no design=)yesyesrefused (SOURCE — same as above)
Categorical observational adjustment (string covariates)yesyesrefused (SOURCE — no design=)yesyesrefused (SOURCE, source.moments.covariate_unavailable — no per-unit covariates)
Categorical nulls with a nondefault missing-value policyyesyesrefused (SOURCE — no design=)default missing-value refusal; YAML does not accept policy overridesdefault missing-value refusal inherited from definitionsrefused (SOURCE — no per-unit covariates)
Observational primary + secondary + guardrail together (role multiplicity)yesyesrefused (SOURCE — no design=)yesyesrefused (SOURCE — same as above)
Quantile inference on tied/rounded outcomesyesyes (unwindowed only — the panel’s per-unit collapse is the same per-unit total mean/ratio already use; a windowed quantile metric can never be declared on the frame path at all, frame.metric.window_days_supported)refused (SOURCE — a quantile metric is refused at the metric-type gate before any frame is read)yesyesrefused (SOURCE, source.frame.quantile_no_moments — a quantile has no moments representation)
Clustered ratio metric with a non-positive control-arm numeratoryesrefused (SOURCE — from_unit_panel(cluster=...) refuses unconditionally; build clustered arm totals with from_unit_summary(cluster=...))refused (SOURCE — a clustered arm-lift ratio has no switchback contrast shape)yesyesrefused (SOURCE, source.moments.cluster_grain — clustered moments transport needs a complete encouragement compliance payload)
Clustered zero relative variance retains its point and any available additive intervalyesrefused (SOURCE — no clustered arm moments; use from_unit_summary)refused (SOURCE — a switchback contrast has no arm-cluster shape)yesyesrefused (SOURCE, source.moments.cluster_grain — clustered moments transport needs a complete encouragement compliance payload)
Analysis.planning_baseline, mean metricyesyesyes (the pilot-fitted SwitchbackBaseline for the switchback_* solvers)yesyesyes
Analysis.planning_baseline, quantile metricyesyes (same unwindowed-only reach as the row above)refused (SOURCE — a quantile metric is refused at the switchback metric-type gate before any frame is read)yesyesrefused (SOURCE, source.frame.quantile_no_moments — no per-unit rows to draw a pilot sample from)
Analysis.planning_baseline with a declared cluster (per-unit mean/var; ICC on the metric’s per-unit score)yesrefused (SOURCE — from_unit_panel takes no cluster)refused (SOURCE — a switchback contrast has no arm cluster)yesyesrefused (SOURCE, source.moments.cluster_grain — clustered moments are not exported)
Analysis.planning_baseline under a declared trigger (triggered-population moments and trigger_rate)refused (SOURCE — from_unit_summary accepts no Experiment declaration, so planning_baseline’s trigger always resolves None; the assigned population is analyzed instead)refused (SOURCE — same as from_unit_summary)refused (SOURCE — same as from_unit_summary)yesyes with the trigger_population and assignment_counts extensions; otherwise refused (SOURCE, analysis.planning_baseline.trigger_evidence_unavailable)refused (SOURCE — same as from_unit_summary)
Logged-policy contrast (estimate_policy_contrast)refused (SOURCE — no Analysis constructor carries a decision trace: the per-decision chosen action, exact logged propensity, logging policy version and reward boundary; the family’s own ingress is LoggedTrace.from_records / from_frame)refused (SOURCE — same)refused (SOURCE — same)refused (SOURCE — same)refused (SOURCE — same)refused (SOURCE — same)
TabularPolicy persistence (model_dump_json / model_validate_json)not comparable (SOURCE — no Analysis constructor carries a policy table; the policy is a LoggedTrace input)not comparable (SOURCE — same)not comparable (SOURCE — same)not comparable (SOURCE — same)not comparable (SOURCE — same)not comparable (SOURCE — same). Within the family: JSON roundtrips string-keyed tables only; any other key type refuses (CONSTRUCTION — JSON object keys are strings, logged_policy.policy.json_context_key); typed keys persist through model_dump(mode="python") or trusted pickle

For from_definitions, pinning preserves dimensions joined by freshness-bearing fact sources. Qualifying non-enrolled events can establish that later data has arrived without entering the enrolled outcome population. In from_unit_panel, natural day labels such as d9 and d10 use numeric order when inferring maturity, matching the same explicitly declared observation bound. Windowed ratios use the earlier numerator/denominator observation bound for totals and completed as-of cohorts. An entirely null component under missing="zero" retains the panel-extent fallback; an explicitly supplied observation_end remains authoritative.

For clustered triggered planning, from_definitions and from_unit_day_artifact retain both populations: assigned mean cluster size determines recruitment; contributing-cluster participation, analyzed size CV, and analyzed-score ICC determine the planning approximation. The artifact needs cluster_identity as well as the trigger and assignment-count extensions. These pilot estimates do not certify future participation or finite-sample coverage. Other constructors cannot supply the declared trigger membership identified in the table. Both clustered trigger routes refuse a pilot whose metric rows omit trigger-eligible control units (SOURCE, analysis.planning_baseline.trigger_metric_population_unavailable). Missing rows can reflect unfinished windows or outcomes undefined even after full follow-up, such as avg_event for a unit with no events. Complete unfinished windows; for structurally undefined outcomes, choose a metric defined for every eligible unit. The adapter does not silently substitute a complete-case population for the declared trigger population.

Mean/Ratio CUPED on the unit panel resolves the declared covariate the same way observational adjustment already does: through the per-unit value unit_frame()/moments(grain="total") join in, constant across a unit’s own rows, refused by name (frame.frame_panel.unit_covariate_varies) when it genuinely varies. A windowed or retention metric still refuses a CUPED covariate (CONSTRUCTION, frame.validation.from_unit_panel): those metrics have no per-unit collapse for that value to attach to; the refusal names from_unit_summary as the route that does. The switchback panel supports neither sequential inference nor CUPED; its contrasts are fixed-horizon. “Asymptotic route” means CUPED is admitted under InferenceSpec(kind="asymptotic_mean"): fitted coefficients use retained joint moments, whose confidence sequence accounts for their estimation; pre-period coefficients use the scalar-mean adjustment. The exact Bernoulli e-process route (AlwaysValid) admits neither form of CUPED. “Checkpoint replay” means from_moments cannot start a sequential process, because it holds no per-unit source to register against, but reproduces an exported checkpoint exactly. On the warehouse the covariate for a ratio metric is the numerator’s pre-period total; a separate denominator covariate is not expressible, and the CUPED guide says why. An encouragement design’s continuous ITT and Bernoulli-uptake compliance, monitored jointly, compose into one sequential family automatically (no explicit registration to declare); the family reports one shared family_guarantee, equal to the weakest regime any member carries — with an asymptotic scalar-mean ITT cell in the family, that value is asymptotic_sequential on every selected row, including the exact Bernoulli compliance cell’s own row, not “exact for one cell, asymptotic for the other.” The from_switchback_panel SOURCE refusal on the encouragement, observational and quantile rows above is structural, not a gap: that constructor accepts only a switchback-shaped frame and an identification: Randomized design, so it never reaches an Encouragement, Observational, or quantile-metric case built for the other five paths.

Parity is asserted, not assumed: identical per-unit data through each path must reproduce the same adjusted moments and interval, and the same retained sequential state byte for byte. That check runs against DuckDB in the ordinary suite and against a live PostgreSQL in the warehouse gate. For the two binary rows it reads one per-unit dataset through the unit panel (under the same two-day window the definitions declare), the definitions reader, the unit-day artifact reader and a replayed portable checkpoint, and requires each to retain the unit-summary oracle’s state, registration and interval exactly.

The informative-prior row uses the existing approximate Normal likelihood on log lift, including for conversion metrics; it is not a binomial posterior or a finite-sample sampling guarantee. The three enumerated prior cases compare posterior bounds, probabilities, metadata and row identity across all five applicable constructors on DuckDB and PostgreSQL. See Priors and Bayesian decisions for binary boundary behavior and the distinction from prior-free sampling.

Binary sequential monitoring (either automatic route) composes as follows, measured on from_unit_summary and pinned per axis. Multiplicity: every in-family secondary on the metric/arm axis — Bernoulli, scalar-mean, or a mix of both — is registered into the family at a nominal per-cell allocation and selected by e-BH at the plan’s q (supported); a selected member’s interval is reinverted at min(q * R / m, nominal_alpha), and the family reports one shared family_guarantee (finite_sample only when every member is exact, asymptotic_sequential when any member is asymptotic). This is unrelated to the registered breakout/segment axis, whose own correction is declared per registration: a continuous (scalar-mean) breakout registers correction="bonferroni" and keeps fixed per-cell Bonferroni (FWER, no reselection); a discrete (Bernoulli) breakout registers correction="bh" and is selected by the same e-BH-at-q family as an ordinary metric/arm secondary, reinverted the same way. This is also unrelated to guardrails, which test at the full plan alpha outside every family. CUPED: a pre-period coefficient and centre declared through InferenceSpec.adjustments is supported on the asymptotic route; an in-experiment coefficient is refused on the exact route (sequential.transform.unpredictable) because it re-weights past increments with information unavailable when they were revealed, a construction this package does not have; the asymptotic route retains the joint moments instead. Clustered assignment: refused before any observation is read (sequential.route.unsupported); no cluster-robust sequential boundary is built. Winsorization: it does not apply to a 0/1 outcome, so a binary metric refuses a percentile or fixed clip at declaration (frame.metric.winsorization_applies_type); on a mean metric a fixed clip is supported and a percentile clip refuses (sequential.transform.unpredictable) because its threshold depends on accumulated data. Quantile metrics: refused on the exact route (definition.inference.always_valid_metric_type) because the Bernoulli law does not describe them; use fixed-horizon inference. Breakouts: InferenceSpec.segments fixes one string-valued dimension and its allowed levels before outcomes on from_unit_summary and from_unit_panel; an explicit registration is optional. Empty levels stay in the roster. Asymptotic segment membership must be determined before assignment and uses the existing fixed-roster Bonferroni construction, not BH. Segment labels are canonical strings on both frame routes: true/false for Booleans and __null__ for missing values, so summary and panel sources label the same units identically. A checkpoint recorded under an earlier raw spelling (True, None) replays unchanged but cannot be continued by a canonically labelled source (COMBINATION — earlier recorded prefix with a different roster spelling; sequential.continuation.rewrite). That is a correction of an earlier inconsistency, not an inherent limitation. from_definitions and from_unit_day_artifact still refuse segmented capture (sequential.route.unsupported). A moments cube can replay retained segment state but has no breakout catalog, so run_breakout refuses (readout.source.dimension). Switchback: refused before any frame is read (source.frame.switchback.plan, reason unsupported_inference) for every sequential kind, automatic or explicit; the switchback contrast has no sequential construction. Continuous intent-to-treat under an encouragement design registers under the public asymptotic mean law exactly as a randomized design does (sequential.route.unsupported still names the Bernoulli route on any other mechanism).

No estimator here is handed a known variance; every one estimates uncertainty from the data. What differs is the approximation each family layers on to get there:

FamilyHow uncertainty is obtained
Conversion and retention arm lift (unadjusted, unit-grain, no informative prior)Exact independent-binomial risk-ratio test inversion (Berger-Boos restricted-nuisance construction) — no delta method, no Normal reference
Unit-grain mean/ratio lift and CUPED, no informative priorDelta-method log-scale uncertainty against a Welch-Satterthwaite t reference, whose degrees of freedom are persisted on the row
Clustered unadjusted arm liftJoint additive/control covariance, Fieller relative set; separate additive uncertainty
SequentialSeparate registered observation-model contract; fixed-horizon approximations do not establish anytime validity
QuantilesInversion of an order-statistic bracket (Woodruff), not a delta method
SwitchbackIndependent unit-cycle orders: default UnitCycleTApproximation (sample variance, at least two units, dof = n_units - 1); a valid prospective UnitCycleVarianceEnvelope is a separate stronger route supporting one unit. Shared schedules: block-t on at least two blocks
Observational adjusted methodsInfluence-function sandwich

Estimating the variance does not by itself forfeit finite-sample validity — under parametric assumptions it need not. But it does mean each family’s interval inherits the quality of its own approximation, and those are not interchangeable. The gap matters most at small n, for heavy-tailed outcomes, and for ratio metrics, where a first-order approximation is furthest from the truth.

Absolute effects on the switchback and observational paths are estimated natively on the additive scale rather than reconstructed from a log-scale estimate, so they carry their own path’s approximation and not an additional one.

Approximate inference paths do not promise finite-sample exactness for a discrete outcome. Eligible fixed-horizon conversion/retention results with reference_kind="binomial" retain a finite-sample guarantee for the relative-risk confidence set and its relative-scale decisions. Their additive abs_lb/abs_ub sidecar, when available, uses a Normal-Wald approximation; that interval is not exact. Asymptotic and empirically qualified methods are labeled as such rather than treated as exact; an unfinished or unsupported capability refuses before producing a result. No result field diagnoses model misspecification. For approximate results, no field certifies calibration. A procedure declares a coarse sampling floor — two units per arm for most arm inference, twenty for a quantile. Cluster methods require estimable independent contributions, not a universal ten-cluster admission rule. A feasibility minimum is not a calibration guarantee.

It is not the estimator’s own feasibility check either, and the two can disagree in both directions. The quantile order-statistic bracket needs a sample that depends on the requested quantile and the alpha it is evaluated at: a median at 95% is feasible on six units per arm, well under the declared twenty, while a 99th percentile needs 368.

The quantile interval stays valid as the recording grid coarsens relative to the sample size, rather than refusing on ties: while an arm’s order-statistic bracket holds a tie and does not yet span a dozen repeated values, the reported interval contains that bracket, so the bracket’s distribution-free coverage carries over whatever the grid (values recorded off it, prices ending in both .99 and .00). The price is width, about 1.1 to 3 times the classical formula’s on coarse grids (milliseconds, seconds or low-value cents at large sample sizes; counts). It has no fixed bound, because a bracket collapsed onto one value has zero classical width, so each widened row states its own ratio in its note.

This feasibility bound applies to a single reported interval, evaluated at one caller alpha. The p-value computation (see below) scans across alpha internally and never surfaces this refusal to the caller: past the infeasible boundary it treats the construction as “does not exclude the null” and stops searching there, rather than raising.

Calibration is metric-specific for approximate methods. Unadjusted unit-grain conversion/retention uses the exact binomial method below at every event count. CUPED and ratio-denominator routes retain their delta approximation. Clustered unadjusted arm lift instead retains joint covariance and relative-set geometry; that representation prevents misleading finite intervals near zero controls, but does not establish finite-sample coverage for rare events or few clusters.

A unit-grain ratio row (unadjusted or CUPED, fixed horizon) carries a ratio_denominator_precision note when either arm’s denominator mean is resolved to a relative standard error sqrt(var_d / n) / d_bar above 0.15. This is an advisory qualification, not a refusal or a corrected interval: the interval is emitted unchanged. The threshold comes from calibration/ratio_precision.py, a seeded coverage grid over sample sizes 20-2000 and lognormal, exponential and gamma denominators with an independent numerator. Every cell covering below 93% at nominal 95% exceeds it in most draws, and no cell covering at least 94.5% does. The note names the arm, the statistic and the threshold. The unchanged interval is still anticonservative under a right-skewed denominator (85.5% coverage at n=20 and 89.6% at n=50 under lognormal(sigma=1.5)). The statistic sees only the denominator’s own precision: a numerator that moves with the denominator covers close to nominal even when flagged, and a numerator independent of it undercovers most. The carried moments end at second order, so the heavy-tail correction and broader small-sample coverage remain unresolved.

The switchback t reference is itself an approximation

Section titled “The switchback t reference is itself an approximation”

Independent unit-cycle orders use UnitCycleTApproximation() by default (equivalent to passing it) unless a valid prospective UnitCycleVarianceEnvelope is supplied; the shared-block reference is separate. Both use the independent contributions’ sample variance. That is not a claim of finite-sample coverage, and the approximation is not confined to discrete outcomes.

Under independent unit-cycle orders, each mean-outcome contribution is inverse-probability weighted, so it takes one of only two values, and those two are asymmetric whenever the declared CT/TC probability is not 0.5. The per-unit deltas are means of such contributions, so the t reference relies on those means being approximately Normal — a condition that is not among the contract’s declared assumptions, which name only no_residual_carryover_after_discarded_steps and independent_units. Coverage therefore degrades with few units, with an assignment probability far from 0.5, and with skewed, heteroskedastic, or high-leverage outcomes, and the declared floor admits as few as two units — the regime where the approximation is weakest. The asymmetric/skew calibration includes a small-N counterexample to nominal coverage. Unit-cycle variance envelopes instead give a finite-sample bound under the declared prospective residual-variance assumption, including at one unit. That assumption must be justified externally; a fitted pilot variance is not a valid envelope. Shared schedules instead apply switchback_block_t with a block_t reference to block contributions; their floor is two independent blocks, even with a one-unit roster. More roster units do not increase the replication count or establish finite-sample coverage.

Cluster references are qualified working approximations

Section titled “Cluster references are qualified working approximations”

Disjoint unadjusted arms retain each arm’s cluster-ratio uncertainty. The relative Fieller set uses a fixed t reference with min(K_treatment - 1, K_control - 1) degrees of freedom; the additive interval separately uses Welch–Satterthwaite. Shared-arm dependence retains the complete response-corrected covariance. Observational IPTW/AIPW/DML use their joint influence-function covariance and an asymptotic Normal reference. Neither a Bessel multiplier nor a generic t choice establishes finite-sample validity for arbitrary small, unbalanced, or skewed clusters. Encouragement LATE and sitewide impacts retain their own component-based references. The cluster-asymptotic justification also needs no cluster to dominate its arm’s variance. Neither a large total row count nor crossing the 40-cluster warning threshold establishes that condition. Concentrated cluster masses can bias the disjoint-arm variance estimate downward; a t reference alone does not correct it.

The clustered encouragement LATE cuts its additive interval at a Welch–Satterthwaite t reference over the two arms’ own cluster-ratio variance components. Measured by tests/calibration/test_cluster_late_welch_df.py (2000 replications per cell over 5/5, 5/25 and 10/50 treatment/control clusters, treatment between-cluster variance 1, 4 and 16 times the control’s, mean cluster size 20; binomial Monte-Carlo standard error 0.5 points at 95% and 0.9 points at 82%): with equal cluster sizes the interval covers the true LATE 94.3–96.0% at nominal 95% in every cell, including 94.55% at 5/25 clusters with 16 times the variance in the small arm, where a pooled t_(K_treatment + K_control - 2) reference over the same components would cover 89.15%. Log-normally dispersed cluster sizes (sigma 0.75) cost about 3–6 points that the reference does not recover: 88.5–92.7% across those nine cells, 89.7% at 5/25 with the 16x ratio, against 82.0% pooled. The first-stage gate withheld at most 4 of 2000 replications in any cell.

relative_confidence_set preserves bounded, disconnected, one-sided, full-real, empty, or unavailable geometry. A zero control mean can leave a confidence set without a finite point. Numeric nulls are not infinities. A genuinely indefinite or unrepresentable joint covariance sets relative_unavailable_reason; additive output survives rather than being replaced by a clipped covariance or fabricated relative precision. Scalar posterior decision probabilities are not recovered from a joint frequentist set.

Representative release checks are diagnostics, not certification of the original stress matrix. Run the preserved cases and paired-width criteria explicitly:

Terminal window
uv run python -m calibration.cluster_diagnostics --output /tmp/cluster-diagnostics --repetitions 4
uv run python -m calibration.cluster_diagnostics --selection full --output /tmp/cluster-stress

Output directories must be new. Budgets are bounded at one hour; started, completed, refused, failed, unfinished, and unstarted work remain distinguishable. The full route preserves original seeds, repetition rules, and paired-width thresholds. Interrupted or diagnostic prefixes never pass those historical gates.

CUPED treats its adjustment coefficient as known

Section titled “CUPED treats its adjustment coefficient as known”

The CUPED standard error conditions on a theta estimated from the same data it adjusts. The uncertainty in that estimate is not propagated, so the reported interval is slightly narrower than one holding the coefficient fixed. Measured by tests/calibration/test_cuped_theta_uncertainty.py (16 cells: covariate correlation 0.1 to 0.9 by 20 to 2000 units per arm, equal allocation, one homogeneous slope, 3000 replications each, intervals cut at the row’s Welch–Satterthwaite t reference; binomial Monte-Carlo standard error 0.4 points): the fitted interval is 1.3–1.4% narrower on the log scale than a known-coefficient comparator at 20 units per arm, 0.5% at 50, 0.1% at 200 and 0.01% at 2000, the same at every covariate correlation, and within Monte-Carlo error of sqrt(1 - 1/(N - 2)) with N the total sample. Coverage of the fitted interval is 93.4–95.3% at nominal 95% across the 16 cells, against 93.8–95.6% for the comparator, and the reported standard error runs 0.94–1.01 times the empirical spread of the log estimate (0.95–0.98 at 20 per arm). The comparator is the same log-scale delta method with the true coefficient substituted, not exact finite-sample inference, so the width ratio measures the residual-variance shrinkage from fitting theta and does not isolate the added sampling variance of the estimate. Unequal allocation and arm-specific slopes are outside that design.

Rare events on an unadjusted conversion/retention arm are estimated exactly, not refused

Section titled “Rare events on an unadjusted conversion/retention arm are estimated exactly, not refused”

An unadjusted (no CUPED, no declared cluster), unit-grain conversion or retention arm pair uses an exact independent-binomial risk-ratio method (Berger-Boos restricted-nuisance test inversion; see increment/estimation/binomial_rr.py), not the log-scale delta method. It handles every event count directly, with no log_se >= 0.5 admission rule and no continuity correction:

An Encouragement design’s ITT on a conversion/retention metric keeps this route: the design’s uptake (first-stage compliance) moments are a different random variable over the same units and do not change the ITT’s own sufficient statistics, so they are stripped before this route reads the arm rather than disqualifying it.

ObservationsPoint estimateConfidence set
Both arms have eventsEmpirical ratioFinite two-sided (or one-sided-plus-unbounded-ceiling) set
Treatment has zero eventsExactly -1 (total loss)Finite two-sided set, lower endpoint exactly -1
Control has zero eventsUnavailable (the empirical ratio’s denominator is zero)A typed, point-less confidence set: finite lower endpoint, genuinely unbounded ceiling
Both arms have zero eventsUnavailableThe full relative-lift support, [-1, +inf)

A row with no finite point (LiftEstimate.lift is None) still carries a typed LiftEstimate.binomial_set (a BinomialConfidenceSet: sufficient counts, the frozen nuisance budget, and finite/unbounded endpoints on the relative-lift scale) and answers stat_sig()/p_value() exactly from it — point availability, confidence-set availability, and reportability are three distinct, independently tracked properties; a point-less row is never a failed one. An unbounded ceiling is represented as upper=None on the set, never as a serialized infinity.

This exact method is fixed-horizon: it does not itself prove validity under optional stopping or repeated peeking (see the sequential-inference limitations elsewhere on this page for that separate guarantee). A conversion/retention request with an informative prior uses the established Normal approximation. A registered sequential specification instead uses its declared raw-observation likelihood and matching sequential inversion, subject to the registered sampling and finalized-window contract.

It also has two further boundaries, both refusals rather than silent degradation:

  • Arm size. Each arm is capped at binomial_rr.MAX_ARM_SIZE (4,000,000). This is a compute-resource applicability boundary, not a statistical one: per-call cost keeps growing with arm size past it, so a call with either arm above the cap refuses immediately (estimation.binomial.arm_too_large_for_exact_enumeration) rather than running an increasingly expensive search. The cap is sized to admit every arm size this method is asked to support today; a workload with a genuinely larger arm would need a higher cap or a closed-form/recurrence tail evaluation (out of scope here).
  • Latency. Cost grows with arm size: measured on commodity hardware (two-sided, alpha=0.05, cold), roughly 1.7s at 100,000 per arm and an extrapolated 6-9s at 1,000,000 (MAX_ARM_SIZE) — down from an unoptimized ~13s and ~60s respectively, a measured 7-8x reduction from tightening the boundary-search iteration budget without weakening the certified interval in any tested regime, including rare events. A readout multiplies this across metrics, arms, and breakout cells. There is no opt-out: every eligible unadjusted conversion/retention contrast takes this route. A further speedup to sub-few-second latency at the largest admitted arm size would need a genuinely different tail-evaluation construction (a closed-form or recurrence update between adjacent risk-ratio candidates); none is implemented today. Reduce the number of eligible contrasts in a single call (narrow metrics=, run large-arm breakouts separately) if latency matters more than exactness at your arm sizes.
  • Extreme alpha. The frozen nuisance tail budget passed to the Clopper-Pearson endpoint solver is min(1e-6, alpha / 32). Below 1e-9 (i.e. alpha < 3.2e-8), SciPy’s iterative beta-quantile solver has demonstrated large relative error against an exact oracle for small n in the validated regime, so the endpoint is refused (estimation.binomial.tail_unrepresentable) rather than certified outside that regime. Ordinary use (including FCR-adjusted alpha) is far above this floor.

CUPED and unit-grain ratio-denominator conversion/retention retain their log-scale guards; their sufficient statistics are not raw Bernoulli count pairs. A clustered unadjusted arm instead uses joint relative inference, without the scalar log-SE admission rule. None of those approximate routes inherits exact-binomial validity. The remaining scalar log-path refusal is:

log-scale relative-lift inference is undefined (zero events, or a signed metric
that needs an absolute-scale estimand)

Signed effects require a suitable scale and reference

Section titled “Signed effects require a suitable scale and reference”

Prior-free adjusted and clustered unadjusted routes preserve additive inference and signed joint-relative sets. A negative control mean is allowed; a zero one can leave no finite ratio point. Ordinary unclustered log-scale arm lift still requires positive means. Absolute-scale estimation is supported directly by observational adjusted methods, encouragement LATE, switchback contrasts, and sitewide impact APIs; Analysis.run(value_scale="absolute") remains an observational selector rather than a general randomized arm-lift option.

A randomized experiment is not shut out, though. Analysis.sitewide() derives an additive effect directly from enrolled-arm moments, so it is a conditional randomized route for a signed quantity. It is conditional on where the analysis came from: a definitions-backed analysis carries the site-volume evidence it needs, and an artifact-backed one qualifies when the published artifact includes a site-volume evidence extension. It also has its own baseline requirements. Otherwise reach for one of the designs above, or compare a transformed quantity that stays positive.

A quantile’s reported interval depends on the alpha you asked for

Section titled “A quantile’s reported interval depends on the alpha you asked for”

The quantile standard error is derived from an order-statistic bracket evaluated at the caller’s alpha, then divided by the corresponding normal critical value. Change the alpha — which multiplicity allocation does — and the same data and the same quantile yield a different standard error and a different reported interval. This is necessary, not a defect: the bracket must be built at the caller’s own alpha for the reported interval to cover at that alpha; a fixed-reference-alpha SE, rescaled to other alphas, was measured to fail coverage on the same grid that validates this construction.

The p-value used to decide significance does not share this dependence: it is computed once, by inverting this same construction (the smallest alpha at which the construction’s own interval excludes the null), so it is a function of the data alone and does not move when a multiplicity allocation changes the alpha assigned to the row. Do not compare quantile standard errors or reported intervals across analyses that allocated different alpha budgets — they are not on the same scale — but the p-value is comparable across them. Fixed-horizon family selection uses that same inversion rather than reconstructing a p-value from the alpha-dependent standard error. A null that is never excluded has p-value exactly one, so every admissible family threshold leaves it unselected.

A quantile metric’s planned variance is a projection from the pilot, not a certified bound

Section titled “A quantile metric’s planned variance is a projection from the pilot, not a certified bound”

Analysis.planning_baseline(metric) builds a QuantileBaseline from the pilot’s own control-arm values; at each candidate sample size the solver applies the runtime’s own half-width rule to a bracket projected from the pilot — on continuous data the pilot’s classical order-statistic standard error scaled by the asymptotic 1/sqrt(n) rate, on rounded data the pilot’s recorded distribution, averaged over where the quantile falls within its recording cell — since resampling at every candidate size the search visits is not possible from a single pilot draw. The projection is close when the candidate size is a modest multiple of the pilot’s own size; extrapolating far past it, especially toward an extreme quantile where the pilot has few points, inherits that region’s own sampling noise and can be off by a larger margin. This is a property of the SOURCE: a finite pilot is one draw, not a certified estimate of the runtime’s exact variance at an arbitrary future size. A bigger or more representative pilot narrows it; planning a design still sizes an experiment, it does not certify a final result. Measured directly (Monte Carlo, 1500 two-arm draws at the planned size): a 90th-percentile metric with a 5,000-unit pilot, 15% relative lift, and target power 0.80 achieved empirical power as low as 0.73 on one draw — a real, several-point shortfall the deterministic projection does not eliminate, not merely simulation noise.

On rounded data the runtime’s interval never narrows below the log gap from the quantile to the next recorded value, however large the sample. Power therefore has a ceiling below one: required_sample_size refuses a target above it with power.quantile_size_search_unreachable, naming in maximum_power the largest power any size reaches (limiting_condition="recording_grid"). This is a property of the SOURCE’s recording resolution; recording the metric more finely raises the ceiling.

Switchback carryover is assumed away, not tested

Section titled “Switchback carryover is assumed away, not tested”

The switchback contract’s identifying assumption is no_residual_carryover_after_discarded_steps. You declare a washout window, and a SwitchbackWindow.carryover_order (validated 0 <= carryover_order < observation_steps, default 0) names how many additional post-washout steps are still assumed contaminated. The retained window is step >= washout_steps + carryover_order, of length observation_steps - carryover_order. Declaring a nonzero order records the assumption precisely and is applied by from_switchback_panel’s runtime reduction; it does not prove that carryover has vanished.

Nothing verifies that washout_steps + carryover_order was long enough. The library validates the logged schedule, but cannot establish that the scheduler was random or that carryover has actually vanished by the retained window’s start. If residual carryover remains past that point, the contrast can be biased and the interval need not cover.

from_switchback_panel retains step >= washout_steps + carryover_order for declared orders 0, 1, and 2 (whenever enough observation steps remain) under both IndependentBernoulliOrder and SharedScheduleOrder; declaring an order only records the assumption; nothing here proves carryover actually vanished by that point.

Switchback inference requires independent units or blocks

Section titled “Switchback inference requires independent units or blocks”

Under independent unit-cycle orders, the estimator averages all cycles within a labelled unit and applies a t test with n_units - 1 degrees of freedom. This accommodates arbitrary serial correlation within a unit, which is the point of the design. It assumes the units are independent of each other, and that each unit’s order was drawn independently.

Several common patterns bear on that assumption. What each one does to the reported interval depends on the randomization law and on which target you are inferring about, and none of it is characterized here:

  • Correlated schedules. If unit orders are not drawn independently — a shared seed, or a balanced or stratified schedule — independent-unit inference is not valid. The direction of the error depends on both the randomization law and how units’ potential outcomes respond to order, not the law alone: when units respond to order in the same direction, shared orders tend to induce positive covariance between contributions and balanced schedules tend to induce negative covariance (leaving the independent-unit standard error conservative); when units respond to order in opposite directions, these signs can reverse.
  • Shared time shocks. Conditional on the outcomes actually realized, an additive shock shared across units does not couple contributions, because the only randomness is each unit’s independent order draw. Whether such a shock matters for a target beyond the periods you ran depends on the mechanism: what makes that broader target random is a treatment effect that varies with the shock, not the shock level on its own. The library does not state which target it addresses.
  • Interference. If one unit’s assignment changes another unit’s outcome, the estimator is no longer targeting the intended contrast, so the point estimate can be biased. This is fundamentally an identification failure, and there is no exposure-mapping facility to address it — but because one unit’s outcome now depends on another unit’s assignment, it can also couple unit contributions and invalidate the independent-unit standard error, so the reported interval is not reliably conservative either.
  • Geographic or network correlation of outcomes, without interference. Nearby units having similar outcomes is not the same as one unit affecting another. On its own it does not bias the point estimate; as with a shared shock, whether it matters depends on the target.

A SharedScheduleOrder sequence declares a coarser randomization: one shared Bernoulli CT/TC order per two-period block, drawn once for a fixed, complete unit roster rather than independently per unit. from_switchback_panel reduces it with a block-level Student-t reference (n_blocks - 1 degrees of freedom): the independent replicate is the block, not the roster unit, so a uniformly cloned roster changes neither the estimate nor its width.

Both laws target the mean-unit retained-window additive effect. Results record observation_steps and retained_steps=observation_steps-carryover_order; the total/conversion estimand labels explicitly name that retained window. A schedule integrity check cannot establish either independent randomization or absence of carryover. The original degenerate all-CT witness (ct=320, tc=0 under a unit-cycle declaration) now refuses before statistics are constructed. Earlier refusal of unsupported shared declarations did not prevent de facto shared data from entering the unit-cycle path. Correctly declared, nondegenerate shared schedules are supported.

Sequential information fractions are unit counts, not Fisher information

Section titled “Sequential information fractions are unit counts, not Fisher information”

Alpha-spending boundaries are computed at information fractions taken as cumulative analyzed-unit fractions of the planned maximum. That equals the true information fraction only when information is proportional to analyzed count: stable allocation, stable variance, comparable independent increments.

Under drifting allocation or drifting variance the spending schedule is mis-timed, which can over- or under-spend error — in exactly the non-stationary conditions that motivate continuous monitoring in the first place.

The sequential certification campaign has not been executed

Section titled “The sequential certification campaign has not been executed”

The public Bernoulli likelihood route has a finite-sample anytime-valid derivation under its registered model, assignment and reveal assumptions. That argument and verification of the deployed numerical implementation are separate obligations. Continuous-mean, adjusted-mean and ratio sequential routes carry asymptotic qualifications; they do not inherit the Bernoulli finite-sample guarantee.

The complete empirical campaign over the registered runtime manifest has not been executed. Its recorded cost is months on twelve cores even after the measured speedups; budgeted profiles deliberately cover only subsets. The presence of this machinery is a specification of the check, not evidence that its cells passed. Complete clustered and unit-cycle research matrices likewise remain uncertified.

Release review for public sequential and difficult-cluster methods instead uses a bounded claim audit, independent numerical oracles and preselected risk-focused calibration, with unchanged per-case thresholds. Those checks can identify concrete failures and establish evidence for their stated cases; they cannot establish universal calibration, resolve unrun regimes, or turn small-cluster/asymptotic approximations into finite-sample guarantees. Incomplete or insufficiently precise checks remain unverified. This narrower release evidence requirement does not claim completion of the full campaigns or change the separate unit-cycle research obligations.

Off-policy evaluation of logged decisions is calibrated only under a fixed logger

Section titled “Off-policy evaluation of logged decisions is calibrated only under a fixed logger”

estimate_policy_contrast evaluates a target policy against a reference policy from a logged decision trace with a trajectory-level self-normalized importance-weighted estimator, cumulative likelihood ratios, and a unit-clustered sandwich interval on a t reference. The interval is asymptotic in the number of units. It runs on no Analysis path (SOURCE: none carries a decision trace); its ingress is LoggedTrace.from_records / from_frame, and the guide states the estimand.

Admission requires every positive registered or recorded logging probability, chosen or not, to meet the declared 0.05 floor (logged_policy.trace.propensity_below_floor). Exact zero support is allowed only where the evaluated policies also assign zero mass. A sub-floor positive probability violates the promised importance-weight bound; it does not imply mathematically unbounded weights. The recorded-law constructor is an alternative to registry reconstruction, not a certificate of logger truthfulness or permission to relabel adaptive data.

A trace logged under more than one policy version is refused (CONSTRUCTION, logged_policy.inference.adaptive_logging_unsupported). Measured before that gate existed, with the same estimator on batch-refit adaptive loggers at horizon 5 under context dependence 0.8 (tests/calibration/test_logged_policy_adaptive.py’s design, 400 replications at 200 units): an epsilon-greedy logger covered 86.4%, with the estimate’s bias within its Monte-Carlo error but the sandwich standard error 27% smaller than the estimate’s realized spread; a Thompson-sampling logger covered 89.6% over the 67% of replications its own propensities kept above the floor; and a contextual Thompson logger at horizon 2, 1000 units, covered 87.2% over the 24% of replications that survived the floor, with a bias of seven Monte-Carlo standard errors in that surviving slice. The variance understatement comes from units coupled through the batch refits and from heavy-tailed cumulative weights; the surviving-slice bias is selection on the outcomes that drove the propensities down. The construction that repairs both (stabilized or adaptively weighted estimators with a martingale variance; Hadad et al. 2021, Zhang, Janson and Murphy 2021) is not implemented, so the refusal names one policy version per trace as the route forward.

Under a fixed logger the interval is nominal from the effective-sample-size floor up. Measured on 26 fixed-logger cells over target divergence, horizon and unit count (9188 emitted intervals; tests/calibration/test_logged_policy_ess_floor.py), coverage by band of the smallest per-index ESS was 0.862 / 0.892 / 0.933 / 0.933 / 0.943 / 0.947 / 0.949 for [2,5), [5,10), [10,20), [20,40), [40,80), [80,160) and [160,inf), with 700 to 2190 intervals per band; [20,40) is the smallest band from which every band is within three Monte-Carlo standard errors of nominal, so ESS_FLOOR = 20 and, since the ESS never exceeds the unit count, the independent-unit floor is 20 as well. The two bands just above the floor sit 1 to 2 points under nominal — within tolerance, not at it — because the sandwich standard error is about 7% too small there; from 80 effective units on it is within 1 point. On the nine fixed-logger cells of the adaptive study (T in {1, 2, 5}, 50 to 1000 units, 160 to 800 replications each) coverage ranged from 0.931 to 0.960, every cell inside its own three-MCSE band; at horizon 5 with 100 units the floor screened 126 of 600 replications and the emitted intervals covered 0.941.

The cumulative-product weighting is what makes the estimator target the dynamic policy value. Measured against a current-step-only comparator at horizon 5, 1000 units, under context dependence 0.8 (truth 0.0275): the current-step estimator is biased by +0.0026, more than three Monte-Carlo standard errors from zero, while the shipped estimator’s bias of +0.0006 stays within its own.

Arm power planning depends on declared alternative-arm variance shapes

Section titled “Arm power planning depends on declared alternative-arm variance shapes”

The three arm solvers evaluate the treatment and control terms at their own means under the alternative. Starting from the control-arm effective variance v, mean-like metrics assume equal absolute variance in the two arms (v_treatment = v). Conversion and retention metrics that the runtime does not decide with the exact binomial test (CUPED, clustered, absorbed-factor, sequential) instead rescale v by the Bernoulli shape at the alternative rate: v_treatment = v * p_treatment * (1 - p_treatment) / (p_control * (1 - p_control)). These are explicit planning assumptions (power_basis="asymptotic"). They do not guarantee that a future data-generating process has either variance shape.

Conversion planning matches the exact binomial decision only within a budget

Section titled “Conversion planning matches the exact binomial decision only within a budget”

Unadjusted, unclustered, fixed-horizon conversion and retention plans report the rejection probability of the runtime’s exact binomial risk-ratio decision. With at most 16,000 retained (control, treatment) cells at the null rate the decision set is replayed exactly (power_basis="exact"); beyond that the replay uses Normal conditional tails (power_basis="approximate"), measured at up to 0.8 percentage points below the runtime’s power (unequal allocation, shifted null) and able to misclassify rare-event count pairs near the tail allocation. The runtime’s own control-arm window and floating-point allowance carry over to both routes. A binomial plan’s call takes seconds rather than milliseconds. Measured on an Apple M3 Pro: exact route at 701 per arm, about 0.8 s for achieved power or MDE and 1.9 s for sizing; approximate route at 10,000 / 20,000 / 35,000 per arm (10% baseline), 1.2 / 2.6 / 3.9 s for achieved power and 6.3 / 15.5 / 22 s for sizing, and at 20,000 per arm with a 50% baseline 6.4 s and 67 s. Every count pair’s decision comes from its own replay, except counts the replay’s first step would settle: those are inferred from a neighbouring count’s margin through the step’s monotonicity in the treatment count, which assumes each computed tail lies within 5e-11 of its exact-arithmetic value. Sizing returns a verified bracket crossing, not a proven global minimum; effect searches exclude earlier effects to within 2e-12 of the target power. Triggered plans use the rounded analyzed counts.

Bounded-metric baselines and implied null/alternative rates must stay strictly positive and at most 1. A requested rate above 1 is refused rather than treated as a conservative approximation. At a fixed sample size, the directed noncentrality can peak before the effect-domain boundary, so a target power can be genuinely unattainable. In that case an achieved-power query still reports power for its supplied effect, but its companion mde_relative is null and mde_unavailable_reason states unattainable, unrepresentable, or numerical_resolution.

The segment-pairwise solvers deliberately retain their baseline-only four-arm approximation. Their numbers should not be compared with the changed arm trio as though the two models were identical.

Sequential MDE inversion recomputes the alternative-arm standard error and boundary at each candidate. Its crossing calculation uses composite eight-point Gauss–Legendre panels. Classical derivative-remainder bounds are propagated through every surviving-density and exit-mass integral; floating arithmetic and normal-tail allowances are added separately. Expected information has its own propagated bound. Observed differences between quadrature levels are diagnostics, not the basis of the enclosure.

Public scalar power and expected-information values require their respective enclosures to meet independent convergence tolerances. An inverse that cannot distinguish the target within the declared quadrature, interval, or resource limits refuses with power.minimum_detectable_effect.numerical_resolution instead of consuming an unresolved midpoint or claiming an effect or unattainability. Information fractions remain the analyzed-unit-count approximation described above.

Switchback planning requires its own model

Section titled “Switchback planning requires its own model”

Ordinary switchback planning derives its reference effect and centered contribution/order-slope covariance from a validated pilot, at the independent unit or shared-block grain. Moment-t power conditions on those fitted estimates; it does not account for their estimation uncertainty or certify finite-sample calibration. Optional prospective envelopes and complete-law research oracles retain separate lower-bound and Monte Carlo meanings. Existing small-sample failures and the unresolved computational-certification work remain; see the switchback guide. Parallel-arm Baseline.icc does not describe within-unit switchback dependence.

Increment checks overlap and post-adjustment balance. It cannot validate conditional ignorability, and no diagnostic on this page substitutes for a defensible identification argument.

Categorical strings are internally encoded with a modal reference and one indicator per remaining training level. High-cardinality columns therefore produce many columns; no automatic pooling or cohort trimming is applied. Encoding and per-level balance diagnostics do not establish overlap. Definitions and artifact ingress preserve their default missing-value refusal for numeric and categorical nulls alike. YAML rejects missing/gate tuning; other policies use a full Observational design on the two dataframe routes.

DML reports a partially-linear slope, not an average treatment effect

Section titled “DML reports a partially-linear slope, not an average treatment effect”

The double machine-learning estimator identifies the slope of the partially linear model, which weights conditional effects by the conditional variance of treatment. That equals the average treatment effect only under homogeneous effects or a constant weight (e(1-e) with one treatment, p_a p_0 / (p_a + p_0) with several). The result labels its estimand plr_slope and carries a note saying so.

A relative DML row divides that slope by the augmented control-arm mean of the analysed population, with their joint influence covariance; it is not mu1 / mu0 - 1. With several treatments, each row keeps its own comparison-specific slope over one shared control mean. Do not read a DML relative lift as a population ATE when you expect effect heterogeneity, and do not compare it against IPTW or AIPW as though all three targeted the same quantity. Its interval is asymptotic: finite-sample calibration of multi-arm and small-cluster designs is not established.

Only untrimmed native-logistic IPTW accounts for fitting its propensity

Section titled “Only untrimmed native-logistic IPTW accounts for fitting its propensity”

The propensity model is fit on the same data used for estimation. Untrimmed IPTW with the default pooled logistic learner includes those estimating equations, and with several treatments the equations of every treatment-versus-control model that shares the control. Generic, pattern-specific and trimmed fits still treat the fitted propensity as known, so their standard error omits the propensity-estimation contribution, with no guarantee of being conservative once trimming is applied. High-capacity or custom learners can overfit assignment and destabilize the weights.

The multi-arm construction couples pairwise treatment-versus-control fits into marginal arm propensities. With the default logistic learner this is consistent for a multinomial logit propensity, but it is not that model’s joint maximum-likelihood fit.

Trimming changes which population you estimated

Section titled “Trimming changes which population you estimated”

When overlap trimming removes units, IPTW and AIPW relabel the estimand to overlap_subpopulation_ate and record the retained population on the result, so the change of target is visible. The standard error conditions on that fitted retained set: it does not include the variability of the trim boundary itself, which depends on estimated propensities.

DML reports plr_slope whether or not trimming occurred, so a DML result alone does not tell you that the population changed. Check the recorded population.

The positivity gate is a fixed threshold with no weight diagnostics

Section titled “The positivity gate is a fixed threshold with no weight diagnostics”

The default identification gate requires propensities within [0.01, 0.99], which caps a raw inverse weight at 100. Passing that gate is not evidence of good overlap. The result reports no effective sample size, no leverage, no weight-tail concentration, and no sensitivity across thresholds.

A handful of near-boundary units can dominate an IPTW or AIPW estimate while the gate still passes, producing an estimate that looks precise and is not stable. If you are relying on weighting, compute weight diagnostics yourself.

Fixed-horizon FDR control assumes a dependence condition that is not checked

Section titled “Fixed-horizon FDR control assumes a dependence condition that is not checked”

Selection across the metric × arm × segment family uses ordinary Benjamini–Hochberg, which controls FDR at the stated level under independence or positive dependence (PRDS). An exploratory family sharing a control arm and correlated outcomes can violate that condition, and there is no Benjamini–Yekutieli option at selection time.

This applies to the fixed-horizon path. Registered sequential inference uses current or explicitly frozen likelihood e-values, whose family validity requires the declared common joint-unit filtration. Distributional assumptions are pre-data declarations; timestamps and observed means do not establish them.

Sequential selected intervals use the same stopped likelihood

Section titled “Sequential selected intervals use the same stopped likelihood”

Sequential families retain every registered cell, including missing or abstaining cells. Selected intervals are reinverted at min(q * R / m, nominal alpha) on the checkpoint used for selection. This selected-set statement is distinct from coverage conditional on selecting one particular metric. Unbounded, full, empty, and numerically abstaining sets remain explicit in the result.

Stopping-date guarantees require the registered reveal contract

Section titled “Stopping-date guarantees require the registered reveal contract”

Native, frame-panel and artifact monitoring require explicit finalized capture. All relevant metrics must be revealed together after their longest required window, in outcome-independent unit order. Previously finalized values and assignments cannot be corrected during continuation. A changed generation or freshness hash is insufficient proof of append. Current as-of readouts expose the retained labeled checkpoint; historical curves require retained per-date checkpoints.

The finite-sample public route uses raw Bernoulli/Beta likelihoods and matching inversion. The public mean, adjusted-mean and ratio routes instead provide asymptotic confidence sequences and family approximations; estimated variance does not preserve the exact e-value martingale argument. Scalar Gaussian NIG and paired-Gaussian NIW kernels are private research diagnostics and do not produce public anytime-valid evidence or decision rows. The filtration qualification is the declared common, outcome-independent joint-unit reveal order; a timestamp or observed mean does not establish it. Primary references are Ville’s maximal inequality and Howard et al.’s time-uniform concentration framework. The bounded release evidence review is separate from these mathematical claims; the full sequential research matrix remains uncertified.

The union across dates is not controlled. “This metric was flagged at least once this week” has inflated FDR and no stated guarantee, and neither does picking the date with the most flags. Act on the current date’s discovery set; do not accumulate flags across days.

Clustered CATE uncertainty is cluster-asymptotic

Section titled “Clustered CATE uncertainty is cluster-asymptotic”

Clustered estimate_cate uses a weighted CR1/HC0 score sandwich and K-1 reference degrees of freedom, with every declared observed cluster counted. It does not implement CR2. Small K, imbalance and high leverage still require scientific calibration; matching the intended sandwich calculation does not establish coverage.

Cluster ARD uses trace(G @ V_JJ) / q as its noise variance, with G the weighted residualized interaction Gram matrix. This is a directional-average precision approximation, not a full correlated-likelihood ARD model. Unclustered ARD retains its existing residual-variance calculation.

Uniformly cloning rows within the same cluster IDs preserves fixed-design point estimates, covariance and cluster ARD precision. Assigning new IDs to the copies declares new independent clusters: under complete doubling the specified sandwich covariance is multiplied by (K-1)/(2*K-1). Invariance to that operation is incompatible with the specified sandwich and literal cluster count; copies that remain dependent must retain their original cluster identity. Learned continuous bases can also change when refitted on duplicated rows.

Exact duplication checks therefore freeze scores and nuisance fits and replicate the entire within-cluster row pattern. A fully refitted pipeline, or a roster-size change that also changes member noise, needs separate behavioral and Monte Carlo evidence. Neither an unchanged independent cluster count nor one replayed p-value establishes null calibration. The existing unclustered calibration figures do not certify clustered AUTOC, Qini, GATES, CLAN or selected policy values.

Targeting validation depends on identification and overlap

Section titled “Targeting validation depends on identification and overlap”

With an Observational design, validate_cate, targeting_rule, and select_targeting_rule use doubly-robust scores; Randomized keeps IPW scores. Observational GATES and the holdout ATE summarize score means, not raw-outcome Welch contrasts; identification still requires conditional ignorability and overlap. When holdout trimming removes units, CateValidation.population == "overlap_subpopulation": AUTOC, Qini, and GATES describe that retained population, not the full one. [estimate_cate][increment.estimate_cate] remains randomized-only (cate.identification.randomized_only); its per-unit point predictions are not doubly robust.

Declared clusters retain unit-level scores and outcomes. cluster_weight="member_count" weights members equally; "equal" gives each retained cluster total weight one and requires cluster IDs. Smooth means use centered cluster contributions with the K/(K-1) scale: disjoint randomized arms use their own cluster counts and Welch reference degrees, overlapping contrasts combine signed contributions within each cluster, and observational DR scores use pooled cluster influence.

Clustered DR nuisance models are fitted once on the training half, with frozen heldout predictions. Ranks, empirical GATES/CLAN cutoffs, and policy values are recomputed in a whole-cluster bootstrap (default [ClusterBootstrap(seed=0, repetitions=999)][increment.ClusterBootstrap]). Randomized pure-arm clusters are resampled within arms; pooled DR and mixed-arm clusters have no arm stratum. Repeated cluster draws are separate sampled instances. Weighted empirical cutoffs keep tied scores together; the clustered AUTOC/Qini integrate the empirical weighted TOC with uniform interpolation within each tied block, making uniform row cloning invariant. Unclustered rank calculations retain their existing discrete definition.

Clustered intervals and tests use bootstrap-t+cluster-jackknife-t. The interval is the smallest interval containing both component intervals; the one-sided p-value is the larger component p-value. Thus a rejection requires both tests to reject, and the reported interval cannot be narrower than either component. This preserves a component’s coverage when that component is valid; it does not prove finite-cluster validity of either component.

For the bootstrap component, use the complete statistic (T), its bootstrap draw (T_b^), and positive covariance reference scales (s,s_b^) to form (R_b^=(T_b^-T)/s_b^). Centering at (T) retains the bootstrap estimate of ratio and cutoff bias; centering at the bootstrap mean would erase it. Dividing by each draw’s own scale retains the dependence between estimation error and uncertainty, which matters for skewed cluster contributions and unequal cluster masses. Even for an ordinary mean of (K) IID clusters, the raw empirical-bootstrap variance is ((K-1)/K) times the usual unbiased variance estimate; more repetitions do not remove that finite-(K) difference. For nonlinear ratios, replicates that omit a large cluster can also have both a large estimation error and a much smaller reference scale. Intervals invert the root tails as ([T-sR^{\mathrm{upper}},T-sR^*{\mathrm{lower}}]); one-sided rank tests compare these same roots with (T/s). The reported se is the standard deviation of the complete bootstrap statistics, including rank and cutoff variation, rather than the conditional reference scale used inside the roots.

For GATES, CLAN, and policy values, the reference is the centered ratio-mean covariance described above, recomputed for each draw’s groups. For ranks, let a tied score block occupy probability interval ((a,b]). Integrating the response weights gives (r_A=(-b\log b+a\log a)/(b-a)) for AUTOC (with (0\log0=0)) and (r_Q=(1-a-b)/2) for Qini. Holding these response weights fixed, the functional (E[r\psi]-E[r]E[\psi]) has influence ((r-E[r])(\psi-E[\psi])-\operatorname{Cov}(r,\psi)). The reference scale sums these signed, member-weighted influences inside each cluster before squaring, with the (K/(K-1)) correction. It is a conditional response covariance, not a claim that estimated ranks or cutoffs are fixed. Inner policy selection similarly uses the cluster covariance of each member’s targeted net-benefit contribution as its reference. All deployed statistics and reference scales are recomputed on every whole-cluster draw; there is no nested bootstrap or extra rank sort for the rank reference.

The second component deletes each observed cluster once and recomputes the complete statistic (T_{(-g)}), including ranks, cutoffs, groups, weights, and policy budgets. It uses the full-estimate-centered cluster jackknife scale [ s_J^2=\frac{K-1}{K}\sum_{g=1}^K (T_{(-g)}-T)^2. ] This is the CV3 construction in MacKinnon, Nielsen and Webb, equation (18). That paper treats regression; applying the construction here requires regularity of the complete nonlinear statistic. For a ratio mean with cluster target-mass share (h_g) and signed contribution (u_g), the exact deletion identity is (T-T_{(-g)}=u_g/(1-h_g)). This captures observed leverage without a fitted multiplier. At equal masses it reduces to the usual cluster-mean standard error. For shared-cluster contrasts, deletion differences combine both sides before squaring, preserving their signed covariance. Estimated cuts and ranks are recomputed, so their deletion changes also enter the scale.

The jackknife interval is (T\pm t_{\nu,\alpha/(2m)}s_J), where (m) is the existing family size and (t) denotes an upper-tail quantile. Its test uses the Student-t survival probability at (T/s_J). The reference has (\nu=\min_h(K_h-1)), using the bootstrap arm/fold strata; without strata it is (K-1). Deletions retain all remaining clusters and their declared target weights, without compensating other clusters in the omitted cluster’s stratum. This unrestricted deletion scale can include composition variation absent from the stratified bootstrap. The smaller stratum reference is conservative relative to using (K-1), but is not an exact small-sample law. This adds (K) full evaluations and no random draws; the bootstrap stream and repetition count are unchanged. reference_df on rank/group/CLAN results describes this jackknife component, while covariance_method describes the bootstrap response reference.

The bootstrap component has a first-order asymptotic justification when the complete cluster bootstrap is consistent and (\sqrt K s) and (\sqrt K s^*) converge to the same positive limit: Slutsky’s theorem then gives the same limiting law for the observed and bootstrap roots. The reference need not equal the full rank/cutoff asymptotic standard error for this argument. It does not establish a higher-order accuracy improvement for these nonlinear statistics. Jackknife-t also requires an asymptotically linear statistic and consistent deletion variance; it is not generally valid for a nonsmooth quantile statistic. Independent clusters, adequate moments, no dominating cluster, and regular population quantiles are needed; stable tied blocks can be evaluated, but a quantile exactly at a jump boundary can be nonregular. The AUTOC logarithmic tail additionally needs an integrable squared response-weight envelope. Training-only nuisance predictions remain frozen. Neither this argument nor a calibration screen establishes distribution-free finite-cluster coverage, particularly with few skewed, high-leverage clusters. See the bootstrap-t inversion and RATE asymptotics for the underlying methods; the latter does not by itself certify this clustered implementation.

Measured behaviour on a null screen with frozen oracle scores (constant effect, independent CLAN covariate), 256 replications and 199 draws per cell: every one of 72 designs stays within its family-aware binomial miss bound for AUTOC, Qini, each GATES interval, the joint GATES family and CLAN, and within the rejection bound for the AUTOC and Qini tests. The designs cross 10, 40 or 200 clusters, equal sizes of 20 or repeating sizes 5/20/100, intracluster correlation 0, 0.2 or 0.5, Gaussian or skewed errors with one high-leverage score per cluster, and both weightings. The tightest design has 10 unequal clusters, skewed errors, member weighting and correlation 0.5: each GATES interval missed 18 of 256 times against a nominal 2.5% (6.4 expected, bound 19). Few unequal clusters therefore still undercover; the screen is a regression check, not a guarantee.

GATES and CLAN retain separate Bonferroni budgets. Tail indices use the exact binary64 input alpha divided rationally by twice the family size, rounded outward on the (B+1) order-statistic grid. Monte Carlo p-values also round upward. A failed bootstrap statistic makes uncertainty unavailable instead of conditioning on successful draws. A finite draw whose reference scale is nonpositive or nonfinite uses the observed positive reference scale in its root denominator. This fallback retains the complete draw; it alone supplies no coverage correction. bootstrap_valid_repetitions counts finite complete bootstrap statistics, independently of their scale availability and of jackknife evaluations. Repeated draws of one source cluster are still distinct instances. The observed response and deletion scales must be positive and finite. An undefined deletion reports estimation.targeting.jackknife_unavailable_replicate; it is never dropped from the sum or replaced by a successful deletion. Small cluster counts, empty groups, constant rankings, zero variance, and inadequate tail resolution retain nullable fields and explicit unavailable_reason codes. Results record cluster count, weighting, method, seed, and requested and valid repetitions. Increasing repetitions resolves smaller tails; it does not remedy insufficient independent clusters.

Deployment grain follows the explicitly declared intervention, independently of uncertainty clustering. A cluster intervention cannot request unit deployment (estimation.targeting.unsupported_unit_deployment). Cluster policies pool member scores and select a score-ordered whole-cluster prefix, breaking ties by canonical ID. Member/equal weighting determines both the budget and the reported population. This prefix is feasible, not a knapsack optimum: unused budget stays unused when the next cluster does not fit. Empty prefixes have no conditional policy effect; zero and one fractions select nobody and everybody. Bootstrap replicates recompute the pooled ranking and budget, counting repeated draws as separate cluster instances while preserving canonical source IDs for score ties. Overlap trimming that retains only part of a cluster is refused for cluster deployment (estimation.targeting.overlap.partial_cluster).

Saved policy JSON retains its fitted scoring basis and coefficients. predict(cols, cluster_ids=...) returns the candidate actions; consult recommendation before deployment. Cluster prediction needs a complete member roster for each deployment cluster and computes the budget over the supplied batch. Unit prediction retains its fitted threshold even with dependence IDs. Neither new-data outcomes nor nuisance refitting enter prediction.

Fraction selection freezes all inner-fold scores before resampling within folds (and arms for randomized designs). The selected inner winner and a policy value reported after the outer validation gate still have no nominal confidence interval; the bootstrap does not remove this selection restriction. A declared cluster intervention defaults to whole-cluster prefix deployment and rejects explicit unit deployment.

The targeting workflow carries no joint error control

Section titled “The targeting workflow carries no joint error control”

Within the targeting workflow, group-effect tests are Bonferroni-corrected inside their own family, and characteristic tests inside theirs. The rank-based tests are not corrected alongside them, and no guarantee spans the whole workflow.

If you scan every interval and test the workflow displays and act on whichever looks strongest, you are not protected at your stated alpha. The per-family scope is real but easy to over-read.

Meta-analysis can be anticonservative at small K

Section titled “Meta-analysis can be anticonservative at small K”

Pooling uses a DerSimonian–Laird heterogeneity estimate followed by an unmodified Hartung–Knapp–Sidik–Jonkman adjustment, with no variance floor. Measured by tests/calibration/test_hksj_small_k.py (10000 replications per cell, unequal segment variances with log-scale standard errors 0.10–0.20, between-segment variance 0, 0.5 and 2 times the average sampling variance; binomial Monte-Carlo standard error 0.2 points at 95% and 0.4 points at 80%): the HKSJ interval covers the random-effects mean 93.7–95.2% at nominal 95% for every K from 2 to 12 and every heterogeneity level. The plug-in random-effects interval covers 80.0% at K = 2, 85.2% at K = 3, 87.2% at K = 4, 88.6% at K = 5, 90.9% at K = 8 and 92.2% at K = 12 when the between-segment variance is twice the average sampling variance, and over-covers (96.0–96.4%) when there is none. HKSJ is the narrower of the two in 12–35% of replications without heterogeneity and 2–13% with it. The degeneracy fallback to the plug-in interval fired in 1 of 10000 replications in two of the K = 2 cells and never otherwise on continuous data. The anticonservatism at small K is real but bounded at about 1.3 points in this design; treat a pooled interval over two or three segments as indicative for that reason, not because it is narrower than plug-in.

Bayesian segment shrinkage places a half-normal prior on the heterogeneity scale, with a configurable default scale of 0.30. This prior remains consequential: a scale suitable for log relative effects may strongly shrink effects measured in dollars or other absolute units. Adaptive integration now resolves posterior mass beyond the old fixed grid for both shrinkage and rollout pricing; it does not remove prior sensitivity. The reported Wald intervals use posterior means and variances, rather than posterior quantiles; they carry no finite-sample frequentist coverage guarantee.

For segment-specific inference, prefer the marginalized segment intervals, which integrate over the heterogeneity posterior instead of plugging in a point estimate. They are a better per-segment answer; they are not a substitute for the pooled interval, which answers a different question.

The read-only SQL gate is a screen, not a sandbox

Section titled “The read-only SQL gate is a screen, not a sandbox”

Every SQL string in a definitions file is admitted as exactly one read-only query. The check parses the string with the declared (or the connection’s) dialect and rejects:

  • more than one statement (definition.sql.statement_count);
  • any statement that is not a query, including INSERT, UPDATE, DELETE, DDL, and DML/DDL nested under a WITH;
  • SELECT ... INTO, FOR UPDATE/FOR SHARE, LOCK IN SHARE MODE, and locking table hints such as T-SQL WITH (UPDLOCK), SERIALIZABLE and REPEATABLEREAD (definition.sql.not_read_only); benign hints such as NOLOCK, READPAST and INDEX(...) are admitted;
  • calls to a fixed list of known side-effecting functions (advisory locks, sleeps, large-object and file writers, dblink, session and query cancellation, and similar).

That list cannot be complete. A side-effecting function outside it, including any user-defined function, is admitted, because syntactically it is an ordinary SELECT. Increment makes no claim that an unlisted function is safe. The connection and its privileges belong to you.

Treat the gate as protection against a mistake in a definitions file, not as a security boundary against a definitions author you do not trust. Definitions are trusted, governed code: review them like any other code that runs against your warehouse, and never load a definitions file from an untrusted source. Connect with read-only, least-privilege warehouse credentials (no write, DDL, or execute grants beyond what the queries need).

Repository code cannot tell you how a downstream application accepts definitions, or what grants your warehouse role actually holds. Before relying on the gate, verify both yourself: that definitions can be changed only through your review process, that the application never builds definitions from end-user input, and that the warehouse role is read-only on the tables in scope (for example by attempting a write with that role and confirming it fails).