Results (`increment.results`)
Sitewide behavior is exposed through Analysis.sitewide, with output values
available from the receive-only results namespace.
Estimate
Section titled “Estimate”EstimateA number with quantified uncertainty.
lb/ub (when set) represent an interval at level confidence.
open_side names an unbounded endpoint of a one-sided interval.
Parameters
Section titled “Parameters”value : float
Point estimate.
lb, ub : float | None
Interval bounds. Closed intervals set both; open intervals leave the
endpoint named by open_side unset.
open_side : {“lower”, “upper”} | None
Unbounded endpoint for a one-sided interval, or None for a closed
or unavailable interval.
level : float | None
Confidence/credible level, e.g. 0.95. Required iff lb/ub are set.
At extreme alpha values this may round to 1.0; alpha retains the
exact allocated noncoverage budget.
alpha : float | None
Allocated noncoverage budget; level is its nominal complement.
Fixed open endpoints use this full quantile tail. Sequential open
endpoints retain the symmetric parent budget, without a fixed-time
tail interpretation. Attained coverage may be higher than nominal.
Unlike level, it remains recoverable when the confidence level
rounds to 1.0.
log_mean, log_se : float | None
Point estimate and standard error on the natural-log scale. Set by
infer_lift to the raw pre-prior statistics, not the posterior
mean/sd - the two differ under an informative prior.
LiftEstimate
Section titled “LiftEstimate”LiftEstimateA lift estimate, carrying the metadata to identify which (metric, method, group) and its persisted inference reference.
value is exp(mu_n) - 1 for scale="log" estimates (under a
near-flat prior this equals treatment_mean / control_mean - 1), or
mu_n directly for scale="linear" estimates, which already carry
the posterior on the relative scale. The CI is the same posterior’s
quantile, back-transformed to this scale.
The Normal posterior is never persisted: with a stored alpha the
mu/sigma recovery from value/lb/alpha/scale is
accurate up to rounding away from a zero critical value; open fixed rows use raw
statistics to recover uncertainty at alpha=0.5. A closed
alpha-less interval falls back to the level tail, which
is exact at ordinary alphas but loses precision as alpha approaches
the representable floor. A mixture posterior (prior_spec set) is not
recoverable, so it persists the prior instead.
reference_kind persists the reference: “normal”, “t”, “sequential”,
“binomial”, or “confidence_set”. Winsor confidence sets persist raw
construction state and endpoint statuses, expose no posterior, and
reinvert through reintervalize(alpha). Their public
confidence_set.qualification is either
pointwise_asymptotic_model_conditioned_v1 (bootstrap candidate) or
uniform_support_conditioned_v1 (rank method); neither is an
unqualified finite-sample promise. Bootstrap p-values invert stored roots;
rank p-values are alpha when the set excludes the null, else one.
dof retains cluster degrees of freedom, so dof=None does not
imply Normal inference: a Welch reference has its own reference_df.
Relative t p-values use this reference; posterior-derived stats refuse it.
inference labels the interval semantics: “fixed” is a single-look
interval; “always_valid”/“asymptotic_mean” are valid at every look.
value carries no stopping adjustment (winner’s curse); the interval
endpoints are the safe summary. Decision-stat methods refuse on any
non-”fixed” inference.
value_scale="absolute" marks a row whose value is not a
relative lift at all - an encouragement design’s additive LATE, or an
observational metric reported additively.
Clustered unadjusted and prior-free adjusted rows preserve additive uncertainty
independently of relative_confidence_set. The joint set can be disconnected,
unbounded, or present without a finite ratio point. Its reference is a qualified
working approximation; relative_unavailable_reason identifies an indefinite
or unrepresentable covariance without discarding additive output.
JointContrastReference
Section titled “JointContrastReference”JointContrastReferenceWorking joint reference for the additive numerator and denominator.
RelativeConfidenceSet
Section titled “RelativeConfidenceSet”RelativeConfidenceSetOutward-enclosed Fieller set under an explicitly approximate joint reference.
LiftEstimates
Section titled “LiftEstimates”LiftEstimateslist[LiftEstimate] with .to_frame() - :meth:~increment. analysis.Analysis.run’s return type.
BreakoutEstimate
Section titled “BreakoutEstimate”BreakoutEstimateA lift estimate for ONE segment (dimension value) of a breakout.
Carries the same identity fields as LiftEstimate (via
_RowIdentity) plus dimension/dimension_value to identify
the segment and source to disambiguate the rare same-name,
different-source case.
reference_kind/reference_df preserve the source row’s sampling
reference through serialization. A Welch t reference can have dof=None;
it must not be reconstructed as a Normal interval.
run_breakout returns one row per (segment, metric, method,
non-control arm) cell, dense - an unestimable cell still gets a row,
with lift=None and excluded naming why (see
:data:ExclusionReason).
low_reliability flags a real estimate whose control or
treatment arm has fewer than reliability_floor units: estimable,
but the Wald interval’s nominal coverage is not trustworthy there.
Call .to_frame() on the :class:BreakoutEstimates this returns
rather than constructing a plain list[BreakoutEstimate].
BreakoutEstimates
Section titled “BreakoutEstimates”BreakoutEstimateslist[BreakoutEstimate] with .to_frame() - :func:run_breakout’s
(and :meth:~increment.analysis.Analysis.run_breakout’s)
return type.
DailyMetricValue
Section titled “DailyMetricValue”DailyMetricValueOne arm’s absolute metric mean (with CI) for ONE day.
The time-series counterpart to a plain arm-level mean - what plots
on a per-day chart of the raw value, as opposed to the relative
lift between arms (see :class:DailyLiftEstimate). Produced by
:func:run_daily directly from one daily_group_summary row.
dimension/dimension_value/source are populated together
when :func:run_daily is given a dimension, and None
otherwise.
ds_basis says what ds indexes: "calendar" (default) is
an observation date; "cohort" means ds is the unit’s own
exposure date, how a retention series is conventionally indexed.
Call .to_frame() on the :class:DailyMetricValues this function
returns rather than constructing a plain list.
DailyMetricValues
Section titled “DailyMetricValues”DailyMetricValueslist[DailyMetricValue] with .to_frame() - :func:run_daily’s
(and :meth:~increment.analysis.Analysis.run_daily/
run_asof’s) return type.
DailyLiftEstimate
Section titled “DailyLiftEstimate”DailyLiftEstimateA relative lift estimate for ONE day of a daily time series.
Carries the same identity fields as :class:BreakoutEstimate and
:class:~increment.estimation.results.LiftEstimate (via
_RowIdentity) plus ds to identify which day’s slice of
moments produced it.
estimand/value_scale/note mirror :class:BreakoutEstimate’s
fields of the same name.
ds retains calendar dates, numeric day indices, or structured string
labels for as-of frame readouts.
low_reliability marks a real estimate whose control or
treatment arm has fewer than reliability_floor units that day.
dimension/dimension_value/source are populated together
when :func:run_daily_lift is given a dimension, and None
otherwise. ds_basis - see :class:DailyMetricValue.
DailyLiftEstimates
Section titled “DailyLiftEstimates”DailyLiftEstimateslist[DailyLiftEstimate] with .to_frame() - :func:run_daily_lift’s
(and :meth:~increment.analysis.Analysis.run_daily_lift/
run_asof_lift’s) return type.
HeterogeneitySummary
Section titled “HeterogeneitySummary”HeterogeneitySummaryOne row per (metric, method, group_id, dimension, source, estimand, value_scale) grouping key: Cochran’s Q, DerSimonian-Laird
tau^2, Higgins-Thompson I^2, and an HKSJ pooled effect, over the
segments declared for that key.
scale is "relative" (log-RR) or "absolute" (risk
difference) - the two frequently disagree, so both ship.
value_scale is the upstream rows’ own value scale; scale is
which of the two heterogeneity passes this row belongs to.
tau2/i2/i2_lb/i2_ub are None whenever
n_excluded_outcome > 0 for this (key, scale) row: an
outcome-based exclusion biases tau^2, so it is suppressed rather than
reported as trustworthy. q/p_value/pooled are not.
HeterogeneitySummaries
Section titled “HeterogeneitySummaries”HeterogeneitySummarieslist[HeterogeneitySummary] with .to_frame() - see
:class:~increment.breakout.estimates.EstimateList.
SegmentEstimate
Section titled “SegmentEstimate”SegmentEstimateOne row per (segment, scale, estimator) for a HeterogeneitySummary
key: two rows per estimable (segment, scale) (estimator="raw"
and "shrunken"), plus an unavailable pair per excluded segment.
excluded is set from an upstream BreakoutEstimate.excluded, a
live segment unusable on this scale only
(excluded="zero_variance"), or a shrunken row withheld because
that scale’s tau posterior could not be integrated within the
numerical budget (excluded="estimation_failed").
raw reuses the segment’s own estimate (for scale="relative",
BreakoutEstimate.lift verbatim). shrunken is the
tau-marginalised posterior estimate, with its own shrink_k.
baseline is the segment’s control-arm absolute mean, recovered
from abs_diff and the raw log-scale lift.log_mean.
value_scale is the upstream rows’ own value scale; scale is
which of the two heterogeneity passes this row belongs to.
role, discovery, family_axes, family_q, and
family_threshold mirror the source BreakoutEstimate’s fields
of the same name verbatim (see there) — a reader of this frame alone
can otherwise not tell a discovery from a non-discovery, or a
multiplicity-corrected interval from an uncorrected one. Both
estimator rows for a segment carry the same source values.
SegmentEstimates
Section titled “SegmentEstimates”SegmentEstimateslist[SegmentEstimate] with .to_frame() - see
:class:~increment.breakout.estimates.EstimateList.
SegmentRolloutResult
Section titled “SegmentRolloutResult”SegmentRolloutResult:func:segment_rollout_recommendation’s return value - two
independent, row-aligned-by-grouping-key result sets.
RolloutRecommendation
Section titled “RolloutRecommendation”RolloutRecommendationOne row per (metric, method, group_id, dimension, source, estimand, value_scale) grouping key: which of that key’s segments to roll out, and
an honest price on doing so.
UNWEIGHTED, and deliberately so: every value field treats each
usable segment as an equal contributor regardless of exposure, so
policy_value is a plain sum over the rolled-out segments, not
traffic-weighted. An exposure-weighted variant was measured and
dropped - its selection-bias correction did not meet the accuracy
bar the unweighted correction is held to.
recommendation carries the estimator’s verdict verbatim:
"rollout" (corrected value is positive), "no_net_benefit"
(evidence cannot demonstrate positive value), or "refuse" (the
offset guard fired - every value field is withheld).
rollout_cost is the relative-lift break-even actually used.
k/n_excluded_design/n_excluded_outcome count how many of
the declared segments the decision could and could not use, so a
reported value can never silently understate its coverage.
RolloutRecommendations
Section titled “RolloutRecommendations”RolloutRecommendationslist[RolloutRecommendation] with .to_frame() - see
:class:~increment.breakout.estimates.EstimateList.
RolloutSegment
Section titled “RolloutSegment”RolloutSegmentOne row per segment declared for a :class:RolloutRecommendation key -
dense: every segment of a key that ran appears, usable or not.
selected is the estimator’s subset membership on a usable
segment, None on one this call could not use - populated even
behind a "refuse" recommendation as evidence about the
selection rule, not a deployable decision.
excluded is an upstream BreakoutEstimate.excluded, or
"zero_variance" for a live segment whose log-scale statistic or
variance is unusable here - the same tag segment_heterogeneity
uses for its own per-scale drop.
estimand/value_scale identify the upstream rows’ own estimand
and value scale; rows are grouped by both, so a family is priced on
its own or not at all. A real LATE family (additive lift, no
log-scale moments) never clears the pricing gate below.
RolloutSegments
Section titled “RolloutSegments”RolloutSegmentslist[RolloutSegment] with .to_frame() - see
:class:~increment.breakout.estimates.EstimateList.
PowerResult
Section titled “PowerResult”PowerResultResult of a power-analysis solver.
Parameters
Section titled “Parameters”n_per_arm : int
Number of units in the treatment arm.
n_total : int
Total N across both arms.
power : float
Planned power at the computed / given sample size, under the model
named by power_basis. For required_sample_size and
achieved_power it describes the SUPPLIED effect; for
minimum_detectable_effect it describes the returned effect’s
implied absolute alternative, which can exceed the target when the
answer is a domain endpoint.
power_basis : {“asymptotic”, “exact”, “approximate”}
How power was computed. "asymptotic": the log-ratio
Normal/noncentral-t planning model (every plan the runtime does not
decide with the exact binomial risk-ratio test). "exact": the
probability that the runtime’s unchanged exact binomial decision
rejects, at the analyzed integer counts (up to at most about 1e-12
of omitted outer count mass). "approximate": the same decision
replayed with Normal conditional tails, for binomial plans whose
exact geometry exceeds the planning cell budget.
mde_relative : float | None
Minimum detectable relative effect on the complier scale, expressed
RELATIVE TO the declared null: (exp(distance) - 1) / baseline.compliance, where distance is the search’s own
log-scale gap from theta0 = log1p(decision.null_lift). At
compliance=1.0 this is exactly exp(distance) - 1; a lower
compliance rescales it up. To recover the implied ABSOLUTE
alternative, compose against the null rather than adding lifts:
expm1(log1p(decision.null_lift) + log1p(mde_relative * baseline.compliance)) / baseline.compliance. Positive for two-sided and one-sided
“greater”; negative for one-sided “less”. None when no minimum
detectable effect exists at the design’s target power — a valid
supplied-effect answer can still be reported then — with the cause
in mde_unavailable_reason.
mde_unavailable_reason : {“unattainable”, “unrepresentable”, “numerical_resolution”} | None
Why mde_relative is None: the target power exceeds every
admissible alternative’s power (unattainable), the admissible
answer has no float64 representation (unrepresentable), or the
search could not resolve it within its numerical limits
(numerical_resolution). None whenever mde_relative is
available; a missing effect always carries a reason.
effective_var : float
Per-unit variance for asymptotic planning, after the effective decision
method’s CUPED / factor-absorption reduction and cluster design effect.
Sensitivity-only CUPED receives no reduction, so this may differ from
the caller’s Baseline.effective_var. Exact and approximate binomial
plans report it as metadata; their power uses event rates and counts.
For a QuantileBaseline, the per-unit variance its pilot’s standard
error implies (see QuantileBaseline.from_control_values).
n_clusters_per_arm : int | None
Randomization clusters needed in the treatment arm,
ceil(n_per_arm / baseline.avg_cluster_size). None under
unit randomization.
n_clusters_total : int | None
Randomization clusters across both arms, ceiled per arm and
summed (not ceiled once on n_total, since a part-cluster in
each arm costs two whole clusters). None under unit
randomization.
n_triggered_per_arm, n_triggered_total : int | None
Units actually entering a triggered analysis
(n_per_arm/n_total times trigger_rate). None when
no trigger_rate was declared; n_per_arm/n_total
always count assigned units.
expected_n_total : int | None
Expected total N a sequential design stops at under the solved-for
effect: E[T] * n_total, where E[T] charges each look’s
first-boundary-exit mass its own information fraction and the
never-crossing remainder the final planned look (see
increment.power.sequential.sequential_expected_information_fraction).
This is the honest economic counterpart to n_total (the
worst-case, never-stops-early size) — early stopping is the
entire argument for monitoring sequentially. None when no
inference spec was given (a fixed-horizon design always runs
to n_total).
inference_to_declare : InferenceSpec | None
The exact runtime InferenceSpec this plan assumed — pass it to
InferenceSpec (or a YAML inference: block) at runtime
declaration so the runtime’s boundary is tuned from the same N
planning assumed. None for a fixed-horizon result.
planned_metric_name, planned_quantile : str | None, float | None
The metric name and quantile level a QuantileBaseline was
built for (Analysis.planning_baseline(metric)), echoed here so
a caller who reuses one metric’s baseline to plan a different
metric sees the mismatch stated in the answer instead of
discovering it, if at all, from a silently mis-sized design.
None for every non-quantile metric.
PowerCurvePoint
Section titled “PowerCurvePoint”PowerCurvePointOne evaluated point in a power or MDE curve.
mde_relative is None with mde_unavailable_reason set when
no minimum detectable effect exists at that row’s size and target
(see PowerResult); frames keep the column numeric with a null.
power_basis names the planning model behind power and
mde_relative, as on PowerResult.
PowerCurve
Section titled “PowerCurve”PowerCurveList-like power-curve result with dict and dataframe conversion.
SRMResult
Section titled “SRMResult”SRMResultResult of an allocation sample-ratio-mismatch check.
fixed_p_value is the ordinary Pearson fixed-look p-value.
log_e_value is the current uniform-Dirichlet mixture evidence
for a predeclared allocation (unavailable when fixed inference
infers equal shares from observed arms). inference states
which quantity controls is_srm.
unassigned_units and mixed_assignment_units are accounting
entries, not arms: they contribute no chi-square degree of freedom
and appear in neither observed nor expected.
mixed_assignment_units counts units observed in more than one
arm, which shrinks every arm’s count symmetrically and so must be
surfaced separately to be seen.
grain is "cluster" when randomization happened over
clusters, so the chi-square belongs there; unit_counts then
carries per-arm unit counts as descriptive context only (cluster
size imbalance earns no degree of freedom). At grain="unit",
unit_counts is empty.
min_expected_count is the smallest per-arm expected count
(the minimum over arms of expected[k] * total), or None when
unavailable. low_expected_count is True when that minimum falls
below 5, the standard Cochran rule of thumb, flagging that the fixed
asymptotic chi-square p-value may be unreliable there.
AllocationBand
Section titled “AllocationBand”AllocationBandA Beta posterior credible band on one arm’s allocation share at one ds.
Purely descriptive: unlike SRMResult, this carries no pass/fail
flag - sample_ratio_mismatch is the actual gate. posterior_a/
posterior_b are the Beta parameters the interval was derived
from, carried so a caller can reconstruct the full posterior (they
cannot be recovered from n/n_total alone without the prior,
a call-site argument).
NotApplicable
Section titled “NotApplicable”NotApplicableA diagnostic check that does not apply to the current design.
Returned in place of a check’s normal result (e.g. SRMResult)
when the design makes the check meaningless - a sample-ratio test
presumes a target randomized allocation an observational design
does not have.
AbsorptionResult
Section titled “AbsorptionResult”AbsorptionResultThe absorbed average treatment effect and its diagnostics.
effect is on the absolute scale (treated mean minus control
mean, factor absorbed). When the effect is homogeneous across
levels this is the ATE; when it varies, effect converges to
the precision-weighted average of per-level effects (weights
n_t * n_c / n_g, Angrist 1998), down-weighting skewed-allocation
levels relative to the unit-weighted ATE.
icc is always the estimated variance-component ratio;
mean_shrinkage is the pooling weight actually applied, pinned
to 0.0 or 1.0 when pooling is forced rather than estimated - the
two can disagree when pooling is not "partial".
ClusterScore
Section titled “ClusterScore”ClusterScoreA pooled score keyed by canonical cluster identity, in canonical ID order.
CateScoreState
Section titled “CateScoreState”CateScoreStateImmutable portable scoring basis, fitted centering and effect coefficients.
CateResult
Section titled “CateResult”CateResultA fitted Lin (2013) interacted regression.
ate is the treatment coefficient at the design centre (the
weighted mean of the fitted treatment effects); se is
its sandwich SE. Without clusters, se_unadjusted is Welch’s SE;
with clusters it uses the same weights and cluster sandwich fitted on
[1, d]. se_reduction reads the width the adjustment bought.
dimension, n_clusters (all observed IDs), immutable vcov and
reference_df are the stored inference contract used by projections.
Cluster intervals use t(K-1); the interaction Wald quadratic uses
F(q, K-1) after division by q. The reference is cluster-asymptotic.
unadjusted_vcov stores the separate two-column fit’s covariance.
Unavailable required uncertainty refuses rather than returning a partial fit.
The state behind cate/contrast/score lives on excluded
fields: a dumped result is report-only, so those need the live
object. ate/se/lb/ub/heterogeneity are always
the unpenalized fit’s; fit_cate(ard=True) only moves the
interaction estimates, in beta_ard (None if ARD did not run).
Pass include_evaluation_population=True to CATE validation or targeting APIs
to retain the immutable evaluation roster and its actual base weights.
The default retains no identifiers. Read the snapshot from
validation.evaluation_population, or from rule.validation.evaluation_population
for fixed and selected rules. Selection captures only the outer evaluation split;
its overlap provenance is independent of the inner selection population.
Equal-cluster weights use retained cluster sizes after overlap trimming, before
policy selection. Snapshot rows, cluster identities, and weights stay aligned
through JSON serialization.
For synthetic validation,
increment.simulate.cluster_dgp.evaluation_policy_truth(rule, population) consumes
a saved rule and its ClusteredCATEResult. It returns exact policy truth
and the retained-population ATE, using the stored IDs and actual base weights.
Missing or mismatched evaluation rosters are refused, not reconstructed.
This is truth for the same retained population, not an independent evaluation batch.
CateEvaluationPopulation
Section titled “CateEvaluationPopulation”CateEvaluationPopulationImmutable roster and weighting provenance for a reported evaluation.
CateValidation
Section titled “CateValidation”CateValidationEverything the held-out half says about a fitted CATE model.
passed is autoc.p_value < alpha; while false, the only
defensible number is the average effect, holdout_ate - the
effect on the same rows the groups and rank tests use. Randomized
sources use a difference in means with a Welch SE; observational
sources use the mean cross-fitted doubly robust score with its SE.
groups/clan intervals are Bonferroni-corrected across their
own family (n_groups group intervals, len(clan) CLAN rows):
a reader scanning every row for the one excluding zero pays the true
familywise error, not the per-row nominal alpha (CDDF 2018).
split_caveat names the single-split limitation this correction
does not address. For declared clusters, top-level uncertainty_method,
reference_df, and unavailable_reason describe the holdout ATE;
each group/rank/CLAN row records its own bootstrap uncertainty. A missing
AUTOC p-value closes the gate. Nuisances are frozen using training rows.
TargetingRule
Section titled “TargetingRule”TargetingRuleA frozen deployment candidate and its honest-split evidence gate.
fraction is the requested budget; achieved_fraction is its realized
holdout share under cluster_weight. Cluster policies pool scores, break
ties by canonical ID and take the longest feasible whole-cluster prefix.
Their threshold is descriptive, never a substitute for the prefix rule.
Unit policies retain their frozen score cutoff.
predict applies this candidate to new columns without fitting again.
recommendation is "target" only when the evidence gate passes;
otherwise it recommends treating everyone alike based on the average effect.
A passing, nonempty candidate reports its conditional policy effect and
uplift as points only: this same holdout gates and reports them, so nominal
intervals would be invalid after selection. Empty candidates retain null
effects and their exact unavailable reason, rather than fabricated zeros.
TargetingSelection
Section titled “TargetingSelection”TargetingSelectionA fraction chosen honestly, and the locked rule’s untouched-test verdict.
inner is the selection table: each grid fraction’s net benefit
E[1{targeted}(tau - cost)] on inner units scored by models that
never saw them, using IPW contributions for randomized sources and
cross-fitted doubly robust scores for observational sources.
selected_fraction is its argmax (ties to the
smaller fraction), locked before the outer test is read. rule
is the ordinary :class:TargetingRule evaluated once on the outer
half. population records when overlap trimming restricts the inner
selection and training population; it is independent of rule.population,
which records trimming on the outer test. seed and the grid are
pre-commitments: rerunning with a new seed until the answer improves is
the failure mode this workflow prevents.
Clustered bootstrap availability and valid-repetition metadata mirror the
selected inner row; outer-test uncertainty remains on rule.
MetricTrend
Section titled “MetricTrend”MetricTrend(table, metric, grain, window, week_start, denominator, start, end, by, alpha)One metric’s calendar trend: a lazy ibis query plus metadata.
Columns (order not guaranteed, only names): metric, grain, window,
period, [dims...], period_complete, n, value, ci_lb, ci_ub.
n is the per-period unit denominator for entity-scoped metrics and
NULL for total/active (an active metric’s count is its value).
ci_lb/ci_ub are NULL for ratio (point estimate only in v1) and for
total/active (no variance concept). grain keeps materialized rows
self-describing across grains. window is the rolling trailing-day
size (total/active only), NULL for every calendar-bucket trend.
period_complete is true once the metric’s own fact has any event at
or past the period’s end (period_end <= max(fact ts)), so a period
can be marked complete up to one partial day early. Fine for date-grain
horizons; a consumer gating on the last complete period of a
still-loading warehouse should ignore it or wait for the next load.
SitewideImpact
Section titled “SitewideImpact”SitewideImpactWhole-site impact of shipping a lift to every enrolled unit.
Parameters
Section titled “Parameters”delta, delta_se : float
Per-unit absolute lift of the target arm and its standard error.
treatment_group : str
group_id of the target arm.
n_control, n_treatment : float
Control and target arm population sizes (float since a cluster
grain population is K * mean cluster size, not an integer).
n_enrolled : float
Every enrolled unit across all arms (N_exp); equals
n_control + n_treatment only when no other arm is enrolled.
other_arm_ids : tuple[str, …]
group_id of every other enrolled non-control arm, whose lift
was netted out of baseline_volume.
site_total_volume : float
The observed site-window total this result was computed against.
baseline_volume : float
Counterfactual site-window total with nobody exposed to treatment.
absolute_impact, absolute_impact_se, absolute_impact_lb, absolute_impact_ub : float
Ship-to-all absolute impact, its SE, and its alpha-level interval.
relative_impact, relative_impact_se, relative_impact_lb, relative_impact_ub : float
Ship-to-all impact as a fraction of baseline_volume, its SE,
and its alpha-level interval.
alpha : float
Two-sided significance level the intervals were built at.
n_clusters : int | None
None at iid grain, else the contrast’s total cluster count.
absolute_dof, relative_dof : float | None
None at iid grain (the Normal reference applies). Otherwise
the dof absolute_impact’s and relative_impact’s own
critical values were cut at - see “Degrees of freedom” in the
module docstring for which reduction each is. They differ from
each other, and both differ from the contrast’s pooled
n_clusters - 2 once another arm is enrolled.
SitewideRatioImpact
Section titled “SitewideRatioImpact”SitewideRatioImpactWhole-site impact of shipping a ratio metric’s lift to every unit.
Ratio metrics need two per-unit lifts (numerator, denominator) instead
of :class:SitewideImpact’s single delta, hence a separate model.
See “The ratio math” in the module docstring for the derivation.
Parameters
Section titled “Parameters”delta_num, delta_num_se : float
Per-unit absolute lift of the target arm’s numerator and its SE.
delta_den, delta_den_se : float
Per-unit absolute lift of the target arm’s denominator and its SE.
delta_cov : float
Covariance of delta_num and delta_den.
treatment_group : str
group_id of the target arm.
n_control, n_treatment : float
Control and target arm population sizes (see
:class:SitewideImpact for why this is a float).
n_enrolled : float
Every enrolled unit across all arms (N_exp).
other_arm_ids : tuple[str, …]
group_id of every other enrolled non-control arm, whose
numerator and denominator lifts were netted out of the baselines.
site_total_numerator, site_total_denominator : float
The observed site-window totals this result was computed against.
baseline_numerator, baseline_denominator, baseline_ratio : float
Counterfactual site-window totals with nobody exposed to
treatment (N0, D0), and their ratio.
shipped_numerator, shipped_denominator, shipped_ratio : float
Ship-to-all site-window totals (N1, D1), and their ratio.
absolute_impact, absolute_impact_se, absolute_impact_lb, absolute_impact_ub : float
Ship-to-all absolute impact (shipped_ratio - baseline_ratio),
its delta-method SE, and its alpha-level Wald interval.
relative_impact, relative_impact_se, relative_impact_lb, relative_impact_ub : float
Ship-to-all impact as a fraction of baseline_ratio, its
delta-method SE, and its alpha-level interval.
alpha : float
Two-sided significance level the interval was built at.
n_clusters, absolute_dof, relative_dof : int | None, float | None, float | None
Same contract as :class:SitewideImpact’s fields of the same
name - unlike the sum-metric’s absolute_dof, this class’s
absolute_dof is always the Satterthwaite reduction, never
the plain pairwise dof (see “Degrees of freedom” in the module
docstring).