Skip to content
Development documentation. The PyPI package predates these APIs. Install from GitHub instead: pip install 'increment @ git+https://github.com/kylejcaron/increment.git' Keep any extras requested by the guide, such as increment[dashboard].

Advanced power (`increment.power`)

Specialized segment and heterogeneity solvers are advanced power entry points:

segment_pairwise_required_sample_size(r_a, r_b, q_a, q_b, baseline_a, procedure, baseline_b, design)

Experiment-wide N needed to detect segment A’s lift differing from segment B’s.

r_a/r_b are the two segments’ relative lifts; q_a/q_b are their shares of the whole experiment (a 10% segment out of ten still has q=0.10 regardless of how many other segments exist). Cost is n_total = n_ATE(delta) * (1/q_a + 1/q_b), where delta = log1p(r_a) - log1p(r_b) is the log-scale contrast (not log1p(r_a - r_b): that errs -27.7% at (0.50, 0.20) and +19.4% at (0.30, -0.10)), and n_ATE(delta) is the total N a standard 50/50-allocation experiment would need to detect delta as a plain ATE. 1/q_a + 1/q_b reduces to the simpler n_ATE(delta)/(q(1-q)) form only when q_a + q_b = 1 (a 2-segment breakout); with more segments the two forms diverge and the simpler one understates N.

n_per_arm/n_total are experiment-wide (summed across every segment, not just A and B). effective_var reflects segment A’s baseline only. baseline_b defaults to baseline_a; design.allocation governs the treatment/control split within each segment. Cluster design effects flow in via each baseline’s effective_var, but PowerResult.n_clusters_* stays None: no single cluster count is meaningful across two possibly-different baselines. Fixed-horizon only; sequential planning for segment contrasts is not yet supported.

ValueError If r_a and r_b give the same relative lift (theta=0), the same refusal required_sample_size makes for a lift exactly at its null boundary. Also raised if the solved-for N is too small for either segment’s 2-arm split (see _segment_arm_sizes).

segment_pairwise_achieved_power(n_per_arm, r_a, r_b, q_a, q_b, baseline_a, procedure, baseline_b, design)

Achieved power for detecting segment A’s lift differing from segment B’s at an experiment-wide treatment-arm size of n_per_arm.

PowerResult.effective_var reflects segment A’s baseline only; it does not summarize baseline_b. Fixed-horizon only; sequential planning for segment contrasts is not yet supported.

ValueError If n_per_arm implies too few units in either segment’s share for a 2-arm split (see _segment_arm_sizes); increase n_per_arm or the smaller segment’s share.

segment_pairwise_minimum_detectable_effect

Section titled “segment_pairwise_minimum_detectable_effect”
segment_pairwise_minimum_detectable_effect(n_per_arm, q_a, q_b, baseline_a, procedure, baseline_b, design)

Smallest segment-A-vs-segment-B difference detectable at n_per_arm.

mde_relative is exp(delta) - 1 for the smallest detectable log-scale contrast delta = log(1+r_A) - log(1+r_B): the smallest detectable ratio (1+r_A)/(1+r_B) - 1, not a lift against a single baseline mean.

PowerResult.effective_var reflects segment A’s baseline only; it does not summarize baseline_b. Fixed-horizon only; sequential planning for segment contrasts is not yet supported.

ValueError If n_per_arm implies too few units in either segment’s share for a 2-arm split (see _segment_arm_sizes); increase n_per_arm or the smaller segment’s share.

joint_q_power_fixed(theta, var, alpha)

Exact power of Cochran’s Q to detect FIXED, named per-segment deviations.

theta : array-like of float Each segment’s true log-scale effect (or any per-segment quantity Q is computed over). Only relative differences matter: a common shift added to every theta_k does not change the result. var : array-like of float Each segment’s sampling variance, same length and order as theta. alpha : float Significance level for Cochran’s Q test (default 0.05).

float Power = P(Q > chi2_crit(K-1, alpha)) under the noncentral chi2(K-1, lambda) distribution Q follows exactly at these theta/var.

Fixed-horizon only; sequential planning for segment contrasts is not yet supported.

Verified against a 100k-rep simulation at a 5-segment unequal-share fixture: predicted 0.0533 vs. empirical 0.0543 (see tests/power/test_core.py).

joint_q_power_random(tau_b, var, alpha)

Power of Cochran’s Q under a random-effects model of segment spread.

Segment effects are modelled as theta_k ~ iid N(mu, tau_b^2): the design-time question “if segments typically differ by about tau_b, what’s my power to detect that?”, vs. joint_q_power_fixed’s “if they differ by exactly these amounts”.

tau_b : float Standard deviation of segment effects around their common mean (same scale as joint_q_power_fixed’s theta). var : array-like of float Each segment’s sampling variance. alpha : float Significance level for Cochran’s Q test (default 0.05).

float Power estimate. Exact when every var entry is equal: Q is then a scaled central chi-square, Q ~ (1 + tau_b^2/v) * chi2(K-1) (verified against a 100k-rep simulation: 0.1078 exact vs. 0.1099 empirical at K=5, v=1.0, tau_b=0.5).

Approximate for unequal ``var``: Q's true distribution is a
generalised (Satterthwaite-type) weighted sum of independent
central chi2(1) variables, not a noncentral chi-square exactly.
This returns a mean-matched noncentral-chi2(K-1, E[lambda])
approximation, ``E[lambda] = tau_b^2 * (sum(w) - sum(w^2)/sum(w))``
(the same building block as ``cochran_q``'s DerSimonian-Laird
denominator). Measured against 60k-rep simulation across 7
configurations (K=5-10, per-segment variance ratios 1x-50x):
relative error within +/-4% up to an 8x ratio, growing to +12.8%
at a 50x ratio - treat as an approximation, not a calibrated
design tool, when segment variances are wildly unequal.

Fixed-horizon only; sequential planning for segment contrasts is not yet supported.