Skip to content
Development documentation. The PyPI package predates these APIs. Install from GitHub instead: pip install 'increment @ git+https://github.com/kylejcaron/increment.git' Keep any extras requested by the guide, such as increment[dashboard].

Conditional effects and targeting

estimate_cate(..., cluster_weight="member_count") uses every member equally. Choose cluster_weight="equal" explicitly to give each observed cluster unit total weight; this requires a declared cluster column. Intervention grain never changes the default weighting. Centering, regression, covariance and ARD all use the selected weights.

Clustered fits use the grouped score sandwich with scale K/(K-1), counting all observed cluster IDs, including singleton and constant-residual clusters. There are no row HC2 divisors or additional row-count corrections. Scalar intervals use t(K-1); the interaction Wald quadratic divided by its dimension q uses F(q, K-1). Without clusters, the existing HC2 and ARD calculations are unchanged.

CateResult.dimension, n_clusters, reference_df, and immutable vcov hold fit-level inference metadata; CATE and contrast projections reuse that state. se_unadjusted uses a separate fit on [1, d] with the same cluster weights and sandwich, stored as immutable unadjusted_vcov. Covariances remain live-object state excluded from the report-only dump; the counts, dimension, reference df and selected weighting are included in reports.

Clustered fits raise InvalidRequestError with stable estimation.cate.* codes for fewer than two clusters (insufficient_clusters), rank deficiency (rank_deficient), invalid covariance (invalid_covariance), directions identified only by one cluster (single_cluster_direction), singular interaction Wald covariance (singular_wald), or unavailable required uncertainty (uncertainty_unavailable). Singular full covariance alone does not invalidate an otherwise supported scalar query. Because the interaction Wald result is required, an unavailable joint test refuses the fit.

Public validation and both policy APIs retain the same cluster_weight and cluster IDs through the honest split and nested selection. Validation reports member counts separately from independent heldout clusters. Numeric nulls and unavailable_reason survive policy JSON round trips; unavailable AUTOC evidence keeps the gate closed. A clustered constant score reports estimation.targeting.degenerate_rank_distribution, rather than manufacturing evidence from repeated members. A supplied randomized training mask that splits a cluster raises estimation.crossfit.cluster_split, as does a split holdout mask.

targeting_rule and select_targeting_rule accept deploy_grain: Literal["unit", "cluster"] | None = None. Only a declared intervention_grain="cluster" defaults to cluster deployment. Dependence clusters alone retain unit policies, including mixed-treatment observational clusters. An explicit unit policy for a cluster intervention refuses before reading data with estimation.targeting.unsupported_unit_deployment.

Cluster deployment averages member scores, orders clusters by descending score and canonical ID, and takes the longest prefix fitting the requested budget. cluster_weight="member_count" budgets members; "equal" budgets clusters. The same weights govern fitting, selection and policy evaluation. The next cluster is never split or skipped to backfill unused capacity. Fractions 0 and 1 select nobody and everybody; an oversized first cluster selects nobody. fraction records the requested share, achieved_fraction the realized share; inner FractionScore rows also retain achieved shares.

CateResult.score(cols, *, cluster_ids=None, deploy_grain=None) returns an aligned unit array in unit mode, or an immutable tuple of ClusterScore(cluster_id, score, member_count) records in canonical ID order in cluster mode. IDs are separate metadata. TargetingRule.predict returns aligned Boolean actions from its saved CateScoreState: unit policies use their frozen cutoff, cluster policies pool and budget the supplied batch and broadcast cluster actions. Supply the complete member roster for each deployment cluster. predict evaluates the candidate; recommendation remains the evidence gate for deploying it. An empty candidate has no conditional policy effect, with reason estimation.targeting.empty_group. A cluster policy’s reported threshold is descriptive; it does not replace its ID tie break and prefix budget.

Policy JSON includes the fitted basis, knots, category levels, centering and coefficients, so TargetingRule.model_validate_json(rule.model_dump_json()) predicts without fitting again. CATE report dumps remain report-only; their separately serializable score_state supplies portable point prediction. Analysis.estimate_cate, validate_cate, targeting_rule and select_targeting_rule delegate to the same source APIs.

Covariate

One pre-exposure column to model effect heterogeneity on.

kind picks the basis: "continuous" standardizes to mean 0 / sd 1; "categorical" one-hot encodes against its modal level.

knots opts a continuous covariate into a piecewise-linear (hinge) basis, free to bend at every knot but still linear in the coefficients. An int places that many knots at the interior quantiles of the fitting column; a tuple gives explicit positions. Continuous-only; opt in per covariate, since each knot spends an interaction column and a Wald degree of freedom.

The basis is additive: knots on spend alongside a categorical country give one spend curve plus a per-country offset, not a differently shaped curve per country.

ClusterBootstrap

Whole-cluster resampling controls shared by every honest-validation entry point.

Bundles the seed and repetition count of the heldout-only whole-cluster bootstrap that :func:validate_cate_arrays, :func:targeting_rule_arrays and :func:select_targeting_rule_arrays use to recompute ranks, empirical GATES/CLAN cutoffs and policy values on every replicate. Immutable: construct once and reuse across calls. seed and repetitions that are not a genuine nonnegative integer and an integer >= 2 — a bool, a float, or a forged/mutated instance — refuse with the coded estimation.targeting.bootstrap_options error before any nuisance model is fit.

estimate_cate(source, metric, control, interact, adjust, alpha, ard, cluster_weight)

Estimate conditional average treatment effects for metric.

Nothing this returns is validated: scoring units with it and reporting the top group’s effect is exactly the in-sample fabrication validate_cate exists to catch - on data with no true heterogeneity, that top quintile reads 2.9x the true effect. Run validate_cate before any subgroup number leaves this function.

source must serve unit-grain rows (CapabilityError otherwise). control names the control group_id; the frame must contain exactly that arm plus one other. interact covariates model effect heterogeneity; adjust covariates enter as main effects only, for precision without a heterogeneity claim. ard shrinks the interaction coefficients by evidence maximization, moving only the scored points - not ate/se/heterogeneity. Returns effects on the metric’s own absolute scale. Declared clusters use a weighted cluster-score sandwich with t(K-1) intervals; otherwise HC2. cluster_weight="member_count" weights units equally (the default); "equal" weights clusters equally and requires a declared cluster. Intervention grain never selects weighting. Cluster rank deficiency, single-cluster-only directions and unavailable uncertainty refuse with InvalidRequestError; inference metadata live on the returned fit. Raises InvalidRequestError (cate.identification.randomized_only) for a source whose design is not randomized; ValueError for an undeclared metric, absent control, or a non-two-arm frame; NotImplementedError for ratio and quantile metrics.

Covariates must be strictly pre-exposure; results do not compose with the package’s default relative lifts (see increment.estimation.cate).

validate_cate(source, metric, control, interact, adjust, n_groups, alpha, cluster_weight, bootstrap, include_evaluation_population)

Fit CATE on one honest partition and validate it on the other.

The gate for every heterogeneity claim: a CATE fit always hands back a winner, so the top group’s in-sample effect is not evidence of anything. Splits the units by a content hash of their id, fits on one half (whole clusters when declared), and reports sorted-group effects, rank tests, and a CLAN profile computed entirely on the other. CateValidation.passed is the verdict; while false, the only defensible number is the average effect.

Randomized sources use raw arm contrasts. Observational sources use cross-fitted doubly robust scores over the design’s declared adjustment set; a numeric adjustment column enters as it is and a string column as modal-reference level indicators fitted inside every nuisance fit (a null level refuses, as this path supports missing="refuse" only). Its overlap gate can refuse or trim the reported population. n_groups sets how many predicted-effect groups to cut the holdout into. alpha is two-sided for every reported interval; the rank tests are one-sided against it, and passed is autoc.p_value < alpha.

Returns holdout-only numbers on the metric’s own absolute scale. Units are keyed by unit_frame’s unit_id (already cast to String), so the same logical unit lands in the same half across runs. There is deliberately no ard= here: this gate’s null size was measured on the unclustered unpenalized score (4.8% against a nominal 5%), and shrinkage is a reporting choice for estimate_cate, not a knob on the test.

Missing or unsupported identification raises cate.identification.unsupported_mechanism. Observational policies other than missing="refuse" or with a non-null max_smd raise cate.identification.unsupported_missing_policy or cate.identification.unsupported_max_smd before reading data. Declared clusters use cluster_weight="member_count" (unit weights) or "equal" (inverse cluster-size weights). Clustered rank, GATES and CLAN intervals contain both a heldout-only whole-cluster bootstrap-t and a delete-one-cluster jackknife-t interval, conditional on training-frozen nuisances; bootstrap (default ClusterBootstrap(seed=0, repetitions=999)) is recorded on results. Missing uncertainty has a nullable numeric field and an unavailable_reason code. n_train and n_holdout count members; n_clusters counts the independent heldout clusters. The unclustered calibration above does not establish calibration of the cluster bootstrap at small cluster counts.

select_targeting_rule(source, metric, control, interact, adjust, fractions, cost_per_treated, n_folds, seed, alpha, cluster_weight, bootstrap, deploy_grain, include_evaluation_population)

Choose the share of units to target on metric, honestly.

:func:targeting_rule demands a pre-committed fraction because choosing the cut after seeing results biases the reported policy value upward by 60-120%. This is the sanctioned way to CHOOSE that fraction: a seeded, arm-stratified half of the units is set aside untouched; on the other half, K-fold out-of-fold scores estimate each grid fraction’s net benefit E[1{targeted} (tau - cost_per_treated)]; the argmax is locked and then evaluated exactly once on the untouched half, through the same gate and policy numbers :func:targeting_rule reports. For observational sources, both the inner selection objective and the untouched outer report use the declared doubly robust score and overlap policy.

fractions, in [0, 1] each, and seed are REQUIRED with no defaults: the grid and the split are pre-commitments. cost_per_treated is in the metric’s own units per treated unit; at the default 0.0 the objective is total benefit, which favors wide fractions whenever the marginal unit’s effect is positive.

Raises exactly what :func:targeting_rule raises, plus ValueError for an invalid grid or a fold too small to hold 2 units per arm. Declared clusters use cluster_weight="member_count" (unit weights) or "equal" (inverse cluster-size weights). Clustered rank, GATES and CLAN intervals contain both a heldout-only whole-cluster bootstrap-t and a delete-one-cluster jackknife-t interval, conditional on training-frozen nuisances; bootstrap (default ClusterBootstrap(seed=0, repetitions=999)) is recorded on results. Deployment defaults only from the declared intervention grain. Cluster policies pool member scores and take the longest feasible whole-cluster prefix, with canonical-ID tie breaks; cluster_weight determines both budget and evaluation mass. Fractions zero/one deploy nobody/everybody. The result stores requested and achieved shares and portable scoring state. predict applies the candidate policy; recommendation is its gate. Missing uncertainty has a nullable numeric field and an unavailable_reason code.

targeting_rule(source, metric, control, interact, adjust, fraction, alpha, cluster_weight, bootstrap, deploy_grain, include_evaluation_population)

Decide whether to target the top fraction of units on metric.

The question a heterogeneity analysis is actually asked is not “which group responded best” but “what rule should I deploy”. Runs validate_cate’s honest split; unless the gate passes, returns recommendation="simple" (treat everyone alike on the average effect). Only a passing gate gets a threshold, a policy value, and an uplift over the average.

fraction is required with no default: pre-committing to the cut is the entire value of this function - choosing it after seeing validate_cate’s group table biases the reported policy value upward by 60-120%. Identification and policy refusals follow validate_cate, including observational adjustment and overlap rules. alpha gates the same autoc.p_value < alpha test. ard= is absent for the same reason it is absent from validate_cate: the gate was calibrated on the unpenalized score.

A failed gate is a result, not an exception: recommendation and the attached validation say why, and the three policy fields are None together so no caller reads an unsupported number. Conditional on a gate that passed by chance, policy_value runs high - unconditionally it is unbiased; treat a barely-passing gate as weak evidence for the magnitude, not just the ranking. Declared clusters use cluster_weight="member_count" (unit weights) or "equal" (inverse cluster-size weights). Clustered rank, GATES and CLAN intervals contain both a heldout-only whole-cluster bootstrap-t and a delete-one-cluster jackknife-t interval, conditional on training-frozen nuisances; bootstrap (default ClusterBootstrap(seed=0, repetitions=999)) is recorded on results. Deployment defaults only from the declared intervention grain. Cluster policies pool member scores and take the longest feasible whole-cluster prefix, with canonical-ID tie breaks; cluster_weight determines both budget and evaluation mass. Fractions zero/one deploy nobody/everybody. The result stores requested and achieved shares and portable scoring state. predict applies the candidate policy; recommendation is its gate. Missing uncertainty has a nullable numeric field and an unavailable_reason code.