Conditional effects and targeting
estimate_cate(..., cluster_weight="member_count") uses every member equally.
Choose cluster_weight="equal" explicitly to give each observed cluster
unit total weight; this requires a declared cluster column. Intervention grain
never changes the default weighting. Centering, regression, covariance and ARD
all use the selected weights.
Clustered fits use the grouped score sandwich with scale K/(K-1), counting
all observed cluster IDs, including singleton and constant-residual clusters.
There are no row HC2 divisors or additional row-count corrections. Scalar intervals
use t(K-1); the interaction Wald quadratic divided by its dimension q uses
F(q, K-1). Without clusters, the existing HC2 and ARD calculations are unchanged.
CateResult.dimension, n_clusters, reference_df, and immutable vcov hold
fit-level inference metadata; CATE and contrast projections reuse that state.
se_unadjusted uses a separate fit on [1, d] with the same cluster weights and
sandwich, stored as immutable unadjusted_vcov. Covariances remain live-object
state excluded from the report-only dump; the counts, dimension, reference df
and selected weighting are included in reports.
Clustered fits raise InvalidRequestError with stable estimation.cate.* codes
for fewer than two clusters (insufficient_clusters), rank deficiency
(rank_deficient), invalid covariance (invalid_covariance), directions identified
only by one cluster (single_cluster_direction), singular interaction Wald
covariance (singular_wald), or unavailable required uncertainty
(uncertainty_unavailable). Singular full covariance alone does not invalidate
an otherwise supported scalar query. Because the interaction Wald result is
required, an unavailable joint test refuses the fit.
Public validation and both policy APIs retain the same cluster_weight and
cluster IDs through the honest split and nested selection. Validation reports
member counts separately from independent heldout clusters. Numeric nulls and
unavailable_reason survive policy JSON round trips; unavailable AUTOC evidence
keeps the gate closed. A clustered constant score reports
estimation.targeting.degenerate_rank_distribution, rather than manufacturing
evidence from repeated members. A supplied randomized training mask that splits
a cluster raises estimation.crossfit.cluster_split, as does a split holdout mask.
targeting_rule and select_targeting_rule accept
deploy_grain: Literal["unit", "cluster"] | None = None. Only a declared
intervention_grain="cluster" defaults to cluster deployment. Dependence clusters
alone retain unit policies, including mixed-treatment observational clusters.
An explicit unit policy for a cluster intervention refuses before reading data
with estimation.targeting.unsupported_unit_deployment.
Cluster deployment averages member scores, orders clusters by descending score
and canonical ID, and takes the longest prefix fitting the requested budget.
cluster_weight="member_count" budgets members; "equal" budgets clusters. The
same weights govern fitting, selection and policy evaluation. The next cluster
is never split or skipped to backfill unused capacity. Fractions 0 and 1
select nobody and everybody; an oversized first cluster selects nobody.
fraction records the requested share, achieved_fraction the realized share;
inner FractionScore rows also retain achieved shares.
CateResult.score(cols, *, cluster_ids=None, deploy_grain=None) returns an aligned
unit array in unit mode, or an immutable tuple of ClusterScore(cluster_id, score, member_count) records in canonical ID order in cluster mode. IDs are separate
metadata. TargetingRule.predict returns aligned Boolean actions from its saved
CateScoreState: unit policies use their frozen cutoff, cluster policies pool
and budget the supplied batch and broadcast cluster actions. Supply the complete
member roster for each deployment cluster. predict evaluates the candidate;
recommendation remains the evidence gate for deploying it. An empty candidate
has no conditional policy effect, with reason estimation.targeting.empty_group.
A cluster policy’s reported threshold is descriptive; it does not replace its
ID tie break and prefix budget.
Policy JSON includes the fitted basis, knots, category levels, centering and
coefficients, so TargetingRule.model_validate_json(rule.model_dump_json())
predicts without fitting again. CATE report dumps remain report-only; their
separately serializable score_state supplies portable point prediction.
Analysis.estimate_cate, validate_cate, targeting_rule and
select_targeting_rule delegate to the same source APIs.
Covariate
Section titled “Covariate”CovariateOne pre-exposure column to model effect heterogeneity on.
kind picks the basis: "continuous" standardizes to mean 0 /
sd 1; "categorical" one-hot encodes against its modal level.
knots opts a continuous covariate into a piecewise-linear
(hinge) basis, free to bend at every knot but still linear in the
coefficients. An int places that many knots at the interior
quantiles of the fitting column; a tuple gives explicit positions.
Continuous-only; opt in per covariate, since each knot spends an
interaction column and a Wald degree of freedom.
The basis is additive: knots on spend alongside a categorical
country give one spend curve plus a per-country offset, not a
differently shaped curve per country.
ClusterBootstrap
Section titled “ClusterBootstrap”ClusterBootstrapWhole-cluster resampling controls shared by every honest-validation entry point.
Bundles the seed and repetition count of the heldout-only whole-cluster
bootstrap that :func:validate_cate_arrays, :func:targeting_rule_arrays
and :func:select_targeting_rule_arrays use to recompute ranks, empirical
GATES/CLAN cutoffs and policy values on every replicate. Immutable:
construct once and reuse across calls. seed and repetitions that
are not a genuine nonnegative integer and an integer >= 2 — a bool, a
float, or a forged/mutated instance — refuse with the coded
estimation.targeting.bootstrap_options error before any nuisance
model is fit.
estimate_cate
Section titled “estimate_cate”estimate_cate(source, metric, control, interact, adjust, alpha, ard, cluster_weight)Estimate conditional average treatment effects for metric.
Nothing this returns is validated: scoring units with it and
reporting the top group’s effect is exactly the in-sample
fabrication validate_cate exists to catch - on data with no true
heterogeneity, that top quintile reads 2.9x the true effect. Run
validate_cate before any subgroup number leaves this function.
source must serve unit-grain rows (CapabilityError
otherwise). control names the control group_id; the frame
must contain exactly that arm plus one other. interact
covariates model effect heterogeneity; adjust covariates enter
as main effects only, for precision without a heterogeneity claim.
ard shrinks the interaction coefficients by evidence
maximization, moving only the scored points - not
ate/se/heterogeneity.
Returns effects on the metric’s own absolute scale. Declared clusters use
a weighted cluster-score sandwich with t(K-1) intervals; otherwise HC2.
cluster_weight="member_count" weights units equally (the default);
"equal" weights clusters equally and requires a declared cluster.
Intervention grain never selects weighting. Cluster rank deficiency,
single-cluster-only directions and unavailable uncertainty refuse with
InvalidRequestError; inference metadata live on the returned fit.
Raises InvalidRequestError (cate.identification.randomized_only)
for a source whose design is not randomized; ValueError for an
undeclared metric, absent control, or a non-two-arm frame;
NotImplementedError for ratio and quantile metrics.
Covariates must be strictly pre-exposure; results do not compose
with the package’s default relative lifts (see
increment.estimation.cate).
validate_cate
Section titled “validate_cate”validate_cate(source, metric, control, interact, adjust, n_groups, alpha, cluster_weight, bootstrap, include_evaluation_population)Fit CATE on one honest partition and validate it on the other.
The gate for every heterogeneity claim: a CATE fit always hands back
a winner, so the top group’s in-sample effect is not evidence of
anything. Splits the units by a content hash of their id, fits on
one half (whole clusters when declared), and reports sorted-group effects,
rank tests, and a CLAN profile computed entirely on the other.
CateValidation.passed is the verdict; while false, the only defensible
number is the average effect.
Randomized sources use raw arm contrasts. Observational sources use
cross-fitted doubly robust scores over the design’s declared adjustment
set; a numeric adjustment column enters as it is and a string column as
modal-reference level indicators fitted inside every nuisance fit (a
null level refuses, as this path supports missing="refuse" only).
Its overlap gate can refuse or trim the
reported population. n_groups sets how many predicted-effect groups
to cut the holdout into. alpha is two-sided for every reported
interval; the rank tests are one-sided against it, and passed is
autoc.p_value < alpha.
Returns holdout-only numbers on the metric’s own absolute scale.
Units are keyed by unit_frame’s unit_id (already cast to
String), so the same logical unit lands in the same half across
runs. There is deliberately no ard= here: this gate’s null size
was measured on the unclustered unpenalized score (4.8% against a nominal 5%),
and shrinkage is a reporting choice for estimate_cate, not a
knob on the test.
Missing or unsupported identification raises
cate.identification.unsupported_mechanism. Observational policies
other than missing="refuse" or with a non-null max_smd raise
cate.identification.unsupported_missing_policy or
cate.identification.unsupported_max_smd before reading data.
Declared clusters use cluster_weight="member_count" (unit weights)
or "equal" (inverse cluster-size weights). Clustered rank, GATES and
CLAN intervals contain both a heldout-only whole-cluster bootstrap-t and a
delete-one-cluster jackknife-t interval, conditional on training-frozen
nuisances; bootstrap (default
ClusterBootstrap(seed=0, repetitions=999)) is recorded on results.
Missing uncertainty has a nullable numeric field and an
unavailable_reason code.
n_train and n_holdout count members; n_clusters counts the
independent heldout clusters. The unclustered calibration above does not
establish calibration of the cluster bootstrap at small cluster counts.
select_targeting_rule
Section titled “select_targeting_rule”select_targeting_rule(source, metric, control, interact, adjust, fractions, cost_per_treated, n_folds, seed, alpha, cluster_weight, bootstrap, deploy_grain, include_evaluation_population)Choose the share of units to target on metric, honestly.
:func:targeting_rule demands a pre-committed fraction because choosing
the cut after seeing results biases the reported policy value upward by
60-120%. This is the sanctioned way to CHOOSE that fraction: a seeded,
arm-stratified half of the units is set aside untouched; on the other
half, K-fold out-of-fold scores estimate each grid fraction’s net
benefit E[1{targeted} (tau - cost_per_treated)]; the argmax is
locked and then evaluated exactly once on the untouched half, through
the same gate and policy numbers :func:targeting_rule reports.
For observational sources, both the inner selection objective and the
untouched outer report use the declared doubly robust score and overlap policy.
fractions, in [0, 1] each, and seed are REQUIRED with no
defaults: the grid and the split are pre-commitments.
cost_per_treated is in the metric’s own units per treated unit;
at the default 0.0 the objective is total benefit, which favors
wide fractions whenever the marginal unit’s effect is positive.
Raises exactly what :func:targeting_rule raises, plus ValueError
for an invalid grid or a fold too small to hold 2 units per arm.
Declared clusters use cluster_weight="member_count" (unit weights)
or "equal" (inverse cluster-size weights). Clustered rank, GATES and
CLAN intervals contain both a heldout-only whole-cluster bootstrap-t and a
delete-one-cluster jackknife-t interval, conditional on training-frozen
nuisances; bootstrap (default
ClusterBootstrap(seed=0, repetitions=999)) is recorded on results.
Deployment defaults only from the declared intervention grain. Cluster
policies pool member scores and take the longest feasible whole-cluster
prefix, with canonical-ID tie breaks; cluster_weight determines both
budget and evaluation mass. Fractions zero/one deploy nobody/everybody.
The result stores requested and achieved shares and portable scoring state.
predict applies the candidate policy; recommendation is its gate.
Missing uncertainty has a nullable numeric field and an
unavailable_reason code.
targeting_rule
Section titled “targeting_rule”targeting_rule(source, metric, control, interact, adjust, fraction, alpha, cluster_weight, bootstrap, deploy_grain, include_evaluation_population)Decide whether to target the top fraction of units on metric.
The question a heterogeneity analysis is actually asked is not
“which group responded best” but “what rule should I deploy”. Runs
validate_cate’s honest split; unless the gate passes, returns
recommendation="simple" (treat everyone alike on the average
effect). Only a passing gate gets a threshold, a policy value, and
an uplift over the average.
fraction is required with no default: pre-committing to the cut
is the entire value of this function - choosing it after seeing
validate_cate’s group table biases the reported policy value
upward by 60-120%. Identification and policy refusals follow
validate_cate, including observational adjustment and overlap rules.
alpha gates the same autoc.p_value < alpha test. ard= is absent for the
same reason it is absent from validate_cate: the gate was
calibrated on the unpenalized score.
A failed gate is a result, not an exception: recommendation and
the attached validation say why, and the three policy fields are
None together so no caller reads an unsupported number.
Conditional on a gate that passed by chance, policy_value runs
high - unconditionally it is unbiased; treat a barely-passing gate
as weak evidence for the magnitude, not just the ranking.
Declared clusters use cluster_weight="member_count" (unit weights)
or "equal" (inverse cluster-size weights). Clustered rank, GATES and
CLAN intervals contain both a heldout-only whole-cluster bootstrap-t and a
delete-one-cluster jackknife-t interval, conditional on training-frozen
nuisances; bootstrap (default
ClusterBootstrap(seed=0, repetitions=999)) is recorded on results.
Deployment defaults only from the declared intervention grain. Cluster
policies pool member scores and take the longest feasible whole-cluster
prefix, with canonical-ID tie breaks; cluster_weight determines both
budget and evaluation mass. Fractions zero/one deploy nobody/everybody.
The result stores requested and achieved shares and portable scoring state.
predict applies the candidate policy; recommendation is its gate.
Missing uncertainty has a nullable numeric field and an
unavailable_reason code.