Logged-policy evaluation
Off-policy contrast compares two registered policies from a decision trace
logged under one fixed policy version. This family uses its own ingress
(LoggedTrace.from_records / from_frame) rather than an Analysis
constructor; see the guide for the estimand, the
assumptions, and the floors.
Both constructors accept either registry= or complete recorded logging laws:
logging_distributions= on records, or logging_distribution_column= on a
frame. Exactly one route is required; the estimator and admission bounds are unchanged.
estimate_policy_contrast
Section titled “estimate_policy_contrast”estimate_policy_contrast(trace, target, reference, alpha)Estimate Delta_T(target, reference) from an admitted trace.
Gates run in order: alpha; one fixed logging law for the whole
trace; policy support at every logged history; the independent-unit
inference floor; the effective-sample-size floor at every decision index
for both policies; a finite positive t critical value; a float64
representation of every reported statistic. Each refusal is a coded
error whose context carries the diagnostics computed so far.
LoggedTrace
Section titled “LoggedTrace”LoggedTraceAdmitted, ordered, complete-horizon decision records.
Construct with :meth:from_records or :meth:from_frame, which admit
either a registered logging policy or a complete recorded distribution.
The model validator re-checks every structural admission rule, so a trace
instance is always ordered by (unit_id, decision_index), has exactly
horizon closed rewards per unit, one fixed candidate set, and one
admitted logging distribution per record.
PolicyRegistry
Section titled “PolicyRegistry”PolicyRegistry(policies)Immutable lookup of registered policies by (policy_id, version).
TabularPolicy
Section titled “TabularPolicy”TabularPolicyA policy whose action distribution is a table over one context value.
probabilities maps the value of context[context_key] to an
action distribution; default applies when the key is absent or its
value is not tabulated. A policy with an empty table and a default
ignores the history entirely (the spec’s reference-policy).
Persistence: Python-mode dumps, copy.deepcopy and (trusted) pickle
keep every key type and rebuild an equivalent policy. JSON dumps
roundtrip only when every table key is a string; any other key type
refuses at serialization with logged_policy.policy.json_context_key
rather than becoming a different key. An absent default stays absent.
PolicyValueContrast
Section titled “PolicyValueContrast”PolicyValueContrastFixed-horizon contrast Delta_T = V_T(target) - V_T(reference) on the reward scale.
ess_by_time is the smaller of the two policies’ effective sample
sizes at each decision index and max_weight_by_time the larger of
their maximum cumulative weights; the per-policy tuples carry both.
The interval is a two-sided Wald interval from the unit-clustered
sandwich standard error with a t reference on n_units - 1
degrees of freedom.