Skip to content
Development documentation. The PyPI package predates these APIs. Install from GitHub instead: pip install 'increment @ git+https://github.com/kylejcaron/increment.git' Keep any extras requested by the guide, such as increment[dashboard].

Dataframe entry points

MetricSpec

One metric to estimate from a dataframe column.

name : str Metric name, and the value column unless value_column is given. value_column : str | None Read metric values from this column instead of name. type : {“mean”, “conversion”, “ratio”, “retention”, “quantile”} Dispatches the variance model: “conversion” is a 0/1 column, “quantile” needs quantile, “retention” needs threshold_days. winsorization : Winsorization | None Optional lower/upper outcome bounds for a mean metric. covariate : str | None Pre-experiment covariate column for CUPED. numerator, denominator : str | None Ratio-metric columns; both required when type="ratio". window_days : int | None Analysis window in days since exposure; None uses full history. Rejected on type="quantile". threshold_days : int | tuple[int, int] | None Retention band: int N is [N, inf); (a, b) is [a, b) (half-open). quantile : float | None The quantile in (0, 1) to estimate; required for type="quantile". missing : {“error”, “zero”, “drop”} Null/NaN policy for the value column: refuse (default), count as 0, or drop those rows from this metric’s moments. "impute" is refused (frame.metric.missing_impute): an outcome is never mean-filled; use "zero" or "drop", or covariate_missing="impute" for a covariate. covariate_missing : {“impute”, “zero”, “error”} Null/NaN policy for covariate: impute the pooled mean (default, warns with the count), treat as 0, or refuse. decision_method : Method | None Decision estimator for this metric. Omitted uses IPTW for an observational design, otherwise unadjusted; run(decision_method=...) overrides it without changing the declared sensitivities. sensitivity_methods : tuple[Method, …] Additional estimators reported after the decision estimator. prior : Normal | None Informative prior for this metric’s run() rows; omitted = inherit the call’s prior=. preferred_direction : {“increase”, “decrease”, “neutral”} | None The metric’s declared favorable side; None (default) means undeclared, not “increase” — synthesise_metric forwards it onto the synthesised Metric only when explicitly set, so an undeclared frame metric never silently reports a favorable direction (see LiftEstimate.preferred_direction).

window_days/threshold_days never raise at construction; band shape is validated once in the models layer. from_unit_summary carries no dates, so it refuses both instead (see CAPABILITY_TABLE).

to_frame(estimates, model, backend)

Convert a sequence of result models (:class:LiftEstimate, :class:BreakoutEstimate, :class:DailyMetricValue, :class:DailyLiftEstimate) to a native backend frame.

Generic over pydantic’s model_fields: an Estimate-typed field flattens into four columns (its name, plus lb/ub/open_side); a binomial_set field (see :class:~increment.estimation.results. BinomialConfidenceSet) flattens into set_lower/set_upper/ set_level — always the row’s confidence-set bounds/level, even for a set-only row with no finite point (<estimate field> and lb/ub stay None there; set_lower/set_upper/ set_level are the row’s ONLY confidence-set representation in that case; see :class:BinomialConfidenceSet); every other field passes through as a scalar column in declaration order. Most callers should use results.to_frame() on a pipeline’s own result rather than calling this function directly.

estimates : Sequence[M] Any sequence of one supported result model. May be empty. model : type[M] | None Which model estimates holds. Required when estimates is empty, since an empty sequence carries no runtime type trace. backend : {“pandas”, “polars”, “pyarrow”} Which native library to build.

IntoDataFrame One row per estimate, columns in the model’s field order, with the Estimate-typed field expanded to <field name>/lb/ub/open_side and a binomial_set field expanded to set_lower/set_upper/ set_level. open_side is "lower"/"upper" for a genuinely unbounded one-sided endpoint, and None both for a closed interval (lb/ub both set) and for an unavailable one (lb/ub/value all None) — distinguish the two by whether <field name> (the point estimate) is None.

impute

Explicit missing-value repairs for dataframe entry paths.

from_unit_summary/from_unit_panel refuse null or NaN metric values by default, since either can silently produce a wrong number (zero-averaged or backend-inconsistent sums) with no warning. This module is the helper tier of the named fixes for callers holding the dataframe; warehouse users instead declare MetricSpec(missing=...), which needs no source mutation.

Helpers are narwhals-based and backend-agnostic, returning the same native frame type that came in, and treat null/NaN identically as “missing” since NaN is a real value (not skipped by sums) on polars/pyarrow but not pandas. Each helper returns (frame, affected_count), so a repair is always visible and loggable.