In modern tabular machine learning, libraries like Polars have shifted dataframe operations from eager in-memory materialization to lazy evaluation graphs. A LazyFrame constructs an abstract syntax tree of transformations, allowing the query engine to apply predicate pushdown, projection pushdown, and streaming execution prior to memory allocation.
However, when wrapping these structures into higher-level feature engineering pipelines (such as skrub in the scikit-learn ecosystem), defensive input validation can silently dismantle these performance guarantees. Below is a retrospective on 5 merged pull requests across skrub and skore resolving lazy evaluation bottlenecks and display serialization.
1. The Eager Evaluation Bottleneck in CheckInput
In skrub, estimators validate tabular inputs via internal checks to verify column types, categorical cardinalities, and missing values. During initial Polars integration, early validation logic performed inspection routines that triggered eager collection of the entire LazyFrame into memory:
# Anti-pattern: Eager collection destroys lazy execution graphs
def check_input(X):
if hasattr(X, "collect"):
df = X.collect() # Forces full memory materialization
validate_schema(df.schema)
return df
When dealing with multi-gigabyte tabular datasets, this eager evaluation forced immediate heap allocation, neutralizing streaming chunk execution and defeating the exact reason users adopted Polars.
In PR #1941, I refactored the CheckInput pipeline to inspect schema metadata natively using lazy schema resolution (X.collect_schema() / X.schema) without collecting data buffers. This kept the query execution graph unmaterialized until downstream fit/transform passes explicitly requested execution.
2. Model Explainability and Display Enhancements in skore
skore is probabl-ai's library for machine learning evaluation, tracking, and diagnostics built on top of scikit-learn. Model explainability displays like ImpurityDecreaseDisplay and CoefficientsDisplay calculate feature importances across ensemble estimators.
For multi-output estimators and complex gradient-boosted trees, raw impurity decreases or coefficients produce multi-dimensional arrays across estimators. In PR #2539 and PR #2552, I implemented aggregate parameters (such as mean, median, and variance pooling), providing flexible single-plot summaries while preserving underlying per-estimator distributions.
3. Matplotlib Figure Legend Export Persistence
In diagnostic reporting, saving interactive diagnostic plots (such as PredictionErrorDisplay) to disk often resulted in clipped or missing external legends depending on the backend used (e.g., Agg vs. inline notebook renderers).
In PR #2530, I restructured legend bounding box management using bbox_inches="tight" and explicit artist inclusion in extra_artists during figure save routines, ensuring that exported figures retain complete annotations regardless of resolution or canvas aspect ratio.
4. Summary of Merged Pull Requests
All contributions were reviewed and merged by core maintainers across the skrub-data and probabl-ai organizations:
| Repository | PR | Title & Core Contribution | Status |
|---|---|---|---|
| skrub-data/skrub | #1941 | FIX: Do not collect LazyFrames in CheckInput | Merged |
| probabl-ai/skore | #2530 | FIX: Ensure figure-level legend is included in PredictionErrorDisplay export | Merged |
| probabl-ai/skore | #2539 | FEAT: Add aggregate parameter to ImpurityDecreaseDisplay | Merged |
| probabl-ai/skore | #2552 | FEAT: Add aggregate parameter to CoefficientsDisplay | Merged |
| skrub-data/skrub | #1940 | DOC: Update gallery examples to load datasets from path | Merged |
| skrub-data/skrub | #1964 | DOC: Update remaining gallery examples to load datasets from paths | Merged |
You can inspect the complete public verification record directly on GitHub: GitHub Search: is:pr is:merged author:MuditAtrey.