Skip to content

Introduced OOFMeanEncoder - #1097

Open
jusspatel wants to merge 3 commits into
feature-engine:mainfrom
jusspatel:main
Open

jusspatel wants to merge 3 commits into
feature-engine:mainfrom
jusspatel:main

Conversation

@jusspatel

Copy link
Copy Markdown

Closes #1050

Standard target encoding (MeanEncoder) replaces each category with the mean target value observed for that
category. When MeanEncoder is fitted and transformed on the same training data, every row contributes its own
target label to its encoded value. For small sample sizes or rare categories with low sample count $k$, each
sample accounts for $1/k$ of its own value (100% when $k=1$). This causes target leakage, resulting in models
memorizing training labels and over-optimistic training performance that collapses on validation or test sets.

This PR introduces **`OOFMeanEncoder`** (Out-of-Fold Mean Encoder) to solve target leakage during training:
- **`fit_transform(X, y)`**: Splits the training dataset into cross-validation folds according to `cv`. For each

fold, target encoding mappings are learned strictly using the remaining folds and applied to the held-out fold.
Consequently, no row's target value is ever used in its own encoding.
- fit(X, y): Computes and stores the target mapping across the full training dataset in encoder_dict_
and y_prior_.
- transform(X): Encodes unseen/test data using the full dataset mapping learned during fit(), matching
standard MeanEncoder behavior at inference time.

---

### Key Features & API Design

- **`cv`**: Supports an integer (default `5`), any scikit-learn CV splitter instance (`KFold`, `StratifiedKFold`,

TimeSeriesSplit, etc.), or an iterable yielding (train, test) index splits.
- smoothing: Supports 0.0 (unmodified means), positive numbers (additive smoothing towards the prior),
or 'auto' (empirical Bayes weighting).
- unseen: Strategy for unseen categories ('encode' with fold/dataset prior, 'ignore' returning np. nan with a warning, or 'raise' raising a ValueError).
- shuffle & random_state: Controls fold randomization when cv is an integer.
- stratify: Enables StratifiedKFold when cv is an integer (for discrete/classification targets).
- Standard feature-engine features: Supports variables, missing_values, ignore_format (enabling
numerical column encoding), and return_empty.
- Preserves pandas structures: Full index, column name, and dtype preservation.

---

### Example Usage

```python
import pandas as pd
from feature_engine.encoding import OOFMeanEncoder

# Training data
X_train = pd.DataFrame({"city": ["London", "London", "Paris", "Paris", "Berlin"]})
y_train = pd.Series([1, 0, 1, 1, 0])

# Initialize encoder
encoder = OOFMeanEncoder(
    variables=["city"],
    cv=3,
    smoothing="auto",
    unseen="encode",
    shuffle=True,
    random_state=42,
)

# Out-of-fold encoding for training (no target leakage)
X_train_oof = encoder.fit_transform(X_train, y_train)

# Inference on unseen/test data using full-dataset statistics
X_test = pd.DataFrame({"city": ["Paris", "Tokyo"]})
X_test_enc = encoder.transform(X_test)

──────

Changes Included

  1. feature_engine/encoding/oof_mean_encoding.py:
    • Implemented OOFMeanEncoder inheriting from CategoricalMethodsMixin and CategoricalInitMixinNA.
    • Comprehensive docstring with math formulas, parameter documentation, and examples.

  2. feature_engine/encoding/init.py:
    • Exposed OOFMeanEncoder in public API and all.

  3. tests/test_encoding/test_oof_mean_encoder.py:
    • 40 dedicated unit tests covering:
    • Initialization parameter validations.
    • Verification that out-of-fold encoding prevents leakage.
    • Custom CV splitters (KFold, StratifiedKFold).
    • Stratification behavior and regression error handling.
    • Additive and 'auto' smoothing.
    • Unseen category strategies ('encode', 'ignore', 'raise').
    • Handling missing values ('raise' vs 'ignore').
    • Numerical encoding (ignore_format=True).
    • Preserving non-standard DataFrame indices.
    • Pipeline integration.
    • inverse_transform behavior.

  4. tests/test_encoding/test_check_estimator_encoders.py:
    • Added OOFMeanEncoder to scikit-learn's check_estimator and feature-engine's check_feature_engine_estimator.

  5. docs/api_doc/encoding/OOFMeanEncoder.rst & index.rst:
    • Added API documentation page and updated the encoder comparison table.

──────

Verification & Testing

• pytest tests/test_encoding/test_oof_mean_encoder.py (40 passed)
• pytest tests/test_encoding/test_check_estimator_encoders.py (52 passed, including sklearn check_estimator)
• Full test suite pytest tests/test_encoding/ (387 passed)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add an out-of-fold target encoder (OOFMeanEncoder)

1 participant