Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1050
Standard target encoding ($k$ , each$1/k$ of its own value (100% when $k=1$ ). This causes target leakage, resulting in models
MeanEncoder) replaces each category with the mean target value observed for thatcategory. When
MeanEncoderis fitted and transformed on the same training data, every row contributes its owntarget label to its encoded value. For small sample sizes or rare categories with low sample count
sample accounts for
memorizing training labels and over-optimistic training performance that collapses on validation or test sets.
fold, target encoding mappings are learned strictly using the remaining folds and applied to the held-out fold.
Consequently, no row's target value is ever used in its own encoding.
-
fit(X, y): Computes and stores the target mapping across the full training dataset inencoder_dict_and
y_prior_.-
transform(X): Encodes unseen/test data using the full dataset mapping learned duringfit(), matchingstandard
MeanEncoderbehavior at inference time.TimeSeriesSplit, etc.), or an iterable yielding(train, test)index splits.-
smoothing: Supports0.0(unmodified means), positive numbers (additive smoothing towards the prior),or
'auto'(empirical Bayes weighting).-
unseen: Strategy for unseen categories ('encode'with fold/dataset prior,'ignore'returningnp. nanwith a warning, or'raise'raising aValueError).-
shuffle&random_state: Controls fold randomization whencvis an integer.-
stratify: EnablesStratifiedKFoldwhencvis an integer (for discrete/classification targets).- Standard feature-engine features: Supports
variables,missing_values,ignore_format(enablingnumerical column encoding), and
return_empty.- Preserves pandas structures: Full index, column name, and dtype preservation.
──────
Changes Included
feature_engine/encoding/oof_mean_encoding.py:
• Implemented OOFMeanEncoder inheriting from CategoricalMethodsMixin and CategoricalInitMixinNA.
• Comprehensive docstring with math formulas, parameter documentation, and examples.
feature_engine/encoding/init.py:
• Exposed OOFMeanEncoder in public API and all.
tests/test_encoding/test_oof_mean_encoder.py:
• 40 dedicated unit tests covering:
• Initialization parameter validations.
• Verification that out-of-fold encoding prevents leakage.
• Custom CV splitters (KFold, StratifiedKFold).
• Stratification behavior and regression error handling.
• Additive and 'auto' smoothing.
• Unseen category strategies ('encode', 'ignore', 'raise').
• Handling missing values ('raise' vs 'ignore').
• Numerical encoding (ignore_format=True).
• Preserving non-standard DataFrame indices.
• Pipeline integration.
• inverse_transform behavior.
tests/test_encoding/test_check_estimator_encoders.py:
• Added OOFMeanEncoder to scikit-learn's check_estimator and feature-engine's check_feature_engine_estimator.
docs/api_doc/encoding/OOFMeanEncoder.rst & index.rst:
• Added API documentation page and updated the encoder comparison table.
──────
Verification & Testing
• pytest tests/test_encoding/test_oof_mean_encoder.py (40 passed)
• pytest tests/test_encoding/test_check_estimator_encoders.py (52 passed, including sklearn check_estimator)
• Full test suite pytest tests/test_encoding/ (387 passed)