A flexible and modular machine learning framework designed to support leakage-free model training through custom cross-validation fold construction
You can install the development version of pipeML from
GitHub with:
# install.packages("pak")
pak::pkg_install("VeraPancaldiLab/pipeML")pipeML is a flexible and leakage-aware machine learning framework for
R designed for predictive modeling in high-dimensional biological data.
The package integrates all key steps of the machine learning workflow —
feature filtering, model training, validation, prediction, and
interpretation — into a single reproducible pipeline.
A key design goal of pipeML is to support fold-aware feature
construction, allowing features that depend on the dataset
(e.g. enrichment scores, correlation-based features, or network-derived
features) to be recomputed within each cross-validation fold. This
prevents information leakage and ensures reliable performance
estimation.
The framework is designed to integrate naturally with R/Bioconductor workflows, making it particularly suitable for omics and biomedical machine learning applications.
Figure 1. General structure of the pipeML machine learning
pipeline.
- Integrated pipeline for feature filtering, model training, validation, prediction, and interpretation
- Custom cross-validation fold construction
- Support for fold-aware feature recomputation
- Prevents information leakage when using dataset-dependent features
- Repeated and stratified k-fold cross-validation
- Leave-one-dataset-out (LODO) evaluation for cross-cohort generalization
- Near-constant and highly correlated features are removed from the
training features (
preprocess = TRUE, the default): once before the cross-validation, or inside each fold for features built by custom fold functions
-
Automatic optimization based on:
- AUROC
- AUPRC
- C-index (survival)
- SHAP values of the selected model, per sample and as global feature importance
- Performance visualization (ROC and PR curves with bootstrap confidence bands, Kaplan-Meier curves by predicted risk group)
- Multi-core support for faster model training and cross-validation
- Users can define custom fold construction functions
- The parameters of the feature construction can be tuned inside the cross-validation, like model hyperparameters
- These functions receive a
bestuneargument after tuning, to rebuild the features on the full training dataset for the final model
For classification tasks, we implemented a diverse set of classification
algorithms that are benchmarked on the fly making extensive use of the R
package caret.
- Bagged classification trees
- Random forests
- C5.0 decision trees
- Regularized logistic regression (elastic net)
- k-nearest neighbors (KNN)
- Classification and regression trees (CART)
- Lasso regression
- Ridge regression
- Support vector machines with linear and radial kernels
- Extreme Gradient Boosting (XGBoost)
For time-to-event outcomes, pipeML implements a unified survival
modeling framework based on the parsnip and workflows ecosystems,
enabling consistent training, hyperparameter tuning, and evaluation
across multiple survival model families.
- Cox proportional hazards model
- Elastic net–regularized Cox regression
- Parametric accelerated failure time (AFT) models
- Conditional inference survival trees
- Bagged CART survival models
- Oblique random survival forests
Below are basic examples showing how to use pipeML. For detailed
tutorials, see Get
started
and the
Articles.
Results (plots, fold files of custom workflows) are written to a
Results/ folder in the working directory.
library(pipeML)
data <- data_example_classification
X <- data[, setdiff(colnames(data), "target")]
y <- data$target
set.seed(123)
train_idx <- caret::createDataPartition(y, p = 0.7, list = FALSE)
X_train <- X[train_idx, ]
X_test <- X[-train_idx, ]
y_train <- y[train_idx]
y_test <- y[-train_idx]res <- compute_features.training.ML(features_train = X_train,
target_var = y_train,
task_type = "classification",
trait.positive = "1",
metric = "AUROC",
k_folds = 5,
n_rep = 10,
ncores = 2)pred = compute_prediction(model = res$Model,
test_data = X_test,
target_var = y_test,
task_type = "classification",
trait.positive = "1")
pred$AUCshap <- compute_shap_values(model_trained = res$Model, task_type = "classification")res <- compute_features.ML(features_train = X_train,
features_test = X_test,
coldata = data,
task_type = "classification",
trait = "target",
trait.positive = "1",
metric = "AUROC",
k_folds = 5,
n_rep = 10,
ncores = 2)If you encounter any problems or have questions about the package, we encourage you to open an issue here. We’ll do our best to assist you!
pipeML was developed by Marcelo
Hurtado in supervision of Vera
Pancaldi and is part of the
Pancaldi team. Currently, Marcelo
is the primary maintainer of this package.
If you use pipeML in a scientific publication, please cite:
Hurtado, M., & Pancaldi, V. (2026). A new pipeline for cross-validation fold-aware machine learning prediction of clinical outcomes addresses hidden data-leakage in omics based ‘predictors’. bioRxiv. https://doi.org/10.64898/2026.03.12.711429
