Mamut
A Python toolkit for tabular classification that keeps preprocessing, hyperparameter search, model selection and validation evidence in one reproducible workflow. It can search across common model families, but the more useful part is what happens around that search: preprocessing stays inside validation, related rows can stay together, final holdout data can remain untouched during selection, and the evidence report can challenge the model that won.
Model search is only part of the experiment
Trying ten classifiers and keeping the largest validation score is easy. The harder questions are whether preprocessing leaked information across folds, whether related observations crossed a split boundary, whether hyperparameters were tuned on the same data used to judge them, whether a simple baseline is just as good, and whether the apparent winner is stable enough to deserve another round of work.
Mamut treats those questions as part of the workflow. A single Mamut object controls candidate selection, preprocessing, Optuna tuning, validation strategy, optional holdout data, final refitting, evidence generation and reporting. Search profiles provide a practical candidate pool, while include_models can lock an experiment to an exact set of estimators.
Preprocessing belongs inside validation
Mamut does not assume that every classifier should receive the same transformed matrix. Linear, kernel and distance-based models use a generic one-hot pipeline. Tree models can skip unnecessary scaling and skew transforms. LightGBM and CatBoost can keep native categorical columns when the configured preprocessing steps are compatible with them.
The important boundary is when those transformations are fitted. During hyperparameter optimization, every cross-validation fold creates and fits a fresh preprocessor on the fold's training rows, then transforms the validation rows with that fitted state. Imputation, encoding, scaling and optional feature transforms therefore stay on the training side of the split instead of being learned once from the full dataset.
The public prediction pipeline keeps the selected preprocessor next to the estimator. It also decodes internally encoded targets back to their original labels, and one-hot paths tolerate previously unseen categories at prediction time.
Selection can be fast or strict
For exploratory work, the default single validation split keeps the cost manageable. For a more careful comparison, selection_strategy="nested_cv" repeats the full tuning procedure inside outer folds over the non-holdout modeling data. Preprocessing and hyperparameter optimization are therefore local to each outer training fold rather than reused from the final fit.
Nested selection also avoids treating tiny score differences as decisive. Candidates within a configurable practical margin are treated as close contenders, then compared by score variability and selection runtime. If observations share a subject, household, session or other unit, group labels can be passed so those related rows stay on the same side of validation and holdout boundaries.
from mamut import Mamut
mamut = Mamut(
search_profile="balanced",
selection_strategy="nested_cv",
holdout_size=0.2,
refit_final_model=True,
random_state=42,
)
mamut.fit(X, y, groups=group_id)
evidence = mamut.generate_evidence(
dataset="holdout",
include_candidate_comparison=False,
)
The evidence layer can disagree with the winner
After fitting, Mamut can run a second layer of checks around the selected model. It screens for obvious leakage risks such as exact target copies, target-like feature names, nearly unique identifier columns, deterministic single-feature mappings and suspicious duplicate patterns. It also evaluates fixed dummy, logistic-regression and random-forest baselines alongside the selected candidate.
Score stability is estimated with repeated stratified cross-validation, or repeated group-disjoint folds when groups are supplied. Each fold refits preprocessing from scratch. The resulting intervals are reported as descriptive stability summaries rather than formal independent confidence intervals.
If a baseline or alternate candidate beats the selected model by a practical margin, the evidence report marks the result as challenged. Critical leakage can block trust in the selection altogether. On a final holdout, Mamut reports a stronger challenger for review without silently replacing the selected model, because choosing a new winner from that holdout would turn the final check into another model-selection set.
From experiment to prediction artifact
Once the validation evidence is accepted, refit_final_model=True fits the selected model on all non-holdout modeling data. With nested selection, the selected model family is retuned on that modeling data before the final fit. Holdout rows remain outside the training path.
The resulting best_model_ is a scikit-learn pipeline that contains preprocessing and the selected estimator, so prediction does not depend on reconstructing notebook state. It can be serialized as a .joblib artifact. evaluate() can also produce an HTML experiment report with validation or holdout metrics, model comparisons, integrity checks, confusion matrices, ROC curves, feature importance and optional SHAP explanations.
Reproducible benchmark campaigns
The repository includes a release benchmark that is deliberately capable of exposing an unimpressive result. In one recorded run on scikit-learn's breast-cancer dataset, the validation-selected random forest scored below a logistic-regression baseline on the holdout, and Mamut marked the selection as challenged. That is the intended behavior of the evidence layer.
A separate Spaceship Titanic harness pushes the same idea further. Each development run writes a result file and manifest containing configuration, data hashes, source commit, branch state, evaluation estimand and relational-overlap audit. A campaign reserves a confirmation partition that can be evaluated once after a candidate is frozen. Submission generation then refits the recorded hyperparameters instead of starting another tuning search. The recorded CatBoost campaign reached a public Kaggle score of 0.80617, with the development and confirmation history preserved alongside it.
Package engineering
Mamut is distributed through PyPI and documented on Read the Docs. The development pipeline installs from a locked uv environment, checks dependency declarations, runs the test suite, executes the documentation notebook, builds Sphinx documentation with warnings treated as errors, builds the package, validates its metadata and smoke-tests the generated wheel.
The repository also runs dependency-health checks and a scheduled security audit. Those are mundane parts of a machine-learning tool, but they matter if an experiment is supposed to remain runnable after the notebook that created it is gone.