mlvern is a Python library for structuring machine learning workflows with consistent dataset handling, experiment tracking, and model management.
It provides a lightweight framework to organize ML projects by separating data processing, experimentation, and evaluation into reproducible units.
https://ml-vern.readthedocs.io/en/latest/
- Dataset registration with fingerprint-based identification
- Metadata tracking for datasets and experiments
- Structured experiment execution workflow
- Model artifact storage and retrieval
- Evaluation tracking and comparison across runs
- Simple prediction interface for trained models
- Utilities for dataset inspection and validation
mlvern is built around the following principles:
- Reproducibility: identical inputs produce identical tracked outputs
- Traceability: datasets, experiments, and models are versioned and linked
- Simplicity: minimal API surface with explicit behavior
- Separation of concerns: data, training, and evaluation are decoupled
- Lightweight structure: avoids unnecessary abstraction layers
pip install mlvernfrom mlvern import Forge
forge = Forge("your_project", "your_base_dir")
forge.init()
dataset_fp, _ = forge.register_dataset(df, "target")
run_id, metrics = forge.run(model, X_train, y_train, X_val, y_val, config, dataset_fp)
from mlvern import ModelComparator
ModelComparator(forge).compare_models([run_id])Python 3.8+, NumPy, Pandas