Training guide
Every command runs from the repository root. uv sync --all-packages first.
MAMBA_SSM_AVAILABLE=0 is recommended unless you have the CUDA Mamba extension
installed: importing mamba_ssm can hang for minutes while it builds kernels.
Classical machine learning
Section titled “Classical machine learning”uv run inmotion train classification \ --data data/processed/dataset.csvRuns EDA, cross-validates a set of scikit-learn classifiers, optionally tunes
them with Optuna, computes feature importance, and writes plots plus a results
CSV. Add --optimize for the Optuna stage and --dashboard to launch
optuna-dashboard.
Deep learning baselines
Section titled “Deep learning baselines”MAMBA_SSM_AVAILABLE=0 uv run inmotion train dl \ --data data/processed/dataset.csv --seed 42 --trials 50Phases, in order, each skippable by flag:
| phase | flag to skip | what it does |
|---|---|---|
| 1 | --no-optuna |
trains each single model on its own GPU (round-robin) |
| 2 | --no-optuna |
Optuna HPO per model |
| 3 | --no-optuna |
NAS |
| 4 | --no-moe |
mixture-of-experts variants |
| 5 | --no-deepstack |
DeepStackEnsemble, a 10-base stack with a learned meta-model |
| — | --no-meta |
metadata-fusion variants |
Checkpoints land in --models-dir, results in --results-dir. An existing
checkpoint is reused unless --force-retrain is given.
World models (SSL)
Section titled “World models (SSL)”# find a configuration firstMAMBA_SSM_AVAILABLE=0 uv run inmotion train exotic \ --model ts_jepa --data data/processed/dataset_augmented3.csv \ --hpo --hpo-trials 50
# pretrain plus fine-tuneMAMBA_SSM_AVAILABLE=0 uv run inmotion train exotic \ --model ts_jepa --data data/processed/dataset_augmented3.csv --seed 42 \ --pretrain-epochs 500 --finetune-epochs 80 --batch-size 512Two-stage by nature: self-supervised pretraining, then supervised fine-tuning.
The JEPA models (t_jepa, ts_jepa, lejepa, cf_jepa) use both stages;
sigreg and the mamba3_* hybrids are directly supervised and take --epochs.
Stage-1 checkpoint selection is driven by a linear-probe MCC, not by the
masked-prediction loss, because the loss does not track representation quality.
--probe-cadence and --probe-patience control that probe.
Resume from a pretrained encoder:
inmotion train exotic --model ts_jepa \ --checkpoint 20-aug/models/exotic/normal-new-ds/ts_jepa_pretrain_best.pt \ --finetune-onlyThe paper’s HPO-best configs
Section titled “The paper’s HPO-best configs”The Optuna databases were deleted, so the winning configurations are re-declared from the documented values and retrained:
MAMBA_SSM_AVAILABLE=0 uv run inmotion train hpo-paper \ --data data/processed/dataset.csv --seed 42Values come from docs/icaisf/hpo_supplementary.md.
Ensembles
Section titled “Ensembles”# greedy selection over existing checkpoints (Caruana forward selection)uv run inmotion evaluate ensemble \ --data data/processed/dataset_augmented3.csv --seed 42
# OOF stacking over every family, with calibrationMAMBA_SSM_AVAILABLE=0 uv run inmotion evaluate mega-ensemble \ --data data/processed/dataset_augmented3.csv --seed 42mega-ensemble combines DL checkpoints, the JEPA family, classical ML and
TabPFN. It needs the member checkpoints present; inmotion model list shows
which are.
On a cluster
Section titled “On a cluster”scripts/slurm/*.sh carry #SBATCH headers:
sbatch scripts/slurm/job_full_train.shscripts/README.md says which script does what and which are superseded.
Credentials come from the environment (env.example), never from the scripts.
Reproducing a run
Section titled “Reproducing a run”Record these four things and the run is repeatable: the command, --seed, the
dataset (inmotion datasets show <id> gives its checksum), and the git commit.
inmotion doctor prints the environment. See docs/reproducibility.md.