Skip to content

Module reference

Generated by scripts/docs/generate_reference.py from the source itself, so it cannot drift from the code. Each entry lists the module’s purpose and its public classes and functions.

Private names (leading underscore) are omitted; they are implementation detail.

src/inmotion/__init__.py - 8 lines

inMotion — WiFi RSSI mobility classification.

Single importable package for the project’s data, model, training, evaluation and CLI surfaces. Legacy top-level modules (dl, ml_classification) are being folded in here; see docs/reports/ for the migration notes.

No public classes or functions; this module provides data or constants consumed elsewhere.

src/inmotion/cli/__init__.py - 73 lines

The inmotion command-line interface.

One package, one module per command group:

model inspect, verify, predict and adopt checkpoints datasets dataset versions, provenance and checksums data build and clean datasets (export, merge, clean, augment) train train models (classification, dl, exotic, hpo-paper) evaluate evaluate and combine models (ensemble, checkpoints, shap, importance) analyze analyse results (seeds, interference) plots regenerate figures from saved results doctor report on the environment and the artifact store

Each command builds a typed RunConfig and calls the pipeline’s run(). The CLI holds no pipeline logic, and the pipelines hold no argument parsing, so either side can change without touching the other.

Public API:

  • func main - [bold]inMotion[/] — WiFi RSSI mobility classification.

src/inmotion/datasets.py - 362 lines

Dataset version registry.

The datasets are committed to git, so their history is the version history. What git cannot tell you is which version is safe to train on: earlier collections were assembled with different device pools, and the project owner reports that older versions carry incorrect class labels. That knowledge is recorded here, next to checksums that prove a file has not silently changed.

Two kinds of information are kept apart on purpose:

  • Measured — sha256, row count, class counts, device count, size, and the commit that first added the file. Regenerated by :func:build_manifest, so it cannot drift.
  • Curated — role, lineage, status and the label audit. Written by hand, because it is project knowledge rather than something derivable from the file.

:func:verify` re-checks the measured side, so a dataset that changed under git’s feet (a re-export written over the same path) is caught.

Public API:

  • class DatasetVersion - One recorded version of one dataset file.
  • func measure - Fill in the measured fields from the file on disk.
  • func build_manifest - Regenerate the whole manifest, measuring every curated version.
  • func write_manifest
  • func load_manifest - Read the manifest, or None when it has not been generated yet.
  • func verify - Check every recorded dataset against its checksum and row count.
  • func resolve - Find a dataset version by id or file name.
  • func canonical - The dataset new training runs should use.

src/inmotion/data/paths.py - 99 lines

Canonical locations for data files, and resolution of dataset references.

The project has three data tiers, matching how a dataset is actually produced:

raw/ Captures straight off the access point (Wavecom .txt dumps). Never edited; the input to inmotion data export. interim/ One CSV per collection session, as written by export. Input to inmotion data merge. processed/ Merged and augmented datasets that training and evaluation consume.

Before this module the tiers all lived in the repository root or in wavecom_files/, and code referred to datasets by bare filename (dataset.csv) from whatever the current directory happened to be.

:func:resolve keeps those bare references working by searching the tiers in order, so an old command line such as --data dataset.csv still finds the file, while new code uses explicit paths.

Public API:

  • func resolve - Resolve a dataset reference to a path.
  • func processed - Path to a processed dataset, by bare filename.
  • func raw - Path to a raw capture, by bare filename.
  • func interim - Path to an interim per-collection CSV, by bare filename.
  • func list_processed - Every file in the processed tier, sorted.
  • func list_raw - Every raw capture, sorted.

src/inmotion/data/preprocessing.py - 185 lines

Reproducible RSSI feature pipeline.

The training code standardised RSSI values with a StandardScaler and then engineered extra channels, but never persisted the scaler. Two different conventions were in use, which is why a saved model could not be re-run without also having the exact training CSV:

inmotion/dl/data_loader.py scaler.fit_transform(X_raw) on the whole file, before the split. inmotion/pipelines/mega_ensemble.py StandardScaler().fit(_raw_X[train_idx]) on the training fold only.

:class:FeaturePipeline records which convention produced a model, stores the fitted statistics next to the checkpoint, and replays them at inference. A pipeline whose statistics were rebuilt from a reference CSV rather than recovered from training is flagged reconstructed so results are never presented as more reproducible than they are.

Public API:

  • class FeaturePipeline - Standardisation + channel expansion for RSSI sequences.

src/inmotion/data/export.py - 130 lines

Public API:

  • func extract_rssi_and_mac_pairs
  • func export_to_csv

src/inmotion/data/merge.py - 133 lines

Merge multiple CSV files into a single dataset.

Public API:

  • func extract_concurrent_noise_path - Extract the concurrent noise path from filename and row label.
  • func merge_csv_files - Merge multiple CSV files into a single dataset.

src/inmotion/data/clean.py - 215 lines

Dataset cleaning audit — find and drop devices/samples that poison training.

Why: on the current data/processed/dataset.csv the absolute RSSI level carries ~3x more information about WHICH DEVICE recorded a sequence than about WHICH movement it is (MI with MAC 0.219 vs MI with label 0.071). Several devices sit at/below chance under cross-device evaluation, i.e. their sequences genuinely do not look like their labels. This script quantifies that per device, drops the poisoned parts, and writes a cleaned CSV plus a full report.

within_mcc : stratified CV inside the device (does it agree with itself?) low -> labels are inconsistent/noisy for this device -> DROP lodo_mcc : leave-one-device-out (train on all other devices, predict this one) low -> device distribution is shifted vs everyone else -> DROP (unless –keep-shifted: keep self-consistent-but-shifted devices)

OOF probabilities from a KNN on the surviving rows: samples the model rejects with high confidence (max prob > –ambiguity-thresh AND wrong class) are reported to _flagged.csv; dropped too with –drop-ambiguous.

MAMBA_SSM_AVAILABLE=0 uv run inmotion data clean -- --data data/processed/dataset.csv
# stricter: also drop ambiguous samples
MAMBA_SSM_AVAILABLE=0 uv run inmotion data clean -- --drop-ambiguous

Public API:

  • func run - Execute the pipeline.

src/inmotion/data/augment.py - 127 lines

Data Augmentation for RSSI sequences.

Two strategies:

  1. Class-Conditioned Noise Augmentation
    • noise=False (clean): inject heavier Gaussian noise to teach model how “clean” data looks when it gets dirty.
    • noise=True (noisy): apply only light jitter so the real signal is not overwhelmed.
  2. Intra-class Mixup λ ~ Beta(α, α), X_new = λ·X₁ + (1-λ)·X₂ (same label only)

All synthetic rows get synthetic=True; originals get synthetic=False.

Public API:

  • func noise_augment - Return a DataFrame of n_copies noisy versions of every row in df.
  • func mixup_augment - Intra-class Mixup: X_new = λ·X₁ + (1−λ)·X₂, same label only.
  • func augment

src/inmotion/data/splits.py - 43 lines

The train/validation/test split convention.

One definition, because every reported number depends on it. The convention is an 80/20 stratified train-test split, then 12.5% of the train portion held out for validation, both with a fixed seed.

This lived in pipelines/mega_ensemble.py and was imported from there by three other modules, which meant a 1,000-line pipeline had to be imported to obtain a fifteen-line function. It is a data concern, so it lives with the data.

Public API:

  • func stratified_indices - Return (train, validation, test) index arrays.

src/inmotion/quant/spec.py - 150 lines

Quantization schemes.

Weight-only quantization: the weights are stored at low precision, activations stay in fp32. That is the right trade here because weights dominate this workload by 9x-2424x over activations (a full inference window is 10 timesteps x 4 channels = 160 bytes), so the entire memory problem is the weights. Keeping activations in fp32 also removes the need for calibration data, which keeps Q4 and Q8 exact and reproducible rather than data-dependent.

Rounding is specified explicitly. NumPy’s np.round rounds half to even, while Rust’s f32::round rounds half away from zero; leaving that implicit would make the Python reference and the Rust kernels disagree on roughly one element in a thousand, which is exactly the kind of drift the golden-vector tests exist to catch.

Public API:

  • class Granularity - How many weights share one scale/zero-point pair.
  • class Rounding - Rounding mode applied before the integer cast.
  • class QuantSpec - A weight-quantization scheme.
  • func resolve - Look up a preset by name, or parse q<N> shorthand.

src/inmotion/quant/reference.py - 203 lines

Normative quantization math, in numpy.

This module is the specification. The Rust kernels in rust/inmotion-core must reproduce it exactly, and tests/test_quant_reference.py pins the properties that make “exactly” checkable:

  1. dequantize is exactly (q - zero_point) * scale in float32;
  2. re-quantizing a dequantized tensor is idempotent (no compounding drift);
  3. rounding ties follow the declared mode, because numpy rounds half to even while Rust rounds half away from zero;
  4. granularity is honoured, so per-group is never silently coarser than asked;
  5. the byte accounting uses the packed size, not one int8 per value.

scale and zero_point are stored in the reduced shape that broadcasts against the codes, so no reader has to re-derive it:

=============== ========================================== ================== granularity reduced shape dequant broadcast =============== ========================================== ================== per_tensor () scalar per_channel (1, .., D, .., 1) with D=shape[axis] direct per_group (*shape[:-1], ceil(D/g)) scale[..., None] =============== ========================================== ==================

Public API:

  • func n_groups_for - Number of groups the last axis is divided into, rounding up.
  • class QuantizedTensor - Integer codes plus the metadata needed to reconstruct the float values.
  • func quantize - Quantize a float tensor into integer codes and per-group metadata.
  • func dequantize - Reconstruct float32 values from codes and metadata.

src/inmotion/quant/packing.py - 145 lines

Bit-packing of integer codes.

Storing a 4-bit code in an int8 wastes half the memory, and this workload is memory-bound, so the packing is not an optimization — it is the point. Q4 must cost half of Q8 and Q2 a quarter of it, or the exercise is pointless.

Codes are biased to unsigned so the byte stream holds no negatives, and written LSB-first: element i occupies bits [i*bits, (i+1)*bits) of the byte stream, so for 4-bit codes element 0 is the low nibble of byte 0 and element 1 the high nibble.

The bias is 2**(bits-1) for a symmetric scheme, whose codes are signed and run [-2**(bits-1), 2**(bits-1)-1]. An asymmetric scheme emits codes that are already unsigned ([0, 2**bits-1]), so its bias is zero; applying the signed bias there would push the top code out of range, which is why q1 could not be exported until this parameter existed.

Biasing means the on-device dequantization stays a single formula with no extra per-weight work: because q = u - bias,

(q - zero_point) * scale == (u - (zero_point + 2**(bits-1))) * scale

so the bias is folded into the zero point once, at export, by :func:fold_bias_into_zero_point.

Public API:

  • func n_bytes_for - Number of bytes needed to pack count codes of bits each.
  • func fold_bias_into_zero_point - Return the zero point that pairs with biased unsigned codes.
  • func code_bias - Offset added to a code to make it unsigned.
  • func pack - Pack int8 codes into a uint8 byte stream, LSB-first.
  • func unpack - Inverse of :func:pack. Returns int8 codes.

src/inmotion/quant/model.py - 272 lines

Weight-only quantization of a model, and honest size accounting.

Only the weights of Linear and Conv1d/Conv2d layers — the tensors with two or more dimensions. Everything else stays fp32:

  • biases, LayerNorm/BatchNorm weights and running statistics are 1-D and a tiny share of the parameters; quantizing them costs accuracy and saves nothing;
  • learned tokens and position embeddings (mask_token, query_tokens, pos_embed) are initialization state, not weights to be approximated.

Activations stay fp32. That is deliberate: activations peak at ~10 KB against 0.57-2.35 MB of weights, so keeping them fp32 costs nothing measurable and removes the need for calibration data.

:func:apply replaces each weight with dequantize(quantize(w)). Running the model afterwards measures exactly what a device does, because the device also reconstructs fp32 weights from codes before the matmul. No forward pass needs to be rewritten, so the measurement cannot drift from the model.

Public API:

  • class TensorRecord - One quantized weight tensor.
  • class QuantReport - Size and error summary for one (model, spec) pair.
  • func quantize_tensor - Quantize one weight tensor.
  • func apply - Return a model whose quantizable weights have been quantized and rebuilt.
  • func apply_cast - Round every parameter and buffer to a narrower float format.

src/inmotion/quant/evaluate.py - 344 lines

Accuracy under quantization, measured on data/processed/dataset.csv.

Every number this module produces is on dataset.csv. It is not on dataset.icaisf.csv, so it is not comparable with the MCC values published for the icaisf variant (lejepa 0.8850, t_jepa 0.8694, ts_jepa 0.8785, cf_jepa 0.8661, sigreg 0.8858). The two splits differ in size and in score, so the numbers are not interchangeable.

Otherwise the harness reuses the same splits and the same feature pipeline the models were trained with:

  • features from DLDataLoader with the model’s own rich flag, because the scaler is fitted inside it and using a different fit would change the inputs;
  • the 80/20 stratified split from inmotion.data.splits, the convention every reported number uses.

The first thing the sweep does is evaluate the unquantized model. If that does not reproduce the recorded MCC, the harness is wrong and every later number is meaningless, so :func:check_baseline fails loudly instead of publishing a table nobody can trace.

Public API:

  • class EvalRow - One (model, precision) measurement.
  • func load_split - Features, labels and test indices, using the training convention.
  • func measure_producer - (quantized model, bits, total bytes, quantized bytes, max weight error).
  • func evaluate - Measure one (model, precision) pair on the test split.
  • func check_baseline - Confirm fp32 reproduces the recorded MCC on the baseline dataset.
  • func sweep - Evaluate every (model, precision) pair and optionally write a CSV.
  • func write_csv - Write the sweep as one row per measurement.
  • func write_json

src/inmotion/quant/export.py - 582 lines

Export a model to a portable artifact a non-Python runtime can execute.

The artifact is a directory:

model.safetensors tensors, quantized ones packed
meta.json architecture, quantization spec, tensor index

ONNX Runtime is a poor fit for an MCU and heavy even for an x86 router, and the graphs here are small and fixed-shape. A flat tensor bag plus a small JSON description is easier to audit and is what embedded runtimes actually do.

Two source conventions are resolved here rather than in every Rust reader:

in_proj_weight is split into separate Q, K and V tensors. PyTorch packs them into one (3*d_model, d_model) matrix; splitting means the Rust side does not have to know that packing convention.

weight_norm is folded into an ordinary weight. It is what broke FX tracing and ONNX export for tcn/deep_stack, and folding removes the problem instead of working around it.

dtype="fp32" writes plain float weights. Use it to verify that the port is correct before introducing quantization, so a mismatch can only be one of the two things.

quant=QuantSpec writes packed integer codes plus scales and zero points. The Rust reader dequantizes exactly as :mod:inmotion.quant.reference does, so parity holds by construction and is asserted by the golden-vector tests.

Public API:

  • class ExportReport - What was written and how big it is.
  • func dequantize_artifact - The float weights an artifact stands for: (u - zero_point) * scale.
  • func rebuild_state_dict - The model’s OWN state dict, from artifact-shaped tensors.
  • func model_with_producer - The model carrying a producer’s quantization in its weights.
  • func resolve_producer - The per-weight quantizer and the quant metadata block for a producer.
  • func export - Write model_name as a portable artifact directory.
  • func write_golden - Write input/output pairs for a non-Python runtime to reproduce.

src/inmotion/models/registry.py - 377 lines

Model registry — the single source of truth for loadable models.

A :class:ModelSpec records everything needed to rebuild a saved model:

factory which architecture constructor to call (see factories.py) arch the keyword arguments for it prefix state-dict wrapping, if the checkpoint was saved from a wrapper params expected parameter count, used as a checksum

Loading is deliberately strict. The previous code called load_state_dict(..., strict=False), which turns an architecture mismatch into a silently random subset of weights — the model still runs and still returns numbers, so the error is invisible. :func:load refuses by default.

Adding a model is data entry, not code. Use adopt (see inmotion model adopt) to infer a spec from a checkpoint’s tensor shapes.

Public API:

  • class ModelSpec - A rebuildable saved model.
  • func spec_names - Names of every registered model.
  • func by_name - Look up a spec by name, optionally overriding its checkpoint path.
  • func all_specs - Every registered spec, including adopted ones.
  • func load - Build spec and load its checkpoint.
  • func param_count - Parameter count of the architecture as built from spec.arch.

src/inmotion/models/factories.py - 751 lines

Architecture factories.

Each factory takes only keyword arguments and is registered in FACTORIES under a stable name. A checkpoint’s architecture is therefore data (an arch dict in the registry or a checkpoint sidecar), not code that has to be reverse-engineered from tensor shapes.

The configurations here were originally recovered by matching saved weight shapes; they are pinned by tests/test_checkpoint_contract.py, which asserts every documented checkpoint still loads with strict=True and matches the parameter count published for it.

Public API:

  • func lejepa - LeJEPA classifier (paper after-HPO winner: d_model=128, ff=256, layers=4).
  • func t_jepa - T-JEPA classifier (paper after-HPO winner: d_model=128, ff=256).
  • func ts_jepa - TS-JEPA classifier (paper: embed_dim=256, layers=4, ff=512).
  • func cf_jepa - CF-JEPA classifier (paper after-HPO: embed 256, layers 3, ff 256).
  • func sigreg - SIGReg classifier wrapped so it returns logits only.
  • func mamba3_cnn - Mamba-3 CNN hybrid.
  • func mamba3_tcn - Mamba-3 TCN hybrid.
  • func mamba3_transformer - Mamba-3 transformer hybrid.
  • func mamba3_multiview - Mamba-3 multi-view hybrid (checkpoint uses d_model=256).
  • func deep_stack - Deep stack ensemble, 10 heterogeneous bases.
  • func build - Instantiate factory with arch keyword arguments.
  • func rnn - Plain RNN classifier.
  • func gru - GRU classifier, optionally with attention pooling.
  • func lstm - LSTM classifier, optionally with attention pooling.
  • func bilstm - Bidirectional LSTM classifier.
  • func tcn - Temporal convolutional network classifier.
  • func mamba - Mamba (selective SSM) classifier.
  • func transformer - Transformer encoder classifier.
  • func cnn - Multi-branch residual CNN classifier.
  • func cnn2d_rnn - 2-D CNN front-end feeding an RNN.
  • func meta_fusion - RNN plus metadata (noise flag, concurrent path) fusion classifier.
  • func voting_ensemble - Averaging ensemble over nested member specs.
  • func stacking_ensemble - Stacking ensemble: base members feed a learned meta-learner.
  • func lejepa_pretrain - Base LeJEPA world model (no classification head).
  • func t_jepa_pretrain - Base T-JEPA world model.
  • func ts_jepa_pretrain - Base TS-JEPA world model.
  • func cf_jepa_pretrain - Base CF-JEPA world model.

src/inmotion/models/checkpoint.py - 166 lines

Checkpoint I/O with a JSON sidecar.

A .pt file alone is not a usable artifact: it does not say which architecture it belongs to, how the inputs were scaled, or what the output classes mean. :class:Sidecar carries that information next to the weights, and :func:save / :func:load keep the two in sync.

Layout written by :func:save::

<dir>/model.pt weights
<dir>/model.json Sidecar
<dir>/preprocessing.json FeaturePipeline metadata
<dir>/preprocessing.npz FeaturePipeline statistics

Loading verifies the sidecar against the weights: the architecture must build, the state dict must load with strict=True, and the parameter count must match. Anything else raises rather than returning a partly-random model.

Public API:

  • class Sidecar - Everything needed to rebuild and interpret a saved model.
  • func save - Write weights, sidecar and (optionally) preprocessing state.
  • func read_sidecar
  • func load - Rebuild the model, its sidecar and (if present) its preprocessing state.
  • func sidecar_for_spec - Build a sidecar from a registry :class:ModelSpec.

src/inmotion/models/adopt.py - 775 lines

Infer a :class:~inmotion.models.checkpoint.Sidecar from an untagged checkpoint.

Existing checkpoints predate the sidecar convention, so their architecture has to be recovered. Two strategies are tried in order:

  1. Signature match. Every registered spec’s module is built and its state-dict key set and shapes are compared with the checkpoint’s. An exact match identifies the architecture with no guessing.
  2. Shape-driven search. For a checkpoint nothing in the registry matches, a bounded per-family grid is searched until a configuration loads with strict=True. Dimensions are read out of the tensor shapes rather than hardcoded.

If neither succeeds, the error lists what was tried and how each candidate differed, so a factory can be added deliberately instead of the model being loaded with strict=False and left partly random.

Checkpoints may or may not carry a wrapper prefix ("c." for the JEPA classifier, "m." for :class:~inmotion.models.wrappers.TupleWrap). Detection works on plain keys; the prefix a factory needs is applied only when checking a candidate, so a raw SIGReg checkpoint and a re-saved wrapped one are both recognised.

Public API:

  • func read_state - Load a state dict, accepting a raw dict or a module checkpoint.
  • func plain_state - Strip any wrapper prefix so detection sees canonical child names.
  • func match_all - Every registered architecture that reproduces plain exactly.
  • func match_registry - First registered architecture that reproduces plain exactly.
  • func search_family - Search the per-family grids for a strictly-loadable configuration.
  • func resolve_plain - Resolve a plain (prefix-stripped) state dict to (factory, arch, prefix).
  • func looks_like_pretrain - True when state is a pre-training bundle rather than a state dict.
  • func describe_pretrain - A one-line description of a pre-training bundle.
  • func spec_for_checkpoint - A :class:ModelSpec for any checkpoint, architecture inferred.
  • func resolve_model - A registered name, or a path to any checkpoint.
  • func infer - Recover a :class:Sidecar for an untagged checkpoint.
  • func infer_many - Adopt every checkpoint in paths, recording failures instead of raising.

src/inmotion/models/wrappers.py - 55 lines

Thin nn.Module adapters and state-dict key helpers.

Some checkpoints were saved from a wrapper module, so their keys carry a prefix that the bare module does not have (or vice versa). These helpers normalise that without the callers needing to know which case they are in.

Public API:

  • class TupleWrap - Expose logits of a model that returns (logits, latents).
  • func strip_prefix - Remove prefix from every key that starts with it.
  • func add_prefix - Prepend prefix to every key.
  • func normalise_state - Return state shaped the way the target module expects.

src/inmotion/cli/__init__.py - 73 lines

The inmotion command-line interface.

One package, one module per command group:

model inspect, verify, predict and adopt checkpoints datasets dataset versions, provenance and checksums data build and clean datasets (export, merge, clean, augment) train train models (classification, dl, exotic, hpo-paper) evaluate evaluate and combine models (ensemble, checkpoints, shap, importance) analyze analyse results (seeds, interference) plots regenerate figures from saved results doctor report on the environment and the artifact store

Each command builds a typed RunConfig and calls the pipeline’s run(). The CLI holds no pipeline logic, and the pipelines hold no argument parsing, so either side can change without touching the other.

Public API:

  • func main - [bold]inMotion[/] — WiFi RSSI mobility classification.

src/inmotion/cli/model.py - 292 lines

inmotion model - inspect, verify, predict and adopt checkpoints.

Public API:

  • func model_list - List every registered model.
  • func model_info - Show the architecture and provenance of one model.
  • func model_verify - Strict-load every model and check it against its recorded parameter count.
  • func model_predict - Run a trained model over a CSV of RSSI windows.
  • func model_adopt - Recover a checkpoint’s architecture and write a sidecar.
  • func model_inventory - Adopt every checkpoint under a directory and report what is recoverable.

src/inmotion/cli/datasets.py - 131 lines

inmotion datasets - dataset versions, provenance and checksums.

Public API:

  • func datasets_list - List recorded dataset versions, newest lineage last.
  • func datasets_show - Show lineage and label provenance for one dataset version.
  • func datasets_verify - Re-check every dataset against its recorded checksum and row count.
  • func datasets_manifest - Regenerate the manifest’s measured fields, or print it.
  • func datasets_resolve - Print the path of a dataset version, warning when it is not trustworthy.

src/inmotion/cli/data.py - 99 lines

inmotion data — build and clean datasets.

Public API:

  • func export - Pivot 10-second captures into one row per route.
  • func merge - Concatenate the interim CSVs and label concurrent-noise rows.
  • func clean - Quantify which devices or samples carry inconsistent labels.
  • func augment - Add class-conditioned noise and mixup rows, flagged synthetic.

src/inmotion/cli/train.py - 277 lines

inmotion train — train models.

Public API:

  • func dl - Train the supervised deep-learning models and their ensembles.
  • func exotic - Pretrain and fine-tune a world model, or train SIGReg / a Mamba-3 hybrid.
  • func classification - Run the classical machine-learning comparison.
  • func hpo_paper - Rebuild the HPO-best models, since the Optuna databases were deleted.

src/inmotion/cli/evaluate.py - 152 lines

inmotion evaluate — evaluate and combine models.

Public API:

  • func ensemble - Combine saved checkpoints with Caruana forward selection.
  • func mega_ensemble - Combine DL, JEPA, classical ML and TabPFN members into one model.
  • func checkpoints - Evaluate every checkpoint in a directory with extended metrics.
  • func shap - Attribute predictions to the 18 feature channels.
  • func feature_importance - Shuffle each feature channel and measure the MCC drop.

src/inmotion/cli/analyze.py - 47 lines

inmotion analyze — analyse results across seeds and routes.

Public API:

  • func seeds - Compute cross-seed statistics and ranking stability.
  • func interference - Quantify how a concurrent device’s route degrades classification.

src/inmotion/cli/plots.py - 99 lines

inmotion plots — regenerate figures from saved results.

Public API:

  • func classical - Rebuild the figures without retraining.
  • func dl - Rebuild the DL figures without retraining.
  • func quantization - Rebuild the six quantization figures without re-measuring anything.

src/inmotion/cli/quantize.py - 429 lines

inmotion quantize — quantize models and measure what it costs.

Public API:

  • func formats - Show each scheme, its bit width and its storage granularity.
  • func show - Report the size, compression and weight error for one scheme.
  • func producers_cmd - Which quantizer made a number matters, so they are named, not implied.
  • func export_cmd - Export an artifact plus the golden vectors a Rust runtime must reproduce.
  • func evaluate - Sweep models x formats and write one row per measurement.
  • func stats - Produce size, accuracy, coverage, latency and per-layer tables.

src/inmotion/cli/doctor.py - 41 lines

inmotion doctor - report on the environment and the artifact store.

Public API:

  • func doctor - Report on the environment and the artifact store.

src/inmotion/pipelines/classification.py - 254 lines

Main script for WiFi Fingerprinting Classification Analysis.

This script runs a comprehensive analysis of various ML classifiers for predicting location classes based on WiFi RSSI fingerprints.

Public API:

  • func save_detailed_classifier_reports - Save detailed per-classifier reports to a CSV file.
  • func run - Execute the pipeline.

src/inmotion/pipelines/dl/__init__.py - 42 lines

DL pipeline: supervised baselines, HPO/NAS, MoE and DeepStack.

Split from a single 1,700-line module into:

config model lists and the typed :class:RunConfig devices GPU pool selection and seeding builders model construction, including from saved HPO parameters workers spawn-safe worker functions used by ProcessPoolExecutor pipeline :func:run, the orchestration

No public classes or functions; this module provides data or constants consumed elsewhere.

src/inmotion/pipelines/dl/config.py - 82 lines

DL pipeline configuration: model lists and run options.

Public API:

  • class RunConfig - Typed options for the DL pipeline.

src/inmotion/pipelines/dl/devices.py - 28 lines

Device pool selection and seeding.

Public API:

  • func set_seed

src/inmotion/pipelines/dl/builders.py - 297 lines

Model construction for the DL pipeline.

No public classes or functions; this module provides data or constants consumed elsewhere.

src/inmotion/pipelines/dl/workers.py - 723 lines

Spawn-safe worker functions for parallel training and HPO.

No public classes or functions; this module provides data or constants consumed elsewhere.

src/inmotion/pipelines/dl/pipeline.py - 626 lines

DL pipeline orchestration.

Public API:

  • func run - Execute the DL pipeline.

src/inmotion/pipelines/exotic/__init__.py - 37 lines

World-model (self-supervised) pipeline.

Split from a single 1,950-line module into modules with one job each:

config the typed :class:RunConfig (54 options) devices device selection and seeding data pretraining and supervised data loading builders the model family and its dispatch table training stage-1 pretraining, linear probing, stage-2 fine-tuning hpo shared Optuna helpers and per-trial logging hpo_jepa the JEPA search (also holds the Pareto machinery) hpo_sigreg the SIGReg search hpo_mamba3 the Mamba-3 search pipeline :func:run, the orchestration

No public classes or functions; this module provides data or constants consumed elsewhere.

src/inmotion/pipelines/exotic/config.py - 78 lines

Options for the world-model pipeline.

Public API:

  • class RunConfig - Typed options for the exotic pipeline.

src/inmotion/pipelines/exotic/devices.py - 26 lines

Device selection and seeding.

Public API:

  • func set_seed
  • func resolve_device

src/inmotion/pipelines/exotic/data.py - 109 lines

Data loading for pretraining and supervised training.

Public API:

  • func load_pretrain_data - Load data for SSL pretraining.
  • func load_supervised_data - Load data for supervised training.

src/inmotion/pipelines/exotic/builders.py - 164 lines

Model construction for the world-model family.

No public classes or functions; this module provides data or constants consumed elsewhere.

src/inmotion/pipelines/exotic/training.py - 456 lines

Stage-1 pretraining, probing and stage-2 fine-tuning.

Public API:

  • func pretrain_jepa - Pretrain a JEPA model (T-JEPA or TS-JEPA).
  • func finetune_jepa - Fine-tune a pretrained JEPA encoder for classification.

src/inmotion/pipelines/exotic/hpo.py - 161 lines

Shared Optuna helpers and per-trial logging.

Public API:

  • class SIGRegWrapper - Wraps SIGRegClassifier so Trainer can use it with compute_loss.

src/inmotion/pipelines/exotic/hpo_jepa.py - 428 lines

Hyperparameter search for the JEPA models.

No public classes or functions; this module provides data or constants consumed elsewhere.

src/inmotion/pipelines/exotic/hpo_sigreg.py - 121 lines

Hyperparameter search for SIGReg.

No public classes or functions; this module provides data or constants consumed elsewhere.

src/inmotion/pipelines/exotic/hpo_mamba3.py - 109 lines

Hyperparameter search for the Mamba-3 hybrids.

No public classes or functions; this module provides data or constants consumed elsewhere.

src/inmotion/pipelines/exotic/pipeline.py - 284 lines

World-model pipeline orchestration.

Public API:

  • func run - Execute the exotic (world-model) pipeline.

src/inmotion/pipelines/ensemble.py - 272 lines

Greedy ensemble selection over existing checkpoints.

Research-backed combiner: Caruana’s forward selection with replacement (“ensemble selection from libraries of models”), which the 2026 TFM-ensembling study recommends as the practical default over plain weighted averaging and stacking. It adds the base whose inclusion maximizes validation accuracy at each step, with replacement, so diverse-but-redundant models can each earn weight without a meta-learner.

Add a model in one line by appending a dict to MODELS: { “name”: “foo”, “rich”: False, # True if it needs 18-channel rich features “build”: lambda: build_foo(), # returns an nn.Module, logits already loaded }

Run: uv run inmotion ensemble –data data/processed/dataset_augmented3.csv –seed 42

Public API:

  • func run - Execute the pipeline.

src/inmotion/pipelines/mega_ensemble.py - 1045 lines

Mega-ensemble across every model family.

Combines DL checkpoints, the JEPA family, classical ML (KNN/CatBoost/GP/LR) and TabPFN via OOF stacking, with temperature calibration and CE-based bagged Caruana selection.

Research-backed combiner pipeline. The combiner choices and their rationale: A. Collect OOF + test probabilities for every member under the same 70/10/20 split B. Per-member temperature calibration on the validation fold C. CE-based bagged Caruana (ensemble selection with replacement) D. Regularized-LR stacker over the calibrated OOF matrix E. Compare: greedy-weighted, logit-avg, rank-avg, entropy-weighted, LR stack

Usage: MAMBA_SSM_AVAILABLE=0 inmotion mega-ensemble –data data/processed/dataset_augmented3.csv –seed 42

Note: set MAMBA_SSM_AVAILABLE=0 to skip the (hang-prone) mamba_ssm CUDA extension import and use the pure-PyTorch fallback for Mamba-based models.

Public API:

  • func rocket_features - MiniRocket-style random convolutional features + PPV pooling.
  • func catch22_features - Hand-crafted time-series statistics per channel (catch22-inspired subset).
  • func fit_temperature - Fit a single temperature scalar to minimize NLL on the calibration set.
  • func run - Execute the pipeline.

src/inmotion/pipelines/hpo_paper.py - 312 lines

Train the paper’s HPO-best models using configs from docs/icaisf/hpo_supplementary.md.

The original Optuna DL DBs (optuna_dl_*.db) were deleted, so the DL pipeline can no longer skip HPO. This script rebuilds each HPO-best model from the paper’s documented config and trains it with the Trainer, saving checkpoints that the mega-ensemble can load directly.

Configs sourced from docs/icaisf/hpo_supplementary.md (best val MCC in parens): HPO GRU 0.9030 HPO TCN 0.8526 HPO Mamba 0.8506 HPO CNN 0.8488 HPO LSTM 0.8388 HPO BiLSTM 0.8410 HPO RNN 0.6402

Usage: MAMBA_SSM_AVAILABLE=0 inmotion hpo-paper –data data/processed/dataset.csv –seed 42 # –models only gru tcn (subset)

Public API:

  • func run - Execute the pipeline.

src/inmotion/evaluation/checkpoints.py - 827 lines

Retroactive DL checkpoint evaluation — produce extended results CSV with per-class metrics.

Usage: uv run inmotion evaluate-checkpoints –seed 42 –data data/processed/dataset.csv

Public API:

  • func set_seed
  • func evaluate_single_checkpoint - Load a checkpoint, build model, evaluate. Returns metrics dict or None.
  • func run - Execute the pipeline.

src/inmotion/evaluation/shap.py - 218 lines

Game-Theoretic SHAP Feature Channel Attribution for WiFi ISAC Models.

Evaluates Shapley Additive exPlanations (SHAP) across the 18-channel feature hierarchy (or the 4-channel base hierarchy) for the time-series world models: TS-JEPA, LeJEPA, CF-JEPA, T-JEPA and SIGReg.

Usage: inmotion shap –data data/processed/dataset.csv –checkpoint backup/ts_jepa_ft_seed42.pt –samples 50 –nsamples 150

Public API:

  • func explain_channels_shap - Compute channel-level Shapley values across test sequences.
  • func run - Execute the pipeline.

src/inmotion/evaluation/feature_importance.py - 612 lines

Permutation feature importance for DL movement classification models.

Computes MCC-drop when each of the 4 engineered channels × 10 timesteps (= 40 features) is independently shuffled, for CNN, TCN, GRU, and Mamba.

DeepStackEnsemble is skipped unless all 8 DS base-variant checkpoints already exist under models/dl/ (it requires 9 pre-trained sub-models).

Output plots (publication-quality PDF, 600 DPI, husl palette): docs/icaisf/paper/images/feature_importance_dl.pdf heatmap docs/icaisf/paper/images/feature_importance_channels_dl.pdf bar chart

Public API:

  • func set_seed
  • func get_device
  • func setup_plot_style - Apply DLVisualizer-compatible plot style.
  • func build_model - Construct a model matching _build_model_by_name in the DL pipeline.
  • func train_or_load - Train model from scratch or load saved checkpoint.
  • func permutation_importance - Return (importance[N_CHANNELS, N_TIMESTEPS], baseline_mcc).
  • func plot_heatmap - Publication-quality heatmap: 4 channels × 10 timesteps, avg across models.
  • func plot_per_model_heatmaps - Multi-panel figure: one 4x10 heatmap per model.
  • func plot_channel_bars - Channel-level bar chart: mean importance +/- std across timesteps & models.
  • func run - Execute the pipeline.

src/inmotion/analysis/seeds.py - 440 lines

Analyze and combine results from multi-seed experiments.

This script reads results from experiments with different seeds, computes statistics, and generates comparison plots.

Public API:

  • func setup_plot_style - Setup publication-ready plot style.
  • func load_seed_results - Load classification results for each seed.
  • func compute_statistics - Compute mean and std statistics across seeds.
  • func plot_metric_comparison_across_seeds - Plot metric comparison for each seed.
  • func plot_metric_variability - Plot metric variability across seeds with error bars.
  • func plot_seed_boxplot - Create boxplot showing metric distribution across seeds.
  • func plot_multi_metric_heatmap - Plot heatmap of mean metrics across seeds.
  • func plot_seed_stability_ranking - Plot how classifier rankings change across seeds.
  • func generate_summary_report - Generate a comprehensive summary report.
  • func run - Execute the pipeline.

src/inmotion/analysis/interference.py - 44 lines

Analyze cross-route interference using the concurrent_noise_path column.

Usage: inmotion analyze-interference –data data/processed/dataset.csv –output-dir plots/interference

Public API:

  • func run - Execute the pipeline.

src/inmotion/viz/regenerate.py - 142 lines

Regenerate plots from saved CSV results.

This script loads classification results from CSV files and regenerates all visualization plots without needing to retrain models.

Usage: uv run inmotion regenerate-plots –results-dir results_3 inmotion regenerate-plots –csv results_3/classification_results.csv –plots-dir plots_regenerated inmotion regenerate-plots –results-dir results_3 –eda –data data/processed/dataset.csv

Public API:

  • func run - Execute the pipeline.

src/inmotion/viz/regenerate_dl.py - 80 lines

Regenerate DL-specific paper plots from extended results CSV.

Usage: inmotion regenerate-dl-plots –results-csv results/dl/dl_detailed_seed42.csv –output-dir plots/dl/42 inmotion regenerate-dl-plots –results-csv results/dl/dl_detailed_seed42.csv –eda –data data/processed/dataset.csv

Public API:

  • func run - Execute the pipeline.

src/inmotion/dl/config.py - 101 lines

DL pipeline configuration.

Public API:

  • class DLConfig

src/inmotion/dl/data_loader.py - 343 lines

Data loading, preprocessing, and PyTorch Dataset/DataLoader utilities.

Public API:

  • class RSSIDataset - PyTorch Dataset for RSSI time-series sequences.
  • class MetaRSSIDataset - Dataset that returns (X, meta, y) triples for metadata-fusion models.
  • class DLDataLoader - Load, preprocess, and split the WiFi fingerprinting dataset.

src/inmotion/dl/training.py - 317 lines

Generic Trainer with WandB logging, early stopping, and L1 regularization.

Public API:

  • class Loggable
  • class TrainResult
  • class Trainer

src/inmotion/dl/optimization.py - 669 lines

Optuna HPO studies — per-model and NAS meta-model.

Public API:

  • func init_optuna_db - Pre-create SQLite schema + studies in the main process to avoid spawn race.
  • func run_hpo_study - Run HPO for a named model type. Returns the finished study.
  • func run_nas_study - NAS: Optuna searches over architectures AND hyperparams.
  • func run_binary_moe_hpo_study - HPO for a single binary MoE expert arch type.
  • func save_optuna_plots
  • func run_moe_multiobjective_study - Multi-objective Optuna HPO for SoftMixtureOfExperts using NSGA-II.

src/inmotion/dl/evaluation.py - 291 lines

Metrics, cross-validation evaluation, and plot generation.

Public API:

  • func compute_metrics
  • func plot_confusion_matrix
  • func plot_training_curves
  • func cross_validate
  • func evaluate_model_on_test
  • func save_results_csv
  • func evaluate_checkpoint_from_disk - Load a checkpoint from disk, reconstruct the model, and evaluate.
  • func metrics_to_extended_row - Convert metrics dict to a flat row dict suitable for CSV storage.

src/inmotion/dl/results.py - 77 lines

Extended results I/O for DL pipeline — detailed per-class metrics + confusion matrices.

Public API:

  • func save_extended_results_csv - Save extended results rows to CSV with all per-class and CM columns.
  • func load_extended_results - Load extended DL results CSV.

src/inmotion/dl/interference.py - 357 lines

Cross-route interference analysis — how concurrent paths affect RSSI signatures.

Analyses the concurrent_noise_path column to measure how strongly each interfering route distorts the RSSI readings of the primary route. Generates publication-quality PDF plots at 600 DPI with large fonts.

Public API:

  • class InterferenceAnalyzer - Analyze how concurrent noise from different routes affects RSSI readings.

src/inmotion/dl/models/__init__.py - 48 lines

DL model exports.

No public classes or functions; this module provides data or constants consumed elsewhere.

Classical machine learning library (inmotion.classification)

Section titled “Classical machine learning library (inmotion.classification)”

src/inmotion/classification/config.py - 115 lines

Configuration module for ML classification pipeline.

Public API:

  • func check_gpu_availability - Check if CUDA GPU is available for ML libraries.
  • class Config - Configuration settings for the ML classification pipeline.

src/inmotion/classification/classifiers.py - 362 lines

Classifier factory with all ML classification algorithms.

Public API:

  • class ClassifierFactory - Factory class to create and manage all classifiers.

src/inmotion/classification/training.py - 407 lines

Training pipeline for ML classification.

Public API:

  • func suppress_stderr - Context manager to suppress stderr output (for C++ library warnings).
  • class ClassifierResult - Results from training a classifier.
  • class TrainingPipeline - Complete training pipeline for all classifiers.

src/inmotion/classification/optimization.py - 325 lines

Optuna hyperparameter optimization module.

Public API:

  • class OptunaOptimizer - Hyperparameter optimization using Optuna.

src/inmotion/classification/visualization.py - 637 lines

Visualization module for ML classification results.

Public API:

  • func load_results_from_csv - Load classification results from CSV files for plot regeneration.
  • class Visualizer - Visualization utilities for ML classification results.

src/inmotion/classification/eda.py - 452 lines

Exploratory Data Analysis module for comprehensive dataset analysis.

Public API:

  • class ExploratoryDataAnalysis - Comprehensive exploratory data analysis for WiFi fingerprinting dataset.

src/inmotion/classification/data_loader.py - 110 lines

Data loading and preprocessing module.

Public API:

  • class DataLoader - Load and preprocess the WiFi fingerprinting dataset.