Skip to content

Promote to main: ml - #80

Merged
hamza-aamer merged 18 commits into
mainfrom
beta
Sep 28, 2026
Merged

hamza-aamer merged 18 commits into
mainfrom
beta

Conversation

@hamza-aamer

Copy link
Copy Markdown

systemwide promotion of beta to main. 18 commits, no divergence (main is 0 ahead of beta).

part of the coordinated cutover tracked in the superproject PR — prod moves to the dataset-store mono-backend and the old koa backend + express auth services are retired.

New model types (auto-registered in the training wizard):
- PyTorch Dense Neural Network (hidden_layers/units, dropout, epochs, lr, batch_size)
- PyTorch 1D Convolutional Neural Network for raw sensor windows
- TorchNeuralNetwork base mirrors the Keras persist/restore contract,
  storing the architecture in the model state so restore can rebuild
  the module before loading weights from the DATASTORE
- New "Raw Time-Series (Sensors only)" feature extraction that drops the
  timestamp column so models only depend on device-obtainable inputs

ExecuTorch export:
- Platforms.EXECUTORCH + dispatch in Pipeline.export (lazy import, the
  service boots without executorch; deploy route returns 400/501 typed
  errors via ExecutorchExportError)
- ExecutorchCompiler composes baked preprocessing (MinMax/Z normalization,
  optional SimpleFeatureExtractor as torch ops) with the trained classifier
  into one torch.export graph, lowered via XNNPACK to model.pte
- Download zip contains model.pte, manifest.json (input shape, sensor
  order, window size/stride, sampling rate, labels), README.md and a
  Kotlin example
- Median is computed via topk (ExecuTorch portable kernels have no
  aten::sort), matching np.median exactly

Model capability plumbing:
- Model.formats persisted at train time; C is only advertised when every
  pipeline step actually implements exportC (get_platforms declarations
  alone are unreliable)

Fixes along the way:
- ZNormalizer never restored mean/std (restore was missing)
- FileSystemDataLoader lacked saveObj/loadObj, so NN persistence only
  worked on S3

Dependencies: executorch==1.0.0 requires numpy>=2/pandas>=2.2.2/
scikit-learn==1.7.1/torch 2.9, so numpy/pandas/h5py pins moved
accordingly; Dockerfile installs CPU torch first (avoids the CUDA
wheel), fetches flatc (needed by the .pte serializer) and rebuilds
lttbc against numpy 2.

Verified: tests/test_torch_executorch.py (10 tests) — persist/restore
round-trips, and numerical parity python pipeline == eager export
module == .pte executed by the ExecuTorch runtime for all three
supported pipeline combinations.
- Precompile the .pte at train time and cache it in the DATASTORE; the
  download route now streams cached bytes instead of re-tracing +
  XNNPACK-lowering on every click. A compile failure at train time just
  drops EXECUTORCH from the advertised formats instead of failing training.
  (buildExecutorchExport split into buildExecutorchPte / assembleExecutorchFiles
  with store/load helpers; falls back to on-the-fly compile for models
  trained before precompilation existed.)
- Run downloadModel in a threadpool so the (fallback) compile and zip
  assembly never block the event loop.
- deploy route: return HTTP 400 for NotImplementedError (e.g. C export
  requested for a classifier without exportC) instead of an opaque 500.
- ExtraFile gained read(), so exporter output is zipped directly;
  downloadModel no longer re-wraps into StringFile (removed a latent
  AttributeError if that wrap were ever dropped).
- Removed dead get_export_formats() hook that looked like it gated export
  but was never called.
- manifest: guard classification_frequency_hint_hz with
  instead of a falsy check.
- tests: cover ExtraFile zipping and the pte store/load/cache-hit paths.
Exercises the seams the running service actually uses, beyond the
in-isolation compiler tests:
- full train -> precompile+cache -> downloadModel -> unzip -> run .pte
  parity, for both raw-CNN/ZNormalizer and dense/features/MinMax pipelines
- Model-doc round-trip: persist each option -> PipelineStepOption ->
  restore -> export still works
- ZNormalizer.restore regression (statistics recovered)
- deploy route status mapping: 200 stream, 400 ExecutorchExportError,
  501 executorch-missing, NotImplementedError -> 400
- downloadModel dispatched off the event loop (run_in_threadpool)
- Pipeline.export platform dispatch routes C vs EXECUTORCH correctly
flatbuffers only ships an x86-64 Linux flatc binary; running it in the
arm64 image failed with 'Exec format error' and broke the arm64 build.
Keep the prebuilt binary on amd64 and compile flatc at the same pinned
version on arm64.
MicroNAS installs triton 2.1.0 (its torch 2.1.1 pin), which is left orphaned
after the torch 2.9.0 upgrade. torch._inductor still detects it via
has_triton_package() and crashes on 'import triton.backends.compiler' (absent
before triton 2.2) during torch.export, so every .pte export/compile failed.
CPU export does not use triton; uninstalling it makes the full ExecuTorch
parity suite pass (14/14).
The PyTorch classifiers export to ExecuTorch (.pte) but only declared
PYTHON in get_platforms(), so the training wizard could not show that they
are mobile-exportable. Add Platforms.EXECUTORCH so the classifier step can
surface it. Actual export eligibility is still decided per-pipeline by
computeFormats; this only advertises the classifier capability.
The windower, feature extractors and normalizers that ExecuTorch export
supports (SampleWindower, Simple/Raw extractors, MinMax/Z normalizers) now
declare Platforms.EXECUTORCH in get_platforms(), matching supportsExecutorch.
This lets the training wizard filter each step's options to the ones valid
for the chosen deployment target. No change to export behavior.
Options now serialize an exportTargets {c, executorch} field computed the
same way the download flow decides formats: a real C download requires
exportC to be implemented, not just a declared C platform (some legacy
options declare C via the old codegen API without an exportC). This lets the
training wizard filter options by deployment target without offering ones
that would then fail to download (e.g. Random Forest / Small Conv NN, which
declare C but produce no C download).
The single-model route passed the pydantic Model straight to json.dumps,
which cannot serialize it, so every fetch 500'd (breaking View live). Convert
to a dict with .dict(by_alias=True) first, exactly as the model-list route
already does.
…, example stride

- Clamp the manifest and Kotlin label arrays to the model's actual output
  count. A configured label with no surviving windows made num_classes <
  len(labels), so the extra names misaligned predictions on-device.
- Add a 1e-8 epsilon to MinMax and Z normalization on both the numpy and the
  baked torch sides (kept identical for export parity), so a constant sensor
  channel no longer yields NaN outputs.
- Make the Kotlin example honour window.stride: it now classifies once per
  stride samples instead of on every sample.
Adds stage/currentEpoch/totalEpochs/progress to the model doc and emits
them during init_train: a coarse stage for every classifier (Loading
data -> Training -> Finalizing) plus per-epoch progress from the shared
PyTorch epoch loop. State is written to Mongo (throttled to one write per
percent) since training runs in a background task and the frontend polls
the model doc from any worker.
…out of bincount

The marker was written 9*10^10, which is (9*10)^10 = 80 in Python (^ is
XOR), not 9e10. Benign only because the assign and filter sites shared
the same wrong value, but a project with >=81 labels would collide a
real label with 80 and silently drop it. Naively changing to 9*10**10
would instead OOM np.bincount (it sizes to the max value).

Introduce UNLABELED_LABEL=-1 and a window_majority_label helper that
computes the per-window majority without passing negatives/huge indices
to bincount. Update the assignment (DataLoader), the filter
(BaseWindower) and all three bincount sites together.
)

New WharModel(TorchNeuralNetwork): a single classifier with an
Architecture dropdown exposing the 17 whar-models neural architectures
(deepsense/global_fusion excluded - they need even/6+ channels and fail
on 3-channel data). build_model calls whar_models.build_model and wraps
it behind a channels-first permute (edge-ml windows are channels-last).
Inherits the training loop, persist/restore and ExecuTorch export.

whar-models installed --no-deps in the Dockerfile (its numpy>=2.2 /
sklearn==1.8 / tsfresh floors, used only by the classical models, would
otherwise break executorch's numpy==2.0.2 pin); einops added to
requirements as the neural models' one light runtime dep. Verified all
17 build + forward + train-step on torch 2.9.0 / numpy 2.0.2 at
WISDM-shape input.
whar-models declares requires-python>=3.11 but the ml image is py3.10;
verified all 17 neural models import/build/train on py3.10 with torch
2.9.0 / numpy 2.0.2, so the pin is conservative. Force the install.
…iers

WHAR Model and TorchCNN1D consume the raw window; block them in preflight
when not paired with 'Raw Time-Series (Sensors only)' so the wizard shows
a clear message instead of training on mis-shaped input.
…#62)

WHAR models (and edge-ml's torch classifiers) often can't lower to mobile
ExecuTorch, leaving them with nothing to download. Add a universal
PyTorch export: trace the composed preprocessing+classifier to TorchScript
(.pt) plus a manifest + README + runnable inference.py, runnable on any
server/desktop. Advertised as the PYTORCH format, but validated at train
time and dropped from formats if the trace fails, so only genuinely
working downloads are ever offered.
… exclusion

Offer all torch architectures in the picker. deepsense (needs an even channel
count) and global_fusion (needs >= 6) were dropped platform-wide; instead the
preflight now blocks them only when the selected data's channel count is
incompatible, with a clear message. Pairs with the wizard hiding them from the
Architecture dropdown on incompatible data.
@hamza-aamer
hamza-aamer merged commit 5a8e2a4 into main Sep 28, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant