Repository navigation
Promote to main: ml - #80
Merged
Merged
Conversation
New model types (auto-registered in the training wizard): - PyTorch Dense Neural Network (hidden_layers/units, dropout, epochs, lr, batch_size) - PyTorch 1D Convolutional Neural Network for raw sensor windows - TorchNeuralNetwork base mirrors the Keras persist/restore contract, storing the architecture in the model state so restore can rebuild the module before loading weights from the DATASTORE - New "Raw Time-Series (Sensors only)" feature extraction that drops the timestamp column so models only depend on device-obtainable inputs ExecuTorch export: - Platforms.EXECUTORCH + dispatch in Pipeline.export (lazy import, the service boots without executorch; deploy route returns 400/501 typed errors via ExecutorchExportError) - ExecutorchCompiler composes baked preprocessing (MinMax/Z normalization, optional SimpleFeatureExtractor as torch ops) with the trained classifier into one torch.export graph, lowered via XNNPACK to model.pte - Download zip contains model.pte, manifest.json (input shape, sensor order, window size/stride, sampling rate, labels), README.md and a Kotlin example - Median is computed via topk (ExecuTorch portable kernels have no aten::sort), matching np.median exactly Model capability plumbing: - Model.formats persisted at train time; C is only advertised when every pipeline step actually implements exportC (get_platforms declarations alone are unreliable) Fixes along the way: - ZNormalizer never restored mean/std (restore was missing) - FileSystemDataLoader lacked saveObj/loadObj, so NN persistence only worked on S3 Dependencies: executorch==1.0.0 requires numpy>=2/pandas>=2.2.2/ scikit-learn==1.7.1/torch 2.9, so numpy/pandas/h5py pins moved accordingly; Dockerfile installs CPU torch first (avoids the CUDA wheel), fetches flatc (needed by the .pte serializer) and rebuilds lttbc against numpy 2. Verified: tests/test_torch_executorch.py (10 tests) — persist/restore round-trips, and numerical parity python pipeline == eager export module == .pte executed by the ExecuTorch runtime for all three supported pipeline combinations.
- Precompile the .pte at train time and cache it in the DATASTORE; the download route now streams cached bytes instead of re-tracing + XNNPACK-lowering on every click. A compile failure at train time just drops EXECUTORCH from the advertised formats instead of failing training. (buildExecutorchExport split into buildExecutorchPte / assembleExecutorchFiles with store/load helpers; falls back to on-the-fly compile for models trained before precompilation existed.) - Run downloadModel in a threadpool so the (fallback) compile and zip assembly never block the event loop. - deploy route: return HTTP 400 for NotImplementedError (e.g. C export requested for a classifier without exportC) instead of an opaque 500. - ExtraFile gained read(), so exporter output is zipped directly; downloadModel no longer re-wraps into StringFile (removed a latent AttributeError if that wrap were ever dropped). - Removed dead get_export_formats() hook that looked like it gated export but was never called. - manifest: guard classification_frequency_hint_hz with instead of a falsy check. - tests: cover ExtraFile zipping and the pte store/load/cache-hit paths.
Exercises the seams the running service actually uses, beyond the in-isolation compiler tests: - full train -> precompile+cache -> downloadModel -> unzip -> run .pte parity, for both raw-CNN/ZNormalizer and dense/features/MinMax pipelines - Model-doc round-trip: persist each option -> PipelineStepOption -> restore -> export still works - ZNormalizer.restore regression (statistics recovered) - deploy route status mapping: 200 stream, 400 ExecutorchExportError, 501 executorch-missing, NotImplementedError -> 400 - downloadModel dispatched off the event loop (run_in_threadpool) - Pipeline.export platform dispatch routes C vs EXECUTORCH correctly
flatbuffers only ships an x86-64 Linux flatc binary; running it in the arm64 image failed with 'Exec format error' and broke the arm64 build. Keep the prebuilt binary on amd64 and compile flatc at the same pinned version on arm64.
MicroNAS installs triton 2.1.0 (its torch 2.1.1 pin), which is left orphaned after the torch 2.9.0 upgrade. torch._inductor still detects it via has_triton_package() and crashes on 'import triton.backends.compiler' (absent before triton 2.2) during torch.export, so every .pte export/compile failed. CPU export does not use triton; uninstalling it makes the full ExecuTorch parity suite pass (14/14).
The PyTorch classifiers export to ExecuTorch (.pte) but only declared PYTHON in get_platforms(), so the training wizard could not show that they are mobile-exportable. Add Platforms.EXECUTORCH so the classifier step can surface it. Actual export eligibility is still decided per-pipeline by computeFormats; this only advertises the classifier capability.
The windower, feature extractors and normalizers that ExecuTorch export supports (SampleWindower, Simple/Raw extractors, MinMax/Z normalizers) now declare Platforms.EXECUTORCH in get_platforms(), matching supportsExecutorch. This lets the training wizard filter each step's options to the ones valid for the chosen deployment target. No change to export behavior.
Options now serialize an exportTargets {c, executorch} field computed the
same way the download flow decides formats: a real C download requires
exportC to be implemented, not just a declared C platform (some legacy
options declare C via the old codegen API without an exportC). This lets the
training wizard filter options by deployment target without offering ones
that would then fail to download (e.g. Random Forest / Small Conv NN, which
declare C but produce no C download).
The single-model route passed the pydantic Model straight to json.dumps, which cannot serialize it, so every fetch 500'd (breaking View live). Convert to a dict with .dict(by_alias=True) first, exactly as the model-list route already does.
…, example stride - Clamp the manifest and Kotlin label arrays to the model's actual output count. A configured label with no surviving windows made num_classes < len(labels), so the extra names misaligned predictions on-device. - Add a 1e-8 epsilon to MinMax and Z normalization on both the numpy and the baked torch sides (kept identical for export parity), so a constant sensor channel no longer yields NaN outputs. - Make the Kotlin example honour window.stride: it now classifies once per stride samples instead of on every sample.
Adds stage/currentEpoch/totalEpochs/progress to the model doc and emits them during init_train: a coarse stage for every classifier (Loading data -> Training -> Finalizing) plus per-epoch progress from the shared PyTorch epoch loop. State is written to Mongo (throttled to one write per percent) since training runs in a background task and the frontend polls the model doc from any worker.
…out of bincount The marker was written 9*10^10, which is (9*10)^10 = 80 in Python (^ is XOR), not 9e10. Benign only because the assign and filter sites shared the same wrong value, but a project with >=81 labels would collide a real label with 80 and silently drop it. Naively changing to 9*10**10 would instead OOM np.bincount (it sizes to the max value). Introduce UNLABELED_LABEL=-1 and a window_majority_label helper that computes the per-window majority without passing negatives/huge indices to bincount. Update the assignment (DataLoader), the filter (BaseWindower) and all three bincount sites together.
) New WharModel(TorchNeuralNetwork): a single classifier with an Architecture dropdown exposing the 17 whar-models neural architectures (deepsense/global_fusion excluded - they need even/6+ channels and fail on 3-channel data). build_model calls whar_models.build_model and wraps it behind a channels-first permute (edge-ml windows are channels-last). Inherits the training loop, persist/restore and ExecuTorch export. whar-models installed --no-deps in the Dockerfile (its numpy>=2.2 / sklearn==1.8 / tsfresh floors, used only by the classical models, would otherwise break executorch's numpy==2.0.2 pin); einops added to requirements as the neural models' one light runtime dep. Verified all 17 build + forward + train-step on torch 2.9.0 / numpy 2.0.2 at WISDM-shape input.
whar-models declares requires-python>=3.11 but the ml image is py3.10; verified all 17 neural models import/build/train on py3.10 with torch 2.9.0 / numpy 2.0.2, so the pin is conservative. Force the install.
…iers WHAR Model and TorchCNN1D consume the raw window; block them in preflight when not paired with 'Raw Time-Series (Sensors only)' so the wizard shows a clear message instead of training on mis-shaped input.
…#62) WHAR models (and edge-ml's torch classifiers) often can't lower to mobile ExecuTorch, leaving them with nothing to download. Add a universal PyTorch export: trace the composed preprocessing+classifier to TorchScript (.pt) plus a manifest + README + runnable inference.py, runnable on any server/desktop. Advertised as the PYTORCH format, but validated at train time and dropped from formats if the trace fails, so only genuinely working downloads are ever offered.
… exclusion Offer all torch architectures in the picker. deepsense (needs an even channel count) and global_fusion (needs >= 6) were dropped platform-wide; instead the preflight now blocks them only when the selected data's channel count is incompatible, with a clear message. Pairs with the wizard hiding them from the Architecture dropdown on incompatible data.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
systemwide promotion of beta to main. 18 commits, no divergence (main is 0 ahead of beta).
part of the coordinated cutover tracked in the superproject PR — prod moves to the dataset-store mono-backend and the old koa backend + express auth services are retired.