Repository navigation
Conversation
Add the language ID role: public header include/transcribe/langid.h (info, label table, session lifecycle, run, result / candidate / timings accessors), TRANSCRIBE_ROLE_LANGID, TRANSCRIBE_ERR_INPUT_TOO_SHORT (21), ABI struct ids 19-23, Arch::langid ops table and its resolve_roles check. The dispatcher (src/transcribe-langid.cpp) owns everything that does not depend on the family: label table validation, the allowed set, the crop to the last max_audio_ms, softmax, allowed_mass, ranking and top-k. Families only return logits. Fixes carried from the langid.cpp review, each covered by tests/langid_dispatch_unit.cpp against a fake arch: - H1: max_audio_ms below 500 ms is rejected at session init, and the minimum is checked on the audio actually scored, after the crop. - H2: only NULL means "all labels"; a non-NULL list with n_allowed <= 0, NULL with n_allowed > 0, or a NULL element is INVALID_ARG. - H3: non-finite PCM is rejected up front; non-finite or miscounted logits fail the run instead of being ranked. - H6: duplicate or empty codes and colliding / dangling aliases fail the load with TRANSCRIBE_ERR_GGUF. No family serves the role yet. Bindings regenerate in a later commit, so the generator / bindgen --check gates are red until then.
Port SpeechBrain's VoxLingua107 ECAPA-TDNN from handy-computer/langid.cpp (96af3a5) into src/arch/ecapa_tdnn/, serving TRANSCRIBE_ROLE_LANGID. - Family-local SpeechBrain log-mel front end (periodic Hamming, 10*log10, 80 dB top-db, per-bin mean), now on run_on_threads; the graph helpers live in graph.cpp. The embed path is dropped. - GGUF contract moves to the shared namespaces: stt.frontend.*, stt.ecapa_tdnn.*, stt.langid.labels.*, stt.variant, uint32 scalars. Tensor names follow the transcribe-quantize rules (.weight / .bias, .conv. for k>1 kernels, frontend.mel_filterbank), so the quant policy needs no ECAPA entries. Metadata-derived dimensions are bounded. - load() builds the label table while the model is still RAII-owned (H4); a duplicate label fails the load (H6, fixture-tested, 0 leaks). - scripts/convert-ecapa_tdnn.py writes F32 only, from the HF cache; its tensors are byte-identical to langid.cpp's GGUF. F16 / Q8_0 come from transcribe-quantize (FAMILY_PRESETS: F16, Q8_0). - Reference env, SpeechBrain dumper and NumPy reference, golden manifest and tolerances (carried over: every C++ stage tensor is bit-identical to langid.cpp's on CPU / 1 thread). validate.py gains an ecapa_tdnn branch: encoder-only reference stage, and a top-1 label gate instead of a transcript (cases set language: null so no hint reaches the model). - transcribe-cli routes LANGID models to a new driver (--allow, --top, --max-audio-ms) with stable `language:` / `candidate:` lines. - Tests: toy-fixture smoke (public API; registered above the shared build cut, so it also runs against shared libtranscribe), front-end unit, real-model smoke (TRANSCRIBE_ECAPA_TDNN_GGUF), CLI smoke. The fixture-generation block in tests/CMakeLists.txt moves above that cut. - Ten FLEURS sample clips (CC-BY-4.0) and the intake JSON; the intake schema and normalize alias table learn hamming_periodic, sentence_mean, encoder-classifier and accuracy. validate.py all --family ecapa_tdnn: 8/8 cases, 17/17 tensors in tolerance, top-1 agrees with SpeechBrain on every case.
Regenerate the low-level layers (_generated.py / _generated.ts /
transcribe.abihash via generate.py, transcribe_sys.rs via cargo xtask
bindgen, Swift pinned hash) and add a hand-written LangIdSession to every
binding, on the existing per-model execution bridge (Python
Model._exclusive, Rust ModelInner::with_compute, Swift Model.withCompute,
TypeScript SessionCore.exclusive):
- Model: roles gains LANGID; langid info, label table (codes, names, alias
lookup) and langid_session(n_threads, max_audio_ms).
- LangIdSession.run(pcm, allowed, top_k) -> result with ranked candidates
(index, code, name, p, p_unrestricted, logit), n_allowed, allowed_mass,
audio_ms; cancellation; timings. Results are copied out under the
compute lock in all four (Rust included).
- allowed: None / nil / null passes NULL ("every label"); an empty list
is rejected by the binding with InvalidArgument, never marshalled to
NULL. The array and every encoded string outlive the native call,
including TypeScript's async worker call.
- TRANSCRIBE_ERR_INPUT_TOO_SHORT maps to InputTooShort /
Error::InputTooShort / .inputTooShort.
Tests: contract checks on the toy ecapa_tdnn fixture (Python generates it
on the fly; Rust / Swift / TypeScript use the C++-built
tests/fixtures/arch_ecapa_tdnn_minimal.gguf and skip without it), top-1
on the FLEURS clips with TRANSCRIBE_SMOKE_LANGID_MODEL, Busy / lock /
empty-allowed / keep-alive checks model-free. docs/bindings.md gains the
LANGID contract section; each README a language ID section.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.