Skip to content

add langid to interface - #193

Open
cjpais wants to merge 8 commits into
mainfrom
langid
Open

cjpais wants to merge 8 commits into
mainfrom
langid

Conversation

@cjpais

@cjpais cjpais commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

No description provided.

cjpais added 8 commits October 5, 2026 14:29
Add the language ID role: public header include/transcribe/langid.h
(info, label table, session lifecycle, run, result / candidate /
timings accessors), TRANSCRIBE_ROLE_LANGID, TRANSCRIBE_ERR_INPUT_TOO_SHORT
(21), ABI struct ids 19-23, Arch::langid ops table and its resolve_roles
check.

The dispatcher (src/transcribe-langid.cpp) owns everything that does not
depend on the family: label table validation, the allowed set, the crop
to the last max_audio_ms, softmax, allowed_mass, ranking and top-k.
Families only return logits.

Fixes carried from the langid.cpp review, each covered by
tests/langid_dispatch_unit.cpp against a fake arch:
- H1: max_audio_ms below 500 ms is rejected at session init, and the
  minimum is checked on the audio actually scored, after the crop.
- H2: only NULL means "all labels"; a non-NULL list with n_allowed <= 0,
  NULL with n_allowed > 0, or a NULL element is INVALID_ARG.
- H3: non-finite PCM is rejected up front; non-finite or miscounted
  logits fail the run instead of being ranked.
- H6: duplicate or empty codes and colliding / dangling aliases fail the
  load with TRANSCRIBE_ERR_GGUF.

No family serves the role yet. Bindings regenerate in a later commit, so
the generator / bindgen --check gates are red until then.
Port SpeechBrain's VoxLingua107 ECAPA-TDNN from handy-computer/langid.cpp
(96af3a5) into src/arch/ecapa_tdnn/, serving TRANSCRIBE_ROLE_LANGID.

- Family-local SpeechBrain log-mel front end (periodic Hamming,
  10*log10, 80 dB top-db, per-bin mean), now on run_on_threads; the
  graph helpers live in graph.cpp. The embed path is dropped.
- GGUF contract moves to the shared namespaces: stt.frontend.*,
  stt.ecapa_tdnn.*, stt.langid.labels.*, stt.variant, uint32 scalars.
  Tensor names follow the transcribe-quantize rules (.weight / .bias,
  .conv. for k>1 kernels, frontend.mel_filterbank), so the quant policy
  needs no ECAPA entries. Metadata-derived dimensions are bounded.
- load() builds the label table while the model is still RAII-owned
  (H4); a duplicate label fails the load (H6, fixture-tested, 0 leaks).
- scripts/convert-ecapa_tdnn.py writes F32 only, from the HF cache; its
  tensors are byte-identical to langid.cpp's GGUF. F16 / Q8_0 come from
  transcribe-quantize (FAMILY_PRESETS: F16, Q8_0).
- Reference env, SpeechBrain dumper and NumPy reference, golden manifest
  and tolerances (carried over: every C++ stage tensor is bit-identical
  to langid.cpp's on CPU / 1 thread). validate.py gains an ecapa_tdnn
  branch: encoder-only reference stage, and a top-1 label gate instead of
  a transcript (cases set language: null so no hint reaches the model).
- transcribe-cli routes LANGID models to a new driver (--allow, --top,
  --max-audio-ms) with stable `language:` / `candidate:` lines.
- Tests: toy-fixture smoke (public API; registered above the shared
  build cut, so it also runs against shared libtranscribe), front-end
  unit, real-model smoke (TRANSCRIBE_ECAPA_TDNN_GGUF), CLI smoke. The
  fixture-generation block in tests/CMakeLists.txt moves above that cut.
- Ten FLEURS sample clips (CC-BY-4.0) and the intake JSON; the intake
  schema and normalize alias table learn hamming_periodic,
  sentence_mean, encoder-classifier and accuracy.

validate.py all --family ecapa_tdnn: 8/8 cases, 17/17 tensors in
tolerance, top-1 agrees with SpeechBrain on every case.
Regenerate the low-level layers (_generated.py / _generated.ts /
transcribe.abihash via generate.py, transcribe_sys.rs via cargo xtask
bindgen, Swift pinned hash) and add a hand-written LangIdSession to every
binding, on the existing per-model execution bridge (Python
Model._exclusive, Rust ModelInner::with_compute, Swift Model.withCompute,
TypeScript SessionCore.exclusive):

- Model: roles gains LANGID; langid info, label table (codes, names, alias
  lookup) and langid_session(n_threads, max_audio_ms).
- LangIdSession.run(pcm, allowed, top_k) -> result with ranked candidates
  (index, code, name, p, p_unrestricted, logit), n_allowed, allowed_mass,
  audio_ms; cancellation; timings. Results are copied out under the
  compute lock in all four (Rust included).
- allowed: None / nil / null passes NULL ("every label"); an empty list
  is rejected by the binding with InvalidArgument, never marshalled to
  NULL. The array and every encoded string outlive the native call,
  including TypeScript's async worker call.
- TRANSCRIBE_ERR_INPUT_TOO_SHORT maps to InputTooShort /
  Error::InputTooShort / .inputTooShort.

Tests: contract checks on the toy ecapa_tdnn fixture (Python generates it
on the fly; Rust / Swift / TypeScript use the C++-built
tests/fixtures/arch_ecapa_tdnn_minimal.gguf and skip without it), top-1
on the FLEURS clips with TRANSCRIBE_SMOKE_LANGID_MODEL, Busy / lock /
empty-allowed / keep-alive checks model-free. docs/bindings.md gains the
LANGID contract section; each README a language ID section.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant