Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
27 commits
Select commit Hold shift + click to select a range
4f2ddeb
langid: LANGID role scaffolding
cjpais Oct 5, 2026
76eb1cf
langid: ecapa_tdnn family (VoxLingua107 ECAPA-TDNN) on the LANGID role
cjpais Oct 5, 2026
6252220
langid: bindings for the LANGID role (Python, Rust, Swift, TypeScript)
cjpais Oct 5, 2026
796545a
wip langid port
cjpais Oct 5, 2026
e8a7745
wip
cjpais Oct 5, 2026
5e7b5cd
kernel work
cjpais Oct 6, 2026
16d46bc
vulkan gpu work
cjpais Oct 6, 2026
d836d01
langid: merge fixes for the ecapa_tdnn CPU/GPU kernel work
cjpais Oct 6, 2026
f563341
langid: drop ecapa_tdnn hand-written CPU kernels
cjpais Oct 7, 2026
113a865
mel: add hamming_periodic window and sentence_mean normalize
cjpais Oct 7, 2026
7ec3311
langid: ecapa_tdnn uses the shared mel frontend
cjpais Oct 7, 2026
00e4aa2
langid: trim ecapa_tdnn validation, dumps and comments
cjpais Oct 7, 2026
106ac48
langid: trim ecapa_tdnn and langid tests
cjpais Oct 7, 2026
2995d13
langid: trim ecapa_tdnn and catalog tooling
cjpais Oct 7, 2026
e4ad30d
langid: trim scripts/langid to the published accuracy path
cjpais Oct 7, 2026
ac5febc
langid: reduce docs to API facts and catalog-rendered tables
cjpais Oct 7, 2026
905eb05
langid: drop two stale notes about langid.cpp parity and CPU kernels
cjpais Oct 7, 2026
144aa7d
langid: restore the 0.05 F32 GPU bound in ecapa_tdnn_real_smoke
cjpais Oct 7, 2026
31f2a1c
mel: skip zero filterbank weights in the scalar mel matmul
cjpais Oct 7, 2026
2fdb5c9
langid: widen Q8_0 weights to F16 on several threads
cjpais Oct 7, 2026
8d00eb4
langid: M4 Max speed cells on ru-long at 2fdb5c95
cjpais Oct 7, 2026
c8b8389
langid: Ryzen 4750U speed cells
cjpais Oct 7, 2026
8ae733e
langid: re-derive two ecapa_tdnn tolerances to cover x86
cjpais Oct 7, 2026
0ce2736
langid: FLEURS accuracy and agreement at 2fdb5c95
cjpais Oct 7, 2026
d492bfd
Merge main into langid
cjpais Oct 8, 2026
31e0c1e
rust: ErrorKind::InputTooShort for the LANGID status
cjpais Oct 8, 2026
2e3ba7d
langid: drop transcribe_langid_candidate::p_unrestricted
cjpais Oct 8, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 9 additions & 2 deletions .claude/skills/porting-1-intake/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,13 +53,20 @@ field semantics.

1. **Reference framework choice.**: The reference implementation should always be the source of truth. If there is not one, this is a red flag and must immediately be told to the human.

2. **Architecture pattern.** One of `encoder-transducer`, `encoder-decoder`, `audio-llm`, `encoder-ctc`. The script's `config.architecture_candidates` is a heuristic starting point. If it doesn't fit, propose a new pattern and have the user accept it based on your research of the architecture.
2. **Architecture pattern.** One of `encoder-transducer`, `encoder-decoder`, `audio-llm`, `encoder-ctc` (`encoder-diarizer` / `encoder-classifier` for the DIARIZE / LANGID roles). The script's `config.architecture_candidates` is a heuristic starting point. If it doesn't fit, propose a new pattern and have the user accept it based on your research of the architecture.

**Role.** ASR (the product is a transcript) or DIARIZE (the product is
who spoke when; see `docs/roles.md`). A DIARIZE port implements
`DiarizeOps` (`src/transcribe-diarize.h`) instead of the ASR hooks, sets
`"role": "diarize"` in its catalog record, and is accepted on DER
(`scripts/diar/`), not WER.
(`scripts/diar/`), not WER. A LANGID port (the product is a language
decision; `docs/langid.md`) implements `LangidOps`
(`src/transcribe-langid.h`) and only returns logits over its label table,
sets `"role": "langid"` with `metric: "accuracy"` / `acc_pct` rows, and
is accepted on top-1 decision parity with the reference over FLEURS
(`scripts/langid/`), not WER. Its golden cases set `"language": null`
and carry the expected label as `expected_language`, so validate.py never
hands the answer to the model.

3. **Acceptance dataset.** Default: LibriSpeech test-clean. Capture any
publisher-reported score in `upstream_benchmarks` when available for
Expand Down
2 changes: 1 addition & 1 deletion .claude/skills/porting-7-wer/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ WER progress:

### Step 1: Acceptance manifest (execute or ask-point)

Read the acceptance dataset and metric from `upstream_benchmarks[0].{dataset, metric}`. Slugify: lowercase, spaces → hyphens. `"LibriSpeech test-clean"` → `librispeech-test-clean`. Resolve to `$MANIFEST`; every later step uses `$MANIFEST`, never a reconstructed path. Confirm `metric` is `wer` or `cer` (`der` for a DIARIZE-role model: score with `scripts/diar/` instead; anything else is out of scope). Do not use any publisher-reported score for pass/fail; the measured Oracle reference score is the gate target.
Read the acceptance dataset and metric from `upstream_benchmarks[0].{dataset, metric}`. Slugify: lowercase, spaces → hyphens. `"LibriSpeech test-clean"` → `librispeech-test-clean`. Resolve to `$MANIFEST`; every later step uses `$MANIFEST`, never a reconstructed path. Confirm `metric` is `wer` or `cer` (`der` for a DIARIZE-role model: score with `scripts/diar/` instead; `accuracy` for a LANGID-role model: run `scripts/langid/run.py` for the reference and every shipped quant, gate the ref dtype with `scripts/langid/compare.py` against the reference, and score with `scripts/langid/score.py`; anything else is out of scope). Do not use any publisher-reported score for pass/fail; the measured Oracle reference score is the gate target.

If the intake's dataset is not covered by `scripts/wer/ingest.py`, extend
that script before running this step.
Expand Down
2 changes: 2 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -59,6 +59,8 @@ __pycache__/

# WER evaluation data + generated working reports.
/samples/wer/
# FLEURS language ID corpus (scripts/langid/ingest.py).
/samples/langid/
# Diarization eval corpora (AMI audio + fetched RTTMs); regenerable via
# scripts/diar/ingest_ami.py + fetch_ami_forced_alignment.py. The committed
# oracle case lives at samples/sortformer-2spk-mix.wav (outside samples/diar/).
Expand Down
8 changes: 8 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,14 @@ C/C++ speech-to-text inference library. Runs diverse STT model families via [GGU
| Streaming Sortformer Diarizer 4spk v2.1 | `diar_streaming_sortformer_4spk-v2.1` | diarize, streaming | [docs/models/diar_streaming_sortformer_4spk-v2.1.md](docs/models/diar_streaming_sortformer_4spk-v2.1.md) |
<!-- /catalog -->

**Language ID models** (no transcription; verified by top-1 decision parity on FLEURS; see [`docs/langid.md`](docs/langid.md)):

<!-- catalog:family-index role=langid -->
| Family | Variants | Available capabilities | Docs |
| --- | --- | --- | --- |
| VoxLingua107 ECAPA-TDNN (language ID) | `lang-id-voxlingua107-ecapa` | language ID (107 languages) | [docs/models/lang-id-voxlingua107-ecapa.md](docs/models/lang-id-voxlingua107-ecapa.md) |
<!-- /catalog -->

Per-variant model cards live under [`docs/models/`](docs/models/).

## Model catalog
Expand Down
19 changes: 19 additions & 0 deletions bindings/python/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -88,6 +88,25 @@ with model.diarize_session() as diarizer:
print(turn.speaker_id, turn.t0_ms, turn.t1_ms)
```

### Language ID

Models whose `model.roles` include `Role.LANGID` (VoxLingua107 ECAPA-TDNN)
identify the spoken language through a langid session. `run()` returns a
`LangIdResult` with candidates ranked by `p`. Codes are the model's own labels
(`"iw"`, `"jw"`; `model.langid_label_index("he")` resolves aliases), so match
`result.code` against the ASR model's `capabilities.languages` yourself.
`allowed=None` scores every label; `allowed_mass` is the unrestricted
probability the allowed set captured. An empty `allowed` list raises
`InvalidArgument`; clips under `model.langid_info.min_audio_ms` raise
`InputTooShort`. Locking, `Busy`, `cancel()` and `close()` work as on
`Session`.

```python
with model.langid_session() as lid:
result = lid.run(pcm, allowed=["en", "de", "fr"], top_k=3)
print(result.code, result.candidates[0].p, result.allowed_mass)
```

## Backends

`Model(backend=...)` applies a backend policy (`"auto"` uses the best
Expand Down
166 changes: 165 additions & 1 deletion bindings/python/src/transcribe_cpp/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,7 @@
BackendError,
Busy,
InputTooLong,
InputTooShort,
InvalidArgument,
ModelFileNotFound,
ModelLoadError,
Expand Down Expand Up @@ -78,6 +79,10 @@
"Session",
"DiarizeSession",
"DiarizeInfo",
"LangIdSession",
"LangIdInfo",
"LangIdResult",
"LangIdCandidate",
"Role",
"Result",
"Segment",
Expand Down Expand Up @@ -120,6 +125,7 @@
"Aborted",
"Busy",
"InputTooLong",
"InputTooShort",
"OutputTruncated",
"OutputRepetition",
"native_version",
Expand Down Expand Up @@ -190,6 +196,8 @@
_Timings = _generated.transcribe_timings
_Segment = _generated.transcribe_segment
_SpeakerSegment = _generated.transcribe_speaker_segment
_LangIdCandidate = _generated.transcribe_langid_candidate
_LangIdResult = _generated.transcribe_langid_result
_Word = _generated.transcribe_word
_Token = _generated.transcribe_token
_StreamParams = _generated.transcribe_stream_params
Expand Down Expand Up @@ -558,10 +566,12 @@ class Capabilities:

class Role(enum.Enum):
"""What a model serves (``Model.roles``): ASR is ``Model.session()``,
DIARIZE is ``Model.diarize_session()``."""
DIARIZE is ``Model.diarize_session()``, LANGID is
``Model.langid_session()``."""

ASR = _generated.TRANSCRIBE_ROLE_ASR
DIARIZE = _generated.TRANSCRIBE_ROLE_DIARIZE
LANGID = _generated.TRANSCRIBE_ROLE_LANGID


@dataclass(frozen=True)
Expand All @@ -570,6 +580,43 @@ class DiarizeInfo:
max_speakers: int # speaker_id is in [1, max_speakers]


@dataclass(frozen=True)
class LangIdInfo:
sample_rate: int
n_labels: int # label indices are [0, n_labels)
min_audio_ms: int # shorter scored audio raises InputTooShort


@dataclass(frozen=True)
class LangIdCandidate:
"""One ranked label. ``code`` is the model's own label ("en", "iw");
``p`` is renormalized over the allowed set."""

index: int
code: str
name: str
p: float
logit: float


@dataclass(frozen=True)
class LangIdResult:
"""``candidates`` are ranked by ``p`` (ties keep label order).
``allowed_mass`` is the unrestricted probability inside the allowed set
(1.0 when unrestricted); a low value means the speech is probably outside
it. ``audio_ms`` is what was scored, after the crop."""

candidates: tuple[LangIdCandidate, ...]
n_allowed: int
allowed_mass: float
audio_ms: int

@property
def code(self) -> str | None:
"""The top candidate's code, or None when there are no candidates."""
return self.candidates[0].code if self.candidates else None


@dataclass(frozen=True)
class SessionLimits:
"""Effective per-session limits (model bound narrowed by session params).
Expand Down Expand Up @@ -1227,6 +1274,38 @@ def diarize_session(self, *, n_threads: int = 0) -> "DiarizeSession":
"""Raises :class:`UnsupportedRole` without the DIARIZE role."""
return DiarizeSession(self, n_threads=n_threads)

@property
def langid_info(self) -> LangIdInfo:
"""Raises :class:`UnsupportedRole` without the LANGID role."""
info = _generated.transcribe_langid_info()
_lib.transcribe_langid_info_init(_byref(info))
_check(_lib.transcribe_langid_get_info(self._h, _byref(info)),
"reading langid info")
return LangIdInfo(sample_rate=info.sample_rate, n_labels=info.n_labels,
min_audio_ms=info.min_audio_ms)

@property
def langid_labels(self) -> tuple[tuple[str, str], ...]:
"""``(code, name)`` per label index. Raises :class:`UnsupportedRole`
without the LANGID role."""
n = self.langid_info.n_labels
return tuple((_lib.transcribe_langid_label_code(self._h, i).decode("utf-8"),
_lib.transcribe_langid_label_name(self._h, i).decode("utf-8"))
for i in range(n))

def langid_label_index(self, code: str) -> int | None:
"""Label index of a code or alias ("he" and "iw" name the same
label), or None. None as well on a model without the LANGID role."""
i = _lib.transcribe_langid_label_index(self._h, code.encode("utf-8"))
return i if i >= 0 else None

def langid_session(self, *, n_threads: int = 0,
max_audio_ms: int = 0) -> "LangIdSession":
"""``max_audio_ms``: longer input is scored on its last
``max_audio_ms`` (0 = 30000). Raises :class:`UnsupportedRole`
without the LANGID role."""
return LangIdSession(self, n_threads=n_threads, max_audio_ms=max_audio_ms)

def close(self) -> None:
"""Free the model. Any session still open on it is closed first —
the C contract forbids freeing a model before its sessions, so this
Expand Down Expand Up @@ -1888,6 +1967,91 @@ def timings(self) -> Timings:
return _timings_from(tm)


class LangIdSession(_SessionBase):
"""A language ID context on a model with the LANGID role. Compute
locking and ``Busy`` rules: see ``Model``."""

_free_fn = "transcribe_langid_session_free"

def __init__(self, model: Model, *, n_threads: int = 0, max_audio_ms: int = 0):
self._model = model # keep the model alive for the session's lifetime
params = _generated.transcribe_langid_session_params()
_lib.transcribe_langid_session_params_init(_byref(params))
params.n_threads = n_threads
params.max_audio_ms = max_audio_ms

handle = ctypes.c_void_p()
_check(_lib.transcribe_langid_session_init(model._h, _byref(params), _byref(handle)),
"opening langid session")
if not handle.value:
raise TranscribeError("langid session init returned a null handle")
self._handle = handle
self._arm_abort(_lib.transcribe_langid_set_abort_callback)

def run(self, pcm: PCMLike, *, allowed: "Sequence[str] | None" = None,
top_k: int = 0) -> LangIdResult:
"""Identify the language of one clip (16 kHz mono float32 PCM).
Longer input than the session's ``max_audio_ms`` is scored on its
tail. ``allowed`` restricts the decision to those codes or aliases;
None means every label, and an empty list is rejected. ``top_k``
keeps the best candidates (0 = every allowed label).

Raises :class:`InputTooShort` below ``LangIdInfo.min_audio_ms``,
:class:`UnsupportedRequest` for an unknown code, :class:`Aborted`
after :meth:`cancel`, and :class:`Busy` if a stream is active on this
model."""
self._cancel.clear() # before the lock wait, as in Session.run()
array, n_samples = _pcm_to_carray(pcm)
params = _generated.transcribe_langid_params()
_lib.transcribe_langid_params_init(_byref(params))
params.top_k = top_k
if allowed is not None:
if isinstance(allowed, (str, bytes)):
raise InvalidArgument("allowed must be a sequence of codes, not a string")
codes = list(allowed)
# An empty list must not reach native as NULL, which means "all".
if not codes:
raise InvalidArgument("allowed is empty; pass None for every label")
if not all(isinstance(c, str) for c in codes):
raise InvalidArgument("allowed entries must be str")
encoded = [c.encode("utf-8") for c in codes]
arr = (ctypes.c_char_p * len(encoded))(*encoded)
params.allowed = ctypes.cast(arr, type(params.allowed))
params.n_allowed = len(encoded)
params._allowed_keepalive = (encoded, arr)
with self._model._exclusive(
"langid_run", busy="a stream is active on this model; "
"finish or drop it before langid run()"):
h = self._h # captured under the lock; close() defers its free
_check(_lib.transcribe_langid_run(h, array, n_samples, _byref(params)),
"transcribe_langid_run")
res = _LangIdResult()
_lib.transcribe_langid_result_init(_byref(res))
_check(_lib.transcribe_langid_get_result(h, _byref(res)),
"transcribe_langid_get_result")
rows = []
for i in range(res.n_candidates):
c = _LangIdCandidate()
_lib.transcribe_langid_candidate_init(_byref(c))
_check(_lib.transcribe_langid_get_candidate(h, i, _byref(c)),
"transcribe_langid_get_candidate")
rows.append(LangIdCandidate(
index=c.index, code=c.code.decode("utf-8"), name=c.name.decode("utf-8"),
p=c.p, logit=c.logit))
return LangIdResult(candidates=tuple(rows), n_allowed=res.n_allowed,
allowed_mass=res.allowed_mass, audio_ms=res.audio_ms)

@property
def timings(self) -> Timings:
"""Load time plus the last run's mel / encode time. Not locked, like
``Session.limits``."""
tm = _Timings()
_lib.transcribe_timings_init(_byref(tm))
_check(_lib.transcribe_langid_get_timings(self._h, _byref(tm)),
"transcribe_langid_get_timings")
return _timings_from(tm)


def transcribe(
model: Model | str | os.PathLike,
pcm: PCMLike,
Expand Down
Loading
Loading