Repository navigation
restructure diarization prep for langid/vad - #187
Merged
Merged
Conversation
Pure move of the ASR C API (sessions, open/close, run, batch, streaming, stream text policy, capabilities, limits, tokenize, timings, result and batch accessors, ASR struct inits) into src/transcribe-asr.cpp. src/transcribe.cpp keeps status, version, ABI sizes, logging, backends, devices, model load/free and model queries. Non-move edits: - api_guard_* move to src/transcribe-api-guard.h (namespace transcribe) - enum_field_raw moves to src/transcribe-abi.h - four ASR log sites call transcribe::log_msg(level, "%s", buf) instead of the library-private emitter (same output) - file header comments, includes, using-declarations; AGENTS.md and two comments point at the new locations No behavior change: exported symbols identical, abihash unchanged, ctest unchanged.
Internal groundwork for per-role sessions; no public API change. - transcribe::SessionCore (src/transcribe-session-core.h) holds the role-agnostic session members: model, n_threads, stage timings, abort callback / poll_abort, compute scratch and release_scratch. transcribe_session derives from it; families keep ctx->sched etc. - SpeakerSegmentEntry is hoisted to transcribe::SpeakerSegmentEntry, with an alias on transcribe_session so existing references compile unchanged. - transcribe_model::roles (internal k_role_asr / k_role_diarize bits). Families that leave it 0 and have ASR hooks default to ASR, so no family changes. - Arch gains a trailing `diarize` ops-table pointer (DiarizeOps is defined with the diarize role). ASR hooks stay flat. - transcribe::resolve_roles runs after every successful load (GGUF and whisper .bin) and fails the load with NOT_IMPLEMENTED when the mask is empty, has unknown bits, or names a role the arch cannot serve. Covered by tests/role_resolve_unit.cpp. Verified: exported symbols and abihash unchanged; ctest green. On CPU, transcribe-cli output matches main (timing lines excluded) for whisper (GGUF, .bin, initial prompt), parakeet (offline, -o, batch, batch-jsonl, cache-aware and buffered streaming), sortformer (speaker segments via batch-jsonl), multitalker --diarize, moonshine streaming, qwen3_asr (vocabulary + prompt), canary --no-pnc, sensevoice --no-itn, gigaam, and cohere with --n-ctx / --kv-type. Real-model ctest: sortformer_stream_ext_unit and the other configured smokes pass; the failures seen are identical on main.
Pure move. main.cpp keeps argument parsing, usage, --list-devices, the log sink and -o setup; asr.cpp gets the ASR JSON helpers, apply_prompting, and the batch and single-file bodies of main(), moved verbatim into transcribe_cli::run_asr_batch / run_asr_file. cli.h holds cli_args and the shared json_escape / write_output_file declarations. Later roles add their own driver file next to asr.cpp. Non-move edits: cli.h, include / using lines, the two function wrappers (each declares the output_ok local main() used to), main()'s two calls, and clang-format reflowing one printf after the de-indent. Verified: --help and --list-devices identical to main; on CPU, output matches main (timing lines excluded) for whisper (GGUF, .bin, initial prompt), parakeet (offline, -o, batch, batch-jsonl, cache-aware and buffered streaming), sortformer (batch-jsonl speaker segments), multitalker --diarize, moonshine streaming, qwen3_asr, canary --no-pnc, sensevoice --no-itn --raw-tokens, gigaam, cohere --n-ctx/--kv-type, and the bad-flag / bad-stream-param error paths.
Bug fix (H5). The host thread pools in the shared mel front end (transcribe-mel.cpp run_threaded), the gigaam mel, and parallel_for_all each hand-rolled spawn / run tid 0 / join: - a thread that failed to launch left already-launched threads joinable when the vector unwound -> std::terminate; - an exception from the calling thread's worker skipped the joins -> std::terminate; - an exception inside a pool thread escaped the thread function -> std::terminate. transcribe::run_on_threads (transcribe-batch-util.h) replaces all three: tid 0 runs on the caller, every launched thread is joined on every path, and the launch error or the lowest tid's worker exception is rethrown after the joins, so the public entry point's api_guard maps it to a status. The vDSP FFT setup in the mel front end is now released by a scope guard, since a worker exception can now unwind past it. Behavior change only on those failure paths (terminate -> error status). tests/run_on_threads_unit.cpp covers launch failure (injected through the Thread template seam), worker exceptions on the caller and pool threads, exception type preservation, lowest-tid ordering, and parallel_for_all propagation. Output on the success path is unchanged: transcribe-cli CPU parity vs main and the real-model ctest results are identical.
Behavior change (D16a, H3). transcribe_run, transcribe_run_batch and transcribe_stream_feed now return TRANSCRIBE_ERR_INVALID_ARG when any input sample is NaN or +-Inf. Before, a NaN went through the front end and could come out as a confident-looking transcript. Silence stays valid. - run: rejected with the other malformed-input checks, before the previous result (scratch slot and batch_results) is touched and before the family hook runs. - run_batch: a non-finite sample in any utterance rejects the whole batch before anything is cleared; the families' batched hooks never see it. (A NULL / empty utterance remains a per-utterance failure.) - stream_feed: rejected before the hook, so the stream stays ACTIVE with its revision unchanged and the caller can keep feeding. transcribe::pcm_is_finite lives in transcribe-abi.h so the upcoming role dispatchers share it. include/transcribe.h documents the rule on all three entry points (abihash unchanged). tests/nonfinite_input_unit.cpp covers NaN / +Inf / -Inf at the start, middle and end, result and stream-state preservation, and silence; it fails before this change. Valid-input behavior unchanged: transcribe-cli CPU parity vs main and the real-model ctest results are identical.
Pin today's behavior of the native compute sites before they move onto a shared helper. New tests/compute_rules.rs (model-gated, skips cleanly without the canaries): - busy_errors_name_the_refused_call: run, run_batch and a second stream on another session of the model each get Error::Busy with their own exact message while a stream is active. - option_validation_precedes_busy_check: a NUL in run options is reported as Error::Nul, not Busy, by run/run_batch/stream. - busy_check_precedes_native_stream_validation: a wrong-family stream extension is refused as Busy while another stream is live, and as InvalidArgument once the lease is free. - failed_stream_begin_does_not_take_the_lease. - ended_stream_never_releases_another_sessions_lease: reset() or drop of a finalized stream leaves another session's lease alone. - run_and_run_batch_across_threads_serialize: runs and batches on sessions of one shared model queue and each get their own results. - cancel_token_is_retained_until_cleared: the session keeps the token's flag alive after the caller drops it, and clear_cancel_token stops it. Also a compile_fail(E0505) doctest on Stream showing a session cannot be dropped while its stream borrows it, so no native free can race a call. (Rustdoc on stable does not check the code; verified by hand that the snippet fails only with E0505.) All new tests pass on the unchanged code (whisper-tiny Q8_0 and moonshine-streaming-tiny Q8_0 canaries).
The per-model compute lock and stream-lease check were open-coded at
seven sites (Session::run, run_batch, stream begin; Stream::feed,
finalize, reset, drop). They now all go through one crate-private
helper:
Session::with_compute(&mut self, refuse_if_streaming: Option<&str>,
f: impl FnOnce(*mut transcribe_session, &mut bool) -> R) -> Result<R>
It takes the model's compute lock (poison is recovered as before), then
if refuse_if_streaming is Some(msg) and a stream holds the lease returns
Error::Busy(msg), else calls f with the raw session handle and the lease
flag under the lock. The closure gets the lease flag so stream begin
can take it, and finalize/reset/drop can clear it only if this stream
still holds it, atomically with their native call.
Behavior-preserving. Each site keeps its exact order and error text:
- run / run_batch: options marshalled and lengths checked before the
lock; Busy "...before run()" / "...before run_batch()"; results are
still materialized after the lock is released.
- stream: Busy "a stream is already active on this model" before the
native begin; the begin status is still checked under the lock and
the lease only set on success.
- feed / finalize / reset / drop: no busy check (the caller holds the
lease); finalize/reset/drop clear the lease only when holds_lease.
Verified (macOS arm64, metal, all model-tier canaries set incl. voxtral):
cargo xtask bindgen --check, cargo fmt --all --check, cargo test -p
transcribe-cpp-sys, cargo test -p transcribe-cpp, the serde test, clippy
-D warnings, and the five examples. Per-test results identical to the
base commit, plus the new compute_rules tests and doctest passing; no
model-tier test skipped.
Pin the rules every native compute call follows today, ahead of moving the model lock behind one helper: - run, runBatch, stream begin/feed/finalize/reset and a dropped Stream's deinit reset all wait on the model-wide runLock (tested by holding the lock and checking the call blocks, then completes once released); - run/runBatch/stream validate string options before taking the lock, and refuse an active stream with .busy only after taking it (exact busy messages pinned, including a second stream on the same session); - the async bridge removes its token after the call, including after an abort; the sync path leaves a caller token installed; - a Session keeps its Model alive, a Stream keeps its Session and Model alive, and an in-flight async run survives the caller dropping refs; - results (Transcript, batch results, stream snapshot/text) are owned copies that later compute does not disturb. Tests only, no library change. All 11 pass on the current code; a quick mutation check (dropping feed's lock, the stream busy check, or moving run's option check under the lock) makes them fail.
Add an internal Model.withCompute(_:) (lock runLock, run the body, unlock via defer; rethrows) and move every native compute site onto it: Session.run, Session.runBatch, Session.stream (begin), Stream.feed, Stream.finalize, Stream.reset and Stream.deinit. No other code takes runLock directly now. The lock stays on Model so every session of one model shares it. Behavior-preserving. Each site keeps its checks in the same order: option NUL checks stay before the lock; the streamActive .busy refusal (run/runBatch/stream only, same messages) is the first step inside the body; finalize still releases the lease before checking the status; reset still reads `state` after the lock is released; deinit still checks holdsLease before locking; results are still copied into Swift values inside the body. reset and deinit used a manual unlock instead of defer; their bodies cannot throw, so defer is equivalent. The async cancel-token bridge stays in Session, outside the lock, as before. Verified with the CI gate (abihash check, macOS xcframework build, swift test with every TRANSCRIBE_SMOKE_* model set, the five examples): 73/73 tests pass, 0 skipped, identical to the run before this change, and the example outputs match.
Add test/compute-rules.test.mjs, pinning the rules every native compute call follows today so the next commit can route run / runBatch / stream begin / feed / finalize through one helper and show nothing moved: - model-wide exclusion: a sibling session's run queues FIFO behind the in-flight one, and only the lock holder is marked in flight; runBatch and finalize use their own in-flight labels - active-stream rejection: run, runBatch and stream begin are refused with Busy (exact "cannot <op>:" prefix) on the stream's own session and on a sibling; disposal is rechecked inside the lock BEFORE the lease, so a queued call on a disposed session sees "disposed", not Busy - feed / finalize do not recheck disposal inside the lock: one queued before session dispose still runs to completion - a non-finite feed is rejected with InvalidArgument and the native stream stays active, but the binding releases the model lease on any non-OK feed status, so a sibling run is no longer refused. Pinned as current behavior, not endorsed - cancel: the abort listener exists only for the call, is removed after success and after abort, is never installed when the call is refused (Busy or disposed), and a later signal-less run is not aborted; a pre-aborted runBatch resolves with per-item Aborted + partials - session or model dispose while run / runBatch / feed is on a worker defers the native free behind it and the call still returns its full result (reads during the call report "in flight" before "disposed") - results are copied out before the lock is released: two back-to-back runs on one session each keep their own text All 13 pass against the current code (3 repeated runs), alongside the existing 52.
Every native compute call now goes through one private Session helper,
#exclusive(gate, body), instead of five inline copies of the lock
sequence. Pure refactor, no behavior change.
What moved:
- #exclusive queues on the model-wide FIFO lock (the Mutex still lives
on TranscribeModel and is shared by all its sessions), then applies the
gate in the old order: disposal recheck, then the stream-lease Busy
check. Each site passes its old settings: run / runBatch / stream begin
recheck disposal and refuse with Busy("run" / "runBatch" / "begin a
stream"); feed / finalize (STREAM_LEASE_HOLDER) do neither, as before.
- The in-flight window (install the abort callback, then set the
in-flight label; on close clear the label, then uninstall) is a
ComputeWindow the body opens around its worker await. The labels stay
"run()", "runBatch()" and "feed()/finalize()". stream begin opens no
window (synchronous native call), and result copy-out stays in the body,
after the window closes and before the lock is released.
- Stream reaches the helper through the existing module-private
SESSION_CONTROL WeakMap (the STREAM_TEARDOWN pattern). Its
enterCompute/leaveCompute entries are replaced by exclusive(); nothing
new is exported.
- #exclusive is not async. Refusals come back as rejected promises and
the body's promise is returned unchanged, so promise settle timing
matches the inline lock bodies.
The one non-literal translation: run/runBatch used to clear the
in-flight label unconditionally, while feed/finalize cleared it only if
it still matched. The window now always uses the matching check. Only the
lock holder sets the label, so at close it always equals the caller's
kind, and the two forms behave the same.
deferFree (stream reset, Session.dispose, Model.dispose) stays as is.
Those are queued frees, not computes, and must not recheck disposal.
Verified against a worktree-built libtranscribe (Metal) with every
TRANSCRIBE_SMOKE_* model set: generate.py --check, version sync, tsc,
npm test (65/65; the 52 baseline tests plus the 13 characterization
tests, the same results as before), the 5 CI examples and the deno smoke.
Mutation checks: swapping the disposed/Busy order, or making feed
recheck disposal, each fails the expected characterization tests.
feed() released the model-wide stream lease on any non-OK status. Since 2013849, native rejects non-finite samples before the family hook and leaves the stream ACTIVE. The binding still freed the lease, so a sibling run or stream could start alongside a live stream, and later feeds ran without holding the lease. The lease now follows the native stream lifecycle. On a failed feed, the binding reads transcribe_stream_get_state, still inside the same exclusive slot after the native feed returns, and releases the lease only if the stream is no longer ACTIVE. - Pre-hook rejections leave the stream ACTIVE and keep the lease. That covers non-finite samples and any other pre-hook INVALID_ARG or NOT_IMPLEMENTED. The stream stays usable and siblings stay Busy. - A hook failure (native FAILED) releases the lease exactly as before. - A feed on a stream that is no longer ACTIVE (e.g. FINISHED) releases as before. That release is a no-op, because finalize already released. - finalize is unchanged and always releases. The feed() comment now describes this, and the README gets one sentence. JS-side checks still stop empty and misaligned PCM, reset or stale wrappers, and disposed sessions before any native call. Tests: the "a rejected feed releases the stream lease (current behavior)" characterization is replaced by two tests. - A NaN feed throws InvalidArgument, and the stream stays active. Sibling run, sibling stream, and a same-session run are all Busy. A following valid feed and finalize succeed, and then a sibling run goes through. - A feed that fails inside the native hook still releases the lease. The test makes the hook fail with an always-firing abort callback, which moonshine-streaming's hook polls, returning ABORTED and leaving the stream FAILED. The callback is installed on the session handle through koffi, test-only. The stream reads "failed" with an Aborted lastStatus, and a sibling run then goes through. The first fails on the previous code. The second fails if the release is removed. Verified with a worktree-built libtranscribe (Metal) and every TRANSCRIBE_SMOKE_* model set: - generate.py --check, version sync, and tsc all pass - npm test passes 66/66 - the 5 CI examples and the deno smoke all pass
Pin the per-call rules the Python binding relies on today before the model-wide compute lock lands, so the next commit can show they are unchanged: - run() / run_batch() on a session with an ACTIVE stream are rejected (native InvalidArgument) and the stream stays usable afterwards (begin-while-active was already covered). - a cancel() requested before run() / run_batch() / stream() is cleared when the call starts and does not abort it. - a Stream keeps its Session (and Model) alive when every other reference is dropped. Verified on the unmodified binding against a build-shared libtranscribe with whisper-tiny Q8_0 and moonshine-streaming-tiny Q8_0: all new tests pass.
BEHAVIOR CHANGE (0.4.0 release notes): compute calls on one Model now serialize. Session.run(), run_batch(), stream() and Stream.feed() / finalize() / reset() take a new model-wide lock, so concurrent calls on sessions of the same model from different threads WAIT for each other instead of racing in native code. Up to 0.3 the binding did not lock: ctypes releases the GIL around every foreign call (CDLL), so such calls overlapped natively and could corrupt decodes or fail Metal command buffers. Callers that serialized runs themselves keep working (their lock simply nests outside this one). Contended calls never raise. Model._exclusive(kind) is the single bridge every compute site goes through. It holds the lock for the native call AND the copy-out of the results, checks the model is still open once the lock is held, and the caller captures its session handle inside the block. - Plain threading.Lock plus an owner thread id, not an RLock. A thread can only re-enter through a callback (the log handler) firing inside a native call. Letting that start a second compute on the same model mid-flight is the race this lock prevents, so a re-entrant call raises TranscribeError instead of deadlocking. The abort callback only reads a threading.Event, so cancel() stays lock-free. - Session.close(), Model.close() and __del__ never wait for the lock. They mark the object closed at once and queue the native free, which runs immediately if the lock is free. Otherwise the current holder runs it just before releasing, so an in-flight call finishes its copy-out on the handle it captured and the free follows it. FIFO order frees sessions before their model, and the deferred session free pins the abort trampoline. This also covers GC running a finalizer on the thread that holds the lock. - The cancel flag is still cleared before the lock wait, so a cancel() issued while a call is queued aborts it as soon as it starts. - Reads (text, snapshot, state, revision, last_status, limits, was_aborted) and model queries stay lock-free. - There is no stream lease: an active stream does not block runs on other sessions between feeds. The docstrings tell callers not to do that. The Model, Session, cancel(), close() and set_log_callback docstrings and the README now describe this instead of "serialize runs yourself". Tests: the new test_compute_lock.py needs no model. It patches the _lib.* native entry points (looked up at call time) with probes on fake handles and covers these cases: - no overlap across two sessions and a stream - a contended call waits rather than raises - the lock is per model - every compute site, including copy-out, runs under the lock - the lock is released after a native error - a cancel issued while queued is kept - session/model close during an in-flight call defers the frees (sessions before model) and does not deadlock - GC on the holder thread defers the free - re-entrant compute raises Ordering uses Events and a lock spy that signals a blocking acquire, not sleeps. With _exclusive swapped for a no-op, 9 of these 11 tests fail; 15 repeated runs of the file all passed. test_concurrent_sessions_on_shared_model (two threads, no caller lock, real whisper-tiny) was an xfail and is now a plain passing test. It was run 5x on Metal and 5x with TRANSCRIBE_BACKEND=cpu. test_rejected_feed_keeps_stream_active checks a real moonshine stream: a NaN feed is rejected before the family hook (InvalidArgument), the stream stays "active", the lock is released, and valid feeds plus finalize still work. The binding has no lease flag, so it has nothing to release on that path. Verified against a build-shared libtranscribe with every model tier enabled (whisper-tiny, moonshine-streaming-tiny, nemotron-3.5 prompted, canary-180m-flash, SenseVoiceSmall, nemotron-speech-streaming, parakeet-unified, Voxtral-Mini-4B-Realtime). Baseline was 176 passed, 1 skipped, 1 xfailed. After: 193 passed, 1 skipped. Every baseline result is unchanged except the xfail, which now passes. The one skip is the speed-dependent cross-thread cancel test, same as baseline. The FFI drift gate and version sync are clean.
Bug fix. The streaming dispatcher marked a stream FAILED only when the family hook returned a non-OK status. If begin / feed / finalize threw, the exception unwound straight to the api_guard: the caller got OOM or BACKEND, but the stream stayed ACTIVE (and last_status unset), so bindings that follow the native state kept treating it as live. A throwing stream_reset skipped the dispatcher's wipe, so reset could leave the stream ACTIVE too. - call_stream_hook wraps the begin / feed / finalize hooks: on an exception it sets FAILED and last_status to the status the guard will return (OOM for bad_alloc, BACKEND otherwise), then rethrows. - reset runs the clear + IDLE + was_aborted wipe even when its hook throws, then rethrows for the guard to log. tests/stream_hook_throw_unit.cpp covers each hook with bad_alloc and a runtime_error (21 failures before this change). Valid-input behavior is unchanged: ctest green and transcribe-cli CPU parity vs main identical.
BEHAVIOR CHANGE (0.4.0 release notes): while a stream is active on a Model, Session.run(), run_batch() and stream() on ANY session of that model now raise the new transcribe_cpp.Busy at once. Before, a run on another session waited for the compute lock and then ran between the stream's feeds, which the C header forbids. A run(), run_batch() or second stream() on the stream's own session also raises Busy now, where it used to raise InvalidArgument from the native dispatcher. Runs that contend with other runs still wait, as before. This matches the stream lease the Rust, Swift and TypeScript bindings already have. Busy is a new TranscribeError subclass in errors.py, exported from the package. It is raised on the Python side only, so status is 0. No native status maps to it, the same as Rust's Error::Busy and TS's Busy. The messages reuse the Rust/Swift wording: - "a stream is active on this model; finish or drop it before run()" - the same with run_batch() - "a stream is already active on this model" (stream begin) How it works: - Model._exclusive(kind, busy=...) checks the lease after it acquires the lock and after the model-closed check. A call that was already queued when a stream began therefore sees that stream's lease. - The lease is _stream_lease.owner on the model: the _StreamLease token of the stream that holds it, or None. It is claimed under the lock only after transcribe_stream_begin succeeds. The Stream is built inside the same locked block. - Each Stream keeps its own token and releases the lease only while it is still the owner. A stream that already ended therefore never clears a lease a later stream took (the per-stream holds_lease of the other bindings). - The stream's own feed() and finalize() take the lock with no busy check. The lease is released by: - finalize(), always, on success or failure. - reset(), and so context-manager exit. - A failed feed(), but only if transcribe_stream_get_state no longer reads ACTIVE. That state is read under the same lock right after the feed. A feed rejected before the family hook (NaN/Inf -> InvalidArgument) leaves the stream ACTIVE and keeps the lease, as in TS commit 8274003. - Session.close(), inside the session's deferred free and right after transcribe_session_free, so no call can start while the stream's native session still exists. Model.close() closes its sessions, and its own deferred free clears the slot too. - Stream.__del__, for a stream that still holds the lease. Like the native frees, it never waits for the lock: GC can run it on the thread that holds the lock. It queues a closure through _free_or_defer that resets the native stream and releases the lease. The closure captures the slot and the token, never the Stream, Session or Model. Under the lock it re-checks that the token is still the owner. While the token is the owner, the session is still open, because session close clears the owner in the same step as its free. So the reset never touches a freed or reused session. Docstrings for Model, Session, Session.stream/run/run_batch/close, Stream, Model.close and the module, the README concurrency note and a test_lifetime comment now describe the lease. Before, they told callers not to run other sessions during a stream. Tests (test_compute_lock.py, model-free, using the fake-handle probes): - Sibling and same-session run / run_batch / stream raise Busy with the exact messages and status 0, and never reach native code. The lease holder keeps feeding and finalizing. - The lease is per model. A failed native begin takes no lease. - The lease is released by finalize, a failed finalize, reset, context-manager exit, GC of an active stream (reset exactly once), session close (a later GC of that stream does no native reset) and model close. - A rejected feed with the stream still ACTIVE keeps the lease, and a valid feed plus finalize then release it. A feed failure that leaves the stream FAILED releases it. - Stream A ends and stream B on another session begins. A failed feed on A, a finalize on A, A.reset(), GC of A and closing A's session then leave B's lease alone. - The busy check runs after the lock is acquired: a run queued behind a gated begin raises Busy if the begin succeeds, and runs if it fails. - An active stream GC'd on the lock-holder thread does not deadlock. It defers its reset and release until that call ends. - A session closed while another call holds the lock is freed before a queued run starts. test_streaming.py (real moonshine-streaming-tiny): - test_second_stream_while_active_rejected and test_run_while_stream_active_rejected now expect Busy on the same session and on a sibling. A sibling run works after finalize. - test_rejected_feed_keeps_stream_active now also checks that sibling run, sibling stream and same-session run are Busy after the NaN feed, and that a sibling run works after finalize. - New: a feed that fails in the hook (pending cancel -> Aborted, state "failed") releases the lease. - New: reset, GC or session close of an active stream frees the model. Mutation checks over test_compute_lock/test_streaming/test_lifetime: - busy check removed: 16 tests fail. - the feed never releases: 2 fail. - the feed always releases (pre-82740039 TS behavior): 2 fail. - the GC release removed: 3 fail. - the GC release run inline instead of deferred: 1 fails. - the session-close release removed: 3 fail. - release without the owner check: 1 fails. test_compute_lock.py passed 15 of 15 repeated runs. Verified against a build-shared libtranscribe built from this tree, with every TRANSCRIBE_SMOKE_* model set (whisper-tiny, moonshine-streaming- tiny, nemotron-3.5 prompted, canary-180m-flash, SenseVoiceSmall, nemotron-speech-streaming, parakeet-unified, Voxtral-Mini-4B-Realtime). Baseline at 2c337e7 was 193 passed, 1 skipped. After: 215 passed, 1 skipped. Every baseline test still passes. The 22 new results are the tests above. The one skip is the same speed-dependent cross-thread cancel test. The FFI drift gate and version sync are clean.
Additive (D2, D5). - transcribe_role (TRANSCRIBE_ROLE_ASR, TRANSCRIBE_ROLE_DIARIZE) and transcribe_model_roles(model): the bitmask a loaded model serves, fixed at load, 0 for NULL. The internal k_role_* bits now take the public values (static_assert'd). - TRANSCRIBE_ERR_UNSUPPORTED_ROLE (20, appended) with a status string. - transcribe_session_init (and so transcribe_open) and transcribe_model_get_capabilities return UNSUPPORTED_ROLE for a model without the ASR role. Capabilities leave the caller's struct untouched rather than returning zeros that read as "supports nothing". Every model that loads today serves ASR, so no current caller sees the new code. Low-level bindings regenerated from the header (Python / TypeScript via generate.py, Rust via cargo xtask bindgen); transcribe.abihash updated. The Swift pinned hash and per-binding status mapping follow in their own commits. tests: role_resolve_unit covers the role checks (session_init leaves *out NULL, capabilities untouched) and the enum/bit agreement; api_smoke covers the new status string and transcribe_model_roles(NULL).
Status 20 now raises a new UnsupportedRole (direct TranscribeError subclass: the model loaded fine but lacks the role the call needs, so it is neither a request option nor a load failure). Session init and capabilities already route through check() -> exceptionForStatus, so they pick it up unchanged. test/errors.test.mjs (no-model tier) pins the full status -> class table against every TRANSCRIBE_ERR_* in _generated.ts, mirroring the Python test_errors.py, plus OK, unknown-status, and OutputRepetition cases.
Status 20 now maps to its own Error::UnsupportedRole variant ("unsupported
role: ...") instead of falling through to Error::Other, and raw_status()
reports it back as 20. Kept separate from Error::Unsupported, which is the
per-request shape group (task / language / timestamps / PNC / ITN).
tests: error.rs unit tests walk every status the linked library names
(until transcribe_status_string returns its unknown text) and check that
each one maps to a typed variant whose raw_status maps back to the same
variant, plus a test that pins the UnsupportedRole message and code.
Additive. include/transcribe/diarize.h: transcribe_diarize_session with get_info / session_init / session_free / set_abort_callback / run / n_segments / get_segment / get_timings, plus info / session_params / params structs, their inits and ABI ids. Streaming is deferred to #175, the first streaming diarizer (plan Q16). TRANSCRIBE_EXT_SLOT_DIARIZE_RUN is appended; diarize.h joins the extensions.h umbrella. src/transcribe-diarize.{h,cpp}: the dispatcher follows the ASR rules (api_guard on compute / ownership entry points, *out NULL on failure, UNSUPPORTED_ROLE for a model without DIARIZE, malformed input / bad extension / family run_validate rejected before the previous result is cleared, non-finite PCM rejected, scratch released after every run). Families only compute: DiarizeOps::run returns frame probabilities and the dispatcher turns them into segments with transcribe::probs_to_segments, the threshold-0.5 per-speaker run builder that used to be sortformer::probs_to_speaker_segments (#175's builder is the same algorithm). The multitalker parakeet bundle now calls the shared one. Sortformer: roles = ASR | DIARIZE. Both session types share one compute path (diarize_pcm over SortformerWork); the DIARIZE run takes the new SFDR kind (transcribe_sortformer_diarize_ext) on the DIARIZE_RUN slot. The ASR path stays until the break commit. tests: diarize_dispatch_unit (fake arch: roles, info, pre-clear rejections, segments, accessors); sortformer_diarize_unit (real model: diarize_run segments byte-identical to the ASR-path transcribe_run at DEFAULT and LOW_LATENCY, bad preset preserves the result, and with a whisper GGUF an ASR session and a diarize session on two threads match their sequential results). ctest green; transcribe-cli CPU parity vs main identical (multitalker --diarize included). Low-level bindings are regenerated with the binding wrappers in a later commit.
transcribe-cli runs the DIARIZE role on a model without ASR, or on any model with --role diarize (--role asr|diarize picks the role on a model serving both; the default stays ASR when available). The shared prelude (WAV load, model load and its metadata lines) is unchanged; right after it, run_asr_file hands a diarize model to run_diarize_file (diarize.cpp), which prints the speaker segments and the realtime line. validate.py and scripts/diar/run_cpp_sortformer.py call the CLI on the Sortformer GGUF and keep working once Sortformer is DIARIZE-only. tests: cli_diarize_smoke (gated on TRANSCRIBE_SORTFORMER_GGUF, reported as skipped otherwise). transcribe-cli CPU parity vs main identical apart from the two new --role usage lines.
Generated only: Python / TypeScript via _generate/generate.py and Rust via cargo xtask bindgen pick up include/transcribe/diarize.h, the DIARIZE_RUN slot, the diarize ABI ids and the Sortformer SFDR extension. transcribe.abihash updated; the Swift pin and the hand-written wrappers follow per binding.
- catalog: optional "role" (asr | diarize, absent = asr); Sortformer is "diarize". render.py's family-index groups README tables by role (role=diarize replaces transcribe=false). format / check / render / db-mapping checks pass. - validate.py: Sortformer's C++ dump runs with --role diarize, and a diarizer run (speaker segments, no text line) no longer needs a transcript. The diarize-API dumps are byte-identical to main's ASR-path dumps on the oracle mix. (The sortformer compare is 1/6 within tolerance against a freshly generated reference on both main and this branch: reference drift, not this change.) - docs/roles.md: what a role is and the rules every role follows. - porting skills: intake records the role; acceptance scores DER for a DIARIZE model.
Model::roles() -> Roles (contains(Role), bits()). Model::diarize_info(),
Model::diarize_session()/_with(DiarizeSessionOptions), and DiarizeSession
with run(pcm, &DiarizeOptions) -> Vec<SpeakerSegment>, timings(), and
set/clear_cancel_token. DiarizeOptions::family takes
DiarizeExtension::Sortformer(SortformerDiarizeOptions { preset }), which
materializes the SFDR ext on the new ExtSlot::DiarizeRun.
The compute lock helper moves to ModelInner::with_compute; Session's
with_compute delegates to it with its handle, and DiarizeSession::run uses
it directly, so diarize runs serialize with every session of the model and
are refused with Busy while a stream holds the lease.
Breaking (0.4.0): Model::capabilities() returns Result and reports the C
status (UnsupportedRole on a model without ASR) instead of zeroed
defaults. ExtSlot gains DiarizeRun.
tests: Busy before native (model-free unit test); Sortformer roles, info,
ext acceptance, diarize run parity with the ASR-path speaker turns at
DEFAULT and LOW_LATENCY, bad input, cancel, model kept alive; whisper
refuses diarize with UnsupportedRole.
Model.roles (Roles option set), Model.diarizeInfo, Model.diarizeSession and DiarizeSession.run returning [SpeakerSegment], with the Sortformer preset on the DIARIZE_RUN slot (DiarizeExtension.sortformer), cancellation tokens and timings. Diarize runs go through Model.withCompute, so they share the model lock and stream lease with sessions. Model.capabilities now throws instead of discarding the native status (.unsupportedRole on a model without ASR); breaking change for 0.4.0. Pin header hash c2979c7e891982ee after auditing diarize.h and the sortformer.h / transcribe.h additions.
model.roles (readonly array of "asr" | "diarize"), model.diarizeInfo
(sampleRate, maxSpeakers), model.createDiarizeSession({ nThreads }), and
DiarizeSession: run(pcm, { family, signal }) -> SpeakerSegment[] (copied
out), timings, dispose() / Symbol.dispose. Both throw UnsupportedRole on a
model without the role; ASR calls already map status 20 the same way.
The Sortformer diarize preset is modeled like the ASR one: a
{ kind: "sortformer_diarize", preset } family extension on the new
"diarize_run" slot (SFDR). Both Sortformer kinds now share one preset
mapper that rejects an unknown preset string instead of silently sending
DEFAULT.
Execution bridge: Session's handle, disposed flag, in-flight mark, compute
admission (#exclusive: model lock, disposed recheck, then stream-lease
Busy), abort install, and deferred free move into a module-private
SessionCore, held in a private field by both Session and DiarizeSession and
parameterized by the role's set_abort_callback and free. Stream still
reaches it through SESSION_CONTROL. The speaker-segment and timings readers
are shared with materialize().
tests: diarize.test.mjs. Model-free: DiarizeSession's public surface.
Whisper: roles ["asr"], diarizeInfo / createDiarizeSession throw
UnsupportedRole, run() rejects the diarize_run extension. Sortformer
(TRANSCRIBE_SMOKE_SORTFORMER_MODEL, local-only): roles, maxSpeakers 4,
diarize rows equal the ASR path's speakerSegments at default and
low_latency, bad preset and wrong-slot extensions rejected, diarize and ASR
runs serialize FIFO on one model (these two use the Sortformer ASR path and
are marked to drop with it), abort listener lifetime and pre-aborted
Aborted, disposed
recheck inside the lock, deferred free on session and model dispose during
an in-flight run. Busy is not covered for diarize: Sortformer cannot begin
a stream, so no model holds a stream lease and serves DIARIZE.
Breaking (0.4.0). Sortformer drops the ASR role: transcribe_session_init / transcribe_open / transcribe_model_get_capabilities return UNSUPPORTED_ROLE for it, and its ASR path (init_context, run, run_validate, the dual-session plumbing) is gone; one SortformerSession serves transcribe_diarize_run. The SFST RUN-slot kind and transcribe_sortformer_stream_ext are removed and the kind is retired in docs/extension-kinds.md (SFDR replaces it). Sortformer no longer sets capabilities or feature bits (capabilities.cpp removed), so TRANSCRIBE_FEATURE_DIARIZATION keeps its ASR meaning only. transcribe-cli routes the DIARIZE-only model to the diarize path by itself, so validate.py (its --role flag dropped again) and scripts/diar/run_cpp_sortformer.py work unchanged; the C++ dumps on the oracle mix are byte-identical to main's. Batch mode stays ASR-only. tests: sortformer_stream_ext_unit removed (its env-preset parity check moves into sortformer_diarize_unit); sortformer_diarize_unit now checks the DIARIZE-only roles, ASR refusal, golden segments at DEFAULT and LOW_LATENCY, ext == TRANSCRIBE_SORTFORMER_STREAM_PRESET env, pre-clear rejection, and cross-role concurrency. Low-level bindings regenerated; the binding-side removals follow per binding. docs: docs/migrating-to-0.4.md; Sortformer model page, HF card and family doc moved to the diarize API; env-var table.
Follows the C break: Sortformer serves only the DIARIZE role and the SFST RUN-slot extension is gone. Remove RunExtension.sortformer and SortformerStreamOptions; SortformerPreset stays for the diarize extension. Pin the reviewed header hash to 30ec88d3b51fc344. tests: drop the temporary Session.run/SFST parity check; testSortformerDiarizes now loads on CPU and checks golden turns on samples/sortformer-2spk-mix.wav at DEFAULT and LOW_LATENCY, roles == [.diarize], and that capabilities and session() throw .unsupportedRole.
Follows the core break: the SFST run-slot extension is gone, so the safe crate drops RunExtension::Sortformer and SortformerStreamOptions. SortformerPreset stays for DiarizeExtension::Sortformer. tests: the ASR-path parity block in tests/diarize.rs is replaced by CPU golden segments on samples/sortformer-2spk-mix.wav (DEFAULT and LOW_LATENCY); role expectations are DIARIZE only, and capabilities() and session() on the Sortformer model return UnsupportedRole. serde coverage drops the removed type.
Follows the core break: Sortformer serves only the DIARIZE role and the RUN-slot SFST extension is gone from the C ABI. Drop the "sortformer" family kind (SortformerStreamOptions, its FAMILY entry, and the transcribe_sortformer_stream_ext_init binding); "sortformer_diarize" keeps the shared preset mapping. README notes the DIARIZE-only model. tests: the ASR-path equality test is replaced by golden segments on samples/sortformer-2spk-mix.wav at DEFAULT and LOW_LATENCY (CPU backend, the goldens are CPU); roles are ["diarize"], and capabilities / createSession throw UnsupportedRole on the Sortformer model. The diarize-vs-ASR exclusivity test is dropped: with two models there is no shared compute lock to order. The retired "sortformer" kind is rejected as an unknown kind, and a run-slot kind (whisper) on diarize_run as InvalidArgument.
Remove SortformerStreamOptions (the SFST run-slot extension retired in the
C API); SortformerDiarizeOptions now owns the preset table, __init__ and
_apply it borrowed. README notes Sortformer serves only Role.DIARIZE.
tests: the temporary DIARIZE-vs-ASR cross-check becomes CPU golden
segments on sortformer-2spk-mix.wav at DEFAULT and LOW_LATENCY; roles are
{Role.DIARIZE}, and Model.capabilities / model.session() raise
UnsupportedRole on the real Sortformer model. test_family_ext drops the
removed class.
No behavior change; exported symbols unchanged. - transcribe-abi.h: TRANSCRIBE_FIELD_END defined once; init_sized (every *_init body) and copy_out_timings; transcribe-session-core.h: copy_out_speaker_segment and ScratchReleaseGuard, shared by the ASR and DIARIZE dispatchers instead of three copies each. - transcribe_model_load_file calls resolve_roles once, in the forwarder. - run_one_inner no longer re-scans PCM for NaN / Inf (both callers did). - DiarizeOps::new_session replaces init_session (the dispatcher sets model and n_threads); the diarize session drops has_result (segments are built into a local and swapped in). - k_role_* aliases removed in favour of TRANSCRIBE_ROLE_*; transcribe_session drops constructor / copy boilerplate SessionCore already provides; the mel run_threaded pass-through is gone. - Unused includes from the split and the Sortformer ASR removal dropped; comments trimmed (history, ASR-only wording on the shared base, the UNSUPPORTED_ROLE and diarize timings docs, stale Sortformer notes).
- transcribe-cli: drop --role (no model serves two roles; a DIARIZE-only model is routed by its role mask, which cli_diarize_smoke now exercises); json_escape / write_output_file move into asr.cpp, their only user; unused includes and stale comments removed. - tests: the five fake-arch units register through one CMake loop; redundant cases dropped (silence, tautological role-bit asserts, dispatcher refusal checks repeated in sortformer_diarize_unit); plan-ID tags and history notes removed from comments. - docs: stale Sortformer ASR-path references fixed; duplicated role / binding notes in migrating-to-0.4.md, roles.md and diarize.h shortened to pointers.
SessionCore.#installAbort's cleanup re-read the handle through the disposed-checking getter, so disposing the session (or model) during a run with a signal turned the finished result into a "disposed" error and leaked the koffi callback and signal listener. Capture the native handle at install; the native free is queued behind the lock, so it is still valid at cleanup.
- SessionCore.exclusive takes busyOp (null = the lease-holding stream) instead of a ComputeGate; drop STREAM_LEASE_HOLDER. - One call(kind, signal, fn, ...args) helper replaces the open()/close() pairs at the five worker-call sites. - Stream reaches SessionCore via SESSION_CONTROL; one assertNotComputing (same messages); drop SessionCore.inFlight. - Inline sortformerPreset; shorten comments and UnsupportedRole's doc. - Tests: drop duplicates of streaming/errors coverage and dead guards. - README: fix a stranded reflow; tighten the diarize and feed paragraphs.
…estored tests, stale Python Stream, CLI diarize -o/--batch, Rust non_exhaustive
- sortformer_diarize_unit: drop test_cross_role_concurrency (two separate models, overlap never forced, needs a second GGUF). - run_on_threads_unit: 4 tests (every tid once, worker exception rethrown after join, launch failure joins and rethrows, parallel_for_all propagates). - nonfinite_input_unit: drop the silence controls and the position loop. - stream_hook_throw_unit: one loop over begin/feed/finalize; drop the feed-after-FAILED tail already covered by stream_dispatch_unit. - role_resolve_unit: table-driven resolve_roles cases; add the diarize-only arch cases (roles=DIARIZE ok, roles=0 rejected). - diarize_dispatch_unit: drop asserts on session internals.
- cancel.rs: drop cancel_token_is_retained_until_cleared (pre-existing behavior, unrelated to roles); file back to its pre-PR form. - streaming.rs: drop option_validation_precedes_busy_check (pins which of two errors wins). - diarize.rs: exact segments live in tests/sortformer_diarize_unit.cpp; the binding test now checks well-formed turns and that LowLatency changes the result (preset reaches native).
test_compute_lock.py keeps 12 tests for the model-wide lock and stream lease: serialization, every compute site locked, cancel while queued, close / model close / GC during an in-flight call, re-entrancy, Busy, failed begin, lease release per ending, ended stream vs a later lease, diarize Busy. Dropped tests that were true by construction (per-model lock/lease), fake twins of real-model tests, internal-ordering pins, and the diarize variants of _SessionBase paths. test_streaming.py: drop the session-close and keep-alive duplicates; the hook-failure lease check moves into the existing cancellation test. test_diarize.py: exact segments live in tests/sortformer_diarize_unit.cpp; check well-formed turns and that low_latency changes the result.
ComputeRulesTests pinned pre-existing locking through internal state and 0.3 s lock-wait timeouts. The two checks worth keeping move to existing files: StreamingTests.testActiveStreamRefusesSiblingCompute (restores the Busy coverage deleted from StreamingTests) and DiarizeTests.testDiarizeRunIsBusyWhileAStreamIsActive. DiarizeTests: exact segments live in tests/sortformer_diarize_unit.cpp; check well-formed turns and that lowLatency changes the result. The status-mapping test compares case names instead of a 20-line switch.
compute-rules.test.mjs pinned pre-existing #exclusive behavior. The two tests for what this PR changed (a pre-hook feed rejection keeps the stream lease; an in-hook feed failure releases it) move to streaming.test.mjs next to the other lease tests. diarize.test.mjs: drop the public-surface, abort-listener and dispose tests (same shared SessionCore path); exact segments live in tests/sortformer_diarize_unit.cpp, so check well-formed turns and that low_latency changes the result. errors.test.mjs: keep the status table and unknown-status tests only.
cjpais
force-pushed
the
roles/pr1-diarize
branch
from
October 4, 2026 02:07
f8bdec4 to
4723ae9
Compare
NairoDorian
pushed a commit
to NairoDorian/transcribe.cpp
that referenced
this pull request
Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.