Skip to content

Release 0.19.1 - #739

Merged
thcp merged 8 commits into
mainfrom
next-release
Oct 2, 2026
Merged

thcp merged 8 commits into
mainfrom
next-release

Conversation

@thcp

@thcp thcp commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Release 0.19.1: everything on next-release since v0.19.0.

Verified on local Windows NVIDIA builds: 0.19.1.dev0 passed 12 of 12 (#727, #737), and 0.19.1.dev1 passed 18 of 18 automated checks including #740, in an isolated profile. uv.lock is unchanged since v0.19.0, so the in-app update works.

Closes #722
Closes #728
Closes #729
Closes #723
Closes #724
Closes #726
Closes #730
Closes #731
Closes #732
Closes #733
Closes #736
Closes #701
Closes #738
Closes #725
Closes #735
Closes #734
Closes #719
Closes #742

Thales and others added 8 commits September 30, 2026 07:23
The lookahead gate compared _scheduledTo, in source seconds, against the
output playhead. With SoundTouch stretching, sources play at 1x and the
worklet buffers the surplus, so below 1x the output playhead lags what
the sources have consumed. The gate overstated its margin by (1 - rate)
seconds per second until the sources ran dry and the next chunk was
scheduled late and skipped. At 0.75x that happened about fifty seconds in.

The gate now measures against ctxTimeToSourceTime, which already follows
the source rate, so it also holds in the tape-effect fallback.

Reported and diagnosed by @goermezer in #701.

Closes #722
Below 1x each WSOLA sequence plays 70 ms of input but advances only
70 * tempo, so consecutive sequences overlap and an attack in the
overlap is heard twice. At 0.75x kicks came out as a ~50 ms flam.

The shared tempo stage now marks attacks on the sample-to-sample
difference, and when the next sequence would replay one it carries
straight on from where the last sequence ended instead. That splices
identical samples, so it is seamless. The extra input is borrowed, at
most 100 ms, and paid back by later sequences stretching slightly more.

Only the shared tempo stage does this. The pitch stages are aligned
sample for sample with the unpitched drums, and borrowing there would
pull the band off the kit.

Measured on a full mix at 0.75x: 93 attacks in, 104 out before, 93 out
after.

Closes #728
WSOLA slows audio by repeating overlapping 82 ms fragments, heard as an
echo on everything sustained. The shared tempo stage now uses the
Signalsmith Stretch WASM core (MIT, vendored from signalsmith-stretch
1.3.2) at 40 ms blocks, chosen by ear against 30, 60 and 120 ms.

- The core is its own worklet module, added before the processor. The
  processor instantiates it asynchronously and swaps it in at the next
  flush, never under audio already playing. If it does not load, WSOLA
  stays the tempo stage.
- The processor reports which stage runs and its latency in two parts.
  Latency is the priming plus the core's input side, divided by the
  tempo, plus its output side; measured end to end it lands within 1 ms.
  Both engines take it from one shared module, tempoStage.js, instead of
  each hard-coding WSOLA's.
- script-src gains 'wasm-unsafe-eval', in the server CSP and the Tauri
  CSP. It allows WebAssembly compilation only, not JS eval. The CSP test
  now pins script-src to exactly 'self' and 'wasm-unsafe-eval'.

On 20 s of a full mix at 0.75x (55 attacks in): WSOLA 60 attacks out
with 8 close pairs, Signalsmith 61 with 2.

Closes #729
Closes #722
Closes #728
Closes #729

Reported and diagnosed by @goermezer in #701. This takes the engine half
of their patch, with two adjustments. The slider half stays open in #701
as a separate design question.

## What changes

The chunk scheduler's lookahead gate now measures against
`ctxTimeToSourceTime(ctx.currentTime)` instead of the output playhead.

- **SoundTouch path:** sources play at 1x, so the gate now counts at 1x.
Below 1x it no longer overstates its margin, and the sources stay fed.
- **Tape-effect fallback:** `ctxTimeToSourceTime` follows `_srcRate()`,
so the gate still counts at the slowed rate there, the same as before.
- **1x:** unchanged when nothing is transposed. With a transpose it
schedules the pipeline latency (a fraction of a second) earlier, which
is harmless.

Differences from the patch in #701:
1. It reuses the existing `ctxTimeToSourceTime` instead of adding a
second clock that assumed 1x, which was wrong for the tape-effect
fallback.
2. It does not subtract the pipeline latency, since that delay sits
after the sources.

## Tests

`tests/js/slowed-scheduling.test.mjs` plays a two minute synthetic track
against a fake clock and checks, every 100 ms, that the scheduled
sources still reach past "now". It covers SoundTouch at 1x, SoundTouch
at 0.75x and tape effect at 0.75x.

- Without the fix, SoundTouch at 0.75x runs dry at 50.1 s.
- With it, all 9 checks pass.

The e2e fixture is 6 seconds long, so no Playwright test can reach this.

## Doubled kicks at slowed speeds (#728)

Found while testing the fix above. Below 1x the stretcher's grains
overlap, and an attack in the overlap is played twice, which turns a
kick into a ~50 ms flam.

- The shared tempo stage marks attacks and, when the next grain would
replay one, carries straight on instead. The splice joins identical
samples, so it is seamless.
- The time that costs is borrowed, at most 100 ms, and paid back over
the next few grains. The whole mix and the click share this stage, so
they stay together. Only the playhead can trail the audio, by up to 100
ms, briefly.
- The pitch stages do not do this. They are aligned sample for sample
with the unpitched drums.

Measured at 0.75x on 30 s of a full mix: 93 attacks in, **104 out
before, 93 after**.

`tests/js/tempo-attacks.test.mjs` stretches 40 synthetic kicks over a
held chord. The old stretcher gives 47 attacks, the new one exactly 40,
and no kick drifts past the borrow limit.

## Signalsmith Stretch as the tempo stage (#729)

Below 1x, WSOLA repeats overlapping 82 ms fragments, which is heard as
an echo on everything sustained, not only drums. The shared tempo stage
now uses the Signalsmith Stretch WASM core (MIT), at 40 ms blocks. That
size was chosen by ear against 30, 60 and the library's 120 ms default.
The 120 ms default softened about a quarter of the drum attacks.

- **Vendored as `static/vendor/signalsmith-stretch.js`**: the loader
from `signalsmith-stretch@1.3.2`, unchanged, with its MIT license. The
package's own AudioWorkletNode is not used, because it ignores the rate
on live input.
- **Falls back to WSOLA** if the core does not load. The swap happens
only at a flush, never under audio already playing.
- **Latency is reported by the processor**, and both engines compute it
from one shared module, `static/js/tempoStage.js`. Measured end to end,
it matches within 1 ms at 44.1 and 48 kHz.
- **CSP:** `script-src` gains `'wasm-unsafe-eval'`, in both the server
and Tauri CSPs. It allows WebAssembly compilation only; JS `eval` stays
blocked. `test_csp.py` now pins `script-src` to exactly those two
sources.
- **CPU:** about 2.5% of one core at 44.1 kHz, offline.

On 20 s of a full mix at 0.75x (55 attacks in): WSOLA 60 attacks out
with 8 close pairs, Signalsmith 61 with 2.

Tests:
- `tests/js/signalsmith-stage.test.mjs`: the swap, the fallback, latency
within 1 ms, and 40 kicks out as 40 with no drift.
- `tests/e2e/tempo-stage.spec.mjs`: in real Chromium, slowed playback
reports the Signalsmith stage. With `'wasm-unsafe-eval'` removed it
fails and reports `wsola`, which also shows the fallback working in a
browser.

Found along the way, not changed here: the WSOLA path's latency as the
engines count it (122 ms) is well under what I measure end to end (about
300 ms at 0.75x). Only the fallback still uses that path.

## Gate

- `node --check` on all of `static/js`, every `tests/js` test, i18n
coverage clean, `uv.lock` untouched.
- Playwright: 411 passed on the final run. `retention-setting` and
`lyrics-align` each failed once in earlier runs and passed on rerun with
and without these changes.
- `pitch-shift.test.mjs` still passes, so the pitch stages are
unaffected.
- pytest: the 8 known machine-local ffmpeg failures only. ruff and
bandit clean.
- Rust: `cargo fmt --check`, `clippy -D warnings` and 116 tests pass
(`tauri.conf.json` changed).
* fix(backend): start on CPU when torch is installed but cannot load

available_torch_devices() caught only ImportError. A torch whose DLLs
fail to load raises OSError instead (WinError 127 from a CUDA DLL left
behind on Windows, #723), which killed startup when the device setting
was "auto" and made every GET /api/settings a 500 otherwise.

The probe now catches any failure, falls back to ["cpu"], logs the
cause once with its traceback, and remembers it. /api/settings carries
it as torch_error, and Settings replaces the device description with a
warning in red: separation cannot run on any device until torch is
repaired, CPU included, so a quiet fallback would hide the real state.

tests/test_torch_unloadable.py installs an import hook that makes
`import torch` raise that OSError; the probe and settings tests fail on
the old code.

Closes #730

* fix(library): favourite a track from its library row

The Now Playing heart was the only way to favourite, and daw.css hides
that whole card below 1460 px. A maximised laptop window is under that,
so there was no way to favourite at all, and no way to un-favourite
from the Favorites view.

- Library and Favorites rows get a heart next to the delete button.
- Every heart goes through toggleFavorite(), which saves, repaints the
  Now Playing heart when it is the open track, and re-renders the list,
  so the hearts cannot disagree with each other or the store.
- Row actions float over the row's right edge on hover or keyboard
  focus. In the flow the hidden delete button took about 29 px from
  every title at rest; a second button would have taken it again.
  Delete was also invisible under keyboard focus before.
- A favourite shows a small heart in its subline at rest.

tests/e2e/favorite-reach.spec.mjs runs at 1280 and 1600 px; all four
fail on the old code.

Closes #724

* fix(desktop): offer older NVIDIA cards a torch build with their kernels

A CUDA 13 driver was offered cu128 alone. cu128 installs torch 2.8,
whose builds start at sm_61 on Windows and sm_70 on Linux
(TORCH_CUDA_ARCH_LIST in pytorch's release/2.8 .ci scripts), so a
GTX 970 (sm_52) installed it, failed verification with "no kernel image
is available", and dropped to CPU.

- A card below compute capability 7 is never offered cu128; it gets
  cu124 then cu118, the torch 2.6 builds, which carry sm_50 on both
  platforms (release/2.6 scripts), and an sm_50 binary runs on 5.x.
- A CUDA 13 driver on a newer card keeps cu128 first, with cu124 and
  cu118 behind it instead of CPU.

New tests: a_pre_volta_card_is_never_offered_cu128 (the old code
returned ["cu128"] for 5.2 under 13.0) and
a_cuda_13_driver_falls_back_to_the_2_6_builds. The published-wheel
check now covers caps 5.2 and 6.1.

Closes #732

* fix(ui): show a track's key once, as one label that fits

The analysis strip said the key three times: "D maj" on the Key card,
"Major" under it, and "Major" again on a Scale card of its own, while
the Key card was too narrow for a longer label.

- keyLabel.js formatKey() turns the stored key and scale into one
  label, "B minor" or "B harmonic minor", from a translated template
  and mode names in all ten languages (plus the European Portuguese
  spelling). Note names stay letters everywhere, as DAWs write them.
- The Scale card, the Major/Minor sub-line and an unused confidence
  label are gone; the confidence ring stays.
- Key takes the Scale card's grid column, so the row keeps the seven
  edges the presence row lines up with. Measured at a 950 px window it
  still has 115 px for a label that can need 155, so the key wraps to a
  second line rather than clip; a short key stays on one line.
- Display only: stored key, scale and key_confidence are unchanged, so
  nothing needs re-analysis.

tests/js/key-label.test.mjs checks every language; the new
tests/e2e/key-card.spec.mjs checks en, de and fr at 1366 and 950 px
for the label and that it is not clipped, and fails on the old code.

Closes #736

* fix(desktop): stop retrying CUDA setup on a GPU no build supports

A CPU result born from a failure is re-probed on every launch (#247),
which is right for a network drop. For a GPU that every offered CUDA
build refuses with "no kernel image is available", it meant downloading
and installing CUDA torch, failing verification and restoring CPU torch
on every launch: six cycles in the #723 reporter's setup.log.

- verify_cuda_torch() now says why it failed; the no-kernel case is
  matched on CUDA's exact wording, so driver, memory or install problems
  are never mistaken for it.
- When every candidate installed and was refused for that reason, setup
  records cuda-unsupported-gpu. Any other mix keeps today's reasons.
- The setup screen treats that reason as settled and says plainly why
  the GPU is not used. A new app version re-runs setup anyway, which is
  when different builds could be on offer.

New tests: a_gpu_every_build_refuses_is_settled_as_unsupported and
only_cudas_own_wording_counts_as_no_kernel_image.

Closes #733

* fix(desktop): put the bundled torch back exactly when CUDA fails

The CUDA install laid the CUDA wheels over the CPU ones, and the restore
laid the CPU wheels back over those, both with --ignore-installed.
Overlaying never removes files only the CUDA wheel has. On Windows torch
loads every DLL in torch/lib, so the leftover c10_cuda.dll of one torch
version failed to link against the restored CPU DLLs, and `import torch`
raised WinError 127 for good while setup reported success (#723).

New module torchswap:
- Before the first real CUDA install the bundled torch is moved aside
  into .stemdeck-torch-backup next to site-packages. What belongs to
  torch is read from pip's RECORD files; on a real install that is only
  torch, functorch, torchgen, torchaudio, torio, torchvision and their
  dist-infos.
- A manifest is written and synced before anything moves, so a crash at
  any point is recoverable; the next setup run restores it first.
- Each candidate installs into clean directories; a failure deletes the
  CUDA build and renames the original back, with no network. Success
  deletes the backup.

After any restore `import torch` is now checked. Without a backup (the
move failed, or the install broke before this fix) the CPU wheels are
reinstalled, and if torch still does not import, everything torch owns
is deleted and they are installed once more into clean directories.
That is what repairs an install already carrying the stray DLL.

Whether a working CUDA build is already present is now asked once,
before anything moves, instead of inside each install.

Eight new tests on real temp directories, including the reporter's case
(no c10_cuda.dll after rollback, c10.dll byte for byte) and a snapshot
interrupted part way.

Closes #731

* fix(analyze): stop reporting minor keys as their relative major

Key detection multiplied each key's profile correlation by how loud its
root was. In a minor song the relative major's root is the minor chord's
third and sits in most of a typical loop's chords, so it is usually
louder: every i-VI-III-VII loop came out as its relative major at up to
100% confidence ("Plug in Baby", B minor, read D major). Harmonic minor
was never reported, and confidence changed with the input's level.

- Keys are ranked by correlation alone, which is level-invariant.
- The keys close to the best (within 0.1) and their relatives are then
  weighed against the song's edges: how well the first and, when the
  whole song was analysed, last 4 s of music fit each key's tonic triad
  and bass note. A song starts and ends on its tonic far more often than
  anywhere else, which is what a whole-song histogram cannot show.
- A minor key is reported as harmonic minor when its raised seventh
  outweighs the flat seventh 1.3 to 1.
- After separation the key is detected again from the stems: harmony
  stems without drums, the bass stem as the bass line, and the whole
  song. Any failure keeps the estimate from the mix.
- Chroma uses 93 ms frames: 4.1 s instead of 8.6 s on a 9.5 min track's
  stems, same score. The stems pass takes about 1 s per minute of audio.

Measured on tests/keyset.py, 33 synthetic songs with known keys (chord
loops over a bass line, 11 progressions in 3 keys), key and scale both
right:
  before   21 of 33 (27 keys right)
  after    30 of 33 from the mix, 30 of 33 from stems
The three misses are vi-IV-I-V, which starts and ends on its minor
chord. A real 9.5 min track (E minor) reads the same before and after.
Already analysed tracks keep their stored key; nothing is re-analysed.

tests/test_key_detection.py scores the set and the reporter's loop, and
fails on the old detector.

Closes #726

* fix(pipeline): count the stems key pass inside post

A separate "key" stage timing broke the timings' order, which
tests/test_identify_regressions.py holds as a contract: new entries may
only follow the existing stages. The pass is short and already inside
post, so it is counted there.

Refs #726

* test(desktop): check the torch rollback on a real torch install

An ignored test, run by hand: copies a real site-packages torch (1.2 GB
here), moves it aside, lays a CUDA build with a stray c10_cuda.dll into
its place, rolls back, and compares every file's SHA-256 with the
original. Passed against the dev venv's torch 2.6.0+cpu, which then
imported and ran a tensor op from the restored copy.

Refs #731

---------

Co-authored-by: Thales <>
* fix(desktop): lift rustls and quick-xml past their advisories

rustls 0.23.40 -> 0.23.45 (RUSTSEC-2026-0285) is reached through reqwest
at runtime: every download and the updater. plist 1.9 -> 1.10.1 moves
quick-xml 0.39.3 -> 0.42.0 (RUSTSEC-2026-0194, -0195), build time only.
Lockfile only: no Cargo.toml or uv.lock change, so the in-app update
still applies. cargo audit now reports no vulnerabilities, only the
existing unmaintained and unsound warnings.

Closes #738

* feat(ui): stack the composer's actions, and keep the card at every width

The top bar's right-hand end follows the layout suggested in #725:
Detect structure sits over the main button, which now reads Extract
stems to match the Extract row it acts on, in all ten languages, and
the Collapse toggles stack in a column beside them. That end of the bar
is about 50px narrower.

The Now Playing card no longer leaves the bar below 1460px, which took
About this song with it and left no way to open it (#735). It gives way
in steps instead, by container query: the artwork and meta line first,
then the title, down to the heart and the About button. The grid gives
the card a 72px floor, so the Extract chips fold into their overflow
button before the card can be squeezed out.

Tests: about-reach.spec.mjs opens About this song at 1024, 1280, 1366
and 1600px, which fails on the old layout below 1460, and checks the
stacking. favorite-reach.spec.mjs no longer expects the card's heart to
be hidden.

Closes #725
Closes #735

* feat(library): keep favourites on the server, so the phone can use them

Favourites lived only in the studio's catalog store, which the phone
cannot reach: it had no heart to press, and its Favorites chip set a
filter nothing read, so it listed every track.

Each job now carries favorite (null until a client says, then true or
false), set by PUT /api/jobs/{id}/favorite with a strict bool. The
studio sends every heart there, adopts the server's value on each sync
and when its window is looked at again, and hands up a heart set before
this change while the server still says null. The phone has a heart on
every row and the chip filters for real.

Tests: test_jobs_favorite.py covers the endpoint (shape, strict body,
404, crafted ids, the record round trip); favorite-sync.spec.mjs covers
both directions and the phone. The e2e helpers stub the writes so no
spec leaves a favourite on the shared backend.

Closes #734

* refactor(lyrics): let the server alone decide which lyrics are a track's

Lyrics were matched to a song in two places: the import's lookup on the
server, and the Lyrics tab's own LRCLIB search in the browser, each
with its own copy of which version is this song by this artist and
which length fits. They were kept equal by hand, and when they drifted
a track could get lyrics at import and none in the tab.

The tab now asks POST /api/jobs/{id}/lyrics/lookup, which runs the
import's lookup, keeps what it finds the way the import does, and
answers as GET .../lyrics then would: 502 when LRCLIB is out of reach,
404 with nothing_known when the track says too little. The band saved
in the artist box goes along, since it lives in the studio's store.
LRCLIB is still only asked when the tab is opened or at import, and a
phone or offline client still shows kept lyrics.

The browser's search, ranking and artist matching are gone from
lyricsLookup.js, and LRCLIB leaves the page's connect-src. sameSong
stays, for lyrics the tab kept before the server held them to the song.

Tests: test_lyrics_api.py covers the endpoint (kept, saved band,
nothing known, another artist, offline, crafted ids); the lyrics e2e
specs answer the lookup with stubLyricsLookup, which keeps lyrics the
way the server does; acceptance E4 answers it as LRCLIB unreachable.

Closes #719

* feat(transport): let the slow speed be set, from 0.50x to 0.99x

Speed was a fixed 0.75x or 1x. Some fills want 0.75x and some only a
nudge, which is what #701 asked for. The slow button now moves a
hundredth at a time when scrolled over, or with the up and down arrow
keys on it, between 0.50x and 0.99x, and plays at it. Its label is the
speed, so the footer keeps its width, 1x stays one press away, and the
slow speed is kept between sessions. The tooltip says how, in all ten
languages.

The 0.75x floor came from WSOLA, whose artefacts made slower speeds
hard to follow (#433). Signalsmith Stretch (#729) holds up further
down, so the floor is 0.5x. A turn of the wheel is applied once it
stops, since each change of rate flushes the tempo stage.

Tests: slow-speed.spec.mjs (wheel, arrow keys, both limits, a reload).

Closes #701

---------

Co-authored-by: Thales <>
urllib3 2.7.0 -> 2.8.0 changes uv.lock, which sends every desktop
install to a full download instead of the one-click update, two days
after 0.19.0 already did. StemDeck's own requests use the standard
library's urllib; urllib3 only serves model and data downloads from
fixed hosts through requests. The entry says when to drop it.

Co-authored-by: Thales <>
0.5x is now one press, next to the slow speed the player sets and 1x,
instead of 25 notches down from 0.75x. The button pressed is the one
lit, so half speed and a slow speed brought down to 0.50x never light
together.

Tests: slow-speed.spec.mjs presses half speed and checks only one
button is lit when the slow speed is also at 0.50x.

Closes #742

Co-authored-by: Thales <>
@thcp
thcp merged commit 939d66e into main Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment