Skip to content

Fix the bugs reported against 0.19.0 - #737

Merged
thcp merged 9 commits into
next-releasefrom
fix/0.19-bugs
Oct 1, 2026
Merged

thcp merged 9 commits into
next-releasefrom
fix/0.19-bugs

Conversation

@thcp

@thcp thcp commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #723
Closes #724
Closes #726
Closes #730
Closes #731
Closes #732
Closes #733
Closes #736

The bugs reported against 0.19.0 by other people, one commit per issue. Found along the way and tracked, but not in this PR: #734 (favourites on the phone UI, needs server storage) and #735 (About this song unreachable below 1460 px).

CUDA setup on Windows (#723, split into #730 to #733)

The reporter's GTX 970 went through a chain of four failures:

  • Older NVIDIA GPUs get a CUDA torch build with no kernels for them #732, wrong build. A CUDA 13 driver was offered cu128 only, which installs torch 2.8. Its Windows builds start at sm_61 and its Linux builds at sm_70 (checked against pytorch's release/2.8 .ci scripts), so a Maxwell card could never pass. Cards below compute capability 7 now get cu124 then cu118 (torch 2.6, which carries sm_50 on both platforms), and newer cards under a CUDA 13 driver fall back to those instead of CPU.
  • Falling back from CUDA to CPU torch leaves torch unimportable on Windows #731, broken rollback (the one that corrupted the install). The CPU restore laid CPU wheels over CUDA ones and never removed CUDA-only files, so a leftover c10_cuda.dll made import torch fail for good. The bundled torch is now moved aside before the first CUDA install (what belongs to torch is read from pip's RECORD files, with a manifest synced to disk first so a crash is recoverable) and renamed back on failure: exact, and no network. After any restore, import torch is checked, and an install already carrying the stray DLL is cleared and reinstalled into clean directories.
  • Backend fails to start when torch cannot be loaded #730, backend crash. The device probe caught only ImportError, so that OSError killed startup and made /api/settings a 500. It now falls back to CPU, logs the cause once, and Settings says in red that separation cannot run until torch is repaired, in all ten languages.
  • A GPU that can never pass CUDA setup is retried on every launch #733, retry loop. A GPU that every offered build refuses with "no kernel image is available" (matched on CUDA's exact wording) is now settled until the next app version, instead of reinstalling CUDA torch on every launch.

Favourites below 1460 px (#724)

The Now Playing heart was the only way to favourite, and that card is hidden below 1460 px. Library and Favorites rows now have a heart next to delete, all hearts go through one toggleFavorite(), and row actions float over the row edge on hover or keyboard focus, which gives titles back the ~29 px the hidden delete button always took. The reporter's other request, removing the music icon, was declined earlier in #636 and is not changed.

Key detection (#726) and the key label (#736)

  • Detection. Keys were ranked by correlation times root loudness, which pulls minor songs to their relative major ("Plug in Baby", B minor, read D major). Now correlation alone picks the close candidates, and the song's first and last seconds (tonic triad and bass note) decide between them. Harmonic minor is reported when the raised seventh dominates. After separation the key is re-detected from the stems, bass stem included.
  • Measured on tests/keyset.py, 33 synthetic songs with known keys, key and scale both right: 21 of 33 before, 30 of 33 after, on the mix and on stems. The three misses are vi-IV-I-V, which starts and ends on its minor chord. The stems pass costs about 1 s per minute of audio.
  • Label. "D maj", "Major" and a separate Scale card saying "Major" again are now one label, "D major" or "B harmonic minor", from a translated template. Note letters stay letters in every language. The Key card takes the Scale card's column and wraps rather than clips at narrow windows. Display only: stored values are unchanged, and existing tracks are not re-analysed.

Tests

  • pytest: test_torch_unloadable.py, test_key_detection.py (with the labelled set)
  • Rust: 8 rollback tests on real temp directories, plus wheel choice and no-kernel-image tests. An ignored test, run by hand, rolled back a real 1.2 GB torch install and compared every file's SHA-256: identical, and it imported and ran afterwards.
  • Playwright: favorite-reach.spec.mjs (1280 and 1600 px), key-card.spec.mjs (en, de, fr at 1366 and 950 px)
  • node: key-label.test.mjs, every language
  • Each new test was run against the old code and fails there.

Gate

  • node tests, JS syntax, i18n coverage clean; uv.lock unchanged
  • Rust: fmt, clippy -D warnings, 128 tests
  • pytest: only the 8 known machine-local ffmpeg failures
  • Playwright: 421 passed
  • Not run: the desktop CUDA path end to end on a real install. It needs a packaged NVIDIA build; see the test notes on this PR.

Verification beyond the tests

  • Real pipeline, key ([Bug]: Scale detection and reporting #726). Three rendered songs uploaded to an isolated dev server and separated by Demucs on an RTX 3080. B minor i-VI-III-VII (the reported case), E harmonic minor and G major all came out right, from the mix and again after the stems pass.
  • Real torch rollback (Falling back from CUDA to CPU torch leaves torch unimportable on Windows #731). A copy of a real 1.2 GB torch 2.6.0+cpu was moved aside, overlaid with a CUDA build carrying a stray c10_cuda.dll, and rolled back. Every file's SHA-256 matched the original, and the restored torch imported and ran.
  • Real first-run GPU setup (Falling back from CUDA to CPU torch leaves torch unimportable on Windows #731, Older NVIDIA GPUs get a CUDA torch build with no kernels for them #732). A Windows NVIDIA package built from this branch, on a fresh data folder: it moved the bundled torch aside (9 entries), installed cu124 into clean directories, verified on the RTX 3080, chose CUDA, and deleted the backup. Only the cu124 dist-infos remained.
  • Acceptance suite on that package, twice, with STEMDECK_ACCEPTANCE_SKIP_LONG=1:
    • Run 1: 31 passed, 2 failed. S2 fails because the local build is unversioned (v0.0.0), so the updater correctly offers v0.19.0. L6 failed after the Demucs worker crashed natively on CUDA ("stack smashing detected") while Whisper was also on the GPU; the job fell back to CPU, where line-by-line re-timing is skipped by design.
    • Run 2: 32 passed, 1 failed (S2 again, same reason). L6 passed and no job fell back to CPU, so the crash did not reproduce. There is no earlier record of it; worth watching.
  • Screenshots of the row hearts at 1280 px and the key card at 1366 and 950 px (en, de) looked right. The confidence ring's "100%" has always been wider than its 20 px circle (since feat(ui-refactor): DAW redesign, library overhaul, sections, analysis, and transport #67); unchanged here.

Not caused by this branch

  • trivy fails on two urllib3 CVEs published today (installed 2.7.0, fixed in 2.8.0). This branch does not touch uv.lock; fixing it means a lock change, which sends desktop users a full download, so it is a separate decision.

Thales added 9 commits September 30, 2026 12:25
available_torch_devices() caught only ImportError. A torch whose DLLs
fail to load raises OSError instead (WinError 127 from a CUDA DLL left
behind on Windows, #723), which killed startup when the device setting
was "auto" and made every GET /api/settings a 500 otherwise.

The probe now catches any failure, falls back to ["cpu"], logs the
cause once with its traceback, and remembers it. /api/settings carries
it as torch_error, and Settings replaces the device description with a
warning in red: separation cannot run on any device until torch is
repaired, CPU included, so a quiet fallback would hide the real state.

tests/test_torch_unloadable.py installs an import hook that makes
`import torch` raise that OSError; the probe and settings tests fail on
the old code.

Closes #730
The Now Playing heart was the only way to favourite, and daw.css hides
that whole card below 1460 px. A maximised laptop window is under that,
so there was no way to favourite at all, and no way to un-favourite
from the Favorites view.

- Library and Favorites rows get a heart next to the delete button.
- Every heart goes through toggleFavorite(), which saves, repaints the
  Now Playing heart when it is the open track, and re-renders the list,
  so the hearts cannot disagree with each other or the store.
- Row actions float over the row's right edge on hover or keyboard
  focus. In the flow the hidden delete button took about 29 px from
  every title at rest; a second button would have taken it again.
  Delete was also invisible under keyboard focus before.
- A favourite shows a small heart in its subline at rest.

tests/e2e/favorite-reach.spec.mjs runs at 1280 and 1600 px; all four
fail on the old code.

Closes #724
A CUDA 13 driver was offered cu128 alone. cu128 installs torch 2.8,
whose builds start at sm_61 on Windows and sm_70 on Linux
(TORCH_CUDA_ARCH_LIST in pytorch's release/2.8 .ci scripts), so a
GTX 970 (sm_52) installed it, failed verification with "no kernel image
is available", and dropped to CPU.

- A card below compute capability 7 is never offered cu128; it gets
  cu124 then cu118, the torch 2.6 builds, which carry sm_50 on both
  platforms (release/2.6 scripts), and an sm_50 binary runs on 5.x.
- A CUDA 13 driver on a newer card keeps cu128 first, with cu124 and
  cu118 behind it instead of CPU.

New tests: a_pre_volta_card_is_never_offered_cu128 (the old code
returned ["cu128"] for 5.2 under 13.0) and
a_cuda_13_driver_falls_back_to_the_2_6_builds. The published-wheel
check now covers caps 5.2 and 6.1.

Closes #732
The analysis strip said the key three times: "D maj" on the Key card,
"Major" under it, and "Major" again on a Scale card of its own, while
the Key card was too narrow for a longer label.

- keyLabel.js formatKey() turns the stored key and scale into one
  label, "B minor" or "B harmonic minor", from a translated template
  and mode names in all ten languages (plus the European Portuguese
  spelling). Note names stay letters everywhere, as DAWs write them.
- The Scale card, the Major/Minor sub-line and an unused confidence
  label are gone; the confidence ring stays.
- Key takes the Scale card's grid column, so the row keeps the seven
  edges the presence row lines up with. Measured at a 950 px window it
  still has 115 px for a label that can need 155, so the key wraps to a
  second line rather than clip; a short key stays on one line.
- Display only: stored key, scale and key_confidence are unchanged, so
  nothing needs re-analysis.

tests/js/key-label.test.mjs checks every language; the new
tests/e2e/key-card.spec.mjs checks en, de and fr at 1366 and 950 px
for the label and that it is not clipped, and fails on the old code.

Closes #736
A CPU result born from a failure is re-probed on every launch (#247),
which is right for a network drop. For a GPU that every offered CUDA
build refuses with "no kernel image is available", it meant downloading
and installing CUDA torch, failing verification and restoring CPU torch
on every launch: six cycles in the #723 reporter's setup.log.

- verify_cuda_torch() now says why it failed; the no-kernel case is
  matched on CUDA's exact wording, so driver, memory or install problems
  are never mistaken for it.
- When every candidate installed and was refused for that reason, setup
  records cuda-unsupported-gpu. Any other mix keeps today's reasons.
- The setup screen treats that reason as settled and says plainly why
  the GPU is not used. A new app version re-runs setup anyway, which is
  when different builds could be on offer.

New tests: a_gpu_every_build_refuses_is_settled_as_unsupported and
only_cudas_own_wording_counts_as_no_kernel_image.

Closes #733
The CUDA install laid the CUDA wheels over the CPU ones, and the restore
laid the CPU wheels back over those, both with --ignore-installed.
Overlaying never removes files only the CUDA wheel has. On Windows torch
loads every DLL in torch/lib, so the leftover c10_cuda.dll of one torch
version failed to link against the restored CPU DLLs, and `import torch`
raised WinError 127 for good while setup reported success (#723).

New module torchswap:
- Before the first real CUDA install the bundled torch is moved aside
  into .stemdeck-torch-backup next to site-packages. What belongs to
  torch is read from pip's RECORD files; on a real install that is only
  torch, functorch, torchgen, torchaudio, torio, torchvision and their
  dist-infos.
- A manifest is written and synced before anything moves, so a crash at
  any point is recoverable; the next setup run restores it first.
- Each candidate installs into clean directories; a failure deletes the
  CUDA build and renames the original back, with no network. Success
  deletes the backup.

After any restore `import torch` is now checked. Without a backup (the
move failed, or the install broke before this fix) the CPU wheels are
reinstalled, and if torch still does not import, everything torch owns
is deleted and they are installed once more into clean directories.
That is what repairs an install already carrying the stray DLL.

Whether a working CUDA build is already present is now asked once,
before anything moves, instead of inside each install.

Eight new tests on real temp directories, including the reporter's case
(no c10_cuda.dll after rollback, c10.dll byte for byte) and a snapshot
interrupted part way.

Closes #731
Key detection multiplied each key's profile correlation by how loud its
root was. In a minor song the relative major's root is the minor chord's
third and sits in most of a typical loop's chords, so it is usually
louder: every i-VI-III-VII loop came out as its relative major at up to
100% confidence ("Plug in Baby", B minor, read D major). Harmonic minor
was never reported, and confidence changed with the input's level.

- Keys are ranked by correlation alone, which is level-invariant.
- The keys close to the best (within 0.1) and their relatives are then
  weighed against the song's edges: how well the first and, when the
  whole song was analysed, last 4 s of music fit each key's tonic triad
  and bass note. A song starts and ends on its tonic far more often than
  anywhere else, which is what a whole-song histogram cannot show.
- A minor key is reported as harmonic minor when its raised seventh
  outweighs the flat seventh 1.3 to 1.
- After separation the key is detected again from the stems: harmony
  stems without drums, the bass stem as the bass line, and the whole
  song. Any failure keeps the estimate from the mix.
- Chroma uses 93 ms frames: 4.1 s instead of 8.6 s on a 9.5 min track's
  stems, same score. The stems pass takes about 1 s per minute of audio.

Measured on tests/keyset.py, 33 synthetic songs with known keys (chord
loops over a bass line, 11 progressions in 3 keys), key and scale both
right:
  before   21 of 33 (27 keys right)
  after    30 of 33 from the mix, 30 of 33 from stems
The three misses are vi-IV-I-V, which starts and ends on its minor
chord. A real 9.5 min track (E minor) reads the same before and after.
Already analysed tracks keep their stored key; nothing is re-analysed.

tests/test_key_detection.py scores the set and the reporter's loop, and
fails on the old detector.

Closes #726
A separate "key" stage timing broke the timings' order, which
tests/test_identify_regressions.py holds as a contract: new entries may
only follow the existing stages. The pass is short and already inside
post, so it is counted there.

Refs #726
An ignored test, run by hand: copies a real site-packages torch (1.2 GB
here), moves it aside, lays a CUDA build with a stray c10_cuda.dll into
its place, rolls back, and compares every file's SHA-256 with the
original. Passed against the dev venv's torch 2.6.0+cpu, which then
imported and ran a tensor op from the restored copy.

Refs #731
@thcp
thcp merged commit d78fa34 into next-release Oct 1, 2026
12 of 13 checks passed
@thcp thcp mentioned this pull request Oct 1, 2026
thcp added a commit that referenced this pull request Oct 2, 2026
Release 0.19.1: everything on `next-release` since v0.19.0.

- #727: slowed playback without stutter, doubled kicks or echo
(Signalsmith Stretch as the tempo stage).
- #737: the bugs reported against 0.19.0 (CUDA setup, torch that cannot
load, favourites in a narrow window, minor keys, the key label).
- #740: the remaining open issues (top bar and Extract stems, the card
at every width, favourites on the phone, the lyrics lookup on the
server, an adjustable slow speed, Rust advisories).
- #741: trivy ignores the two urllib3 CVEs until the lock next moves.
- #743: a 0.5x button beside the slow speed.

Verified on local Windows NVIDIA builds: 0.19.1.dev0 passed 12 of 12
(#727, #737), and 0.19.1.dev1 passed 18 of 18 automated checks including
#740, in an isolated profile. `uv.lock` is unchanged since v0.19.0, so
the in-app update works.

Closes #722
Closes #728
Closes #729
Closes #723
Closes #724
Closes #726
Closes #730
Closes #731
Closes #732
Closes #733
Closes #736
Closes #701
Closes #738
Closes #725
Closes #735
Closes #734
Closes #719
Closes #742
@thcp
thcp deleted the fix/0.19-bugs branch October 5, 2026 08:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant