Conversation
…g stream The Mac app gates its held-out empty-room check on `calibrated_presence_evidence` carrying `ruview.calibration.calibrated-presence-evidence.v2`. No shipped server has ever emitted that key, so the check fails on its first iteration with "No fresh calibrated presence result matched this room's current model receipt", `bootstrap/promote` is never called, and `bootstrap_baseline::store` -- the only path that persists a calibration -- never runs. A completed 62-minute field model is therefore lost on the next restart. Emit the evidence bound to the active model receipt. `None` unless an explicit calibration is fresh and the field model scores the observation without a heuristic fallback, so a present value is always calibrated. Extract the field-model scoring history into `scoring_history()` and reuse it for both `person_count_at` and the evidence, so the published evidence cannot disagree with the count derived from the same model. Co-Authored-By: Ruflo & AQE Claude-Session: https://claude.ai/code/session_01Qh7eir6C2U5CSwD7GR43Zv
…fusals Assert the exact wire schema and every identity field the Mac app matches against its model receipt, plus that evidence agrees with the server's own person count. Pair the refusals (no receipt, bootstrap-only authority, expired model) with the positive case so an absent value can never be read as proof when the path is simply broken. Co-Authored-By: Ruflo & AQE Claude-Session: https://claude.ai/code/session_01Qh7eir6C2U5CSwD7GR43Zv
… baseline
A room binding may name several radios, and every one of them was admitted
into the field model. That model is single-link: one baseline, one set of
amplitude offsets. Averaging two radios' offsets into it flattens the
eigenstructure and leaves only a scalar energy threshold behind.
Measured on three ESP32-C6 nodes:
4-node binding baseline_eigenvalue_count 0, residual threshold 18.90,
empty room reported an occupant in 1162/1162 frames
1-node binding baseline_eigenvalue_count 5, residual threshold 3.65,
empty room absent 604/604, occupied present 616/616
`person_count_at` already refused to score a non-bound radio, calling it
"deterministic false occupancy from hardware-specific amplitude offsets" --
but only when exactly one node was bound. With more it fell through to the
shared mixed history, so the guard never fired in the case it described.
Restrict the baseline write, not the admission. Every bound radio is still
admitted and recorded as a contributor: `calibration_source_nodes_missing`
refuses to finalize until each one proves it is live on the frozen grid, so
gating them out of the feed path entirely deadlocks the capture -- measured
on hardware as 43,120 frames over 51 minutes stuck in `collecting` with
missing_source_node_ids=[12,13,14]. Only `maybe_feed_calibration` is
restricted to the grid-bound radio, and scoring follows that same radio
whatever the size of the room's node set.
This is the narrowing `bootstrap_baseline::store` already applies when it
persists `vec![binding.source_node_id]` as the frozen model source.
Co-Authored-By: Ruflo & AQE
Claude-Session: https://claude.ai/code/session_01Qh7eir6C2U5CSwD7GR43Zv
…rse minority `accept_grid` locks each node onto the densest grid it has seen and rejects sparser frames from the feature path -- on an ESP32-C6 the ~16% HT 64-bin minority that arrives alongside HE-SU 256-bin. `select_calibration_grid` ordered candidates by `max_gap_s` first, so it preferred whichever grid looked temporally smoothest. A sparse trickle at a metronomic cadence has a smaller worst-case gap than a busy stream with one scheduling hiccup, so selection kept choosing the grid admission is designed to discard. Measured on four ESP32-C6 nodes: a capture bound 64sc on node 11, every node then locked onto 256sc, the bound grid went stale with no frame for five minutes, frame_count froze at 10,997 and `calibration_grid_is_fresh` failed. The capture could never finalize -- `collecting` for 858 s with both minimums long since met, no error surfaced, and a flickering UI as the client retried. The same rig bound 256sc on a different capture and finalized with five baseline eigenvalues. Order by subcarrier count first so selection agrees with admission, then fall back to the previous gap/rate/ppdu ordering to choose among equally dense candidates. Co-Authored-By: Ruflo & AQE Claude-Session: https://claude.ai/code/session_01Qh7eir6C2U5CSwD7GR43Zv
This was referenced Sep 15, 2026
Closed
ruvnet
added a commit
that referenced
this pull request
Sep 16, 2026
calibration_status unconditionally overwrote status with "none" whenever active was false, which also erased the legitimate "expired" status for a calibration that ran long ago and has since expired -- losing the distinction between "never calibrated" and "was calibrated, now expired". Fixes the pre-existing failure in calibration_expiry_tests::long_running_expiry_is_not_reported_active_for_runtime_or_bootstrap, called out as a known, unrelated gap in PR #1936. Co-Authored-By: claude-flow <ruv@ruv.net>
oga35767-eng
approved these changes
Sep 26, 2026
|
👍 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Calibration: make the empty-room baseline save path work on real hardware
Three defects found while driving the Mac Catalyst app against three to four ESP32-C6
nodes, 2026-09-14/15. Each is measured on hardware, each has a test, and each test was
mutation-checked — reverting the fix fails it.
Base:
2c0abcc8(PR #1903 head, which carries the room-identity receipt this builds on).Touches one file:
wifi-densepose-sensing-server/src/main.rs.1.
feat(calibration): emit calibrated presence evidence v2 on the sensing streamThe app gates its held-out empty check on
calibrated_presence_evidence(
ruview.calibration.calibrated-presence-evidence.v2). No shipped server has ever emittedit — measured: 156 frames, 0 carrying the key. So the check bailed on iteration one,
bootstrap/promotewas never called, andbootstrap_baseline::store— the only path thatpersists a calibration — never ran. A 62-minute, 41,138-frame model was lost on restart.
Emit it, bound to the active model receipt.
Noneunless an explicit calibration is freshand the field model scores without a heuristic fallback, so a present value is always
calibrated. The scoring history is extracted into
scoring_history()and shared withperson_count_at, so published evidence cannot disagree with the server's own count.Verified live: 604/604 frames carried it after the fix, and correctly 0/429 while the
model was stale.
2.
fix(calibration): only the grid-bound radio may write the single-link baselineEvery bound radio was fed into the single-link field model. Averaging several radios'
amplitude offsets into one baseline flattens the eigenstructure.
person_count_atalready called this "deterministic false occupancy from hardware-specificamplitude offsets" — but only engaged when exactly one node was bound, so it never fired in
the case it described.
Restricts the baseline write, not the admission. Every bound radio still registers as a
contributor, because
calibration_source_nodes_missingrefuses to finalize until each oneproves it is live. Gating the whole feed path deadlocks the capture — measured: 43,120
frames over 51 minutes stuck in
collecting. There is a regression test for that exactdeadlock.
3.
fix(calibration): bind the grid the radio keeps emitting, not the sparse minorityaccept_gridlocks each node onto the densest grid and rejects the ~16% HT 64-bin minoritya C6 emits alongside HE-SU 256-bin.
select_calibration_gridordered bymax_gap_sfirst —and a sparse metronomic trickle has a smaller gap than a busy stream with one hiccup. So
selection kept binding the grid admission discards.
Measured: bound 64sc on node 11; every node then locked onto 256sc; the bound grid went
stale,
frame_countfroze at 10,997, and the capture hung incollectingfor 858 s withboth minimums met, no error surfaced, and a flickering UI as the client retried. Both
captures that bound 256sc produced a usable model.
Orders by subcarrier count first, then the existing gap/rate/ppdu ordering.
Same defect class as #1919, measured independently on C6.
Verification
316 passed, 1 failed— the single failure iscalibration_expiry_tests:: long_running_expiry_is_not_reported_active_for_runtime_or_bootstrap, pre-existing on the2c0abcc8base and untouched here (it reportsnonewhere the test expectsexpired).It is unrelated to these changes and worth its own fix.
Each new test was mutation-checked. One of them initially passed against the unfixed
comparator — the test data did not discriminate — and was rewritten until reverting the fix
failed it with
left: 64, right: 256.Verified on hardware (2026-09-15, four ESP32-C6 nodes, empty room)
One capture exercising all three fixes together:
Evidence stream: 1,415 of 1,468 frames carried
ruview.calibration.calibrated-presence-evidence.v2, and correctly 0 of 429 while themodel was stale.
bootstrap/promotenow runs to completion and returns honest counts:Not claimed
bootstrap_baseline::storehas never executed. Itno longer fails because the path is broken -- every gate works and reports truthfully --
but because the detector does not meet its own quality bar: 7 of 12 samples read empty
where 10 are required. That is up from 0 of 12 before these fixes.
same empty room, sampled two minutes apart with nobody entering: false positives went
from 107/1385 (7.7%) to 364/1415 (25.7%).
inference_methodwasfield_model_perturbation_energy_v1in all 1,415 frames, becauseestimate_occupancyisthe
#[cfg(not(feature = "eigenvalue"))]stub in the shipped build. All occupancy is onescalar against one frozen threshold. Tracked separately -- it needs a decision alongside
the Windows build, not a flag flip.
🤖 Generated with Ruflo & AQE