Skip to content

Calibration: make the empty-room baseline save path work on real hardware - #1936

Merged
ruvnet merged 4 commits into
codex/calibration-process-identityfrom
fix/calibrated-presence-evidence-v2
Sep 16, 2026
Merged

ruvnet merged 4 commits into
codex/calibration-process-identityfrom
fix/calibrated-presence-evidence-v2

Conversation

@proffesor-for-testing

@proffesor-for-testing proffesor-for-testing commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Calibration: make the empty-room baseline save path work on real hardware

Three defects found while driving the Mac Catalyst app against three to four ESP32-C6
nodes, 2026-09-14/15. Each is measured on hardware, each has a test, and each test was
mutation-checked — reverting the fix fails it.

Base: 2c0abcc8 (PR #1903 head, which carries the room-identity receipt this builds on).
Touches one file: wifi-densepose-sensing-server/src/main.rs.


1. feat(calibration): emit calibrated presence evidence v2 on the sensing stream

The app gates its held-out empty check on calibrated_presence_evidence
(ruview.calibration.calibrated-presence-evidence.v2). No shipped server has ever emitted
it — measured: 156 frames, 0 carrying the key. So the check bailed on iteration one,
bootstrap/promote was never called, and bootstrap_baseline::store — the only path that
persists a calibration — never ran. A 62-minute, 41,138-frame model was lost on restart.

Emit it, bound to the active model receipt. None unless an explicit calibration is fresh
and the field model scores without a heuristic fallback, so a present value is always
calibrated. The scoring history is extracted into scoring_history() and shared with
person_count_at, so published evidence cannot disagree with the server's own count.

Verified live: 604/604 frames carried it after the fix, and correctly 0/429 while the
model was stale.

2. fix(calibration): only the grid-bound radio may write the single-link baseline

Every bound radio was fed into the single-link field model. Averaging several radios'
amplitude offsets into one baseline flattens the eigenstructure.

bound nodes eigenvalues threshold empty room
4 0 18.90 present 1162/1162
1 5 3.65 absent 604/604 (occupied: present 616/616)

person_count_at already called this "deterministic false occupancy from hardware-specific
amplitude offsets" — but only engaged when exactly one node was bound, so it never fired in
the case it described.

Restricts the baseline write, not the admission. Every bound radio still registers as a
contributor, because calibration_source_nodes_missing refuses to finalize until each one
proves it is live. Gating the whole feed path deadlocks the capture — measured: 43,120
frames over 51 minutes stuck in collecting. There is a regression test for that exact
deadlock.

3. fix(calibration): bind the grid the radio keeps emitting, not the sparse minority

accept_grid locks each node onto the densest grid and rejects the ~16% HT 64-bin minority
a C6 emits alongside HE-SU 256-bin. select_calibration_grid ordered by max_gap_s first —
and a sparse metronomic trickle has a smaller gap than a busy stream with one hiccup. So
selection kept binding the grid admission discards.

Measured: bound 64sc on node 11; every node then locked onto 256sc; the bound grid went
stale, frame_count froze at 10,997, and the capture hung in collecting for 858 s with
both minimums met, no error surfaced, and a flickering UI as the client retried. Both
captures that bound 256sc produced a usable model.

Orders by subcarrier count first, then the existing gap/rate/ppdu ordering.

Same defect class as #1919, measured independently on C6.


Verification

316 passed, 1 failed — the single failure is calibration_expiry_tests:: long_running_expiry_is_not_reported_active_for_runtime_or_bootstrap, pre-existing on the
2c0abcc8 base
and untouched here (it reports none where the test expects expired).
It is unrelated to these changes and worth its own fix.

Each new test was mutation-checked. One of them initially passed against the unfixed
comparator — the test data did not discriminate — and was rewritten until reverting the fix
failed it with left: 64, right: 256.

Verified on hardware (2026-09-15, four ESP32-C6 nodes, empty room)

One capture exercising all three fixes together:

source_node_ids            [11, 12, 13, 14]   four nodes bound
source_grid                256sc, ppdu_type 1 grid fix picked it, active for all 889 s
baseline_eigenvalue_count  7                  every previous 4-node capture gave 0
variance_explained         0.1493             previous: 0.0816 / 0.0966
frame_count                15,169
missing_source_node_ids    []                 throughout -- no deadlock

Evidence stream: 1,415 of 1,468 frames carried
ruview.calibration.calibrated-presence-evidence.v2, and correctly 0 of 429 while the
model was stale.

bootstrap/promote now runs to completion and returns honest counts:

sample_count 12   stale 0   empty 7   vital_signs 1   -> needs empty >= 10, vitals 0

Not claimed

  • The baseline is still not stored. bootstrap_baseline::store has never executed. It
    no longer fails because the path is broken -- every gate works and reports truthfully --
    but because the detector does not meet its own quality bar: 7 of 12 samples read empty
    where 10 are required. That is up from 0 of 12 before these fixes.
  • The detector is unstable, and these fixes do not address it. The same model in the
    same empty room, sampled two minutes apart with nobody entering: false positives went
    from 107/1385 (7.7%) to 364/1415 (25.7%).
  • The seven baseline eigenvalues are unused. inference_method was
    field_model_perturbation_energy_v1 in all 1,415 frames, because estimate_occupancy is
    the #[cfg(not(feature = "eigenvalue"))] stub in the shipped build. All occupancy is one
    scalar against one frozen threshold. Tracked separately -- it needs a decision alongside
    the Windows build, not a flag flip.

🤖 Generated with Ruflo & AQE

…g stream

The Mac app gates its held-out empty-room check on
`calibrated_presence_evidence` carrying
`ruview.calibration.calibrated-presence-evidence.v2`. No shipped server has
ever emitted that key, so the check fails on its first iteration with "No
fresh calibrated presence result matched this room's current model receipt",
`bootstrap/promote` is never called, and `bootstrap_baseline::store` -- the
only path that persists a calibration -- never runs. A completed 62-minute
field model is therefore lost on the next restart.

Emit the evidence bound to the active model receipt. `None` unless an
explicit calibration is fresh and the field model scores the observation
without a heuristic fallback, so a present value is always calibrated.

Extract the field-model scoring history into `scoring_history()` and reuse it
for both `person_count_at` and the evidence, so the published evidence cannot
disagree with the count derived from the same model.

Co-Authored-By: Ruflo & AQE
Claude-Session: https://claude.ai/code/session_01Qh7eir6C2U5CSwD7GR43Zv
…fusals

Assert the exact wire schema and every identity field the Mac app matches
against its model receipt, plus that evidence agrees with the server's own
person count. Pair the refusals (no receipt, bootstrap-only authority,
expired model) with the positive case so an absent value can never be read as
proof when the path is simply broken.

Co-Authored-By: Ruflo & AQE
Claude-Session: https://claude.ai/code/session_01Qh7eir6C2U5CSwD7GR43Zv
… baseline

A room binding may name several radios, and every one of them was admitted
into the field model. That model is single-link: one baseline, one set of
amplitude offsets. Averaging two radios' offsets into it flattens the
eigenstructure and leaves only a scalar energy threshold behind.

Measured on three ESP32-C6 nodes:

  4-node binding  baseline_eigenvalue_count 0, residual threshold 18.90,
                  empty room reported an occupant in 1162/1162 frames
  1-node binding  baseline_eigenvalue_count 5, residual threshold 3.65,
                  empty room absent 604/604, occupied present 616/616

`person_count_at` already refused to score a non-bound radio, calling it
"deterministic false occupancy from hardware-specific amplitude offsets" --
but only when exactly one node was bound. With more it fell through to the
shared mixed history, so the guard never fired in the case it described.

Restrict the baseline write, not the admission. Every bound radio is still
admitted and recorded as a contributor: `calibration_source_nodes_missing`
refuses to finalize until each one proves it is live on the frozen grid, so
gating them out of the feed path entirely deadlocks the capture -- measured
on hardware as 43,120 frames over 51 minutes stuck in `collecting` with
missing_source_node_ids=[12,13,14]. Only `maybe_feed_calibration` is
restricted to the grid-bound radio, and scoring follows that same radio
whatever the size of the room's node set.

This is the narrowing `bootstrap_baseline::store` already applies when it
persists `vec![binding.source_node_id]` as the frozen model source.

Co-Authored-By: Ruflo & AQE
Claude-Session: https://claude.ai/code/session_01Qh7eir6C2U5CSwD7GR43Zv
…rse minority

`accept_grid` locks each node onto the densest grid it has seen and rejects
sparser frames from the feature path -- on an ESP32-C6 the ~16% HT 64-bin
minority that arrives alongside HE-SU 256-bin. `select_calibration_grid`
ordered candidates by `max_gap_s` first, so it preferred whichever grid
looked temporally smoothest. A sparse trickle at a metronomic cadence has a
smaller worst-case gap than a busy stream with one scheduling hiccup, so
selection kept choosing the grid admission is designed to discard.

Measured on four ESP32-C6 nodes: a capture bound 64sc on node 11, every node
then locked onto 256sc, the bound grid went stale with no frame for five
minutes, frame_count froze at 10,997 and `calibration_grid_is_fresh` failed.
The capture could never finalize -- `collecting` for 858 s with both minimums
long since met, no error surfaced, and a flickering UI as the client retried.
The same rig bound 256sc on a different capture and finalized with five
baseline eigenvalues.

Order by subcarrier count first so selection agrees with admission, then fall
back to the previous gap/rate/ppdu ordering to choose among equally dense
candidates.

Co-Authored-By: Ruflo & AQE
Claude-Session: https://claude.ai/code/session_01Qh7eir6C2U5CSwD7GR43Zv
ruvnet added a commit that referenced this pull request Sep 16, 2026
calibration_status unconditionally overwrote status with "none" whenever
active was false, which also erased the legitimate "expired" status for
a calibration that ran long ago and has since expired -- losing the
distinction between "never calibrated" and "was calibrated, now expired".

Fixes the pre-existing failure in
calibration_expiry_tests::long_running_expiry_is_not_reported_active_for_runtime_or_bootstrap,
called out as a known, unrelated gap in PR #1936.

Co-Authored-By: claude-flow <ruv@ruv.net>
@ruvnet
ruvnet merged commit 686e4ba into codex/calibration-process-identity Sep 16, 2026
4 checks passed
@oga35767-eng

Copy link
Copy Markdown

👍

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants