Skip to content

Pad Gaussian tree-cull nodes per-subtree, not per-asset - #1250

Closed
untoldengine wants to merge 2 commits into
developfrom
bugfix/headset_rotation
Closed

untoldengine wants to merge 2 commits into
developfrom
bugfix/headset_rotation

Conversation

@untoldengine

Copy link
Copy Markdown
Owner

Summary

  • Fixes GaussianChunkTreeCull's per-node padding: every node in a chunked
    .untoldgs asset's cluster tree was padded by the whole asset's single
    largest splat scale (no per-node subtree maximum was stored), so one
    outlier splat — a background "sky" splat, common in unbounded outdoor
    captures — inflated every node's box enough to defeat pruning entirely.
    Confirmed on a real garden capture: padding came out to ~1093m, and the
    tree pre-filter pruned 0 of 814 chunks across 20 consecutive frames of a
    stationary view. Bumps .untoldgs to version 4 (48 → 52 byte tree nodes)
    to add UntoldGSTreeNode.maxLogScaleMax, computed once per node at cook
    time; older files are cleanly rejected with .unsupportedVersion rather
    than misread, per the format's existing "a version bump means re-bake"
    contract. Confirmed on-device: pruning went from 0/814 to 94/814 on the
    same view.
  • Adds diagnostics that made the above findable: os_signpost coverage for
    the full visionOS frame loop (XRRenderFrame/XRQueryNextFrame/
    XRQueryDrawable/XRFramePacingWait/XRCommandBufferWait) and the
    previously-unwired Gaussian preprocess stage (GaussianDepth) and
    per-chunk pager tick (GaussianPagerTick), plus two new [Gaussian] log
    fields: chunkCull=passed/failed and the asset's largest chunk log-scale
    (maxLogScaleMax=).

Context

Investigating a fast-rotation stutter on large Gaussian splat scenes
(#1248). This tree-cull bug is real and fixed here, but on-device tracing
after the fix (with the new signposts) showed it wasn't the dominant cost:
frame.queryDrawable() accounts for 71-77% of frame time in both
before/after traces, independent of Gaussian pipeline cost, which stayed
flat at ~4-5% throughout. That's a compositor-level pacing issue, tracked
separately in #1249 — not resolved by this PR.

Test plan

  • swift build / xcodebuild -destination generic/platform=visionOS — both clean
  • swiftformat --lint clean against project version (0.60.1)
  • UntoldGSFormatTests, UntoldGSCookerTests, UntoldGSCookerEquivalenceTests,
    GaussianChunkTreeCullTests, GaussianChunkCullMathTests — all passing
    (golden CRC32/SHA256 hashes and version assertions updated for the
    52-byte node / version-4 format change)
  • Full UntoldEngineRenderTests Gaussian suite (208 tests) passing end to end
  • Verified on-device: tree pruning 0/814 → 94/814 chunks on the same
    stationary view, before/after

…il, tree-scale stat

Investigating a fast-rotation stutter on large Gaussian splat scenes needed
visibility the engine didn't have. Adds os_signpost coverage for the full
visionOS frame loop (XRRenderFrame/XRQueryNextFrame/XRQueryDrawable/
XRFramePacingWait/XRCommandBufferWait) and the previously-uninstrumented
Gaussian preprocess stage (GaussianDepth) and per-chunk pager tick
(GaussianPagerTick), plus two new [Gaussian] log fields: chunkCull=
passed/failed counts from the per-chunk frustum/HZB test, and the asset's
largest chunk log-scale (maxLogScaleMax) to flag outlier splats.

This instrumentation is what surfaced two real findings: the tree pre-filter
providing zero pruning on an asset with a large background splat, and the
compositor's queryDrawable() call dominating frame time independent of the
Gaussian pipeline's own (small) cost.
…at scale

GaussianChunkTreeCull padded every node in a chunked .untoldgs asset's
cluster tree by the whole asset's single largest splat scale, since a
node's own subtree maximum wasn't stored. One outlier splat (a background
"sky" splat is common in unbounded outdoor captures) padded every node by
its scale, regardless of which subtree actually held it.

On a real garden capture this padded every node by exp(5.745) * 3.5 ≈
1093m, large enough that the tree pre-filter pruned nothing at all,
confirmed at 814/814 chunks reaching the per-chunk cull every frame across
20 consecutive frames of a stationary view.

Bumps .untoldgs to version 4 (48 -> 52 byte tree nodes) to add
UntoldGSTreeNode.maxLogScaleMax, computed once at cook time in
UntoldGSWriter.buildTree from the same per-node scan that already finds
each node's AABB. GaussianChunkTreeCull now pads each node by its own
value instead of one shared figure. Older .untoldgs files are rejected
with .unsupportedVersion rather than misread against the new layout, per
the format's existing "a version bump means re-bake" contract.

Confirmed on-device: tree pruning went from 0/814 to 94/814 chunks pruned
on the same view that showed no pruning before.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant