Pad Gaussian tree-cull nodes per-subtree, not per-asset - #1250
Closed
untoldengine wants to merge 2 commits into
Closed
untoldengine wants to merge 2 commits into
untoldengine wants to merge 2 commits into
Conversation
…il, tree-scale stat Investigating a fast-rotation stutter on large Gaussian splat scenes needed visibility the engine didn't have. Adds os_signpost coverage for the full visionOS frame loop (XRRenderFrame/XRQueryNextFrame/XRQueryDrawable/ XRFramePacingWait/XRCommandBufferWait) and the previously-uninstrumented Gaussian preprocess stage (GaussianDepth) and per-chunk pager tick (GaussianPagerTick), plus two new [Gaussian] log fields: chunkCull= passed/failed counts from the per-chunk frustum/HZB test, and the asset's largest chunk log-scale (maxLogScaleMax) to flag outlier splats. This instrumentation is what surfaced two real findings: the tree pre-filter providing zero pruning on an asset with a large background splat, and the compositor's queryDrawable() call dominating frame time independent of the Gaussian pipeline's own (small) cost.
…at scale GaussianChunkTreeCull padded every node in a chunked .untoldgs asset's cluster tree by the whole asset's single largest splat scale, since a node's own subtree maximum wasn't stored. One outlier splat (a background "sky" splat is common in unbounded outdoor captures) padded every node by its scale, regardless of which subtree actually held it. On a real garden capture this padded every node by exp(5.745) * 3.5 ≈ 1093m, large enough that the tree pre-filter pruned nothing at all, confirmed at 814/814 chunks reaching the per-chunk cull every frame across 20 consecutive frames of a stationary view. Bumps .untoldgs to version 4 (48 -> 52 byte tree nodes) to add UntoldGSTreeNode.maxLogScaleMax, computed once at cook time in UntoldGSWriter.buildTree from the same per-node scan that already finds each node's AABB. GaussianChunkTreeCull now pads each node by its own value instead of one shared figure. Older .untoldgs files are rejected with .unsupportedVersion rather than misread against the new layout, per the format's existing "a version bump means re-bake" contract. Confirmed on-device: tree pruning went from 0/814 to 94/814 chunks pruned on the same view that showed no pruning before.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
GaussianChunkTreeCull's per-node padding: every node in a chunked.untoldgsasset's cluster tree was padded by the whole asset's singlelargest splat scale (no per-node subtree maximum was stored), so one
outlier splat — a background "sky" splat, common in unbounded outdoor
captures — inflated every node's box enough to defeat pruning entirely.
Confirmed on a real garden capture: padding came out to ~1093m, and the
tree pre-filter pruned 0 of 814 chunks across 20 consecutive frames of a
stationary view. Bumps
.untoldgsto version 4 (48 → 52 byte tree nodes)to add
UntoldGSTreeNode.maxLogScaleMax, computed once per node at cooktime; older files are cleanly rejected with
.unsupportedVersionratherthan misread, per the format's existing "a version bump means re-bake"
contract. Confirmed on-device: pruning went from 0/814 to 94/814 on the
same view.
os_signpostcoverage forthe full visionOS frame loop (
XRRenderFrame/XRQueryNextFrame/XRQueryDrawable/XRFramePacingWait/XRCommandBufferWait) and thepreviously-unwired Gaussian preprocess stage (
GaussianDepth) andper-chunk pager tick (
GaussianPagerTick), plus two new[Gaussian]logfields:
chunkCull=passed/failedand the asset's largest chunk log-scale(
maxLogScaleMax=).Context
Investigating a fast-rotation stutter on large Gaussian splat scenes
(#1248). This tree-cull bug is real and fixed here, but on-device tracing
after the fix (with the new signposts) showed it wasn't the dominant cost:
frame.queryDrawable()accounts for 71-77% of frame time in bothbefore/after traces, independent of Gaussian pipeline cost, which stayed
flat at ~4-5% throughout. That's a compositor-level pacing issue, tracked
separately in #1249 — not resolved by this PR.
Test plan
swift build/xcodebuild -destination generic/platform=visionOS— both cleanswiftformat --lintclean against project version (0.60.1)UntoldGSFormatTests,UntoldGSCookerTests,UntoldGSCookerEquivalenceTests,GaussianChunkTreeCullTests,GaussianChunkCullMathTests— all passing(golden CRC32/SHA256 hashes and version assertions updated for the
52-byte node / version-4 format change)
UntoldEngineRenderTestsGaussian suite (208 tests) passing end to endstationary view, before/after