Skip to content

Vectorize contact solving and reuse world step storage - #101

Open
edwardgushchin wants to merge 5 commits into
ikpil:mainfrom
edwardgushchin:codex/reuse-physics-step-buffers
Open

edwardgushchin wants to merge 5 commits into
ikpil:mainfrom
edwardgushchin:codex/reuse-physics-step-buffers

Conversation

@edwardgushchin

@edwardgushchin edwardgushchin commented Sep 24, 2026 •

Copy link
Copy Markdown

Active physics steps allocate transient solver/task storage, and wide contact math processes four scalar lanes. This PR collects all portable downstream Box2D.NET runtime patches accumulated since tag 3.1.654, adapts them to current upstream main (154bacd0fddc0e72e484d3061cdc76aa2d842994), and adds regression tests and a reproducible dense workload.

It includes the original per-world buffer reuse in this PR and both changes from #102. #102 remains open; its static revolute/wheel fixes are also included here.

Changes

  • Eight-lane Vector256<float> contact storage/arithmetic, gather/scatter, indices and layout validation. Scalar NaN/signed-zero selection and separate multiply/add remain intact; the existing Vector<float> overloads are unchanged.
  • Skip absent second-point normal/friction work and empty overflow stages; directly execute serial solver stages.
  • Reuse step contexts, graph block arrays, the newer upstream preparation-span arrays and worker-0 context.
  • Preserve distinct spare reference objects during swap removal; reuse prepared contact/island slots and retain sleeping-set/island capacities until world destruction, including dormant sets.
  • Reuse upstream parallel-for task storage, with independent storage for nested calls. Keep upstream's scheduler and stable worker indices. Capture constraint/callback failures, wake peers and join started work before reporting errors; world destruction remains possible after a fatal step.
  • Preserve all four target frameworks. The netstandard2.1 path uses scalar eight-lane storage and C# 12 static method-group caching, with no new package dependencies.

The complete patch inventory and report maps every portable vendored delta to its current upstream implementation. Visibility/nullable/XML-warning embedding adaptations and engine-owned scene/contact/interpolation code require downstream types and are excluded explicitly. Upstream's existing B2WorldDef.capacity covers expected body/contact capacity preparation.

Measurements

Ryzen 7 5700X, Linux x64, .NET 10.0.1, Release with DOTNET_TieredCompilation=0. Same 1,536 always-awake circles, 144 Hz, four substeps, 1,600 warmup and 1,024 measured steps. Three fresh-process trials per configuration, alternated within each trial; medians of trial means and p99 values:

Library Workers Mean ms p99 ms Managed bytes, all threads, per sampled run
Upstream main 154bacd 1 6.905 7.612 2,090,040
This PR 1 0.997 1.153 0
This PR, built-in scheduler 4 0.465 0.594 0

All nine trials produce the identical complete body-state SHA-256:
C2554173F6A5D78AD49163D4BBF0909906B31A92518A795A2FEC21A63E22CEDD.
Raw JSON trials are committed beside the benchmark. The host desktop load was uncontrolled. These are kernel measurements, not rendered FPS or a native Box2D comparison.

DOTNET_TieredCompilation=0 dotnet run --project tools/Box2D.NET.Benchmark -c Release -f net10.0 -- --dense 1
DOTNET_TieredCompilation=0 dotnet run --project tools/Box2D.NET.Benchmark -c Release -f net10.0 -- --dense 4

Verification and boundaries

  • Release build succeeds for netstandard2.1, net8.0, net9.0, net10.0.
  • Full NUnit suite: 208/208 in Release on .NET 8 and .NET 10, and 208/208 in Debug on .NET 10.
  • AVX-disabled regression/determinism suite: 12/12; the AVX-disabled dense run preserves the complete state hash and zero all-thread allocation.
  • Directly loaded scalar netstandard2.1 assembly on .NET 10: 12/12 of the same regression/determinism tests, including warmed static revolute/wheel allocation checks.
  • Existing falling-hinge golden values remain unchanged across worker counts: sleep step 328 and hash 0x26E08AEE.
  • .NET 9 runtime execution and other CPUs/operating systems remain unverified locally. Upstream's workflow currently limits PR base branches to pr/**, so a PR targeting main does not automatically receive its checks.

Low-level API/layout change: B2FloatW becomes 32 bytes and wide constraint indices become eight elements. On .NET 8+, X/Y/Z/W are ref-returning properties over vector storage, while scalar fallback retains fields. Direct reads/writes and the four-argument constructor remain usable, but field reflection/ABI and wide-structure layout require consumers of these public low-level types to rebuild. The existing two-argument b2DestroySolverSet entry point is retained. Buffers retain peak capacity for a world's lifetime; new topology/nested jobs can allocate. Resuming a partly solved world after a fatal error is not a recovery contract.

@edwardgushchin edwardgushchin changed the title Reuse per-world solver buffers during active steps Vectorize contact solving and reuse world step storage Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant