Skip to content
A2R-LabPublic

About

A GPU accelerated library for computing rigid body dynamics with analytical gradients

Resources

Code of conduct

Contributing

Security policy

Stars

17 stars

Watchers

0 watching

Forks

Repository files navigation

GRiD

CI docs license python agent-ready All Contributors

A GPU-accelerated library for robot dynamics, kinematics, and collisions, with analytical derivatives and Hessians for supported numerical operations.

The GRiD package ecosystem: a user's URDF goes through URDFParser to the code generator (built on GLASS) and RBDReference, producing CUDA C++ with NumPy, JAX, and PyTorch wrappers; benchmarks and tests, backed by pytest-gpu-proof and external oracles, produce validated outputs and performance benchmarks.

GRiD turns a URDF into optimized, per-robot CUDA C++ for rigid-body dynamics, kinematics, their analytical first- and second-order derivatives and a trajectory-optimization plant layer, then hands you that code three ways — a numpy handle, a jax.jit-able FFI surface, or torch.autograd-aware ops — from one content-addressed .so cache. One CUDA block per problem, batched, bit-deterministic and thread-count invariant; the same model and API rebuild for each target architecture (one artifact per sm_XX), with runtime memory adaptation from embedded Jetson class devices to desktop GPUs. Website: https://a2r-lab.org/GRiD/.

GRiD builds on our URDFParser, RBDReference, and GLASS packages (URDF parsing, Pinocchio-validated reference dynamics, and GPU linear algebra), together with its own bundled code generator. Using its scripts, users can easily generate and test optimized rigid body dynamics CUDA C++ code for their URDF files.

Status: 0.5.0 is the first packaged release and is alpha software. APIs may still change between minor versions; see CHANGELOG.md for what has landed since the paper and what is known to be missing. The original ICRA 2022 paper describes the implementation preserved in the archival robot-acceleration/GRiD repository, not the feature set or performance of this release. See the project website for the overview. Collision routines use the generated CUDA interface; numerical Python interface coverage is documented separately.

I want to…

Task Start here
Call GRiD from Python (numpy/JAX/torch) grid_rbd.load_robot("robot.urdf", backend=...) — Python wrappers docs · agent guide
Generate CUDA for a new robot grid-generate config/robot_assets/iiwa14.urdf — see Quick Start below
Fit a humanoid build in RAM fast robot setup (algorithm_list=, enable_mujoco_kernels=False)
Add an algorithm adding an algorithm
Run tests / fix a red receipt CI job CUDA validation + test/run_gpu_proof.sh --help
Benchmark benchmarks
Debug a CUDA-vs-numpy mismatch docs/agent_debugging_guide.md — the bug-class bible
Get MuJoCo/mjx-convention I/O handle.mujoco.<method>(...) — values AND derivatives/second-order
Everything else How do I…? on the docs site

Start-here track: examples/README.md routes the four usage tracks — the examples/notebooks/ Python-bindings tour (01-quickstart → 07-inline-cuda), the runnable bindings/examples/ scripts, codegen scripts, and hand-written-CUDA walkthroughs.

A checkout contains submodules: run git submodule update --init --recursive after cloning without --recursive. A PyPI install bundles them.

Quick Start

Install from PyPI (Linux x86-64, Python 3.10–3.12; no GPU or compiler needed to install — robot CUDA is compiled by register_robot using the nvcc on your PATH):

pip install grid-rbd            # code generator + numpy backend + grid-generate / grid-spherize CLIs
pip install "grid-rbd[jax]"     # + JAX FFI surface   (install a CUDA jax wheel yourself)
pip install "grid-rbd[torch]"   # + torch backend     (install a CUDA torch wheel yourself)
pip install "grid-rbd[all]"     # jax + torch

Or work from a checkout (creates a local venv, editable install, registers the CLIs):

git clone --recursive https://github.com/A2R-Lab/GRiD.git && cd GRiD
bash install/base_install.sh
source .venv/bin/activate

Generate CUDA code for your robot (ten ready-to-use URDFs ship in config/robot_assets/ — iiwa14, go2, fr3, g1, h1_2, …):

# Via the installed CLI (works on a clean base install):
grid-generate config/robot_assets/iiwa14.urdf                # arm, fixed base
grid-generate config/robot_assets/go2.urdf -f                # quadruped, floating base
grid-generate path/to/robot.urdf [-t EE_JOINT_NAME] [-n NAMESPACE] [-f] [--algorithm-list LIST] [-o OUT.cuh]

# Or via a hardcoded zero-config example (these two pull their URDFs from the
# robot_descriptions package — a DEV dependency; install install/requirements-dev.txt first):
python examples/codegen/generate_iiwa14.py       # iiwa14 fixed base
python examples/codegen/generate_go2_floating.py # Go2 floating base

Validate and debug:

# Print CPU reference values for all algorithms:
python examples/codegen/print_reference_values.py path/to/robot.urdf

# Compile and run the CUDA print kernel (requires nvcc):
python examples/codegen/print_grid.py path/to/robot.urdf

Write your own CUDA kernel against the generated header:

# Step-by-step walkthrough + compiling/validated example kernels:
#   examples/cuda/README.md   (and examples/cuda/wrapper_types.md)
bash examples/cuda/build_and_validate.sh   # generate → nvcc → run → validate

Requires a C++17-capable host compiler (e.g., g++ ≥ 7, clang++ ≥ 5). The benchmark and codegen runtime compile with -std=c++17 — needed for inline variables in the bench common header. With the [torch] extra the per-robot .so follows torch's ATen requirement (-std=c++20 from torch 2.14; needs CUDA 12+ and g++ ≥ 10).

Usage

  • grid-generate PATH_TO_URDF — generate grid.cuh; add -d for full debug mode, -f for floating base, -t JOINT_NAME to target a specific end-effector joint
  • python examples/codegen/print_reference_values.py PATH_TO_URDF — print CPU reference values for all algorithms to validate CUDA output
  • python examples/codegen/print_grid.py PATH_TO_URDF — compile and run the CUDA print kernel against the generated header

Floating-Base Conventions

Floating-base parsing and the Python reference path now accept a public floating-base convention flag. The default is Pinocchio-compatible:

  • floating_base_convention="pinocchio": q = [x, y, z, qx, qy, qz, qw], v = [vx, vy, vz, wx, wy, wz]
  • floating_base_convention="legacy": q = [x, y, z, qw, qx, qy, qz], v = [wx, wy, wz, vx, vy, vz]

GRiD normalizes both public conventions into one shared internal floating-base representation, so the code generator and RBDReference stay consistent under the hood while callers can choose the input/output ordering they need.

Developer Testing

Contributor-facing test workflows (floating-convention regression suite, CUDA equivalence env overrides, shared-memory targets) moved to CONTRIBUTING.md; the receipt/verification policy lives in the CUDA validation guide.

Current Support

GRiD supports open-chain and tree models with revolute, prismatic, fixed, floating, mimic, and additional joint types. Support is operation-specific; see the joint and algorithm restrictions and backend inventory.

GRiD implements the full modern rigid-body-dynamics stack: RNEA / CRBA / ABA / Minv / forward dynamics; analytical first-order gradients (ID + FD, incl. external-force gradients); the second-order derivatives (IDSVA-SO both frames with a codegen-time dispatcher, FDSVA-SO); the kinematics family (EE pose / Jacobian / Hessian, general-frame frame_jacobian/J̇/OSC inertia, runtime multi-EE targets); integrators + integrator gradients; the centroidal family (CoM, CCRBA, dccrba, CMM time-variation, Coriolis matrix, energy/ID regressors); the inertial-parameter (π) family (the joint-torque regressor Y with its analytic gradient ∂Y/∂(q,v) and the FD parameter gradient ∂q̈/∂π); contact-frame wrench mapping (contact_fext / register_robot(contact_frames=...)); runtime tool/payload welding (attach_tool/tool_fext); runtime multi-target positions; the collision family (two-tier config_free); a trajectory-optimization grid_plant cost/step layer; and runtime-mutable inertia/transform/joint-dynamics tables. The complete per-algorithm catalog with citations and per-feature detail lives in the CUDA support status page.

RBDReference additionally provides numpy reference oracles — validated against Pinocchio — for generalized gravity, nonlinear effects, kinetic/potential/mechanical energy, the Coriolis matrix, the centroidal quantities (CoM, CoM Jacobian, CCRBA, centroidal momentum) and their derivatives (the analytic dccrba ∂A/∂q tensor — replacing the prior finite-difference oracle — and cmm_time_variation Ȧ), the inverse-dynamics and kinetic/potential-energy regressors, the general-frame Jacobian / J̇ / OSC inertia described above, and the plant/cost/barrier layer above.

Dual-surface equivalence. Every algorithm exists on two surfaces that are tested for numerical agreement: the RBDReference numpy implementation (the oracle, checked against Pinocchio) and the generated CUDA C++ kernels (checked against that same numpy reference). This keeps the GPU codegen honest against an independent, Pinocchio-validated baseline.

Mimic-joint support: per-robot gating is now essentially eliminated. Non-gradient algorithms (RNEA, forward dynamics, ABA, CRBA, …) work for robots with mimic joints, and every gradient emits a correct mimic-reduced result on both the fixed and floating base: inverse_dynamics_gradient/forward_dynamics_gradient, end_effector_pose_gradient/end_effector_pose_hessian, the second-order idsva_so/fdsva_so, the external-force gradients (f_ext_gradient), and the integrator gradients. The centroidal family — com, ccrba, energy, and the centroidal derivatives dccrba/cmm_time_variation — now also runs on mimic robots (the per-body Jacobian and per-unit motion columns carry the mimic multiplier α, validated against the mimic-aware reference). dccrba/cmm_time_variation additionally run on big floating-base robots (e.g. g1/h1_2-floating) via the sweep-pool spill path. No algorithm raises NotImplementedError for mimic robots anymore.

Additional algorithms and features are in development. If you have a particular algorithm or feature in mind please let us know by posting a GitHub issue. We'd also love your collaboration in implementing the Python reference implementation of any algorithm you'd like implemented!

Repo map

Directory Owns Entry doc
grid_codegen/ the code-generation engine: emits grid.cuh AND the checked-in generated binding regions, all driven by the abi_specs.py table codegen architecture
bindings/ the grid-rbd Python package (numpy/jax/torch handles over a cached per-robot .so) bindings/README.md · agent guide
external/ the peer-product submodules: GLASS (GPU linear algebra), RBDReference (Pinocchio-validated numpy oracle), URDFParser each submodule's README
examples/ the start-here track: notebooks/ (Python tour), codegen/, cuda/ examples/README.md
test/ pytest suites + the split-suite/receipt machinery (run_split_suite.py, run_gpu_proof.sh, compile_sched.py) CUDA validation
config/ ten sample URDFs (robot_assets/) + tuned per-GPU launch configs (launch_configs/) + autotune_robot.sh config/robot_assets/URDF_SOURCES.md
docs/ the Sphinx site (source/) + agent_debugging_guide.md (the bug-class bible) docs site
install/ install scripts (base_install.sh, developer_install.sh) + requirements files installation guide

C++ API

The generated external interface has three layers: *_device (algorithm-specific buffer and scratch contract, with placement owned by the device function), *_kernel (global entry point with batched timestep loop), and the host wrapper (CPU launcher with H↔D copies). Internal *_inner helpers support composition without repeating shared setup. See the codegen architecture docs for the rationale and concrete signatures.

Python API (grid-rbd)

For Python users the grid-rbd package (in bindings/) wraps the per-robot codegen behind a register-then-run UX with numpy, jax, and torch backends. It ships in the same grid-rbd distribution as the code generator, so pip install grid-rbd (or pip install -e . from a checkout, what install/base_install.sh runs) installs the codegen toolkit and the grid_rbd wrapper together. The base install is minimal; pick a backend extra for the surface you want:

pip install grid-rbd            # base: numpy backend only
pip install "grid-rbd[jax]"     # + JAX FFI surface
pip install "grid-rbd[torch]"   # + torch backend (CUDA wheel matching your GPU arch)
pip install "grid-rbd[all]"     # jax + torch

See the install matrix in bindings/README.md for what each extra unlocks (and the torch CUDA-wheel note).

import grid_rbd

# numpy (default), jax, or torch; urdf_string= also accepted instead of urdf_path
handle = grid_rbd.register_robot("iiwa14", urdf_path="iiwa.urdf", backend="torch")

qdd = handle.forward_dynamics(q, qd, u)   # autograd-aware torch.Tensor
qdd.sum().backward()                      # gradients flow to q, qd, u

The torch backend exposes autograd-aware inverse_dynamics / forward_dynamics / aba / integrator (analytic backward passes) plus CUDA-Graphs capture, and the handle also surfaces the grid_plant cost/barrier methods. inverse_dynamics (alias rnea) takes an optional qdd= (the autograd gradient is qdd-aware, returning the correct ∂τ/∂(q,q̇) including the ∂(M·q̈)/∂q term), and all three backends expose the value ops coriolis_matrix, kinetic_energy_regressor, potential_energy_regressor, dccrba, and cmm_time_variation (forward-only on jax/torch). The π-regressor family (inverse_dynamics_regressor, the differentiable inverse_dynamics_wrt_params/forward_dynamics_wrt_params, and the forward_dynamics_parameter_gradient ∂q̈/∂π), runtime tool welding (attach_tool/tool_fext, via enable_tool=True), and multi-contact wrench mapping (contact_fext, via register_robot(contact_frames=[...])) are bound as well. For true fp64 compute build with register_robot(..., dtype="float64") (its own cache entry); allow_fp64=True is only the numpy handle's fp64-in/fp64-out convenience cast on an fp32 build (ignored when dtype="float64"). See bindings/README.md and the Python wrappers docs.

Citing GRiD

To cite GRiD in your research, please use the following bibtex for our paper "GRiD: GPU-Accelerated Rigid Body Dynamics with Analytical Gradients":

@inproceedings{plancher2022grid,
  title={GRiD: GPU-Accelerated Rigid Body Dynamics with Analytical Gradients}, 
  author={Brian Plancher and Sabrina M. Neuman and Radhika Ghosal and Scott Kuindersma and Vijay Janapa Reddi},
  booktitle={IEEE International Conference on Robotics and Automation (ICRA)}, 
  year={2022}, 
  month={May}
}

Performance

Release measurements from the 27 September 2026 run on one NVIDIA RTX 5090 with an Intel Core Ultra 9 285K cover RNEA, its analytical gradient (∇RNEA), and its analytical Hessian (∇²RNEA) on iiwa14 (fixed base, 7 velocities), go2 (floating base, 18), and G1 (floating base, 35) at batch sizes 16–1024. The release measurements page gives the method, every timing boundary, and the caveats; the benchmark harness reproduces the collection.

Core-operation speedups against seven baseline modes, with timing boundaries and fp64 exceptions labeled.

Ratios are baseline time divided by GRiD time; above 1× favors GRiD. Each column names its timing boundary: GRiD host calls including copies against the CPU libraries, and GRiD compute-only calls against the GPU libraries' resident calls. * marks cells where the evaluated baseline path required fp64 and ~ a side whose run means span more than 1.5×. Colors are clipped at 100×.

Clustered GRiD, Pinocchio and MuJoCo timing bars on three robots; Hessians compare GRiD with Pinocchio's standard API only.

Microseconds per complete batch on a log axis. GRiD's bar splits into its CUDA compute-only call, the GPU–CPU I/O increment, and the JAX wrapper increment. These are differences of measured call times, not isolated measurements of each component.

Call wall times for RNEA, its gradient and its Hessian on three robots: a no-I/O group (CUDA Device, PyTorch resident, JAX resident) beside a host-call group (C++ Host, NumPy, PyTorch, JAX); Python host bars are solid to the allocate-once call and hatched up to the default call.

Call wall times through each API boundary: native CUDA, the C++ host call, NumPy, PyTorch, and JAX. Pick the boundary your application uses. For the Python surfaces the solid bar is the call with its buffers allocated once and reused (measured 2 October 2026), and the hatched cap reaches the default call, which allocates its output every time. With reused buffers, NumPy and PyTorch land within a few percent of the C++ host call on large outputs.

Installation

The Quick Start above covers the common-case install. For CUDA Toolkit setup, developer dependencies (Pinocchio, robot_descriptions, benchmarks), and Docker, see the full installation guide.

Troubleshooting

Bench harness nvcc hangs on floating-base kernels (sm_8x)

On Ampere (sm_86 / CUDA 12.6) the bench harness can wedge nvcc / ptxas at 100 % CPU when compiling heavy floating-base GRiD harnesses. Pass --ptxas-opt-level 2 to test/benchmarks/run_multi_version.py — it forwards -Xptxas -O2 to floating-base compiles only. Blackwell (sm_120) does not hit this. Typical user code that includes grid.cuh and calls the batch host wrappers (e.g. grid::forward_dynamics<T>(...)) does not trigger the hang — it's specific to the timing-bench template surface.

Contributing

Contributions welcome — see CONTRIBUTING.md for the workflow (and CLAUDE.md for the repo conventions AI agents and humans both follow).

Contributors

Brian Plancher
Brian Plancher

Zachary Pestrikov
Zachary Pestrikov

Kwamena A
Kwamena A

Ann Li
Ann Li

Cael Yasutake
Cael Yasutake

pruyontrarakk
pruyontrarakk

About

A GPU accelerated library for computing rigid body dynamics with analytical gradients

Resources

Code of conduct

Contributing

Security policy

Stars

17 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages