Skip to content

Boot-crash triage tooling: map raw crash EIP back to Java method (3 proposals for later evaluation) #652

Description

@LSantha

Problem statement

When a JNode boot image crashes (page fault / panic in int_die_halt), the only attribution available is a raw native register dump, e.g.:

int  : 0000000E  Error: 00000004  CR2  : 00000000  CR3  : 00001000
EIP  : 0018E74B  CS   : 0000001B  FLAGS: 00013206  CR0  : 80000011
EAX  : 00000000  EBX  : 00000004  ECX  : 00000001  EDX  : 02ED2130
STACK: 00000000 00000000 00000006 00000000 00000000 00000518
CODE(0018E73B ): 00 FF 75 EC FF 75 F0 FF 97 8C 05 00 00 8B 45 EC
CODE(0018E74B ): 8B 00 3B 45 F0 0F 86 E6 FF FF FF 8B 45 EC 52 8B ...

EIP is a native address inside the boot image. Nothing maps it back to the executing Java method / bytecode, so every boot-crash triage today is manual archaeology:

  1. Guess candidate methods from the boot phase visible in the log (e.g. crash right after MemoryBlockManager init → early-init path).
  2. On the host, regenerate per-method codegen with the boot-image-identical pipeline and eyeball-match the faulting instruction bytes (above: 8B 00 = mov eax,[eax] with EAX=0, i.e. a null read) and surrounding shape.
  3. Iterate until something matches — 30+ minutes per crash, and short byte patterns occur in many methods, so attribution stays uncertain.

This issue asks for tooling that answers "which Java method + bytecode faulted?" from a crash dump. The solution will be designed against master (whatever shape the tree has at implementation time); three candidate approaches are sketched below and a later evaluation should pick one (or a combination).

Proposal A — Capture KDB's own Java-frame resolution at the crash

Idea. KDB already resolves Java frames itself: its T command prints the current thread plus full Java stack (core/src/core/org/jnode/vm/scheduler/KernelDebugger.java, serial input via core/src/native/x86/unsafe.asm:readKdbInput, port config in core/src/native/x86/jnode.asm). A triage helper would send T (and t/W for context) at/after the crash and record the resolved frames — no host-side guessing at all.

Details / prerequisites.

  • KDB input only works when UART1 is a pipe server (--uartmode1 server <socket>); in the common file logging mode only the automatic panic dump is captured. Single-client pipe: needs a persistent drainer/multiplexer while attached, otherwise an undrained pipe stalls the boot (every UART1 byte blocks until read).
  • KDB services input during reschedule(); a CPU-bound or fully wedged guest may starve it, and it is unverified whether the panic path still pumps KDB input — needs a spike: kdb on (default via kdb boot flag), reproduce a fault, send T.
  • Output parsing is straightforward (<currentthread .../> + frame lines).

Pros: zero build changes; uses the kernel's own (authoritative) maps; works for any compiler backend.
Cons: needs UART replumbing + drainer discipline; useless if the fault path can't service input; post-mortem only if someone/something sends T in time (could be automated by a host-side watchdog on the UART stream).
Acceptance: from a reproduced boot crash, one command yields the faulting Java method + bytecode index, cross-checked against the EIP bytes.

Proposal B — Host-side crash-dump analyzer (no guest/build changes)

Idea. A host script (e.g. local/kdb_eip.sh <eip> [code-bytes]) that, given the EIP and the CODE(...) bytes from the panic dump, regenerates boot-image-identical codegen for candidate methods (existing BinDump --resolve flow in local/l2boot-tools/) and ranks matches by byte-pattern + surrounding-shape similarity.

Details.

  • Pure host-side; no guest reconfiguration, no image changes; works on old logs.
  • Search space must be bounded (whole-image regeneration per crash is too slow): seed candidates from boot-phase markers in the log, then widen on miss.
  • Short patterns (8B 00) are ambiguous, so ranking needs context windows (prefix/suffix bytes, register setup like the preceding 8B 45 EC) and should report confidence, not a single answer.

Pros: works retroactively on any saved KDB log; no risk to boot/runtime behavior.
Cons: heuristic — can misattribute; slow on wide searches; still requires the analyst to pick candidates when it is unsure.
Acceptance: on the recorded EIP 0018E74B crash, the tool lists the true method in its top-3 with the matching offset, without manual candidate selection.

Proposal C — Emit a native-address→method map at image-build time

Idea. Teach the image build (compiler backends and/or BootImageBuilder in builder/) to write a symbol-map artifact next to the ISO (e.g. jnode-x86-lite.map: native range → class/method/signature/bytecode-range per compiled method). Crash triage then becomes a trivial range lookup, exact and instant.

Details.

  • Must be opt-in / side-artifact only: zero bytes of difference in the shipped image when disabled (or a documented, negligible header addition).
  • Map must cover all backends that can appear in an image (stub, L1A, L2) and survive any post-link relocation applied by the builder (record final linked addresses, not pre-link offsets).
  • Format should be line-oriented text (grep/awk-friendly), one range per line, with a version header; large images ⇒ tens of thousands of methods, keep it compact.
  • Consumers: the host-side analyzer from Proposal B degrades to an exact lookup when a map is present.

Pros: exact attribution, O(log n), permanent value for every future crash; also useful for profilers/debuggers (JDWP work).
Cons: touches the builder/compiler — highest implementation risk and review burden; map faithfulness must be tested (stale-map misattribution is worse than none: embed image checksum + build id in the header and verify on load).
Acceptance: build with map enabled → addresses from a forced crash resolve to method+bytecode via lookup; build without the flag is bit-identical to today; a stale map is detected and rejected by checksum.

Evaluation guidance (for later)

Criterion A (KDB capture) B (host analyzer) C (build map)
Exactness exact (kernel maps) heuristic exact
Works on old logs no yes no (needs new build)
Guest/build changes UART plumbing + drainer none builder patch
Failure modes input starvation at panic misattribution stale-map risk
Effort small spike + helper small script medium, needs tests

Suggested path: spike A at the next reproduced crash (cheap, may solve it outright); implement B as a stopgap helper regardless; do C only if boot-crash triage becomes a recurring bottleneck. A+C compose well (map as cross-check for KDB frames).

Context

Raised during L2-backend differential testing (mauve corpus green; remaining work is boot-image crashes with raw-EIP-only dumps). Environment: JNode 0.2.9-dev x86, VirtualBox 7.2.6, UART1 file-mode KDB log + UART2 serial shell, host Ubuntu + JDK 1.6. No code changes proposed in this issue — tooling/design decision only.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/vmJVM internals: JIT compilers, GC, VM magic, type system.kind/featureNew feature or enhancement request.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions