Repository navigation
Performance strategy: TypeScript first, native helpers only when measured #155
Description
Activity
- addedarea:performancePerformance and renderingPerformance and renderingtype:researchResearch or investigation taskResearch or investigation task
on Sep 27, 2026 Research findings (moving v0.6 branch not benchmarked)
Likely hot paths. Per-keystroke input layout/highlighting and suggestion ranking; completion currently launches a zsh helper per query; ANSI parsing creates per-character strings/cells and calls string-width; wrapping/transcript presentation scans/allocates rows; archive snapshot/restore serializes large cell arrays; IPC JSON framing and detached buffering can become costly at large output volumes. Chroma/gradient work is future and has no evidence yet. These are hypotheses from code paths, not measured bottlenecks. No repeatable performance harness was found in the targeted repository inspection.
TypeScript remains a sound default. Node handles UI, protocol framing, session coordination, storage orchestration, lexical highlighting and moderate collections adequately. First measure and improve algorithmic complexity, avoid redundant layout/render work, bound/cap transcript and suggestion data, use streaming/backpressure, cache stable results, and move shell lookups to persistent helpers where valid. For SQLite question in this issue, compare supported built-in node:sqlite with an addon only using actual query/write/packaging requirements; do not assume SQLite needs native project code.
Measurements to add before optimization: reproducible fixtures (short and long commands, 10k/100k history entries, 100k/1m transcript cells, ANSI-heavy output and Unicode); p50/p95 keystroke-to-frame and suggestion latency; parser/wrap throughput; startup and first-prompt time; CPU over 30–60s sustained output; peak/RSS growth and retained transcript size; IPC throughput, event-loop delay and backpressure. Pin Node/hardware, warm/cold runs, memory cap and repeat counts. Do not use moving v0.6 measurements to draw stable conclusions.
Suggested native trigger (decision thresholds, not current performance claims): only consider a native helper when a representative production-scale profile shows a single isolated computation consumes >20% of end-to-end latency or >25% sustained CPU, or causes >100ms p95 interactive stalls after TS algorithm/caching fixes; AND a prototype provides >=2x speedup with meaningful user-visible improvement, stable ABI/fallback, and total install/distribution cost acceptable across supported OS/architectures. Treat >250ms p95 keystroke-to-render, >100ms p95 suggestions, repeated >100ms event-loop stalls, or unbounded memory growth at 100k+ transcript lines as investigation triggers first, not automatic native-code permission.
If proven, keep a coarse boundary around pure CPU work over typed arrays/strings (e.g. ANSI batch tokenization or fuzzy rank over a bounded candidate set); never put session state, PTY protocol, shell bootstrap, or rendering ownership in native code. Rust/N-API has mature ecosystem but ABI/prebuild and cross-platform release costs; Zig/native Node addon have similar distribution/maintenance costs; WASM improves portability but has interop/copy costs and may not outperform optimized JS. No candidate is recommended before measurements.
#193 was preserved and updated by merging current dev into its existing branch, then used as the base of the v0.8 stack. Node 22/26 CI passed on the updated head. Measurements now cover completion filtering, editor/ScreenPlan, chunked IPC decoding, 10k/100k history queries/imports, native navigation and large transcript wrap/serialization.
On macOS arm64 Node 26.8.1: 500 completion filtering p50/p95 0.06/0.13 ms; 100k structured history query 11.38/11.70 ms; 100k import indexing 107.08/121.41 ms total, yielding every 1,024 records; 100k directory rank 6.56/9.26 ms. Cold native capture is asynchronous and above the warm-query budget. Full 100k transcript wrap and JSON serialization remain expensive; no claim of end-to-end startup/socket optimization.
The Node >=22 floor spans built-in SQLite availability/flag changes, and synchronous database operations alone would not solve responsiveness. v0.8 therefore derives an in-memory command index from journals/imports, uses background journal projection and cooperative yielding, and persists only private deletion IDs. No database binding/helper was added. Broader storage/startup/IPC strategy remains open; #155 was not added to milestone #8.
See the updated docs/performance-benchmarks.md on the cumulative #241 branch for measurement methodology and limits.
Implemented on the v0.16 branch (PR #303,
feature/v016-platform-portability-agents). Strategy recorded and evidence in place:docs/performance-benchmarks.md,scripts/benchmarks.ts, and budgets asserted in tests. NMSh stays TypeScript-first. It was part of the integrated build the maintainer physically validated in the earlier v0.16 QA pass; closing with the implementation.
Goal
Record the performance strategy and set up the evidence needed to apply it.
Decision
The current architecture can implement the following in TypeScript: renderer, editor, provider orchestration, Unix sockets, session service/client (#125), persistence, Kitty/OSC protocols, SQLite bindings, command intelligence.
Research / acceptance
node:sqlitevs a native binding) for Rich history search and metadata #143/NMSh Native Suggestions v2 #138.Dependencies / related
Related #138, #143, #125, #20.