Skip to content

[FFL-3347] Benchmark mobile flag evaluation, hydration, and native tracking - #1453

Draft
sameerank wants to merge 6 commits into
developfrom
sameerank/mobile-flags-benchmarks
Draft

sameerank wants to merge 6 commits into
developfrom
sameerank/mobile-flags-benchmarks

Conversation

@sameerank

@sameerank sameerank commented Sep 28, 2026 •

Copy link
Copy Markdown

What does this PR do?

Adds benchmark-only flags experiments for FFL-3347, supporting the mobile offline initialization and dynamic-context RFC.

Companion native prototypes: iOS #3239 and Android #3931.

Decode Evaluate Reads
JS JS Local JS
Native JS Local JS after a ProtoJSON handoff
JS Native Sync and async native calls after a ProtoJSON handoff
Native Native Sync and async native calls

Uses published @datadog/flagging-core@3.1.1, the SwiftProtobuf 1.38.1 prototype, and the Kotlin/protobuf-java 3.25.5 prototype. Synthetic configurations contain 10, 100, and 1,000 flags. Four hot boolean flags exercise static, membership, compound numeric, and MD5-split evaluation. Values, metadata, context A -> B -> A, and configuration replacement are correctness gates.

Both platforms include:

  • Saved configuration: an asynchronous native file read through the first correct result in JS. Compare bytes -> JS decode/evaluate, native decode -> ProtoJSON -> JS evaluate, and native decode/evaluate -> result. This uses a running app and warm OS cache, not cold startup.
  • Native tracking: JS evaluation followed by the real DdFlags.trackEvaluation bridge. Test no tracking, disabled loggers, exposures, evaluations, and both, with repeated/changing contexts and uninterrupted/yielding bursts.
  • JS batching: compare no tracking, the existing bridge, a one-record wrapper control, and a 25-record/50 ms buffer. Test 10 and 100 evaluations/second, 20-call bursts, and a 500-call burst in both mode orders.
  • Event verification: a loopback HTTP collector verifies actual exposure deduplication and evaluation aggregation. Validators check totals, individual targeting keys for batching, and rejected or incomplete calls.
  • Published evidence: 43 successful reports, measured source snapshots, checksums, and summary JSON. Reviewers can verify saved results without running a simulator. Build logs, binaries, device identifiers, system-property dumps, and failed diagnostic runs are excluded from the export. Report bytes are unchanged.

Motivation

The RFC needs evidence for parsing, evaluation, configuration ownership, and tracking delivery. Baseline measurements cover setup, first read, repeated reads, throughput, and controls. Follow-ups cover saved configuration and native tracking without customer data or external intake.

Setup starts with bytes in memory and excludes loading, networking, storage, and fixture generation. The saved-configuration experiment measures file access separately. Crossed placements use one ProtoJSON handoff, not the fastest possible JSI or shared-memory implementation.

Evidence and Interpretation

Main evidence compares matching Release/Hermes experiments on the iOS simulator and Android emulator. Compare alternatives within each platform, not absolute iOS-versus-Android times. The iPhone run is supplementary; it ended at serious thermal state.

  • At 100 flags, repeated JS reads took about 6-7 us p50 on both emulators. Native-direct evaluation was faster, but RN reads through the synchronous bridge took about 18.3 us on iOS and 14.2-14.6 us on Android.
  • At 1,000 flags, the first result from saved bytes took about 30 ms through JS on both platforms. Native parsing and evaluation returned it sooner. JS parsing can block the JS thread even after an asynchronous file read.
  • Tracking event totals matched in both mode orders on both platforms. Caller timing varied, so it does not reliably rank individual loggers. These experiments do not compare native tracking with a JS-only logger.
  • Batching improved bursts and reduced measured JS work at 100 evaluations/second. At 10/second, it did not reduce calls or consistently reduce work, and delayed submission by about 66 ms. No conclusion about production queue limits, crash reliability, or optimal batch size follows from these tests.
  • The existing bridge makes one trackEvaluation call per evaluation, including the mode with both loggers enabled. It does not measure independently registered hooks, which could make separate calls. Repeat measurements with the final public API.

The RFC proposes local JS evaluation for RN-only use, shared native ownership when RN and native readers need one configuration, and optional native tracking. It favors native tracking for reuse and maintenance, not a measured speed or reliability advantage over JS logging.

Detailed results and limits: iOS, Android, and batching.

Additional Notes

  • No shipping provider API or tracking behavior changes. The non-benchmark fixes add missing Foundation imports to native test mocks so CI can compile them.
  • Native prototypes are opt-in through DD_FLAGS_PROTOTYPE_PATH and DD_FLAGS_KOTLIN_PROTOTYPE_PATH. The baseline bypasses Datadog initialization. Tracking runs use synthetic data, loopback intake, disabled RUM, small/frequent uploads, and a one-second aggregation interval. These are benchmark settings, not SDK defaults.
  • The benchmark Metro resolver uses public CommonJS exports so core and the handoff adapter share the same Protobuf-ES decoder and UTF-8 fallback on Hermes.
  • Setup and run instructions cover both platforms. Archived manifests identify the measured code, including changes not committed at measurement time. Later documentation, validator, and CI fixes are not new performance runs.
  • Native prototypes implement only the tested rules. Full conformance, cold startup, shipping package size, native queues, RUM attribution, and device energy/memory-pressure behavior remain unvalidated. Benchmarks do not replace correctness tests of the final provider lifecycle.

Validation

  • 57 focused Jest tests across nine suites pass. Targeted TypeScript, Prettier, and git diff --check pass.
  • Both ARM64 Release benchmark builds succeeded for the recorded batching runs. No new simulator build or performance run was required for the later CI import/header fixes.
  • Placement runs contain 900,000 timed reads each. Each platform has 810 saved-configuration samples and 60,000 timed reads across the five tracking modes and two orders.
  • All 16 batching runs passed, with 88,320 timed reads. Each tracking run emitted exactly 2,772 exposures and 5,544 represented evaluations. Baselines emitted no events.
  • All 43 exported reports pass their applicable correctness and event checks. Manifest-backed report and source hashes were verified. Reprocessing the 16 archived batching reports reproduces the saved summary exactly.
  • Earlier companion validation: 22 Swift tests passed with one opt-in host benchmark skipped; eight Kotlin tests passed. This update does not claim a new full native SDK test run.
  • The pre-existing benchmark ESLint configuration reports Environment key "jest/globals" is unknown; targeted TypeScript/Prettier and repository CI are separate checks.

Review checklist (to be filled by reviewers)

  • Feature or bugfix MUST have appropriate tests
  • Make sure you discussed the feature or bugfix with the maintaining team in an Issue
  • Make sure each commit and the PR mention the Issue number
  • If this PR is auto-generated, please make sure also to manually update the code related to the change

@sameerank sameerank changed the title [FFL-3347] Benchmark JS and native flag parsing and evaluation [FFL-3347] Benchmark mobile flag evaluation, hydration, and native tracking Sep 29, 2026
@datadog-prod-us1-5

datadog-prod-us1-5 Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Tests

✅ All CI checks and tests passed.

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 2633f58 | Docs | View more details | Give us feedback!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant