Skip to content

FFL-3347: Add Kotlin rules evaluation benchmark prototype - #3931

Draft
sameerank wants to merge 1 commit into
developfrom
sameerank/FFL-3347-kotlin-rules-evaluator-prototype
Draft

sameerank wants to merge 1 commit into
developfrom
sameerank/FFL-3347-kotlin-rules-evaluator-prototype

Conversation

@sameerank

@sameerank sameerank commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

What does this PR do?

Adds a standalone, benchmark-only Kotlin rules evaluator under prototypes/rules-evaluation. This provides the Android counterpart to the Swift prototype so we can compare parsing and evaluation in JavaScript versus native code from React Native.

  • Parses client UFC protobuf and ProtoJSON using generated protobuf-java models.
  • Implements the same boolean benchmark workloads as Swift: static values, string membership, compound numeric targeting, and MD5 percentage splits.
  • Supports configuration replacement, changing evaluation contexts, native-direct timing controls, and warm persisted-configuration loading.
  • Checks both decoders against 128 synthetic reference cases generated by @datadog/flagging-core@3.1.1, plus targeted correctness and error-handling tests.

This is not a production provider or a complete UFC evaluator. It adds no shipping SDK API and is not included in the Android SDK's module graph. The RN adapter owns bridge serialization, scheduling, and platform reporting; this module owns parsing, evaluation, and benchmark operations.

Motivation

FFL-3347: provide Android evidence for the mobile SDK building blocks and benchmarks RFC.

We want to separate the cost of decoding, evaluation, RN transport, and loading cached configuration before choosing where each responsibility belongs. This prototype supports six decode/evaluation placements and three cached-loading paths through the RN harness.

Companion work:

The Android adapter, runner, native-tracking and batching experiments, and detailed results are published in RN commit c7b19737. Kotlin unit tests can run independently of that harness.

Additional Notes

Validation

  • Eight focused Kotlin tests pass, including both decoders, A/B/A context changes, configuration replacement, failed-install preservation, unsupported and cyclic rules, targeting errors, persisted loading, and timing-control checksums.
  • The RN Release/Hermes ARM64 APK built successfully with this prototype; 28 RN harness tests also passed.
  • On a hardware-accelerated Android 15 / API 35 ARM64 emulator: six-placement smoke test, two full runs totaling 1.8 million timed reads, and 810 persisted-loading samples passed their correctness gates.
  • git diff --check passes. The full shipping Android SDK build/lint suite was not run; its root tasks do not exercise this standalone project.

To run the focused tests with JDK 17 and Gradle 8.10.2:

gradle -p prototypes/rules-evaluation --no-daemon --max-workers=2 test

The RN Gradle 8.12 wrapper can also run the standalone project with -p. This experiment does not use the Android repository's root Gradle 9 wrapper. See the prototype README for dependencies, fixture provenance, and companion setup.

Initial Findings

For the 100-flag configuration, repeated-read p50 was 5.96-6.00 us for JS decoding/evaluation, 14.21-14.62 us for native decoding/evaluation over the synchronous bridge, and 188.04-199.63 us over the asynchronous bridge. Ranges summarize two runs, each using the median of five per-repetition percentiles; they are not confidence intervals.

The native-direct evaluator control was faster (0.750-0.875 us p50), but excludes RN transport and uses preconverted inputs. These findings support keeping repeated RN reads in JS, not a claim that the JS evaluator itself is faster. Native-only cached loading reached the first result sooner; native decoding followed by a ProtoJSON transfer to JS was slower than transferring bytes and decoding in JS.

Limits

  • Emulator measurements are not physical-device latency or energy guarantees. No physical Android device was used.
  • The evaluator supports a boolean workload subset, not full UFC conformance. Unsupported features are explicit errors; regex, semantic versions, obfuscated targeting, and other variation types are out of scope.
  • The full Java protobuf runtime supports the ProtoJSON comparison. This is not a production dependency or APK-size recommendation.
  • Persisted loading uses a warm OS file cache, not cold app startup. This Kotlin prototype does not add tracking. The RN companion measures the existing Android tracking implementation and compares per-evaluation delivery with JS batching.
  • The RN Release build used New Architecture/Hermes with R8 disabled. ProtoJSON is one transfer format, not a lower bound for optimized binary, typed-object, or JSI designs.
  • Published evidence includes 43 successful reports across both platforms, source snapshots, checksums, and summary JSON. Report bytes are unchanged. Exported manifests omit device identifiers and system-property dumps; build logs, binaries, and failed diagnostic runs are excluded. Reproduction commands and Android results are available in the RN companion.

Review checklist (to be filled by reviewers)

  • Feature or bugfix MUST have appropriate tests (unit, integration, e2e)
  • Make sure you discussed the feature or bugfix with the maintaining team in an Issue
  • Make sure each commit and the PR mention the Issue number (cf the CONTRIBUTING doc)

@datadog-prod-us1-3

datadog-prod-us1-3 Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Tests

✅ All CI checks and tests passed. Datadog automation helped this PR pass.

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🔄 Datadog retried 1 test - 1 passed on retry View in Datadog

🎯 Code Coverage (details)
• Patch Coverage: 100.00%
• Overall Coverage: 71.75% (-0.02%)

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 9a463ce | Docs | View more details | Give us feedback!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant