feat: integrate E2EE remote-control foundation - #394
Draft
Lokesh7025 wants to merge 138 commits into
Draft
Lokesh7025 wants to merge 138 commits into
Lokesh7025 wants to merge 138 commits into
Conversation
Dependency ReviewThe following issues were found:
|
8 tasks
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Serialize lifecycle transitions across processes, recover authenticated ready and pre-join phone state, and restrict destructive cleanup to provably pristine initialization state. Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Reconcile the Session 50 Node and browser scope, pairing transcript, durable ownership, and browser persistence constraints. Defer native mobile bindings to Phase 13 and keep browser pairing disabled until rollback protection is approved. Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Signed-off-by: VishnuM049 <vishnu.muthiah04@gmail.com>
Wire protocol 32 adds remote.status, a local-only method reporting the remote host's pairing phase, relay connection, whether the paired phone holds a relay route, the endpoint's witness state, the latest failure, and the log path. /remote status in the terminal shows it. The deployment-test host now: - writes an owner-only, size-bounded remote.log, since a daemon started by the terminal discards its output, and logs errors with their code and HTTP status; - records host faults, not requests the daemon refused, as the latest failure; - probes its relay every 30 seconds; - lets a reply wait out a short relay reconnect instead of dropping it; - ignores repeated pair activations once the bridge serves the device; - retries restoring a paired session with backoff when the network or the witness is briefly unreachable, instead of giving up until restart; - makes /remote wait for a restore still in progress, so a replaced session can never reconnect a second daemon route that fights the new one on the relay. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
The page probes the relay on a heartbeat and as soon as it becomes visible or regains its network, so a socket that died while the phone slept is replaced right away. After any reconnect an open conversation subscribes again after its last acknowledged cursor, replaying only what was missed in one round trip; a cursor the daemon no longer knows falls back to reopening the session, and a session that failed to open while the daemon was away is opened again once it returns. The page says when the daemon is offline, and concurrent refreshes share one session list request. A newer tab for the same pairing asks the older one to let go of the endpoint and its lock, and waits briefly for it, instead of failing with lifecycle_busy. Pairing confirms the activation by getting an answer from the daemon, and startup failures offer a Retry button. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
The barrier scenario tracks which step it is in, but the test worker reduced every failure to a bare code, so the flaky Chrome and Safari failures only reported "The E2EE operation failed safely". Scenario operations now return their failure message, the barrier scenario keeps the binding's code in it, and the page rejects with both. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Two ways a prompt from the phone was lost, both found by the new failure-injection E2E: - Reopening a thread after a reconnect cleared the session while it waited for the old subscription to close, so a prompt sent in that window was silently dropped. The thread now takes over the view before any round trip and closes the old subscription in the background. - After a daemon restart the resent session.send reached the daemon before the page reopened the session, and was refused with unknown_session. A refused send never ran, so the page reopens the session and sends the prompt once more. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
A phone-emulated Chrome pairs with a real sandboxed daemon through a local copy of the deployment-test stack: the deployment-test control plane with its witness, the Elixir relay, and one HTTPS origin that serves the phone page and forwards to both, like the stack's CloudFront. The runner builds the Node and browser bindings pinned to a witness generated for the run and restores any existing deployment-test builds afterwards. Each scenario injects one failure and sends a prompt through it: dropped connections, a silent phone socket, a silent daemon socket, an offline period, a frozen page, a daemon restart, a relay restart, and a second tab taking over. Every prompt must reach the model exactly once and every reply must appear exactly once, within a bound. The run records reply times after each failure and keeps the stack and page logs when it fails. CI runs it as "Remote phone E2E" when remote-control code changes. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
The daemon authority doc still said the module was not wired to the runtime, the implementation briefs still pointed at draft PR #413, and the hosted-path evidence said no witness was deployed. Describe the deployment-test wiring and its limits, and record the phone evidence from the AWS stack and the local failure E2E. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
"editing and /quit recover from a stale shutdown status" failed about half the time with three shutdown attempts instead of two. The test injects one stale status, but /quit could also race the "hello" turn: a turn that is still settling changes the daemon's status revision, so the real shutdown was refused a second time and the TUI correctly retried again. Wait for the reply and an idle daemon before /quit, so only the injected stale status is exercised. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Until now every phone shared one device ID from the stack configuration and presented one static relay possession proof carried in the pairing link, so a copied link was a working credential and two phones collided on one route. Each /remote pairing now mints a fresh UUIDv7 device ID and a 32-byte enrollment secret. The daemon invites the ID with only the secret's SHA-256 digest; the version 2 pairing link carries the secret instead of the proof. The phone creates a non-extractable ECDSA P-256 key in WebCrypto, keeps it in IndexedDB, enrolls its public half once, and signs every relay admission over the ticket and connection nonce. The control plane verifies that signature and issues tickets only for enrolled, unrevoked devices. Enrollment binds exactly one key within a 10 minute window, so a copied link cannot enroll a second browser. Pairing again revokes the previous device in the control plane and in the daemon's authority store. A refused ticket now surfaces as route_forbidden and is not retried, so a locked-out phone says so at once. Device records persist in the stack's DynamoDB table. The daemon keeps its static deployment-test proof; production gates stay closed. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
The deployment-test witness kept its three replicas in process memory, so every control-plane restart or redeploy emptied them and every paired phone had to pair again. Replica records and high-water journals now live in two DynamoDB tables. A record is split into a head item and one item per ledger event, operation, retained response, and used recovery read, so it never meets the 400 KB item limit. Each step rereads the head consistently and writes only its new entries plus the next head in one transaction conditioned on the revision it read. History entries are write-once; the journal table takes only conditional first writes, and the task role cannot update or delete in it. A replica that restarts over durable state cannot recover all lineages at once, since fresh signed heads exist only for endpoints that are reading. WitnessReplica.resumeDurable and recoverLineage apply the existing restart checks one lineage at a time: until a lineage's head matches fresh heads from both other replicas the replica does not vote for it, and the only other request it admits is the registration of a lineage its store and journal have never held. The deployment-test gateway recovers a lineage on its first fresh read, which endpoints make before every mutation. Advances and catch-up no longer drop the record of used recovery reads. A control-plane restart also exposed two daemon paths that lost work while the witness was briefly unavailable. The bridge dropped a delivery batch the witness refused, so the phone never saw the reply; it now holds such batches and sends them, in order, once the witness recovers. The host dropped a request refused the same way, and the device resends only when its connection changes; the host now receives it again once the witness settles, a bounded number of times. Failed recovery attempts are logged, and the control plane logs witness failures with their stage and cause. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
A witness step reread the record head before computing even when this process had just written that revision, and every journal append read the entry and the latest sequence before its conditional put. A writing step now computes from the cached record and relies on the transaction's existing revision condition; a conflict drops the cache and reruns the step on a fresh load. A step that writes nothing still confirms the head revision with a consistent read before answering. A journal append right after the newest sequence this process saw is one conditional put, and a refused put falls back to the existing compare. An advance goes from 36 DynamoDB calls to 21 and a fresh read from 9 to 6 across the three replicas. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
The phone page showed a turn only once it finished. Model output went into one status line, tool calls were bare names, and every event rebuilt the transcript. An open session now renders with the desktop client's transcript renderer, reconciled by event ID and coalesced to one update per frame. The reply in progress streams in place with its thinking and running tools, the composer shows the turn's stage and elapsed time, and a sent prompt shows at once until the daemon records it. The page gains a header with a connection indicator, session cards that refresh on session changes, a growing composer, keyboard and safe-area handling, and the phone's light or dark theme. The renderer loads as its own chunk while pairing runs, and the view's cursor is acknowledged after events pause rather than every half second. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
The stack answered only on its CloudFront host name. With domain_name set, Terraform issues a us-east-1 certificate validated in the parent Route 53 zone, adds alias records, and hands clients a relay URL on that host, which the phone page's same-origin CSP needs. The bare host opens the phone page, and the CloudFront host name redirects the page there. deploy.sh defaults domain_name to remote.observal.io. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
A paired phone can now answer an ask_user_question interaction instead of waiting for someone at the computer. session.interaction.respond is allowed with the steer scope, and the daemon refuses it from a remote device for every interaction kind other than user_question, so MCP tool approvals and elicitations stay on the computer until the approve_within_policy scope in docs/architecture/remote-permission-authorization.md exists. The phone renders the questionnaire card inline, shows "Waiting for your answer" while one is pending (or "Waiting for approval on the computer" for kinds it cannot answer), and sends the answer through the encrypted session. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
The full pairing link is 671 characters, which makes a version 18 QR code about 50 terminal rows tall. The daemon now seals the full link's fragment with AES-256-GCM under a fresh key, parks the ciphertext on the control plane under a random 128-bit ID for at most 15 minutes, and shows https://<host>/remote/#p=<id>.<key> instead: about 100 characters and a version 5 code. The key never leaves the fragment, so the control plane stores only ciphertext it cannot open. If parking fails the daemon shows the full link as before. - protocol: publish and fetch requests for sealed links - control plane: PairingLinkService with in-memory and DynamoDB stores; publishing needs the account credential, fetching needs none - sdk: sealRemotePairingLink, parseShortRemotePairingFragment, openRemotePairingLink, fetchRemotePairingLink - phone: resolves a short fragment before pairing, and explains an expired link - tui: /remote draws the code in a framed panel with what to do beside it, falling back to the stacked layout in narrower terminals Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
The phone page can sign in through a Cognito user pool that federates Google (auth.remote.observal.io), and pairing links then leave out the account token. Until now every link carried the stack's account credential, which could call every control-plane route, including the daemon's. - control plane: a principal can carry the phone scope. Phone principals reach only pairing, device enrollment, device relay tickets, and the witness; the daemon's routes answer 403 scope_forbidden and daemon relay tickets 403 forbidden_route. - aws control plane: CognitoPhoneAuthenticator verifies the pool's RS256 access tokens (issuer, token_use access, client_id) and grants the phone scope. The deployment-test runtime accepts them beside the static token when AXL_TEST_PHONE_ISSUER and AXL_TEST_PHONE_CLIENT_ID are set. - page: the authorization code flow with PKCE. The pairing fragment waits in session storage across the redirect, the refresh token stays in local storage, access tokens are refreshed five minutes before expiry and handed to the witness binding. - daemon: phoneSignIn in the deployment-test config drops `t=` from links. - infra: cognito.tf (user pool, Google IdP with the secret from SSM, public PKCE client, custom domain with its certificate and alias records), the page CSP allows the token endpoint, remote-page.sh publishes sign-in.json, remote-config.sh sets phoneSignIn. - e2e: a fake user pool signs every phone in, so all scenarios pair without an account token in the link. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
… stack Adds a production control-plane runtime with no account credential: Cognito access tokens from the daemon client carry the account and the phone client the phone scope, only members of the remote group are accepted, and each daemon installation registers its own P-256 key, whose signature admits its relay tickets. Relay tickets need the installation to belong to the caller. The relay accepts production mode, and Terraform gains the daemon client, the remote group, and a control_plane_mode switch that defaults to deployment-test. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
A daemon in WSL has no v1 envelope-key store: WSL is the headless Linux case the storage RFC leaves unsupported. This adds one for that target. Records keep the native Windows store model (one file per key and lifecycle, exclusive temporary file, sync, rename, directory sync, the same activation and reconciliation) on the Linux side. Only the sealing crosses to Windows: axl-dpapi-helper.exe, started through WSL interop as the signed-in Windows user, seals each record with nested machine-scope then user-scope DPAPI and reports the SID it binds to. The helper is stateless and sees one bounded value per request, never a path or a record identity. docs/architecture/remote-production-wsl.md amends the storage RFC for this target and describes the production remote path this is the first piece of: accounts, a per-account control plane, a three-region witness, and an opt-in that keeps it off for everyone else. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Adds the hosted-wsl build of the Node binding: the production binding plus one constructor, hostedWslDaemonEndpoint(root, helper, account, installation, session), whose envelope keys are sealed by Windows DPAPI through the Windows user's axl-dpapi-helper.exe and whose endpoint verifies certificates against build-pinned replica trust (AXL_E2EE_HOSTED_TRUST_FILE). check-abi keeps it out of the production artifact, and CI lints the feature. The architecture note now describes the account opt-in group, installation keys, and the enablement conditions as built. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
| if (!(cause instanceof DynamoError)) status = 500; | ||
| } | ||
| response.writeHead(status, { "content-type": "application/x-amz-json-1.0" }); | ||
| response.end(JSON.stringify(body)); |
On SIGTERM the control plane kept serving kept-alive connections until it was killed, so a witness step could be cut between its replica writes and leave one replica behind the other two, which lineage recovery then refuses. Requests already received now finish, each connection closes after its response, and the process exits once none is left. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
axl remote login signs the daemon in with Google through the stack's user pool, creates the installation's P-256 key, seals the refresh token and key with Windows DPAPI through the helper, and registers the key. The remote host now takes its credential, possession proof, and endpoint from settings, so the daemon runs the production host for a stored account and the deployment-test host only when AXL_REMOTE_DEPLOYMENT_TEST is set. /remote is always listed and says how to set remote access up. The phone E2E runs in production mode by default. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
The deployment-test smoke test pairs with the static test token, which production refuses, so deploy.sh failed after a successful production deploy. In production mode it now checks that the control plane reports production health. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
#457 merged two terraform applies with a literal \n where a line continuation belonged, which passes a stray argument to terraform. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
Skills written for other agents list allowed-tools as YAML items. Axl accepted only the spec's single string, and one skill it could not parse failed every session.create. A list of non-empty strings is now joined into the space-delimited form. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
…erlap connect CloudFront PriceClass_100 routed India through Europe on every witness call, and every page asset (including the 3 MB wasm) was re-downloaded on each open. Use PriceClass_200, cache hashed assets for good and revalidate the rest, start the encryption module and the first relay ticket early, and subscribe before resuming a session that is already open. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
The DynamoDB witness storage cloned and re-encoded a lineage's entire append-only record on every transact and read, so each witness call cost more the longer a pairing was used. Freeze the cached record and hand it out as is, and encode only new or replaced entries. The phone also acknowledges its cursor every 10 s and not while a request it waits on is in flight, and the control plane gets a full vCPU. Signed-off-by: Lokesh <lokeshselvam7025@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
RC integrates end-to-end encrypted remote control for Axl: a phone or browser device controls a local daemon through a hosted relay, while the relay and control plane only ever see ciphertext. This PR merges RC into
mainonce the code is complete, reviewed, and fail-closed for production.Turning remote control on for users is deliberately not part of this merge. Each platform gets its own reviewed enablement change once its storage, witness, packaging, and runtime evidence gates pass, as
production-e2ee-storage-and-rollback.mdandsession-60-windows-integration.mdrequire.What RC contains
device <-> daemontopology). Cross-language framing and internal-contract fixtures.observe,steer), hosted and local scope intersection, terminal revocation, a bounded audit log, and an explicit RPC allowlist (packages/daemon/src/remote-rpc.ts).--unsafedaemons grant onlyobserveremotely.axl-e2ee-mls-pq-v1(hybrid X-Wing suite), durable transactional persistence, the pairing transcript and lifecycle, epoch update barriers, and the rollback-witness protocol with a three-replica quorum.AXL_REMOTE_DEPLOYMENT_TESTset,/remotein the TUI prints a QR code. A phone browser pairs with the local daemon over the relay, lists and opens sessions, sends and queues prompts, streams replies, and stops turns./remote statusshows the pairing, relay, phone, witness, and latest error, and the daemon writes a remote log file.infra/aws/hosted-path-test). The control plane, relay, rollback witness, and phone page sit behind one CloudFront origin with a strict CSP.What stays fail-closed or out of scope
productionStorageReadyis false, and production endpoint constructors returnunsupported_platform,secure_store_unavailable, orrollback_anchor_unavailable.remote-permission-authorization.mdis specified only)steer; exposing them to users needs its own enablement review.How was this tested?
Run on RC
39cc7e3(116 commits ahead ofmain, 0 behind):pnpm check: 1213 tests, 1207 passed, 6 platform skips, 0 failures. Also formatting, lint, type-check, boundaries, and generated files.pnpm relay:check: 17 relay tests, 0 failures. Also formatting, warnings-as-errors compilation, Credo, Dialyzer, and the dependency audit (no vulnerabilities).CI runs:
Remote phone E2E (
pnpm --filter @axl/web test:remote-e2e, in CI): a phone-emulated Chrome pairs with a real sandboxed daemon through a local copy of the deployment-test stack. The stack includes the real control plane, witness, relay, and page. The test then sends a prompt through each of these failures:It asserts that every prompt reaches the model exactly once and every reply appears exactly once.
Live on the AWS deployment-test stack from a phone-sized browser: pairing, sessions, prompts and streamed replies, recovery after a daemon restart, tab handoff, and
/remote status(docs/evidence/hosted-path-deployment-test.md).Learning
The relay stays independent of the cryptographic message formats, so transport, authorization, and delivery were built and tested against opaque bytes while the endpoint cryptography developed separately. Most of the phone flakiness came from the connection lifecycle rather than the cryptography: silent sockets, suspended pages, duplicate tabs, and the order in which a restarted daemon sees resent requests. The failure-injection E2E now pins each of those down.
Checklist
REUSE.toml.Signed-off-bytrailer.Licenses
openmls, withopenmls_basic_credential,openmls_memory_storage,openmls_traits0.6.0)js)packages/e2ee/DEPENDENCIES.mdrecords the full Rust closure, checksums, and provenance.AI assistance
gpt-5.6-sol.