feat: show live MCP server status and reconnect from /mcp - #579
Merged
nikita-ashihmin merged 22 commits intoOct 2, 2026
Merged
nikita-ashihmin merged 22 commits into
nikita-ashihmin merged 22 commits into
Conversation
nikita-ashihmin
force-pushed
the
nikita.ashikhmin/mcp-status-command
branch
from
October 1, 2026 18:05
6656288 to
615ee24
Compare
/mcp now shows the live status of each MCP server in a list, not in a table. /mcp reconnect reloads the MCP configuration and shows the status after the servers start. The list shows the reconnect hint only when a server failed, was cancelled, needs authentication, or did not start. The hint says to restart the servers that are not connected. The list groups the servers by the status. The groups with a problem come first, and the disabled servers share one line. A table rendered badly in AIR. The narrow columns broke the words, and the details cells showed the raw error chains. So each item now shows the tool count or one short error. The cleanup removes the wrapper messages, the numeric codes, the Rust type paths, and the repeated messages. It keeps at most two sentences. An "e.g." or an "i.e." does not end a sentence. The cleanup removes only a real type path, such as crate::module::Type<T>. A lone "<" and an IPv6 URL stay. A segment that repeats and extends an earlier segment replaces it. The cut never splits a surrogate pair, and an empty error shows no detail. A server that needs authentication, or a server that is not logged in, shows "not signed in". A status that the adapter does not know shows as unknown. A server name gets a backtick fence that is longer than the backticks in the name. An empty name shows as "(unnamed)". The reload errors in the notes use the same cleanup. The status list gets the session thread id, because Codex fills runtimeStatus only for a thread. Without the thread id, every status is unknown. The adapter reads all pages of the list. When Codex does not know the thread, the adapter lists the servers without the thread and does not wait for the startup. Other list errors are not retried. The reconnect sends config/mcpServer/reload. Codex applies the reload to every loaded thread. It restarts the servers that failed, stopped, or changed, and a healthy connection stays as it is. The protocol has no request to reconnect one server, so /mcp reconnect <server> reloads all servers and tells the user so. The wait list comes from the status list after the reload, so a removed server does not hold the wait until the timeout. When a server needs authentication and the client supports URL elicitation, the adapter starts the OAuth sign-in and reloads again after a successful sign-in. A failed reload skips the sign-in. The notes show as a list under a "Reconnect:" label, above the status list. A reload error shows on one line, shortened, and without inline code. The MCP startup state is now kept per thread. A startup event of one thread no longer ends the wait of another session, and an event without a thread id counts for every thread. A wait must name its thread. The session startup reports use the session thread too. A startup wait removes its waiter at the timeout or at an abort. A startup wait and a sign-in wait reject when the Codex connection ends. The exit of the Codex process disposes the connection, and the stdio reader never fires a close event. So the waits end on onDispose and on onClose. A test kills a real child process to check this path. The command gets the abort signal of the prompt. After a cancel, it sends no update. A cancel also stops the sign-in: the sign-in closes the dialog with elicitation/complete and drops its completion wait. The adapter routes mcpServer/oauthLogin/completed once, to a waiter list by server name. Before, each sign-in replaced the notification handler, so two sign-ins at the same time lost one completion, and a declined sign-in left its handler. Now a decline or an abort removes its waiter. A session close aborts the startup wait and the sign-in of the startup report, and forgets the startup states of the session thread. A fork publishes no startup report, so it aborts its wait after the optional startup timeout. The scenario baseline stays a recording of the adapter before the AIR extensions. A new allowed difference covers the new /mcp command entry. The /mcp code is in src/mcp/. McpCommand runs the command, McpServerStatusList reads the status pages, McpStatusMarkdown formats the markdown list, and McpServerSignIn signs in to a server. McpStartupTracker keeps the startup states and the startup waits, and McpOauthCompletions keeps the sign-in waits. CodexAppServerClient only feeds them the notifications and the connection end. Doc fix: docs/mcp-startup-await-timeout.md said that, with the timeout field, the wait ends at the timeout. The wait ends when the startup completes or at the timeout, whichever comes first.
nikita-ashihmin
force-pushed
the
nikita.ashikhmin/mcp-status-command
branch
from
October 1, 2026 19:42
615ee24 to
97b6bf5
Compare
A skill command now carries _meta.jetbrains.air.skillPath with the path that skills/list reports. AIR opens this file on a click on the skill chip in the transcript. Other clients get no metadata.
The compaction wait ended only on the close event of the connection. The exit of the Codex process disposes the connection and fires no close event. So the wait never ended when Codex exited during a compaction. The wait now also ends on the dispose event, as the MCP waits do.
McpSessionStartup now owns the pending startup of each session, the startup wait with the timeout from _meta, the sign-in, and the report. CodexAcpServer keeps only the calls. The behaviour does not change.
The AIR keys table and the presentation hints now state _meta.jetbrains.air.skillPath. The test also checks that an AIR client gets the path.
Cap the raw error at 2000 characters, so that a long error does not block the event loop. Keep an IPv6 address, a quoted value, and a short non-Rust path such as Foo::Bar. Keep an HTTP status. Drop only a JSON-RPC code. Do not end a sentence after a list number or etc. Put a server name with a line break on one line.
Make the requested servers of publish required and drop the unused unfiltered branch. Return the timeout race promise itself, so that the caller resumes at the same microtask as before.
Note in the docs that skillPath is an OS-native absolute path, not a URI.
A session load now requests the next item page before it sends the current page. It reads at most one page ahead, so at most two pages are in memory. It does not read a page after the resume boundary. Each page keeps its 60 s timeout, and the close check stops the load as before. A turn page search still reads one page at a time. Measured with a local Codex app-server and real sessions, median of 4 runs: - 180 MB session: the load response goes from 486 ms to 382 ms. - 287 MB session: the load response goes from 566 ms to 442 ms. - 2.1 GB session: the load response goes from 5043 ms to 4572 ms. - 2.1 GB session with AIR capabilities: from 4276 ms to 3696 ms. The time to the first history update does not change. The history updates are byte-equal to the updates before the change. The peak adapter memory grows by up to about 165 MB on the 2.1 GB session.
The fallback title is the whole text of the first prompt. A large prompt made a 25023-character title, and the IntelliJ client ignored it. The history title, the thread name and the session/list title had no limit too. All four paths now use one normalizeSessionTitle with the same limit as claude-agent-acp.
A tool call that starts with a final status keeps only its small fields. So the adapter no longer stringifies its content, raw input, raw output and _meta. A history replay starts most tool calls this way. Load of a 2.1 GB session, 8 runs, median (min): - AIR client: 2952 ms (2827) before, 2641 ms (2595) after. - other client: 4881 ms (4659) before, 4259 ms (4072) after. - adapter CPU: 2.88 s before, 2.56 s after (AIR). A 180 MB session does not change (about 380 ms). The session/update stream is byte-equal to the parent commit for both sessions.
The SDK writer encoded each message with TextEncoder and wrote it through a web stream adapter. The adapter now writes each line as a string to the Node stream. A write waits for drain when the stream buffer is full, so the backpressure stays. After an error, an end or a close of stdout, each write rejects, so the ACP connection closes as before. The process still exits when stdin closes. Load time, 8 runs, median (min), with the parent commit first: - 2.1 GB session, AIR client: 2543 ms (2519), then 2326 ms (2251). - 2.1 GB session, other client: 4105 ms (4055), then 3678 ms (3598). - 180 MB session, AIR client: 367 ms (360), then 349 ms (336). Adapter CPU for the 2.1 GB session goes from 3.94 s to 3.35 s. The session/update stream is byte-equal to the parent commit.
The account read of session/load waits for the app-server workspace routing discovery. That is a network call to the ChatGPT backend. The load now starts the read and thread/resume at the same time, and reads the account once. The auth state of the session uses that read. A login with the default auth request keeps the old order, because it can change the provider of the resume. When the agent needs a login, the load still fails with auth_required and closes the resumed thread. The auth error also wins when thread/resume fails. The client exposes readAuthRequirement instead of authRequired. The tests mock the new method. Load time, 8 runs, median, with the parent commit first: - 2.1 GB session, AIR client: 2295 ms, then 2242 ms. - 2.1 GB session, other client: 3618 ms, then 3598 ms. - 180 MB session: no change (about 345 ms). The gain is the time of the account read, which is short with a fast backend. The session/update stream is byte-equal to the parent commit.
Codex 0.159 account/read runs a workspace routing discovery for a ChatGPT login. That is a network call to the ChatGPT backend. When the backend is not reachable, account/read fails with "workspace routing discovery failed". When the backend does not answer, it fails after 15 s with "workspace routing discovery timed out". session/new, session/resume, session/load and session/list then failed before Codex opened the thread. Such an error does not mean "not logged in". isAccountReadUnavailableError matches the internal error code and the whole text of these errors. A refused token (401) and other errors stay errors, as before. On such an error the session open continues, and the adapter pushes no auth status. AIR keeps the last status, because AIR shows a login request for a none status. The adapter does not log out. A missing login fails the resume or the first turn with the real error. session/new and session/resume now read the account once, so a slow backend costs one wait, not two. session/load waits for the account read at most 1 s after its start when thread/resume is done. Only the routing discovery is slow, and it runs only for an account that the app-server has. A read that ends later sets the account of the session and pushes the status. If it finds that the agent needs a login, the adapter pushes none. Checked with an isolated CODEX_HOME and a 180 MB session: - unresolvable backend: the load failed in 0.3 s. Now it answers in 0.33 s with the full history. - backend that does not answer: the load failed after about 15 s. Now it answers in 1.26 s with the full history. - unresolvable backend: session/new failed. Now it opens the session. No case pushes an auth status.
A late account read of session/load that fails with a revoked token (401) now pushes the none status at once. A missing workspace and a missing account id count as the same auth failure. Before, the adapter read the account again, got the same error, and only logged it. When another client changed the login during the read, the adapter reads the account again. Any other late error is logged as an error and pushes nothing. A late answer applies only while its load is the current open of the session. A close or a new load of the session drops it. An auth status that came after the read started also drops it. A failed thread/resume now waits for the account read only for the rest of the 1 s grace time. Before, it waited for the whole read, up to 15 s. The last known account now follows every successful read and every account/updated, and a logout clears it. The first auth status read pushes none for a revoked token. authentication/status answers chat-gpt with the last known email when the read is unavailable. It answers unauthenticated for a revoked token. authenticate with chat-gpt or chat-gpt-device-code keeps a stored login when the read is unavailable. It starts a new login for a revoked token. The unavailable set no longer has the two "cancelled" texts. The app-server never closes the semaphore of the first one, and the second one comes from a shutdown. The KDoc says that a 403 and a 5xx of accounts/check also give "discovery failed". It also says that the 15 s timeout covers the config load and the token refresh. The grace tests use fake timers. New tests cover a late 401, a late other error, a close and a second load before the late answer, a failed resume with a pending read, and refreshAuthState with an unavailable read.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Cancellation during the initial status request can still trigger a global MCP reload and late output.
Review effort: Balanced
Findings: 1
Open (1)
What changed in this PR
Adds live MCP status reporting, reconnect support, OAuth recovery, and thread-aware startup tracking.
Changes:
- Implements paginated
/mcpstatus and reconnect workflows. - Adds cancellable startup/OAuth tracking and connection-disposal handling.
- Updates AIR skill metadata, documentation, tests, and snapshots.
| File(s) | Description |
|---|---|
src/mcp/McpCommand.ts |
Implements /mcp status and reconnect. |
src/mcp/McpStatusMarkdown.ts |
Formats grouped MCP status output. |
src/mcp/McpStartupTracker.ts |
Tracks thread-scoped startup events. |
src/mcp/McpSessionStartup.ts |
Manages startup reporting and sign-in. |
src/mcp/McpServerStatusList.ts |
Reads all status pages. |
src/mcp/McpServerSignIn.ts |
Implements URL-based OAuth sign-in. |
src/mcp/McpOauthCompletions.ts |
Routes concurrent OAuth completions. |
src/CodexAcpServer.ts |
Wires MCP lifecycle handling. |
src/CodexThreadErrors.ts |
Recognizes unknown status-list threads. |
src/CodexCommands.ts |
Exposes enhanced /mcp and skill metadata. |
src/CodexAppServerClient.ts |
Adds reload and lifecycle event routing. |
src/CodexAcpClient.ts |
Exposes the new MCP APIs. |
src/AirExtension.ts |
Defines AIR skill-path metadata. |
src/__tests__/McpStartupTracker.test.ts |
Tests startup and OAuth tracking. |
src/__tests__/McpStatusMarkdown.test.ts |
Tests status and error formatting. |
src/__tests__/CodexACPAgent/mcp-command.test.ts |
Tests /mcp workflows. |
src/__tests__/CodexACPAgent/session-fork.test.ts |
Verifies fork startup cleanup. |
src/__tests__/CodexACPAgent/session-close.test.ts |
Verifies close-time cleanup. |
src/__tests__/CodexACPAgent/mcp-session.test.ts |
Covers real MCP reconnect behavior. |
src/__tests__/CodexACPAgent/mcp-config-merge.test.ts |
Updates MCP output assertions. |
src/__tests__/CodexACPAgent/elicitation-events.test.ts |
Updates sign-in assertions. |
src/__tests__/CodexACPAgent/compact-command-lifecycle.test.ts |
Tests disposal during compaction. |
src/__tests__/CodexACPAgent/CodexAcpClient.test.ts |
Tests MCP and AIR skill integration. |
src/__tests__/acp-test-utils.ts |
Adds disposal support to mocks. |
src/__tests__/scenarios/client-profiles.test.ts |
Allows the /mcp command change. |
src/__tests__/scenarios/baseline.ts |
Applies MCP compatibility transforms. |
src/__tests__/CodexACPAgent/data/mcp-command-*.md |
Adds MCP command snapshots. |
src/__tests__/CodexACPAgent/data/*.json |
Updates command and skill snapshots. |
src/__tests__/scenarios/data/air/*.jsonl |
Updates AIR command recordings. |
README.md |
Documents /mcp reconnect. |
docs/mcp-startup-await-timeout.md |
Documents startup reporting behavior. |
docs/air-extensions.md |
Documents command and skill metadata changes. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| } | ||
|
|
||
| private async reconnect(sessionState: SessionState, serverName: string | null, signal?: AbortSignal): Promise<string | null> { | ||
| const before = await this.listServers(sessionState); |
AIR ignores an agent session title that is longer than 160 characters. It does not cut such a title. The limit of 256 characters made AIR drop a long fallback title, a long thread name and a long history title. The adapter now cuts every session title to 160 characters with an ellipsis.
The tracker stored the thread id of a startup event as it came. When Codex omits the field, the key was undefined, not null. Such an event matched no thread and was not global, so the startup wait did not end. The tracker now stores a missing thread id as null.
This reverts commit eb1e005. The adapter keeps the limit of 256 characters. AIR will accept a longer agent title instead.
Codex cannot reconnect one MCP server. config/mcpServer/reload reloads all servers. The adapter accepted a server name only to check it and to warn that the reload applies to all servers. /mcp reconnect now accepts no argument. An argument gets the same answer as an unknown subcommand. The command hint is now "[reconnect]", and the README shows /mcp reconnect.
The cleanup now collapses the whitespace, removes the wrapper segments and the Rust type paths, keeps two sentences, and cuts the text to 160 characters. One regular expression finds the wrapper segments, and one finds a type path with three or more parts. The cleanup no longer parses the generic arguments, removes repeated segments, or skips the sentence ends after e.g., i.e., etc. and list numbers. The ijproxy error, the startup errors and the HTTP status text do not change. Some texts change: - The Glean transport error gives "Send message error Transport Auth: not logged in". Before, it gave "Send message error: not logged in". - "JSON-RPC error: 401: Unauthorized" gives "401: Unauthorized". Before, it gave "Unauthorized". - A repeated segment stays, such as "timeout: timeout after 30s". - A two-part type path in brackets stays, such as "[rmcp::Transport<A>]". - A sentence can end after e.g. or etc. The cleanup still reads at most 2000 characters of the raw text.
raceMcpStartupTimeout and settledWithin had the same behaviour. Both end at the value or at the timeout, and both reject only for a rejection before the timeout. settledWithin moves to StdUtils, and the MCP session startup uses it. The wait of /mcp reconnect stays as it is. It aborts the startup wait at the timeout, so that the tracker releases the waiter.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Summary
/mcplisted the configured servers without a live status, becausemcpServerStatus/listwas called without athreadIdandruntimeStatuswas always null. It also read only the first page. The user of an ACP client could not restart a failed server./mcpasks for the status of the session thread, reads all pages, and shows a summary line and a list grouped by status: failed, needs authentication, connecting, connected, and disabled on one line. A failed server shows a short cleaned error without the wrapper chain and Rust type paths. A list is used and not a table, because a table breaks words in narrow columns and raw errors fill the cells./mcp reconnect [server]callsconfig/mcpServer/reload, waits for the servers of this thread to start, and shows the new table. The reload applies to every loaded thread. Codex restarts the servers that failed, stopped, or changed, and keeps a healthy connection. The protocol has no reconnect for one server, so the output says that all servers reload.mcpcommand has a new description and an input hint.Fixes on the way
onDispose, notonClose, on a process exit.docs/mcp-startup-await-timeout.mdnow describes what the adapter really sends for a startup failure.Layout
The MCP logic lives in
src/mcp/:McpCommand,McpStatusMarkdown,McpServerStatusList,McpServerSignIn,McpStartupTracker,McpOauthCompletions. The main files keep only the wiring. The startup tracking that was inCodexAppServerClientmoved toMcpStartupTracker.Tests
mcp-command.test.ts,McpStatusMarkdown.test.ts(real error samples) andMcpStartupTracker.test.ts, including a real child-process exit.mcp-session.test.tscovers the reconnect.withFeatureChangesinbaseline.tsand the compatibility rule indocs/air-extensions.mdallow the newmcpcommand entry.npm run typecheck,npm run build,npm testpass.