Guided connect setup, and notices that speak to the person in the thread - #794
Merged
Merged
Conversation
robzolkos
force-pushed
the
guided-connect-setup
branch
from
September 29, 2026 00:20
88bfe83 to
6874547
Compare
robzolkos
marked this pull request as ready for review
September 30, 2026 03:27
The connector's automatic notices were written for whoever runs it. They used terms like "task", "redispatch", "attempt" and "worker", and ended with a shell command. The person who reads them is the one who mentioned the agent. They now say plainly what happened and what that person can do, and keep the ids on a closing Ref line: - "I'm not set up to work in this project yet, so I haven't started on this", instead of a holding notice about trust and projects. - "Working on this as of 14:05 UTC", dated, because a notice must never claim the present. - For work that stopped: interrupted, out of time, or went wrong. Then "Mention me again to try again" when mentioning again could help. - Nothing, when the agent already replied to explain why it failed. Mentioning it again won't change a refusal. connect status still shows the failure. Notices posted by earlier versions are still recognised and retracted. A worker that can't start (Claude Code missing, too old or logged out) used to show up only as a notice. Mentioning the agent again gave the same notice every time. Now: - connect doctor and the connector's startup try the worker the way the connector runs it: it starts, it knows every flag the connector passes, and it's logged in. - After two start failures in a row, the dispatcher stops taking work. It rechecks every minute and backs off to an hour. - A task that never started says so: "I couldn't start on this: something's wrong on the computer I run on. The person who runs me needs to check it."
Setting up an agent took a separate sign-in as yourself to name the
operator, project ids nobody knows, and output that read like a log.
Plain `basecamp connect setup` is now a guided, resumable conversation:
- It connects this computer to the agent if it isn't already
("✓ Connected this computer as Rob Zolkos (agent)").
- It reads the agent's owner from Basecamp and makes them the operator.
For a personal agent that's the only choice: "Rob Zolkos (you) is the
only person who can give it work."
- It checks the worker starts ("✓ Claude Code 2.1.283 is ready").
- It serves every project the agent is in, by name. An agent in none
gets told how to add it, and setup exits cleanly.
- It offers to start the connector whenever you log in. If the service
can't start, it says what to run instead.
- When everything passes it says so, instead of listing ids and paths.
Setup with flags, doctor, service install and --json keep their
detailed output. `auth agent connect` and guided setup share one
connection routine. The guided connection identifies itself to Basecamp
as "basecamp connect".
Guided setup ended by offering to "start your agent now, and whenever you log in", which installed the systemd service. The service isn't ready: it runs the agent in the home directory, where it may change files without asking; it freezes PATH at install; it is Linux-only; and it has not been tried on a real machine. Setup no longer offers it or touches systemd. When the agent isn't running, the summary says to start it in the folder it should work in, that it can change files there without asking, and gives the one command. `connect service` itself is unchanged here.
robzolkos
force-pushed
the
guided-connect-setup
branch
from
September 30, 2026 13:47
6874547 to
a104bea
Compare
This was referenced Sep 30, 2026
Two starts in a row that never ran hold new work, but the dispatcher asked about the hold once a pass. A start that fails at once gives its slot back, so the rest of the batch still launched: with two slots and five queued records, all five were started and answered "I couldn't start on this". Each launch now asks whether work is held.
…ests
The worker's preflight only ran as the re-check of a hold, and a hold took
two failed starts to happen. A connector started with Claude Code logged
out reported itself running and answered the first two requests with "I
couldn't start on this" before it said anything was wrong.
The dispatcher now runs the preflight before its first pass. A worker that
isn't ready holds new work from the start: the connector keeps listening,
says why in its log and in connect status ("Claude Code isn't ready — …"),
and takes work once a check passes. A preflight that passes, or a driver
without one, changes nothing.
…ofile Guided setup switches to the agent profile when it isn't already active, and applying a profile overwrites the account and base URL with the profile's own. Unlike root, it didn't re-apply the environment and flags afterwards, so `connect setup --account 777` against a profile bound to 999 quietly worked in 999 instead of refusing, as it does with -P. The overrides are re-applied, in root's order.
A setup being resumed kept its connect.json and skipped setup's checks, scope among them. A profile reconnected read-only since then (auth agent connect --scope read) was reported as set up and ready to mention, though it couldn't reply or acknowledge. The resumed setup now runs setup's scope check and refuses such a credential with the way to reconnect.
Guided setup checked the worker through its spawn preflight, and the acp driver has none, so a setup on the acp driver skipped the worker check altogether: with its pinned adapter missing, setup still said the agent was set up. It now falls back to doctor's check of where the driver finds its worker, and stops on a failure with its fix.
When a second failed start held new work, the dispatcher asked the worker's preflight why with the lock released. A start that worked in the meantime cleared the hold and recorded the connector running, and then the late reason recorded it as not taking work, which stayed in connect status while work was being taken. The connection's status is now written under a lock of its own, and each writer asks the hold again first.
A request was classed as never started from its own record alone: not picked up, outcome unknown, and the attempt failed. In a conversation where the worker had already picked up and answered an earlier request, a follow-up it failed to get to was reported as "I couldn't start on this: something's wrong on the computer I run on", with no suggestion to mention the agent again. A request only never started now if nothing in its settlement was picked up; the follow-up reads as unfinished, with the retry.
…ilding, and say when nothing is served - On a platform the connector doesn't run on, guided setup went ahead and connected, replacing the secret another computer might be running the agent on, then gave a command that refuses to start. It now refuses first, as the connector does. - Rebuilding a connect.json that didn't fit moved the file aside after the question, without reading it again: a file another setup had repaired meanwhile was moved aside. It's read again under the lock, and a file that now fits is resumed. - A setup serving no projects printed an empty "Works in:" and suggested mentioning the agent "in one of those projects". It now says it works in no projects and how to add one.
A start whose session couldn't even be prepared (its directory, its token socket) settled as a start that ran nothing, like a driver's refusal, but wasn't counted toward the hold. With the sessions directory unusable, every queued request was started, failed, retried and blocked. It now counts, and two in a row hold new work.
A held dispatcher checks its worker in the background. If a start that worked cleared that hold while the check ran, and two failures then made a new one, the old check coming back "fine" cleared the new hold, and new work went to a worker that had just failed twice. Holds are now numbered as they're made, and a check, or a reason, for a hold that is no longer the one standing decides nothing.
…ken program missing - A probe that timed out killed only the program it ran. A launcher that starts the real worker as a child and waits left that child running, and repeated checks piled them up. Probes now run as a process group of their own, and the group is ended when the probe's context is (Unix; off Unix the connector runs no workers). - A program on PATH whose interpreter or loader was gone fails its exec with ENOENT, the same error as a missing program, and was reported as not on PATH. It's now missing only when the program itself isn't there; otherwise it's a program that can't start, with the error.
Any settlement that picked its request up cleared the start-failure hold. A task launched before the worker broke, and finishing after two later launches had failed and held new work, cleared that hold, and new work went to a worker that was still broken. Launches are now numbered in the order they're made; a start that worked clears only failures from launches made at or after it.
Ending the probe's process group only on timeout missed a launcher that starts a child, prints its answer and exits: the probe returned and the child ran on, one more with every recheck. The probe's group is now ended whenever a probe returns.
…as it is - Root checks the base URL of the profile it starts with; guided setup then switched to the agent profile without checking its, and an insecure base_url there panicked in the SDK. It's now checked as root checks it, and refused with where to correct it. - The setup help still said guided setup "offers to keep the connector running" and makes "a connect.json that cannot be used again". It now says what it does: writes connect.json, offers to redo one that can't be used, and ends with how to start the connector in its folder.
A request the worker never reported on has an unknown outcome, not an unfinished one: the worker may have done the work, and even replied, then stopped before it said so. The notice said "I stopped before I finished this. Mention me again to try again", and following it could do the work twice. For an unknown outcome the notice now says it "may not have finished this" (stopped, interrupted, or ran out of time) and "If I didn't, mention me again to try again", so the reader looks first. A request that failed keeps "Mention me again to try again".
Sessions run with --setting-sources "", but the login probe didn't, so a user setting that blanks ANTHROPIC_API_KEY made preflight report a logged-out worker, and hold all work, though the worker itself would run. The probe now asks `claude --setting-sources "" auth status`; the option goes before the subcommand, which doesn't take it.
The start record kept only its earliest failure, so a task launched between two failures, still running when the second held work, cleared the hold when it finished, and queued work went to the broken worker. The record now keeps each failure's launch: a start that worked drops the failures launched before it, and leaves any launched after it, and their hold, standing.
release freed the worker's slot before the failure was noted, so a dispatch pass between the two saw capacity and no hold, and could launch a third request into a worker that had just failed twice. Whether the worker started is now noted inside release, while its slot is still taken.
The probe's cleanup ran after it had been reaped and signaled its pid's group unconditionally, so a pid the kernel had since given to another group leader would have had that group killed. The probe's identity is now recorded, by the kernel's start time, while it runs, and cleanup signals through signalRecordedGroup, which leaves a reused pid alone and still ends children a launcher left behind.
Guided setup reads the agent, asks about its projects, then hands the answers to setup, which reads the credential again. If the profile was connected to another agent while a question was open, setup saved the answers for that agent while the summary named the first. Guided setup now passes the agent and account it showed, and setup refuses anything else.
Guided setup read the lock's metadata and called any live pid running, but the metadata outlives a crash and its pid can be reused, so it could say "Running: yes" and leave out how to start the agent. The holder now counts only if its process started before the lock was taken; one the kernel gave the pid to later started after. The lock itself is never taken, so a starting connector can't find it held.
…le it could The context watcher can fire after Wait has reaped the probe, and a launcher can exit before it is recorded. Kill a running leader only while its start time matches; once the leader is gone, kill only leftovers of a group whose id nobody holds. A probe that never started has nothing to end.
The worker has started since it launched, so it neither holds work nor counts towards a hold.
A bare `connect setup` may pick another profile the second time.
Another setup's file may trust more people than guided setup told the person; it is left as that setup wrote it.
A file another command replaced meanwhile may be a finished setup for another agent; it is left in place.
A worker may reply and not report it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
These fixes came out of running
basecamp connectend to end against a local Basecamp: first from nothing, then by working in threads with real Claude Code. There are two commits.The connector speaks to the person in the thread
Cards:
https://app.basecamp.com/2914079/buckets/48699913/card_tables/cards/10349162082
https://app.basecamp.com/2914079/buckets/48699913/card_tables/cards/10349584227
https://app.basecamp.com/2914079/buckets/48699913/card_tables/cards/10349174102
Lifecycle notices speak to whoever mentioned the agent, not to the operator. They say what happened and what that person can do. The ids move to a closing
Ref …line, and there are no shell commands.lifecycle_legacy_test.go).No notice when the agent already replied to explain a failure. Mentioning it again won't change a refusal.
connect statusstill shows the failure.Worker preflight.
connect doctorand the connector's startup run the worker the way the connector runs it. They check it starts, that it knows every flag passed (so "too old" gets caught), and that it's logged in (claude auth status, which makes no model call).The dispatcher pauses after two start failures in a row. It rechecks every minute and backs off to an hour. A task that never started says the computer needs checking. It doesn't suggest mentioning the agent again, because that would fail the same way.
Guided
basecamp connect setupCards:
Plain
basecamp connect setupis now a resumable conversation:/my/profile.json(boss) and makes them the operator. For a personal agent that's the only choice.basecamp connect -P <profile>). It doesn't offer the background service, which isn't ready: it would run the agent in the home directory. Card: https://app.basecamp.com/2914079/buckets/48699913/card_tables/cards/10356300850When everything passes it prints one line, instead of the ids and paths.
Setup with flags,
doctorand--jsonare unchanged.connect serviceitself isn't changed here. Hiding it until it's ready is a separate PR.auth agent connectand guided setup share one connection routine (connectAgentProfile).This depends on the Basecamp side for
bosson the agent's profile and for agents listing their own projects: basecamp/bc3#13514. Against a Basecamp without them, guided setup stops and says which flag to pass (--operator-profileor--operator, and--serve <project-id>) instead of guessing.Rebased on
mainafter #783, #784 and #785 (one conflict inconnect_service.go, resolved by keeping both changes). Checked end to end against a local Basecamp at today'smaster(which includes bc3 #13514 and #13573):doctor,auth statusandmeon the agent's profile;Changes after an adversarial review
The connector:
connect status, and takes work once a check passes.Preflight probes:
Guided setup:
--accountorBASECAMP_*overrides when it switches to the agent profile.doctordoes.connect.jsonunder its lock before rebuilding.Notices:
Also:
basecamp connect setup -P <profile>).connect setup --helptext describes what guided setup now does.