Skip to content

Guided connect setup, and notices that speak to the person in the thread - #794

Merged
robzolkos merged 31 commits into
mainfrom
guided-connect-setup
Sep 30, 2026
Merged

robzolkos merged 31 commits into
mainfrom
guided-connect-setup

Conversation

@robzolkos

@robzolkos robzolkos commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

These fixes came out of running basecamp connect end to end against a local Basecamp: first from nothing, then by working in threads with real Claude Code. There are two commits.

The connector speaks to the person in the thread

Cards:

  • https://app.basecamp.com/2914079/buckets/48699913/card_tables/cards/10349162082

  • https://app.basecamp.com/2914079/buckets/48699913/card_tables/cards/10349584227

  • https://app.basecamp.com/2914079/buckets/48699913/card_tables/cards/10349174102

  • Lifecycle notices speak to whoever mentioned the agent, not to the operator. They say what happened and what that person can do. The ids move to a closing Ref … line, and there are no shell commands.

    • The still-running notice is dated ("Working on this as of 14:05 UTC"), because a notice must never claim the present.
    • Notices from earlier versions are still recognised and retracted (lifecycle_legacy_test.go).
  • No notice when the agent already replied to explain a failure. Mentioning it again won't change a refusal. connect status still shows the failure.

  • Worker preflight. connect doctor and the connector's startup run the worker the way the connector runs it. They check it starts, that it knows every flag passed (so "too old" gets caught), and that it's logged in (claude auth status, which makes no model call).

  • The dispatcher pauses after two start failures in a row. It rechecks every minute and backs off to an hour. A task that never started says the computer needs checking. It doesn't suggest mentioning the agent again, because that would fail the same way.

Guided basecamp connect setup

Cards:

Plain basecamp connect setup is now a resumable conversation:

  1. It connects this computer if it isn't already.
  2. It reads the agent's owner from /my/profile.json (boss) and makes them the operator. For a personal agent that's the only choice.
  3. It checks the worker.
  4. It serves every project the agent is in, by name.
  5. It ends by saying how to start the agent: in the folder it should work in, left running (basecamp connect -P <profile>). It doesn't offer the background service, which isn't ready: it would run the agent in the home directory. Card: https://app.basecamp.com/2914079/buckets/48699913/card_tables/cards/10356300850

When everything passes it prints one line, instead of the ids and paths.

Setup with flags, doctor and --json are unchanged. connect service itself isn't changed here. Hiding it until it's ready is a separate PR. auth agent connect and guided setup share one connection routine (connectAgentProfile).

This depends on the Basecamp side for boss on the agent's profile and for agents listing their own projects: basecamp/bc3#13514. Against a Basecamp without them, guided setup stops and says which flag to pass (--operator-profile or --operator, and --serve <project-id>) instead of guessing.

Rebased on main after #783, #784 and #785 (one conflict in connect_service.go, resolved by keeping both changes). Checked end to end against a local Basecamp at today's master (which includes bc3 #13514 and #13573):

  • setup from nothing, with "Create my agent" on the approval page;
  • reconnecting after a disconnect;
  • moving to a second computer;
  • doctor, auth status and me on the agent's profile;
  • someone outside internal testing.

Changes after an adversarial review

The connector:

  • Checks the worker when it starts, not only after two failed requests. A worker that isn't ready holds new work from the start, says why in connect status, and takes work once a check passes.
  • Holds new work from the next launch in the same batch, not the next pass.
  • Counts a session that can't be prepared toward the hold.
  • Numbers each hold, so a late check or reason can't clear or overwrite a newer one.

Preflight probes:

  • They run in a process group of their own, so a timeout ends the launcher's children too.
  • A program whose interpreter is gone is reported as "can't start", not "not on PATH".

Guided setup:

  • Refuses unsupported platforms before connecting, so it never replaces a working computer's secret for a connector that can't run.
  • Keeps an explicit --account or BASECAMP_* overrides when it switches to the agent profile.
  • When resuming, refuses a read-only credential.
  • Checks an ACP worker the way doctor does.
  • Re-reads connect.json under its lock before rebuilding.
  • Says so when no projects are served.

Notices:

  • A follow-up lost by a worker that had already picked up an earlier request reads as unfinished, with the retry. It isn't "I couldn't start on this".
  • Wording change: an unknown outcome is no longer described as unfinished. The worker may have done the work, even replied, and stopped before saying so. It now reads "I stopped and may not have finished this." (or "interrupted" / "ran out of time") and "If I didn't, mention me again to try again.", so nobody repeats work that was already done. A request that failed keeps "Mention me again to try again."

Also:

  • A task launched before the worker broke no longer clears the hold made by failures after it.
  • A probe's children are ended even when the launcher exits first.
  • A probe's cleanup signals only the probe's own group, even after its pid has been reused.
  • A start failure that settles after a later launch worked isn't counted towards a hold.
  • Guided setup's retry instructions name the profile it was working on (basecamp connect setup -P <profile>).
  • Guided setup doesn't save over a connect.json that another setup wrote while it was asking.
  • Guided setup moves a broken connect.json aside only if it's still the file it asked about.
  • A request that succeeded without reporting a reply says "I finished this, but didn't report a reply."
  • Guided setup checks the agent profile's base URL as root does, instead of panicking on an insecure one.
  • The connect setup --help text describes what guided setup now does.

Copilot AI balanced review requested due to automatic review settings September 29, 2026 00:05
@github-actions github-actions Bot added commands CLI command implementations tests Tests (unit and e2e) skills Agent skills auth OAuth authentication labels Sep 29, 2026
@robzolkos
robzolkos marked this pull request as ready for review September 30, 2026 03:27
The connector's automatic notices were written for whoever runs it. They
used terms like "task", "redispatch", "attempt" and "worker", and ended
with a shell command. The person who reads them is the one who mentioned
the agent. They now say plainly what happened and what that person can
do, and keep the ids on a closing Ref line:

- "I'm not set up to work in this project yet, so I haven't started on
  this", instead of a holding notice about trust and projects.
- "Working on this as of 14:05 UTC", dated, because a notice must never
  claim the present.
- For work that stopped: interrupted, out of time, or went wrong. Then
  "Mention me again to try again" when mentioning again could help.
- Nothing, when the agent already replied to explain why it failed.
  Mentioning it again won't change a refusal. connect status still shows
  the failure.

Notices posted by earlier versions are still recognised and retracted.

A worker that can't start (Claude Code missing, too old or logged out)
used to show up only as a notice. Mentioning the agent again gave the
same notice every time. Now:

- connect doctor and the connector's startup try the worker the way the
  connector runs it: it starts, it knows every flag the connector passes,
  and it's logged in.
- After two start failures in a row, the dispatcher stops taking work.
  It rechecks every minute and backs off to an hour.
- A task that never started says so: "I couldn't start on this:
  something's wrong on the computer I run on. The person who runs me
  needs to check it."
Setting up an agent took a separate sign-in as yourself to name the
operator, project ids nobody knows, and output that read like a log.
Plain `basecamp connect setup` is now a guided, resumable conversation:

- It connects this computer to the agent if it isn't already
  ("✓ Connected this computer as Rob Zolkos (agent)").
- It reads the agent's owner from Basecamp and makes them the operator.
  For a personal agent that's the only choice: "Rob Zolkos (you) is the
  only person who can give it work."
- It checks the worker starts ("✓ Claude Code 2.1.283 is ready").
- It serves every project the agent is in, by name. An agent in none
  gets told how to add it, and setup exits cleanly.
- It offers to start the connector whenever you log in. If the service
  can't start, it says what to run instead.
- When everything passes it says so, instead of listing ids and paths.

Setup with flags, doctor, service install and --json keep their
detailed output. `auth agent connect` and guided setup share one
connection routine. The guided connection identifies itself to Basecamp
as "basecamp connect".
Guided setup ended by offering to "start your agent now, and whenever you
log in", which installed the systemd service. The service isn't ready: it
runs the agent in the home directory, where it may change files without
asking; it freezes PATH at install; it is Linux-only; and it has not been
tried on a real machine.

Setup no longer offers it or touches systemd. When the agent isn't
running, the summary says to start it in the folder it should work in,
that it can change files there without asking, and gives the one command.
`connect service` itself is unchanged here.
Two starts in a row that never ran hold new work, but the dispatcher asked
about the hold once a pass. A start that fails at once gives its slot back,
so the rest of the batch still launched: with two slots and five queued
records, all five were started and answered "I couldn't start on this".
Each launch now asks whether work is held.
…ests

The worker's preflight only ran as the re-check of a hold, and a hold took
two failed starts to happen. A connector started with Claude Code logged
out reported itself running and answered the first two requests with "I
couldn't start on this" before it said anything was wrong.

The dispatcher now runs the preflight before its first pass. A worker that
isn't ready holds new work from the start: the connector keeps listening,
says why in its log and in connect status ("Claude Code isn't ready — …"),
and takes work once a check passes. A preflight that passes, or a driver
without one, changes nothing.
…ofile

Guided setup switches to the agent profile when it isn't already active,
and applying a profile overwrites the account and base URL with the
profile's own. Unlike root, it didn't re-apply the environment and flags
afterwards, so `connect setup --account 777` against a profile bound to
999 quietly worked in 999 instead of refusing, as it does with -P. The
overrides are re-applied, in root's order.
A setup being resumed kept its connect.json and skipped setup's checks,
scope among them. A profile reconnected read-only since then (auth agent
connect --scope read) was reported as set up and ready to mention, though
it couldn't reply or acknowledge. The resumed setup now runs setup's scope
check and refuses such a credential with the way to reconnect.
Guided setup checked the worker through its spawn preflight, and the acp
driver has none, so a setup on the acp driver skipped the worker check
altogether: with its pinned adapter missing, setup still said the agent
was set up. It now falls back to doctor's check of where the driver finds
its worker, and stops on a failure with its fix.
When a second failed start held new work, the dispatcher asked the
worker's preflight why with the lock released. A start that worked in the
meantime cleared the hold and recorded the connector running, and then the
late reason recorded it as not taking work, which stayed in connect status
while work was being taken. The connection's status is now written under a
lock of its own, and each writer asks the hold again first.
A request was classed as never started from its own record alone: not
picked up, outcome unknown, and the attempt failed. In a conversation
where the worker had already picked up and answered an earlier request,
a follow-up it failed to get to was reported as "I couldn't start on
this: something's wrong on the computer I run on", with no suggestion to
mention the agent again. A request only never started now if nothing in
its settlement was picked up; the follow-up reads as unfinished, with
the retry.
…ilding, and say when nothing is served

- On a platform the connector doesn't run on, guided setup went ahead and
  connected, replacing the secret another computer might be running the
  agent on, then gave a command that refuses to start. It now refuses
  first, as the connector does.
- Rebuilding a connect.json that didn't fit moved the file aside after the
  question, without reading it again: a file another setup had repaired
  meanwhile was moved aside. It's read again under the lock, and a file
  that now fits is resumed.
- A setup serving no projects printed an empty "Works in:" and suggested
  mentioning the agent "in one of those projects". It now says it works
  in no projects and how to add one.
A start whose session couldn't even be prepared (its directory, its token
socket) settled as a start that ran nothing, like a driver's refusal, but
wasn't counted toward the hold. With the sessions directory unusable, every
queued request was started, failed, retried and blocked. It now counts, and
two in a row hold new work.
A held dispatcher checks its worker in the background. If a start that
worked cleared that hold while the check ran, and two failures then made a
new one, the old check coming back "fine" cleared the new hold, and new
work went to a worker that had just failed twice. Holds are now numbered as
they're made, and a check, or a reason, for a hold that is no longer the
one standing decides nothing.
…ken program missing

- A probe that timed out killed only the program it ran. A launcher that
  starts the real worker as a child and waits left that child running, and
  repeated checks piled them up. Probes now run as a process group of their
  own, and the group is ended when the probe's context is (Unix; off Unix
  the connector runs no workers).
- A program on PATH whose interpreter or loader was gone fails its exec with
  ENOENT, the same error as a missing program, and was reported as not on
  PATH. It's now missing only when the program itself isn't there; otherwise
  it's a program that can't start, with the error.
Any settlement that picked its request up cleared the start-failure hold.
A task launched before the worker broke, and finishing after two later
launches had failed and held new work, cleared that hold, and new work
went to a worker that was still broken. Launches are now numbered in the
order they're made; a start that worked clears only failures from launches
made at or after it.
Ending the probe's process group only on timeout missed a launcher that
starts a child, prints its answer and exits: the probe returned and the
child ran on, one more with every recheck. The probe's group is now ended
whenever a probe returns.
…as it is

- Root checks the base URL of the profile it starts with; guided setup then
  switched to the agent profile without checking its, and an insecure
  base_url there panicked in the SDK. It's now checked as root checks it,
  and refused with where to correct it.
- The setup help still said guided setup "offers to keep the connector
  running" and makes "a connect.json that cannot be used again". It now
  says what it does: writes connect.json, offers to redo one that can't be
  used, and ends with how to start the connector in its folder.
A request the worker never reported on has an unknown outcome, not an
unfinished one: the worker may have done the work, and even replied, then
stopped before it said so. The notice said "I stopped before I finished
this. Mention me again to try again", and following it could do the work
twice.

For an unknown outcome the notice now says it "may not have finished
this" (stopped, interrupted, or ran out of time) and "If I didn't, mention
me again to try again", so the reader looks first. A request that failed
keeps "Mention me again to try again".
Sessions run with --setting-sources "", but the login probe didn't, so
a user setting that blanks ANTHROPIC_API_KEY made preflight report a
logged-out worker, and hold all work, though the worker itself would
run. The probe now asks `claude --setting-sources "" auth status`; the
option goes before the subcommand, which doesn't take it.
The start record kept only its earliest failure, so a task launched
between two failures, still running when the second held work, cleared
the hold when it finished, and queued work went to the broken worker.
The record now keeps each failure's launch: a start that worked drops
the failures launched before it, and leaves any launched after it, and
their hold, standing.
release freed the worker's slot before the failure was noted, so a
dispatch pass between the two saw capacity and no hold, and could launch
a third request into a worker that had just failed twice. Whether the
worker started is now noted inside release, while its slot is still
taken.
The probe's cleanup ran after it had been reaped and signaled its pid's
group unconditionally, so a pid the kernel had since given to another
group leader would have had that group killed. The probe's identity is
now recorded, by the kernel's start time, while it runs, and cleanup
signals through signalRecordedGroup, which leaves a reused pid alone and
still ends children a launcher left behind.
Guided setup reads the agent, asks about its projects, then hands the
answers to setup, which reads the credential again. If the profile was
connected to another agent while a question was open, setup saved the
answers for that agent while the summary named the first. Guided setup
now passes the agent and account it showed, and setup refuses anything
else.
Guided setup read the lock's metadata and called any live pid running,
but the metadata outlives a crash and its pid can be reused, so it could
say "Running: yes" and leave out how to start the agent. The holder now
counts only if its process started before the lock was taken; one the
kernel gave the pid to later started after. The lock itself is never
taken, so a starting connector can't find it held.
…le it could

The context watcher can fire after Wait has reaped the probe, and a launcher can
exit before it is recorded. Kill a running leader only while its start time
matches; once the leader is gone, kill only leftovers of a group whose id nobody
holds. A probe that never started has nothing to end.
The worker has started since it launched, so it neither holds work nor counts
towards a hold.
A bare `connect setup` may pick another profile the second time.
Another setup's file may trust more people than guided setup told the person;
it is left as that setup wrote it.
A file another command replaced meanwhile may be a finished setup for another
agent; it is left in place.
@robzolkos
robzolkos merged commit 61adfbf into main Sep 30, 2026
31 checks passed
@robzolkos
robzolkos deleted the guided-connect-setup branch September 30, 2026 19:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

auth OAuth authentication commands CLI command implementations skills Agent skills tests Tests (unit and e2e)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant