Skip to content

MCP: agents spawn real top-level chats or side chats on any device - #7

Draft
katulevskiy wants to merge 15 commits into
mainfrom
agent-spawned-chats
Draft

katulevskiy wants to merge 15 commits into
mainfrom
agent-spawned-chats

Conversation

@katulevskiy

Copy link
Copy Markdown
Owner

Mirror of zeronsh#671, opened on this fork to run CI. Not for merging.

Summary

An agent in a Zeron chat can now create a real top-level chat, one that
appears in the Sessions sidebar on every device just like a chat the user
created. It can still create a side chat (today's behaviour). Either kind
can run on any online device connected to the account, and the chosen
device's engine runs it.

  • MCP create_chat / create_chats take kind: side (placement under
    the parent, hidden from the sidebar; still the default inside a chat, so
    existing callers behave the same) or chat (top-level, listed normally).
  • Provenance without hiding. A new additive, serde-defaulted field,
    Chat.spawnedByChatId, records whose agent created a chat. Both kinds set
    it. parentChatId keeps meaning "placement". list_chats { spawned_by }
    lists every chat a chat spawned. Summaries carry kind, parentChatId and
    spawnedByChatId.
  • Device targeting. Pass device (id or name) and/or project (on any
    device). Each host is checked before anything is written: it must be an
    execution host (an engine) and online, and errors list the valid hosts or
    that device's projects. Harness and model are checked against the host's
    own catalog through targetDeviceId. list_devices reports online,
    executionHost and each device's projects. list_harnesses and
    list_models take a device. worktree: true has the host run the first
    turn in a fresh isolated worktree.
  • Running them across devices. The first prompt, send_message,
    wait_for_turn, read_chat, interrupt_chat and respond_to_input use
    the same synced doc/session paths as the desktop composer. This was
    verified end to end on two engines.
  • Loop guards. A top-level spawned chat may spawn more chats (a side chat
    still cannot), within these limits: provenance depth MAX_SPAWN_DEPTH = 3
    (checked in the MCP and re-checked by the engine on createChat; a looping
    chain counts as over the limit), at most 32 live spawned chats per spawner,
    and at most 32 creations per minute per MCP server.
  • Desktop. Spawned top-level chats get a subtle "Spawned by " mark in
    the sidebar (tooltip; clicking it opens the spawner). The spawner's
    transcript shows a compact card for each chat it created, and the explorer
    footer's Chats section lists them too. Tool outputs never enter the
    doc, so the fold keeps just the created chat ids on the create call
    (additive createdChatIds), which makes the link exact on every device.
    read_chat shows them as [created: …]. See Screenshots.
  • Mobile core. SessionRow carries spawnedByChatId and
    spawnedByTitle. iOS and Android list these chats normally. Swift
    bindings are regenerated, and the only change is the two new fields.
  • Docs. docs/mcp.md covers kinds, provenance, who may create chats,
    device targeting and orchestration examples, with a mermaid diagram.

Motivation

Until now, create_chat always recorded the calling chat as parentChatId,
which hides the new chat from the sidebar. That suits a quick sub-task. It
does not suit work the user should own: a long training run on a GPU box, or
a feature branch that a worker takes over. In those cases the agent should be
able to start "a chat as if the user had created it", on whichever device
has the hardware, repo or harness for the job, and the user should still see
which chat started it.

flowchart LR
    U[User] -->|creates| C["Coordinator chat<br/>laptop (depth 0)"]
    C -->|"create_chat kind: chat<br/>device: GPU box"| W["Top-level chat<br/>GPU box (depth 1)<br/>in every sidebar,<br/>Spawned by Coordinator"]
    C -->|"create_chat kind: side<br/>device: GPU box"| S["Side chat<br/>GPU box<br/>under Coordinator"]
    W -->|"create_chat (either kind)"| W2["depth 2"]
    W2 --> W3["depth 3: runs,<br/>cannot spawn"]
    classDef side fill:#eee,stroke:#999
    class S side
Loading

Screenshots

The review fixture (crates/ui/examples/agent-chats-fixture.rs) is rendered on a temporary headless Hyprland output. A coordinator chat's create_chats call made one top-level chat on "GPU box" and one side chat.

Sidebar: the spawned top-level chat lists under Sessions with an agent glyph. Its tooltip names the spawner, and clicking it opens the spawner.

Sidebar with a spawned chat

Spawner transcript: the create call renders as one compact card per created chat, showing kind, live title, host device and live status. Clicking a card opens that chat.

Spawner transcript cards

Overview (sidebar mark plus cards):

Overview

Light variants: docs/media/agent-chats/*-light.png.

Testing

Unit and integration tests:

  • zeron-mcp (25 tests): both kinds with placement and provenance; kind/parent validation; device targeting (project on a remote device, a device by name, device+project scoping, and a same-named project resolving to this device); the host's harness/model catalog through targetDeviceId; offline, phone-viewer and unknown targets refused with valid hosts listed; list_devices/list_projects/list_harnesses { device }; the depth guard (a spawned top-level chat spawns until depth 3; a spawned side chat never); list_chats { spawned_by } vs { parent }; the live fan-out cap under a concurrent batch; the per-server rate limit; worktree; explicit mock; a remote reply that syncs after the session row; [created: …] in read_chat.
  • zeron-engine: create_chat_records_provenance_and_guards_spawn_depth (Mutate createChat records both links, the engine-side depth backstop, WatchChats carries the field, forks drop provenance). Also passing: lib (370), side_chats, workspace_sync, projectless, device_routing and local_first.
  • zeron-doc: registry sync, update, restart and legacy-row serde for spawnedByChatId; the fold keeps only created ids (every naming scheme, but not for errors or other servers); createdChatIds survives the streaming writer and snapshots.
  • zeron-proto: additive serde for spawnedByChatId; spawn_depth (chains, dangling ids, loops); create-call detection and result parsing.
  • zeron-client: spawned top-level chats list normally with spawnedByTitle; spawned side chats stay children.
  • zeron-ui lib (1382 tests): visible_chats, the spawned-by label and target helper, the explorer Chats rows, card matching (kept ids preferred; timing fallback for old transcripts), and transcript grouping.
  • cargo check --workspace --all-targets is clean. Clippy reports no new warnings on the lines this PR touches (the crates' existing warnings are untouched).

Two devices on real processes (scripts/e2e-agent-chats.sh: a wrangler dev edge in AUTH_MODE=dev plus two headless engines, driven through real zeron mcp stdio servers). All steps passed:

PASS: both engines list the other as an online execution host
PASS: list_harnesses { device: B } answers from B through targetDeviceId
PASS: agent on A spawned a top-level chat on B; B's engine ran the prompt
PASS: agent on A spawned a side chat on B; B's engine ran the prompt
PASS: engine A: list_chats {spawned_by} + read_chat see both spawned chats
PASS: engine B: list_chats {spawned_by} + read_chat see both spawned chats
PASS: send_message … wait round-trips to B, attributed to the coordinator
PASS: interrupt_chat queues on the remote chat; wait_for_turn reads B's session row
PASS: invalid targets and kind/parent combinations are refused with guidance
PASS: spawned top-level chats spawn further chats across devices until depth 3
PASS: a real opencode agent (opencode/big-pickle) on A spawned a top-level chat on B that ran
PASS: agent-spawned chats across two devices

With ZERON_E2E_AGENT_MODEL=opencode/big-pickle (a free OpenCode Zen model), a real agent in a chat on A called the engine-injected zeron_create_chat itself. Its transcript on A reads zeron_create_chat [created: <id>], and the new chat ran on B with kind: chat and spawnedByChatId = the agent's chat.

The first e2e run found two gaps, both fixed here:

  • A completed wait on a remote chat could return an empty reply, because the idle session row synced before the chat doc.
  • A large create_chat future overflowed a 2 MB stack in a feature-unified test build.

Timing-sensitive tests under load (not caused by this PR):

  • local_first: the headless stop and sign-out tests failed twice on "headless IPC did not start" (a 5 s boot budget) at load average 48. Rerun at load 16, they pass 11/11.
  • In the UI lib, one flaky test per full run, a different one each time; each passes alone. Final run: 1382/1382.

Not verified: iOS/Android app UI (only the core view model and the regenerated Swift bindings were checked); clicks in a live desktop app (the screenshots come from the fixture); harnesses other than OpenCode and the mock making the call.

Compatibility

  • spawnedByChatId is additive and serde-defaulted on the proto, the
    registry/workspace rows and the Mutate createChat params. Older rows read
    as user-created. An older engine ignores the field, so the chat just looks
    user-made there. It is never written when unset.
  • The create_chat default is unchanged inside a chat (side, with the
    caller as parent). Without a calling chat and without a parent, the
    default is chat, which matches today's result (no parent). The only new
    refusals: parent together with kind: "chat"; kind: "side" with nothing
    to hang it under; offline or phone-viewer target devices, which were
    previously accepted and never ran; and the depth, fan-out and rate limits.
  • Forks never inherit provenance.
  • createdChatIds on tool parts is additive in the session doc. Older
    readers ignore it, and transcripts written without it fall back to
    timing-based matching on desktop.
  • Mobile FFI: SessionRow gains two optional fields (Swift bindings are
    regenerated and committed).

Overlaps


Chat.spawnedByChatId is an additive, serde-defaulted field carried on the
registry and workspace chat rows next to parentChatId. parentChatId stays
placement (a side chat, hidden from the sidebar); spawnedByChatId is
provenance (whose agent created the chat), so an agent-spawned top-level
chat can list in the sidebar like a user-created one and still say who
started it.

The engine's Mutate createChat accepts spawnedByChatId
(WorkspaceHost::create_chat_linked) and refuses provenance chains deeper
than zeron_proto::MAX_SPAWN_DEPTH as a backstop to the MCP server's own
check. A fork is the user's own side chat and never inherits provenance.
create_chat/create_chats gain kind: side (default inside a chat, backward
compatible) hangs the chat under its parent; chat makes a real top-level
chat that shows in the Sessions sidebar on every device. Both record the
calling chat as spawnedByChatId, and list_chats { spawned_by } finds them.

Targets are validated up front: project and/or device resolve to an online
execution host (engines stamp capabilities; phone viewers and stale
heartbeats are refused with the valid hosts listed), with harness and model
checked on that host through targetDeviceId. list_devices reports online,
executionHost and each device's projects; list_harnesses/list_models take a
device. worktree: true asks the host to run the first turn in a fresh
isolated worktree.

Spawned top-level chats may spawn further chats, bounded by provenance depth
(MAX_SPAWN_DEPTH), a live fan-out cap per spawner and a per-server creation
rate.
SessionRow gains spawnedByChatId and spawnedByTitle (the spawner's display
title, None when its row is gone), so iOS and Android list agent-spawned
top-level chats like any session and can say who started them. Spawned
side chats stay under their parent's children. Swift bindings regenerated;
the only change is the two new record fields.
A chat an agent created with create_chat { kind: "chat" } is a real
top-level session, so it lists in Sessions like one the user started.
Its row now carries a small agent glyph after the title whose tooltip
names the spawner ("Spawned by <title>", or "Spawned by an agent" once
the spawner is gone) and whose click opens the spawner. Side chats carry
no marker: their placement under the parent already says where they
came from.

The label/target resolution is a pure helper (spawned_by_affordance)
with unit tests, and visible_chats gains a test pinning that provenance
alone never hides a chat while placement still does.
The footer's Chats section listed only side chats (parentChatId), so a
coordinator that fanned work out with create_chat { kind: "chat" } had
no place to find its workers besides scanning the sidebar. It now also
lists the top-level chats whose spawnedByChatId is this chat, merged in
the same most-recent-first order; the fingerprint hashes the new
top-level flag. Opening such a row selects it in the main view, the way
its own sidebar row does, while side chats keep docking beside the
parent.
For a chat hosted on another device the idle session row and the chat doc
sync separately, so a create_chat/send_message wait could settle before the
reply reached this engine's replica and report an empty reply. After a
completed send, await_turn now re-reads the transcript for up to 20s until
the new assistant message has landed.

An explicit harness: mock is honoured when the host has the test rig (it is
never a default and stays hidden from pickers), which lets the new
scripts/e2e-agent-chats.sh drive real zeron mcp servers against two headless
engines and a wrangler dev edge: an agent on A spawns a top-level chat and a
side chat on B, both engines agree through list_chats/read_chat, follow-ups,
interrupts, validation and the depth guard work across devices.
A Zeron MCP create_chat/create_chats call in the spawner's transcript
now renders as the spawn chip's card, one per created chat: "Chat" or
"Side chat", the live title, the host device when it is not this one,
and the sidebar's status glyph and word. Clicking a card opens the chat
the way its home surface does: a top-level chat is selected like its
sidebar row, a side chat docks beside its parent like a footer row.
While the call runs (or if it failed) the ordinary chip stays.

Detection is tolerant of every harness's MCP naming: Mcp { server,
tool } from Claude (mcp__zeron__...), Pi (zeron_...), Codex and Cursor,
and Unknown names such as zeron_create_chat, zeron/create_chat or
"create_chat (zeron MCP Server)" from OpenCode and ACP. The created ids
are read from the tool result (chatId, or results[].result.chatId for
the batch) when the transcript carries it. The doc fold keeps tool
outputs out of synced docs, so the usual path recovers them from the
chats' own provenance instead: every chat whose spawnedByChatId is this
chat belongs to the last message that started before it was created,
split across that message's calls in order. Ambiguous cases (two
batches in one message) stay on the plain chip rather than guess.

subagent_chip's card is factored into link_chip so both chips share it.
ZERON_E2E_AGENT_MODEL (e.g. opencode/big-pickle, a free OpenCode Zen model)
runs a real harness in a chat on engine A with the engine-injected Zeron MCP
server; the model itself calls create_chat to start a top-level chat on B,
which B's engine runs.
docs/mcp.md now covers kind side vs chat, the parentChatId/spawnedByChatId
split, who may create chats (depth, fan-out and rate limits), running chats
on any online execution host, orchestration examples with a sequence
diagram, and the two-device e2e recipe.
create_chat now resolves the host, reads its catalog, delivers and can wait
inline; in a feature-unified test build that future overflowed a 2 MB
test-thread stack. Boxing the create/send (and batch) futures in the
dispatcher keeps every tool call's stack small, in tests and in zeron mcp.
agent-chats-fixture renders the real shell over a demo state: a
coordinator whose transcript holds a Zeron create_chats call that made a
top-level chat on another device ("GPU box") and a side chat, with the
spawned chat in the sidebar. The tool call carries no output, like a
synced doc, so the cards resolve through provenance.
ZERON_AGENT_CHATS_HOVER hovers a window point so the sidebar tooltip can
be captured; ZERON_PALETTE_LIGHT selects the light appearance.

docs/media/agent-chats holds the sidebar marker (with its tooltip) and
the spawner's transcript cards, dark and light, captured on a temporary
headless Hyprland output.
Tool outputs never enter the session doc, so a spawner's transcript could
only guess which chats its create_chat calls made (by creation time), and
parallel calls in one message could swap. The fold now keeps just the
created chat ids from a successful Zeron create_chat/create_chats result on
the tool part (additive createdChatIds; the output itself stays out), and
render_parts carries them into the doc.

Detection and result parsing move to zeron_proto::created_chats so the doc
fold and the desktop share one definition. The desktop card prefers the
kept ids and falls back to timing only for older transcripts; read_chat
shows them as [created: …] so an orchestrator can follow its chats.
The agent-chats fixture's create_chats part now carries the created chat
ids the fold keeps, so the screenshots exercise the exact link path rather
than the timing fallback. overview.png shows the sidebar mark and the
spawner's cards together.
@contribution-manager-katulevskiy contribution-manager-katulevskiy Bot added the review:needs-triage Managed by Contribution Manager label Oct 6, 2026
@contribution-manager-katulevskiy

Copy link
Copy Markdown

Contribution Manager

needs triage · medium risk · untriaged

Open review workspace · Revision a0573e84

Next actions:

  • Direction decision needed
  • Scope, risk, and platform triage needed
  • Author has marked this PR as draft
  • Independent code review needed
  • Required check: core-tests (missing)

This summary is maintained by the review app. Decisions and test evidence are recorded in the workspace; the app does not merge PRs.

@contribution-manager-katulevskiy contribution-manager-katulevskiy Bot added cm:inbox Contribution Manager: waiting for triage and removed review:needs-triage Managed by Contribution Manager labels Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cm:inbox Contribution Manager: waiting for triage

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant