Skip to content

Workflows: Starlark scripts orchestrate child agent chats, engine half (stacked on #708) - #712

Open
katulevskiy wants to merge 36 commits into
zeronsh:mainfrom
katulevskiy:workflows-engine
Open

katulevskiy wants to merge 36 commits into
zeronsh:mainfrom
katulevskiy:workflows-engine

Conversation

@katulevskiy

@katulevskiy katulevskiy commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #708 (goal mode, which is stacked on #707). Until those merge, this PR's diff also contains their commits; review git log goal-mode..workflows-engine (24 commits).

Summary

Dynamic workflows, the engine half. The agent writes a Starlark script that orchestrates many child agent chats (phases, parallel fan-out, typed results, shell gates, reports, artifacts). The user approves it, the engine runs it in the background with a journal, resume and throttling, and the result is delivered to the parent chat. State, events, RPCs and MCP tools are designed for the workflow card and run pane (next PR), mobile, and saved workflows.

This PR adds no card UI: the transcript shows approval, started/completed/stopped marker rows and the result message only.

Screenshots

Captured from the real desktop app against the mock harness (ZERON_MOCK_WORKFLOW=1; the child agents are scripted, not real models). The approval prompt is the stock question UI, and the dedicated dialog comes in the card PR.

Approval: phases, agents, literal commands, limits, script

Approval question

Running

Workflow started marker

Completed, and the result delivered to the parent chat

Workflow completed

Stopped by a provider error (resumable)

Workflow stopped

Design

  • zeron-workflow (new, pure crate): Starlark runtime (starlark pinned, pure Rust: blake3 uses its pure feature so aarch64 needs no C compiler), static analysis with path:line:col diagnostics producing a WorkflowGraph, and an idempotent pure reducer. No tokio. Wire and state types live in zeron-proto so clients never link the interpreter (checked: no starlark under zeron-mobile, zeron-client, zeron-doc).
  • Language: scripts define main(args). Hermetic (no clock, random, env or IO beyond the API), no while, no recursion, so scripts always terminate. agent()/.ask() return handles dispatched immediately; .result() joins; failures are values (ok/error), not exceptions. run(cmd, args) needs a literal command (shown at approval and re-checked at run time). Also phase, pmap, files/git read-only queries, report, artifact.*, and schema.* helpers for typed results.
  • Replay: host effects are journaled by (call site, ordinal), so a re-run answers from the journal; resume pins the script hash and checks input hashes.
  • Ask layer extended (from Goal mode: /goal keeps a chat working until a verifier chat confirms (stacked on #707) #708): persistent actors (hidden child chats reused across asks) and an actor-only escalate tool. A workflow may mix harnesses per actor.
  • Scheduler: per-actor FIFO, a global cap, a per-model gate moved by an AIMD governor, provider-fault redrive with backoff, deterministic provider errors stop the run (resumable), stall notice, budgets.
  • State and sync: event-sourced run state projected into meta.workflowRuns (one small Loro map entry per run/actor/node, host-only writer, coalesced to one write per 250 ms). Full results stay in a device-local JSONL journal. Measured on a 200-node, 20-actor run with instant agents: 1805 events, 6 doc writes, ~107 KB of deltas.
  • Approval: an ordinary input question on the live turn carrying the graph in an additive meta field. Scripts with problems are rejected before anything is created.
  • Delivery: exactly one machine-origin result message in the parent chat through the normal queue. Goals defer verification while a workflow of that chat runs.
  • MCP: workflow_guide, start_workflow, get_workflow_run, list_workflow_runs, stop_workflow_run, resume_workflow_run, resolve_workflow_question; actor-only submit_result and escalate. The authoring guide (docs/workflow-guide.md) is embedded and its worked example is executed in a test.
  • Security: hermetic interpreter, literal commands, user approval by default, cwd confinement, safe-alphabet ids, escaped untrusted text, children never widen permissions.
  • Prompt text adapted from ZCode (Apache-2.0); attribution in THIRD_PARTY_NOTICES.md. Details: docs/workflows.md.

Testing

  • zeron-workflow 60 (reducer, analysis, runtime limits and determinism, the guide's example against a fake host), zeron-proto 74, zeron-mcp 36, zeron-doc workflow tests.
  • zeron-engine: lib 429; workflows 20 (scheduling, journal/resume, approval, delivery, restart, budgets, escalation, goal deferral, artifacts); workflow_e2e 2 (real child chats and the real approval question); ask_child 16; goal_mode 15.
  • cargo check --workspace --all-targets passes; cargo check -p zeron-workflow --target aarch64-unknown-linux-musl passes; clippy clean with -D warnings on the new crate.

Known limitations / deviations

  • agent(isolate="worktree") is not built; parallel writers share the workspace (documented, with guidance in the authoring guide).
  • start_workflow blocks until the user answers; a harness that times out the tool call (e.g. Codex's default) denies the run. Documented.
  • A resume creates a new run (resumedFrom) that replays the old journal.
  • Escalation waits poll (at most ~45 s per call).

Not verified

  • Windows, macOS, iOS and Android builds (only the aarch64-musl check of the new crates was run).
  • A run with a real model and real child harnesses (Claude/Codex): the demo uses scripted agents.
  • Pre-existing, unrelated failure: m5_repos_diffs_terminals::diff_capture_tracked_untracked_and_checksum fails on the base branch too.

Claude TodoWrite, OpenCode, ACP plans and Codex update_plan expose an
in-progress step that TodoItem flattened into done/not-done. Add a serde-
defaulted status (written only for in-progress, so existing docs are byte-
identical and old builds/edge ignore it; unknown values derive from done),
map it in every normalizer, and wire Codex turn/plan/updated into the live
plan chip.
The agent's checklist is now a tray stacked above the queue/composer: a
collapsed header (Todo 2/5 plus the current item), an expanded list with
check / active / pending glyphs, and a 3-item focus window with N earlier /
N later fold rows for lists over 6. Open state is remembered per chat for the
app run, a finished list tidies itself to a compact done state once the turn
is idle, and every control is keyboard reachable with tooltips. The
transcript chip's detail marks the in-progress item [~].
ZERON_MOCK_TODO walks an 8-item list through every panel state for live
checks; docs/todo-panel.md records the data model, behaviour and limits.
Child ask (engine::ask): hidden child chat under the requester, run-scoped
submit_result validated against a caller-supplied JSON Schema (boon), repair
rounds, nudge, typed failures, cancellation and archive-on-exit, plus a
FakeAsk backend. Goal mode: additive session-doc goal state, goal commands on
the command plane, machine-origin queue/message origin, pure state machine and
prompts, doc-host controller with restart recovery, and tests.
The expanded goal list was cut off by the todo tray tucked over it. Every
expanded list now keeps a bottom clearance equal to the tuck overlap (the todo
tray too, a one-line change in todo_panel.rs), the goal list's height share
shrinks as trays stack below it, and it follows the newest round. Retook the
goal-mode screenshots and added a stacked-trays one.
The in-process verifier entry was dropped when the verifier returned, but the
verdict was applied only after taking the controller lock. A tick (turn end,
queue watcher, status change) in that window saw 'verifying' with no verifier
and judged the same round again. Reserve the entry under the lock together with
the Verifying ledger write, release it under the lock after the verdict is
applied, and make a second start for the same (goal, round) a no-op.
@katulevskiy

Copy link
Copy Markdown
Contributor Author

Force-pushed: rebased the stack onto current main (80b946b, v0.2.101). CI was failing on the stacked PRs because the merged Cargo.lock had zeron-workflow at 0.2.100 while the workspace is now 0.2.101, so every --locked job stopped at once. The lockfile now differs from main only by our own packages. Also adapted to upstream's new zeron-voice crate and SettingsSection::Voice (lock and shell conflicts), and the dictation test helper for the new UserInputQuestion.meta field. Verified locally with --locked: zeron-workflow 60, workflows 20, workflow_e2e 2, goal_mode 16, ask_child 16.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant