Workflows: Starlark scripts orchestrate child agent chats, engine half (stacked on #708) - #712
Open
katulevskiy wants to merge 36 commits into
Open
katulevskiy wants to merge 36 commits into
katulevskiy wants to merge 36 commits into
Conversation
katulevskiy
force-pushed
the
workflows-engine
branch
from
October 1, 2026 19:04
fb0791d to
8c64132
Compare
Claude TodoWrite, OpenCode, ACP plans and Codex update_plan expose an in-progress step that TodoItem flattened into done/not-done. Add a serde- defaulted status (written only for in-progress, so existing docs are byte- identical and old builds/edge ignore it; unknown values derive from done), map it in every normalizer, and wire Codex turn/plan/updated into the live plan chip.
The agent's checklist is now a tray stacked above the queue/composer: a collapsed header (Todo 2/5 plus the current item), an expanded list with check / active / pending glyphs, and a 3-item focus window with N earlier / N later fold rows for lists over 6. Open state is remembered per chat for the app run, a finished list tidies itself to a compact done state once the turn is idle, and every control is keyboard reachable with tooltips. The transcript chip's detail marks the in-progress item [~].
ZERON_MOCK_TODO walks an 8-item list through every panel state for live checks; docs/todo-panel.md records the data model, behaviour and limits.
Child ask (engine::ask): hidden child chat under the requester, run-scoped submit_result validated against a caller-supplied JSON Schema (boon), repair rounds, nudge, typed failures, cancellation and archive-on-exit, plus a FakeAsk backend. Goal mode: additive session-doc goal state, goal commands on the command plane, machine-origin queue/message origin, pure state machine and prompts, doc-host controller with restart recovery, and tests.
The expanded goal list was cut off by the todo tray tucked over it. Every expanded list now keeps a bottom clearance equal to the tuck overlap (the todo tray too, a one-line change in todo_panel.rs), the goal list's height share shrinks as trays stack below it, and it follows the newest round. Retook the goal-mode screenshots and added a stacked-trays one.
The in-process verifier entry was dropped when the verifier returned, but the verdict was applied only after taking the controller lock. A tick (turn end, queue watcher, status change) in that window saw 'verifying' with no verifier and judged the same round again. Reserve the entry under the lock together with the Verifying ledger write, release it under the lock after the verdict is applied, and make a second start for the same (goal, round) a no-op.
…erInputQuestion.meta
…ysis and host seam
…PC and command plane
… reason in the end marker
…g; steadier tests; demo host ignores workflow commands
katulevskiy
force-pushed
the
workflows-engine
branch
from
October 2, 2026 00:04
8c64132 to
c05800f
Compare
Contributor
Author
|
Force-pushed: rebased the stack onto current |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #708 (goal mode, which is stacked on #707). Until those merge, this PR's diff also contains their commits; review
git log goal-mode..workflows-engine(24 commits).Summary
Dynamic workflows, the engine half. The agent writes a Starlark script that orchestrates many child agent chats (phases, parallel fan-out, typed results, shell gates, reports, artifacts). The user approves it, the engine runs it in the background with a journal, resume and throttling, and the result is delivered to the parent chat. State, events, RPCs and MCP tools are designed for the workflow card and run pane (next PR), mobile, and saved workflows.
This PR adds no card UI: the transcript shows approval, started/completed/stopped marker rows and the result message only.
Screenshots
Captured from the real desktop app against the mock harness (
ZERON_MOCK_WORKFLOW=1; the child agents are scripted, not real models). The approval prompt is the stock question UI, and the dedicated dialog comes in the card PR.Approval: phases, agents, literal commands, limits, script
Running
Completed, and the result delivered to the parent chat
Stopped by a provider error (resumable)
Design
zeron-workflow(new, pure crate): Starlark runtime (starlarkpinned, pure Rust:blake3uses itspurefeature so aarch64 needs no C compiler), static analysis withpath:line:coldiagnostics producing aWorkflowGraph, and an idempotent pure reducer. No tokio. Wire and state types live inzeron-protoso clients never link the interpreter (checked: nostarlarkunderzeron-mobile,zeron-client,zeron-doc).main(args). Hermetic (no clock, random, env or IO beyond the API), nowhile, no recursion, so scripts always terminate.agent()/.ask()return handles dispatched immediately;.result()joins; failures are values (ok/error), not exceptions.run(cmd, args)needs a literal command (shown at approval and re-checked at run time). Alsophase,pmap,files/gitread-only queries,report,artifact.*, andschema.*helpers for typed results.(call site, ordinal), so a re-run answers from the journal; resume pins the script hash and checks input hashes.escalatetool. A workflow may mix harnesses per actor.meta.workflowRuns(one small Loro map entry per run/actor/node, host-only writer, coalesced to one write per 250 ms). Full results stay in a device-local JSONL journal. Measured on a 200-node, 20-actor run with instant agents: 1805 events, 6 doc writes, ~107 KB of deltas.metafield. Scripts with problems are rejected before anything is created.workflow_guide,start_workflow,get_workflow_run,list_workflow_runs,stop_workflow_run,resume_workflow_run,resolve_workflow_question; actor-onlysubmit_resultandescalate. The authoring guide (docs/workflow-guide.md) is embedded and its worked example is executed in a test.THIRD_PARTY_NOTICES.md. Details:docs/workflows.md.Testing
zeron-workflow60 (reducer, analysis, runtime limits and determinism, the guide's example against a fake host),zeron-proto74,zeron-mcp36,zeron-docworkflow tests.zeron-engine: lib 429;workflows20 (scheduling, journal/resume, approval, delivery, restart, budgets, escalation, goal deferral, artifacts);workflow_e2e2 (real child chats and the real approval question);ask_child16;goal_mode15.cargo check --workspace --all-targetspasses;cargo check -p zeron-workflow --target aarch64-unknown-linux-muslpasses; clippy clean with-D warningson the new crate.Known limitations / deviations
agent(isolate="worktree")is not built; parallel writers share the workspace (documented, with guidance in the authoring guide).start_workflowblocks until the user answers; a harness that times out the tool call (e.g. Codex's default) denies the run. Documented.resumedFrom) that replays the old journal.Not verified
m5_repos_diffs_terminals::diff_capture_tracked_untracked_and_checksumfails on the base branch too.