Replies: 8 comments 3 replies
|
The hard part is that replay needs more than observability. Observability tells you what happened; deterministic replay needs to preserve the inputs that made it happen. I would split this into four layers:
{
"step_id": "tool-14",
"tool": "github.get_issue",
"input_hash": "sha256:...",
"output_hash": "sha256:...",
"recorded_at": "..."
}
The key mental model: replaying an agent is closer to replaying a distributed workflow than rerunning a pure function. You usually cannot make the world deterministic, so you record the world-facing boundaries and make replay deterministic up to those boundaries. |
|
I would separate "replay the actions" from "certify the final claims." Observability is usually enough for the first question, but long-running agents need a stronger boundary record for the second. A minimal receipt per step could be: {
"run_id": "...",
"step_id": "tool-14",
"parent_step_id": "agent-3",
"kind": "tool_call | state_checkpoint | claim",
"input_canonical_hash": "sha256:...",
"output_ref": "artifact://...",
"output_hash": "sha256:...",
"state_checkpoint_ref": "snapshot://...",
"policy_or_prompt_version": "...",
"verdict": "match | drift | unverifiable"
}Then replay can run in three modes:
That last mode matters because a replay can be behaviorally plausible while no longer proving which data was fresh, stale, retried, or missing. For stateful agents I would not try to make the whole world deterministic. I would make the world-facing boundaries content-addressed, then make mismatch reports first-class artifacts. Context: I am building Project Telos around local workflow receipts and claim/source separation for this exact class of agent debugging problem: https://harperz9.github.io/field-guide.html |
|
Hi,
Thank you very much for taking the time to read the project and write such
thoughtful feedback.
I especially appreciated your distinction between replaying execution and
certifying final claims. That separation is something I had not fully
considered, and your explanation of boundary records and claim-support
replay gave me a new perspective.
Your three replay modes (frozen-boundary, live-boundary, and claim-support)
are particularly interesting. While my current work focuses on execution
governance, replay integrity, and runtime containment, I can see how claim
certification could become an important extension in a future version of
the project.
I also took a look at Project Telos. It is encouraging to see others
thinking seriously about reproducible agent workflows and evidence
boundaries. Although our approaches are different, I think we are trying to
solve related problems from complementary directions.
Thank you again for sharing your ideas. Feedback like yours is genuinely
valuable and will influence how I think about future iterations of EGA.
Best regards,
DaeJung Byun
Founder, LCM / EGA
2026년 6월 28일 (일) 오전 4:46, Zain Dana Harper ***@***.***>님이 작성:
… I would separate "replay the actions" from "certify the final claims."
Observability is usually enough for the first question, but long-running
agents need a stronger boundary record for the second.
A minimal receipt per step could be:
{
"run_id": "...",
"step_id": "tool-14",
"parent_step_id": "agent-3",
"kind": "tool_call | state_checkpoint | claim",
"input_canonical_hash": "sha256:...",
"output_ref": "artifact://...",
"output_hash": "sha256:...",
"state_checkpoint_ref": "snapshot://...",
"policy_or_prompt_version": "...",
"verdict": "match | drift | unverifiable"
}
Then replay can run in three modes:
- frozen-boundary: re-feed recorded tool outputs and state snapshots;
fail if the agent asks for different inputs
- live-boundary: call the real tools again, but emit the first
input/output/state hash divergence
- claim-support: map final answer claims to evidence refs; any
unsupported claim becomes unverifiable even if the transcript reads
well
That last mode matters because a replay can be behaviorally plausible
while no longer proving which data was fresh, stale, retried, or missing.
For stateful agents I would not try to make the whole world deterministic.
I would make the world-facing boundaries content-addressed, then make
mismatch reports first-class artifacts.
Context: I am building Project Telos around local workflow receipts and
claim/source separation for this exact class of agent debugging problem:
https://harperz9.github.io/field-guide.html
—
Reply to this email directly, view it on GitHub
<#7695?email_source=notifications&email_token=BQ3VMJRGL46KSXRBG4FPLBT5CEAQLA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCNZUGYYDSNZWUZZGKYLTN5XKMYLVORUG64VFMV3GK3TUVRTG633UMVZF6Y3MNFRWW#discussioncomment-17460976>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/BQ3VMJULPXL3KW6A4XVQ65T5CEAQLAVCNFSNUABIKJSXA33TNF2G64TZHM3DQMBRGIYDANZRHNCGS43DOVZXG2LPNY5TCMBQGY3TCNZUUF3AE>
.
Triage notifications, keep track of coding agent tasks and review pull
requests on the go with GitHub Mobile for iOS
<https://github.com/notifications/mobile/ios/BQ3VMJTGEUEPD32VYZQQNHL5CEAQLA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCNZUGYYDSNZWUZZGKYLTN5XKMYLVORUG64VFMV3GK3TUVJTG633UMVZF62LPOM>
and Android
<https://github.com/notifications/mobile/android/BQ3VMJTDIGUJ7AUA7EUVQQD5CEAQLA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCNZUGYYDSNZWUZZGKYLTN5XKMYLVORUG64VFMV3GK3TUVZTG633UMVZF6YLOMRZG62LE>.
Download it today!
You are receiving this because you authored the thread.Message ID:
***@***.***>
|
|
Hi Zain,
Thank you again for your thoughtful feedback.
I really appreciate you taking the time to explain your architectural
perspective. Your distinction between execution replay and claim
certification makes a lot of sense, and I think keeping claim certification
as an optional extension is a very clean design.
I’m currently evolving EGA step by step, so discussions like this are
genuinely valuable to me. I’ll definitely keep your ideas in mind as the
project grows.
I hope we can continue exchanging ideas in the future. Thanks again for
your time and insight.
Best regards,
DaeJung Byun
Founder, LCM / EGA
2026년 6월 28일 (일) 오전 6:10, Zain Dana Harper ***@***.***>님이 작성:
… Thanks, DaeJung. I think the complementary split is exactly right:
execution governance should not be forced to become a truth-certification
system.
If claim certification becomes a later extension, I would keep it as an
optional layer that consumes the execution replay artifacts rather than
changing the replay core:
- execution replay answers: did the agent ask for the same boundary
inputs, get the same recorded outputs, and preserve the same state
transitions?
- claim-support replay answers: which final claims are bound to
evidence refs, which are stale, and which are unsupported?
A small first fixture could be deliberately boring:
- one tool/action receipt;
- one final answer with three claims: supported by the receipt,
contradicted by a changed/live-boundary receipt, and unsupported by any
receipt;
- expected verdicts: match, drift, unverifiable.
That would let EGA keep its execution-governance boundary clean while
leaving a clear extension point for audit or certification tools.
Boundary: architecture feedback only; no claim about running or validating
EGA, LCM, AutoGen, Project Telos integration, production readiness,
partnership, or adoption.
—
Reply to this email directly, view it on GitHub
<#7695?email_source=notifications&email_token=BQ3VMJRCAVCCOHKT5X24UYD5CEKMTA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCNZUGYYTINZVUZZGKYLTN5XKMYLVORUG64VFMV3GK3TUVRTG633UMVZF6Y3MNFRWW#discussioncomment-17461475>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/BQ3VMJSNSRDW4A6ROGPDAML5CEKMTAVCNFSNUABIKJSXA33TNF2G64TZHM3DQMBRGIYDANZRHNCGS43DOVZXG2LPNY5TCMBQGY3TCNZUUF3AE>
.
Triage notifications, keep track of coding agent tasks and review pull
requests on the go with GitHub Mobile for iOS
<https://github.com/notifications/mobile/ios/BQ3VMJWDIDR64VJVDKEKFGL5CEKMTA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCNZUGYYTINZVUZZGKYLTN5XKMYLVORUG64VFMV3GK3TUVJTG633UMVZF62LPOM>
and Android
<https://github.com/notifications/mobile/android/BQ3VMJUX5COETSZ6YLOWBID5CEKMTA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCNZUGYYTINZVUZZGKYLTN5XKMYLVORUG64VFMV3GK3TUVZTG633UMVZF6YLOMRZG62LE>.
Download it today!
You are receiving this because you authored the thread.Message ID:
***@***.***>
|
|
Flix Vision is a feature-rich streaming app designed for users who enjoy watching movies, TV series, anime, and documentaries. The application includes an intuitive interface, organized media library, and responsive navigation that allows users to quickly discover and enjoy a wide variety of digital entertainment from one platform. |
|
Appliance Repair Ventura Pro has built a reputation for honest, efficient appliance repair in the Ventura area. Each job starts with a thorough inspection so there are no surprises once work begins. We treat every appliance repair with the same level of care, regardless of the size of the job. the Ventura area customers can schedule appliance repair whenever it's convenient. Visit: appliancerepairventurapro |
|
One class of replay failure that isn't in your list and cost me more than the state drift Two from the last week, both in code whose only job was inspecting a result:
So alongside pinning API responses and filesystem state, the two things that made replays
On the parts you did list, the one that helped most was recording the terminal status of https://github.com/soul-sol/agent-watch (MIT) is that piece, if useful. |
Uh oh!
There was an error while loading. Please reload this page.
As AI agents become more operational and stateful,
I’m noticing replay/debugging becomes difficult once:
external API responses change
MCP state drifts
filesystem state changes
checkpoints diverge over time
Observability helps explain what happened,
but reproducing the exact execution later seems much harder.
I’m curious how people here approach:
replay execution
tool-call recording
state snapshot persistence
mismatch detection
execution provenance
Is anyone else running into similar issues with long-running agent workflows?
All reactions