Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -304,6 +304,16 @@ LANGSMITH_PROJECT=decepticon
# Set this to 1 to remove --allow-blocking and reinstate the fatal-
# on-sync behavior (useful when debugging async issues in your own code).
# LANGGRAPH_STRICT_ASYNC=1
#
# Forward-compatible flag, currently a no-op through `langgraph dev`. The inmem
# runtime re-pickles its ENTIRE checkpoint state to /app/.langgraph_api/*.pckl
# every 10s regardless of idle state, no dirty check, no pruning (measured:
# 198 MB rewritten every ~10s, 26.9 MiB/s). This var is meant to disable that,
# but langgraph-cli 0.4.29 strips the config key and run_server overwrites the
# var in the worker to "false", so it has no effect today. The real fix is the
# 1 GiB tmpfs on /app/.langgraph_api in docker-compose.yml (writes land in RAM,
# zero disk wear). See docs/adr/0013.
# LANGGRAPH_DISABLE_FILE_PERSISTENCE=true

# --- C2 Framework / Specialist workloads ---
# ADR-0006 (v1.1.8+): specialist workloads (c2-sliver, c2-havoc, ad,
Expand Down
15 changes: 15 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,21 @@ workflow.
heartbeat writes (~every 10s), which with `--allow-blocking` wedged the event
loop so `:2024` never served and `decepticon start`'s "assistant loaded" check
timed out. Hot-reload is a dev-only convenience not wanted on an engagement box.
- **langgraph dev idle disk writes moved off the NVMe** (`docker-compose.yml`,
`.env.example`). The in-memory runtime re-pickled its entire checkpoint state
to `/app/.langgraph_api/*.pckl` every 10s with no dirty check and no pruning;
measured on VM 150 the `.langgraph_checkpoint.3.pckl` blob reached 198 MB
rewritten every ~10s (26.9 MiB/s, 32.3 TB in 18 days, 0.44%/day of NVMe wear).
The runtime's own disable switch (`LANGGRAPH_DISABLE_FILE_PERSISTENCE` /
`disable_persistence`) is a no-op through `langgraph dev` in langgraph-cli
0.4.29 (the config validator strips the key and `run_server` overwrites the env
var in the worker), so the fix is a 1 GiB size-capped tmpfs on
`/app/.langgraph_api`: the runtime still re-pickles every 10s but into RAM, and
the container's block-IO write measured 0 B (vs 277 MB without it). The
forward-compatible flag is still set for the day upstream honors it. Trade-off:
langgraph run history no longer survives a container restart (it already did
not survive a recreate, and engagement evidence persists separately). Guarded
by `tests/unit/ops/test_langgraph_persistence_config.py`. See `docs/adr/0013`.

### Removed

Expand Down
34 changes: 34 additions & 0 deletions docker-compose.yml
Original file line number Diff line number Diff line change
Expand Up @@ -400,7 +400,41 @@ services:
# BENCHMARK_MODE flows in via `env_file: .env` above — no need to
# re-declare it here. Set BENCHMARK_MODE=1 in .env to engage CTF/XBOW
# benchmark mode (rule overrides + challenge context inject).
#
# Forward-compatible flag (currently a no-op through `langgraph dev`; the
# tmpfs mount below is the load-bearing fix). `langgraph dev` always runs
# the inmem runtime (it hardcodes :memory: and never reads DATABASE_URI),
# whose _flush_loop re-pickles the ENTIRE checkpoint state to
# /app/.langgraph_api/*.pckl every 10s unconditionally, no dirty check, no
# pruning. Measured on VM 150: .langgraph_checkpoint.3.pckl reached 198 MB
# and was rewritten every ~10s (26.9 MiB/s, 32.3 TB in 18 days, 0.44%/day
# of NVMe wear). This env var is read at import by _persistence.py and the
# checkpoint memory backend, BUT langgraph-cli 0.4.29 defeats it two ways:
# validate_config_file() strips unknown keys (so langgraph.json cannot pass
# disable_persistence) and run_server() then overwrites this env var in the
# worker process to str(disable_persistence).lower() = "false" via
# patch_environment (langgraph_api/cli.py). So it does not disable writes
# today; it is set for the day upstream honors it, and does no harm. See
# docs/adr/0013.
- LANGGRAPH_DISABLE_FILE_PERSISTENCE=${LANGGRAPH_DISABLE_FILE_PERSISTENCE:-true}
volumes:
# THE FIX: back /app/.langgraph_api with a size-capped tmpfs so the
# runtime's unavoidable 10s re-pickling lands in RAM, never the host NVMe.
# Verified on this box: with this mount the container's block-IO write
# stays at 0 B while the pickle is still rewritten every 10s; without it,
# block-IO write climbs continuously. Capped at 1 GiB (~5x the observed
# 198 MB) so a runaway blob fails loud with ENOSPC (which crashes the
# daemon flush thread — _flush_loop has no try/except — stopping further
# writes) instead of silently wearing the disk. Raise the size for very
# long engagements. Contents are lost on restart, which changes nothing:
# the stack already ran :memory: (state ephemeral across restarts) and
# .langgraph_api was never volume-mounted (lost on every recreate); real
# engagement evidence persists via EventLogMiddleware + the workspace mount.
# See docs/adr/0013.
- type: tmpfs
target: /app/.langgraph_api
tmpfs:
size: 1073741824
# Mount the OAuth credential files read-only so the LLM factory can
# check file presence + valid JSON before adding a method to the
# fallback chain. Without this the factory would trust the
Expand Down
124 changes: 124 additions & 0 deletions docs/adr/0013-disable-langgraph-dev-file-persistence.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,124 @@
# 0013. Keep langgraph-dev's idle pickle writes off the host disk with a tmpfs

- **Status:** Proposed
- **Date:** 2026-09-27
- **Deciders:** Exploitacious
- **Related:** docker-compose.yml (`langgraph` service), `.env.example`, `packages/decepticon/tests/unit/ops/test_langgraph_persistence_config.py`

## Context

The stack runs the LangGraph server as `langgraph dev` (a compose service,
built from `containers/langgraph.Dockerfile`). `langgraph dev` always selects
the in-memory runtime: it hardcodes `__database_uri__=":memory:"` and
`runtime_edition="inmem"` and never reads the process `DATABASE_URI`, so setting
`DATABASE_URI` has no effect on which runtime loads (`langgraph_cli` 0.4.29
`cli.py` `dev()` -> `run_server()`; `langgraph_api` 0.10.0 `cli.py`).

The in-memory runtime keeps its checkpoint state in three dicts (`storage`,
`writes`, `blobs`) mirrored to disk under `/app/.langgraph_api/` as
`.langgraph_checkpoint.{1,2,3}.pckl` plus `.langgraph_ops.pckl`. A daemon thread
(`langgraph_runtime_inmem` 0.30.0 `_persistence.py`, `_flush_interval = 10`,
`_flush_loop`) calls `PersistentDict.sync()` on every registered store every 10
seconds unconditionally. `sync()` (`langgraph_checkpoint` 4.1.1
`langgraph/checkpoint/memory/__init__.py`) re-pickles the entire dict and
atomically replaces the file. There is no dirty flag anywhere in the class, so
the whole state is rewritten every interval whether or not anything changed, and
the `blobs` dict (`.langgraph_checkpoint.3.pckl`) is only pruned by an explicit
`delete_thread()`, never by TTL or the flush loop, so it grows unbounded.

Measured on VM 150 (idle): `.3.pckl` reached 198 MB, rewritten roughly every 10
seconds even with zero activity: 26.9 MiB/s, 32.3 TB over 18 days, driving the
homelab's only NVMe at 0.44%/day. `/app/.langgraph_api` was not volume-mounted,
so those writes hit the container's writable layer on the host disk.

The obvious fix would be to disable the file persistence. The runtime does have
a switch, `LANGGRAPH_DISABLE_FILE_PERSISTENCE` (read at import by
`_persistence.py` and the checkpoint memory backend), settable in principle via
`"disable_persistence": true` in `langgraph.json`. We verified on this box that
**neither route works through `langgraph dev` in the pinned versions**:

- `langgraph_cli` 0.4.29 loads config through `validate_config_file()`, whose
schema does not include `disable_persistence`; the key is silently stripped
(confirmed: the validated config's keys do not contain it), so
`config_json.get("disable_persistence", False)` is always `False`.
- `run_server()` then builds a `to_patch` env dict with
`LANGGRAPH_DISABLE_FILE_PERSISTENCE=str(disable_persistence).lower()` = `false`
and applies it with `patch_environment()`; the loaded-env merge explicitly
refuses to overwrite keys already in `to_patch`. So a container-level env var,
a `.env` entry, and the config key are all overridden to `false` in the worker
process that does the writing (observed directly: the worker subprocess env
showed `LANGGRAPH_DISABLE_FILE_PERSISTENCE=false` even when the container was
launched with it set to `true`, and the `.pckl` files were still written).

So there is no supported way to stop the writes from inside `langgraph dev` at
these versions. There is also no knob for the 10s interval (`_flush_interval` is
a hardcoded module constant).

## Decision

Keep the writes off the host NVMe by backing `/app/.langgraph_api` with a
size-capped `tmpfs` in the compose `langgraph` service (1 GiB). The runtime
still re-pickles every 10s, but into RAM, so the disk sees nothing.

Also set `LANGGRAPH_DISABLE_FILE_PERSISTENCE=true` in the compose env as a
forward-compatible signal: it is a no-op today (overridden as described above),
does no harm, and will disable the writes for real if a future upstream bump
honors it. `langgraph.json` is left unchanged (an unrecognized key is stripped
today and could be rejected by a stricter future validator).

Verified on this box (single image build, run idle ~60s each):
- Without the tmpfs: `.langgraph_ops.pckl` (8.4 MB after one seeded run) was
rewritten every 10s and the container's block-IO **write** climbed to 277 MB.
- With the tmpfs: the same file was still rewritten every 10s (into RAM), the
tmpfs held ~8 MB, and the container's block-IO **write stayed at 0 B**.

## Consequences

- **Easier:** idle host-disk writes from this service are zero; NVMe wear from
it stops. Guarded by a static test on the shipped compose
(`test_langgraph_persistence_config.py`) that fails without the tmpfs.
- **Harder:** nothing operationally.
- **Given up:** LangGraph run/checkpoint history no longer survives a container
restart (the tmpfs is cleared). Near-zero loss in practice: the stack already
ran `:memory:` (ephemeral across restarts), `/app/.langgraph_api` was never
volume-mounted (state was already lost on every recreate/rebuild), and
engagement evidence persists independently via `EventLogMiddleware`
(`events.jsonl`) and the workspace mount. In-run state (sub-agent handoffs
within one server process) is unaffected: it lives in the same in-memory dicts.
- **Residual cost (not solved here):** the runtime still burns CPU and RAM
bandwidth re-pickling the full state every 10s; only the *disk* cost is
removed. Eliminating the re-pickle needs an upstream fix (see below).
- **Cap behavior:** the tmpfs is capped at 1 GiB (~5x the observed 198 MB). If
the state exceeds it, `sync()` fails with `ENOSPC`; `_flush_loop` has no
`try/except`, so its daemon thread dies loudly and further persistence stops,
with no disk wear. Raise the size for very long engagements.

## Alternatives considered

- **Disable file persistence via `disable_persistence` / the env var.** Rejected
as the fix (kept only as a forward-compatible signal): proven non-functional
through `langgraph dev` in langgraph-cli 0.4.29 / langgraph-api 0.10.0, because
the config validator strips the key and `run_server` overwrites the env var in
the worker (evidence above). This is an upstream defect; the clean long-term
fix is to get upstream to honor the switch (or to plumb it correctly), at which
point the flag we already set takes over from the tmpfs.
- **A real Postgres checkpointer on the stack's existing Postgres.** Rejected:
`langgraph dev` cannot use it; the Postgres runtime needs
`LANGGRAPH_RUNTIME_EDITION=postgres`, the separate `langgraph-runtime-postgres`
package, Redis, and a migrations path, i.e. abandoning `langgraph dev`. Far
larger and riskier than the disk problem warrants; the box does not need
durable cross-restart run history.
- **Throttle the flush interval.** Rejected: `_flush_interval` is a hardcoded
module constant with no env or config knob, and a longer interval still
re-pickles the full unbounded blob.
- **A custom container entrypoint that calls `run_server(disable_persistence=True)`
directly, bypassing the broken config path.** Rejected: it would have to
replicate `langgraph dev`'s config loading, graph resolution, and the fork's
`--no-reload` / `--allow-blocking` handling, and would re-break on every
upstream bump. The tmpfs is version-independent.

## See also

- [/CHANGELOG.md](../../CHANGELOG.md) - the fork's Unreleased entry for this fix.
- `docs/adr/0006-agent-driven-container-lifecycle.md` - the broader
compose/runtime lifecycle this service lives in.
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
"""Regression guard: the shipped langgraph service must not write to the host disk.

`langgraph dev` always runs the in-memory runtime, whose flush loop re-pickles
the ENTIRE checkpoint state to /app/.langgraph_api/*.pckl every 10s with no
dirty check and no pruning (measured on VM 150: 198 MB rewritten every ~10s,
26.9 MiB/s, 32.3 TB in 18 days). The runtime's disable switch is a no-op through
`langgraph dev` in the pinned langgraph-cli (see docs/adr/0013), so the fix is a
size-capped tmpfs on /app/.langgraph_api: the writes land in RAM, not the NVMe.

This test fails on the old config (no tmpfs floor) and passes once the tmpfs is
in place. It is a static config check, no container is started. The tmpfs
assertion is the load-bearing guard; the env-var assertion records the
forward-compatible flag that will take over if upstream honors it.
"""

from __future__ import annotations

import re
from pathlib import Path

import yaml

_REPO_ROOT = Path(__file__).resolve().parents[5]
_COMPOSE = _REPO_ROOT / "docker-compose.yml"

_DISABLE_KEY = "LANGGRAPH_DISABLE_FILE_PERSISTENCE"
_API_DIR = "/app/.langgraph_api"


def _langgraph_service() -> dict:
compose = yaml.safe_load(_COMPOSE.read_text(encoding="utf-8"))
services = compose["services"]
assert "langgraph" in services, "compose has no langgraph service"
return services["langgraph"]


def _env_default_truthy(value: str) -> bool:
"""True if a compose env value resolves to a truthy default.

Accepts a bare ``true`` or a ``${VAR:-true}`` interpolation whose default
is truthy.
"""
value = value.strip()
if value.lower() == "true":
return True
m = re.fullmatch(r"\$\{[^:}]+:-([^}]*)\}", value)
return bool(m) and m.group(1).strip().lower() == "true"


def test_compose_mounts_size_capped_tmpfs_floor() -> None:
"""THE fix: /app/.langgraph_api is a size-capped tmpfs so the runtime's
unavoidable 10s re-pickling hits RAM (capped), never the host NVMe."""
volumes = _langgraph_service().get("volumes", [])
tmpfs_mounts = [
v
for v in volumes
if isinstance(v, dict) and v.get("type") == "tmpfs" and v.get("target") == _API_DIR
]
assert tmpfs_mounts, (
f"compose langgraph service must back {_API_DIR} with a tmpfs so writes "
"never reach the host disk (ADR-0013)."
)
size = tmpfs_mounts[0].get("tmpfs", {}).get("size")
assert isinstance(size, int) and size > 0, (
f"the {_API_DIR} tmpfs must be size-capped so a runaway blob fails loud "
f"with ENOSPC in RAM instead of wearing the disk; got size={size!r}."
)


def test_compose_sets_forward_compat_disable_flag() -> None:
"""Records the forward-compatible disable flag. It is a no-op through
`langgraph dev` today (ADR-0013) but must stay set so it takes over if a
future upstream bump honors it."""
env = _langgraph_service().get("environment", [])
# compose environment is a list of "KEY=VALUE" strings here.
pairs = dict(item.split("=", 1) for item in env if "=" in item)
assert _DISABLE_KEY in pairs, f"compose langgraph service must set {_DISABLE_KEY} (ADR-0013)."
assert _env_default_truthy(pairs[_DISABLE_KEY]), (
f"{_DISABLE_KEY} must default to true; got {pairs[_DISABLE_KEY]!r} (ADR-0013)."
)
Loading