From 12140320fbef97c7478914516bd152fa03b669a7 Mon Sep 17 00:00:00 2001 From: Alex Ivantsov Date: Sat, 26 Sep 2026 23:44:51 -0400 Subject: [PATCH 1/2] fix(langgraph): keep dev-runtime idle pickle writes off the host NVMe `langgraph dev` always runs the in-memory runtime (it hardcodes :memory: and never reads DATABASE_URI). That runtime's daemon flush loop (langgraph_runtime_inmem 0.30.0 _persistence.py, _flush_interval=10) calls PersistentDict.sync() on every store every 10s unconditionally, and sync() (langgraph_checkpoint 4.1.1 memory backend) re-pickles the ENTIRE dict and atomically replaces the file with no dirty check and no pruning. The blobs dict (.langgraph_checkpoint.3.pckl) is only pruned by explicit delete_thread(). Measured on VM 150 (idle): the .3.pckl blob reached 198 MB, rewritten every ~10s = 26.9 MiB/s, 32.3 TB in 18 days, 0.44%/day of NVMe wear. The runtime's own disable switch does NOT work through `langgraph dev` in the pinned versions (verified on this box, see docs/adr/0013): langgraph-cli 0.4.29 validate_config_file() strips the unknown "disable_persistence" key, and run_server() then overwrites LANGGRAPH_DISABLE_FILE_PERSISTENCE to "false" in the worker via patch_environment. A container env var and a .env entry are overridden the same way. There is also no knob for the 10s interval. Fix: back /app/.langgraph_api with a 1 GiB size-capped tmpfs in the compose langgraph service. The runtime still re-pickles every 10s, but into RAM, so the NVMe sees nothing. Verified idle for ~60s each on one image build: without the tmpfs the container block-IO write climbed to 277 MB; with it, block-IO write stayed 0 B while the pickle was still rewritten every 10s. Also set LANGGRAPH_DISABLE_FILE_PERSISTENCE=true as a forward-compatible no-op that takes over if upstream honors it. langgraph.json left unchanged (unknown key is stripped now and could be rejected by a stricter future validator). Rejected: the disable switch (proven non-functional here); a Postgres checkpointer (unreachable from `langgraph dev`); throttling the interval (hardcoded constant); a custom entrypoint replicating dev (fragile across bumps). Trade-off: langgraph run history no longer survives a container restart (the tmpfs clears). It already did not survive a recreate (/app/.langgraph_api was never volume-mounted) and the stack already ran :memory:; engagement evidence persists separately. Residual: the runtime still burns CPU/RAM re-pickling; only the disk cost is removed, pending an upstream fix. Guard: tests/unit/ops/test_langgraph_persistence_config.py asserts the shipped compose mounts the size-capped tmpfs and sets the forward-compat flag; it fails on the old config. --- .env.example | 10 ++ CHANGELOG.md | 15 +++ docker-compose.yml | 34 +++++ ...-disable-langgraph-dev-file-persistence.md | 124 ++++++++++++++++++ .../ops/test_langgraph_persistence_config.py | 85 ++++++++++++ 5 files changed, 268 insertions(+) create mode 100644 docs/adr/0013-disable-langgraph-dev-file-persistence.md create mode 100644 packages/decepticon/tests/unit/ops/test_langgraph_persistence_config.py diff --git a/.env.example b/.env.example index 7a676f722..b01b18c00 100644 --- a/.env.example +++ b/.env.example @@ -304,6 +304,16 @@ LANGSMITH_PROJECT=decepticon # Set this to 1 to remove --allow-blocking and reinstate the fatal- # on-sync behavior (useful when debugging async issues in your own code). # LANGGRAPH_STRICT_ASYNC=1 +# +# Forward-compatible flag, currently a no-op through `langgraph dev`. The inmem +# runtime re-pickles its ENTIRE checkpoint state to /app/.langgraph_api/*.pckl +# every 10s regardless of idle state, no dirty check, no pruning (measured: +# 198 MB rewritten every ~10s, 26.9 MiB/s). This var is meant to disable that, +# but langgraph-cli 0.4.29 strips the config key and run_server overwrites the +# var in the worker to "false", so it has no effect today. The real fix is the +# 1 GiB tmpfs on /app/.langgraph_api in docker-compose.yml (writes land in RAM, +# zero disk wear). See docs/adr/0013. +# LANGGRAPH_DISABLE_FILE_PERSISTENCE=true # --- C2 Framework / Specialist workloads --- # ADR-0006 (v1.1.8+): specialist workloads (c2-sliver, c2-havoc, ad, diff --git a/CHANGELOG.md b/CHANGELOG.md index 2a83fadcb..c2babfadc 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -47,6 +47,21 @@ workflow. heartbeat writes (~every 10s), which with `--allow-blocking` wedged the event loop so `:2024` never served and `decepticon start`'s "assistant loaded" check timed out. Hot-reload is a dev-only convenience not wanted on an engagement box. +- **langgraph dev idle disk writes moved off the NVMe** (`docker-compose.yml`, + `.env.example`). The in-memory runtime re-pickled its entire checkpoint state + to `/app/.langgraph_api/*.pckl` every 10s with no dirty check and no pruning; + measured on VM 150 the `.langgraph_checkpoint.3.pckl` blob reached 198 MB + rewritten every ~10s (26.9 MiB/s, 32.3 TB in 18 days, 0.44%/day of NVMe wear). + The runtime's own disable switch (`LANGGRAPH_DISABLE_FILE_PERSISTENCE` / + `disable_persistence`) is a no-op through `langgraph dev` in langgraph-cli + 0.4.29 (the config validator strips the key and `run_server` overwrites the env + var in the worker), so the fix is a 1 GiB size-capped tmpfs on + `/app/.langgraph_api`: the runtime still re-pickles every 10s but into RAM, and + the container's block-IO write measured 0 B (vs 277 MB without it). The + forward-compatible flag is still set for the day upstream honors it. Trade-off: + langgraph run history no longer survives a container restart (it already did + not survive a recreate, and engagement evidence persists separately). Guarded + by `tests/unit/ops/test_langgraph_persistence_config.py`. See `docs/adr/0013`. ### Removed diff --git a/docker-compose.yml b/docker-compose.yml index 34f8fc6f2..1d6244ccf 100644 --- a/docker-compose.yml +++ b/docker-compose.yml @@ -400,7 +400,41 @@ services: # BENCHMARK_MODE flows in via `env_file: .env` above — no need to # re-declare it here. Set BENCHMARK_MODE=1 in .env to engage CTF/XBOW # benchmark mode (rule overrides + challenge context inject). + # + # Forward-compatible flag (currently a no-op through `langgraph dev`; the + # tmpfs mount below is the load-bearing fix). `langgraph dev` always runs + # the inmem runtime (it hardcodes :memory: and never reads DATABASE_URI), + # whose _flush_loop re-pickles the ENTIRE checkpoint state to + # /app/.langgraph_api/*.pckl every 10s unconditionally, no dirty check, no + # pruning. Measured on VM 150: .langgraph_checkpoint.3.pckl reached 198 MB + # and was rewritten every ~10s (26.9 MiB/s, 32.3 TB in 18 days, 0.44%/day + # of NVMe wear). This env var is read at import by _persistence.py and the + # checkpoint memory backend, BUT langgraph-cli 0.4.29 defeats it two ways: + # validate_config_file() strips unknown keys (so langgraph.json cannot pass + # disable_persistence) and run_server() then overwrites this env var in the + # worker process to str(disable_persistence).lower() = "false" via + # patch_environment (langgraph_api/cli.py). So it does not disable writes + # today; it is set for the day upstream honors it, and does no harm. See + # docs/adr/0013. + - LANGGRAPH_DISABLE_FILE_PERSISTENCE=${LANGGRAPH_DISABLE_FILE_PERSISTENCE:-true} volumes: + # THE FIX: back /app/.langgraph_api with a size-capped tmpfs so the + # runtime's unavoidable 10s re-pickling lands in RAM, never the host NVMe. + # Verified on this box: with this mount the container's block-IO write + # stays at 0 B while the pickle is still rewritten every 10s; without it, + # block-IO write climbs continuously. Capped at 1 GiB (~5x the observed + # 198 MB) so a runaway blob fails loud with ENOSPC (which crashes the + # daemon flush thread — _flush_loop has no try/except — stopping further + # writes) instead of silently wearing the disk. Raise the size for very + # long engagements. Contents are lost on restart, which changes nothing: + # the stack already ran :memory: (state ephemeral across restarts) and + # .langgraph_api was never volume-mounted (lost on every recreate); real + # engagement evidence persists via EventLogMiddleware + the workspace mount. + # See docs/adr/0013. + - type: tmpfs + target: /app/.langgraph_api + tmpfs: + size: 1073741824 # Mount the OAuth credential files read-only so the LLM factory can # check file presence + valid JSON before adding a method to the # fallback chain. Without this the factory would trust the diff --git a/docs/adr/0013-disable-langgraph-dev-file-persistence.md b/docs/adr/0013-disable-langgraph-dev-file-persistence.md new file mode 100644 index 000000000..b38d4f839 --- /dev/null +++ b/docs/adr/0013-disable-langgraph-dev-file-persistence.md @@ -0,0 +1,124 @@ +# 0013. Keep langgraph-dev's idle pickle writes off the host disk with a tmpfs + +- **Status:** Proposed +- **Date:** 2026-09-27 +- **Deciders:** Exploitacious +- **Related:** docker-compose.yml (`langgraph` service), `.env.example`, `packages/decepticon/tests/unit/ops/test_langgraph_persistence_config.py` + +## Context + +The stack runs the LangGraph server as `langgraph dev` (a compose service, +built from `containers/langgraph.Dockerfile`). `langgraph dev` always selects +the in-memory runtime: it hardcodes `__database_uri__=":memory:"` and +`runtime_edition="inmem"` and never reads the process `DATABASE_URI`, so setting +`DATABASE_URI` has no effect on which runtime loads (`langgraph_cli` 0.4.29 +`cli.py` `dev()` -> `run_server()`; `langgraph_api` 0.10.0 `cli.py`). + +The in-memory runtime keeps its checkpoint state in three dicts (`storage`, +`writes`, `blobs`) mirrored to disk under `/app/.langgraph_api/` as +`.langgraph_checkpoint.{1,2,3}.pckl` plus `.langgraph_ops.pckl`. A daemon thread +(`langgraph_runtime_inmem` 0.30.0 `_persistence.py`, `_flush_interval = 10`, +`_flush_loop`) calls `PersistentDict.sync()` on every registered store every 10 +seconds unconditionally. `sync()` (`langgraph_checkpoint` 4.1.1 +`langgraph/checkpoint/memory/__init__.py`) re-pickles the entire dict and +atomically replaces the file. There is no dirty flag anywhere in the class, so +the whole state is rewritten every interval whether or not anything changed, and +the `blobs` dict (`.langgraph_checkpoint.3.pckl`) is only pruned by an explicit +`delete_thread()`, never by TTL or the flush loop, so it grows unbounded. + +Measured on VM 150 (idle): `.3.pckl` reached 198 MB, rewritten roughly every 10 +seconds even with zero activity: 26.9 MiB/s, 32.3 TB over 18 days, driving the +homelab's only NVMe at 0.44%/day. `/app/.langgraph_api` was not volume-mounted, +so those writes hit the container's writable layer on the host disk. + +The obvious fix would be to disable the file persistence. The runtime does have +a switch, `LANGGRAPH_DISABLE_FILE_PERSISTENCE` (read at import by +`_persistence.py` and the checkpoint memory backend), settable in principle via +`"disable_persistence": true` in `langgraph.json`. We verified on this box that +**neither route works through `langgraph dev` in the pinned versions**: + +- `langgraph_cli` 0.4.29 loads config through `validate_config_file()`, whose + schema does not include `disable_persistence`; the key is silently stripped + (confirmed: the validated config's keys do not contain it), so + `config_json.get("disable_persistence", False)` is always `False`. +- `run_server()` then builds a `to_patch` env dict with + `LANGGRAPH_DISABLE_FILE_PERSISTENCE=str(disable_persistence).lower()` = `false` + and applies it with `patch_environment()`; the loaded-env merge explicitly + refuses to overwrite keys already in `to_patch`. So a container-level env var, + a `.env` entry, and the config key are all overridden to `false` in the worker + process that does the writing (observed directly: the worker subprocess env + showed `LANGGRAPH_DISABLE_FILE_PERSISTENCE=false` even when the container was + launched with it set to `true`, and the `.pckl` files were still written). + +So there is no supported way to stop the writes from inside `langgraph dev` at +these versions. There is also no knob for the 10s interval (`_flush_interval` is +a hardcoded module constant). + +## Decision + +Keep the writes off the host NVMe by backing `/app/.langgraph_api` with a +size-capped `tmpfs` in the compose `langgraph` service (1 GiB). The runtime +still re-pickles every 10s, but into RAM, so the disk sees nothing. + +Also set `LANGGRAPH_DISABLE_FILE_PERSISTENCE=true` in the compose env as a +forward-compatible signal: it is a no-op today (overridden as described above), +does no harm, and will disable the writes for real if a future upstream bump +honors it. `langgraph.json` is left unchanged (an unrecognized key is stripped +today and could be rejected by a stricter future validator). + +Verified on this box (single image build, run idle ~60s each): +- Without the tmpfs: `.langgraph_ops.pckl` (8.4 MB after one seeded run) was + rewritten every 10s and the container's block-IO **write** climbed to 277 MB. +- With the tmpfs: the same file was still rewritten every 10s (into RAM), the + tmpfs held ~8 MB, and the container's block-IO **write stayed at 0 B**. + +## Consequences + +- **Easier:** idle host-disk writes from this service are zero; NVMe wear from + it stops. Guarded by a static test on the shipped compose + (`test_langgraph_persistence_config.py`) that fails without the tmpfs. +- **Harder:** nothing operationally. +- **Given up:** LangGraph run/checkpoint history no longer survives a container + restart (the tmpfs is cleared). Near-zero loss in practice: the stack already + ran `:memory:` (ephemeral across restarts), `/app/.langgraph_api` was never + volume-mounted (state was already lost on every recreate/rebuild), and + engagement evidence persists independently via `EventLogMiddleware` + (`events.jsonl`) and the workspace mount. In-run state (sub-agent handoffs + within one server process) is unaffected: it lives in the same in-memory dicts. +- **Residual cost (not solved here):** the runtime still burns CPU and RAM + bandwidth re-pickling the full state every 10s; only the *disk* cost is + removed. Eliminating the re-pickle needs an upstream fix (see below). +- **Cap behavior:** the tmpfs is capped at 1 GiB (~5x the observed 198 MB). If + the state exceeds it, `sync()` fails with `ENOSPC`; `_flush_loop` has no + `try/except`, so its daemon thread dies loudly and further persistence stops, + with no disk wear. Raise the size for very long engagements. + +## Alternatives considered + +- **Disable file persistence via `disable_persistence` / the env var.** Rejected + as the fix (kept only as a forward-compatible signal): proven non-functional + through `langgraph dev` in langgraph-cli 0.4.29 / langgraph-api 0.10.0, because + the config validator strips the key and `run_server` overwrites the env var in + the worker (evidence above). This is an upstream defect; the clean long-term + fix is to get upstream to honor the switch (or to plumb it correctly), at which + point the flag we already set takes over from the tmpfs. +- **A real Postgres checkpointer on the stack's existing Postgres.** Rejected: + `langgraph dev` cannot use it; the Postgres runtime needs + `LANGGRAPH_RUNTIME_EDITION=postgres`, the separate `langgraph-runtime-postgres` + package, Redis, and a migrations path, i.e. abandoning `langgraph dev`. Far + larger and riskier than the disk problem warrants; the box does not need + durable cross-restart run history. +- **Throttle the flush interval.** Rejected: `_flush_interval` is a hardcoded + module constant with no env or config knob, and a longer interval still + re-pickles the full unbounded blob. +- **A custom container entrypoint that calls `run_server(disable_persistence=True)` + directly, bypassing the broken config path.** Rejected: it would have to + replicate `langgraph dev`'s config loading, graph resolution, and the fork's + `--no-reload` / `--allow-blocking` handling, and would re-break on every + upstream bump. The tmpfs is version-independent. + +## See also + +- [/CHANGELOG.md](../../CHANGELOG.md) - the fork's Unreleased entry for this fix. +- `docs/adr/0006-agent-driven-container-lifecycle.md` - the broader + compose/runtime lifecycle this service lives in. diff --git a/packages/decepticon/tests/unit/ops/test_langgraph_persistence_config.py b/packages/decepticon/tests/unit/ops/test_langgraph_persistence_config.py new file mode 100644 index 000000000..e6b72e6a2 --- /dev/null +++ b/packages/decepticon/tests/unit/ops/test_langgraph_persistence_config.py @@ -0,0 +1,85 @@ +"""Regression guard: the shipped langgraph service must not write to the host disk. + +`langgraph dev` always runs the in-memory runtime, whose flush loop re-pickles +the ENTIRE checkpoint state to /app/.langgraph_api/*.pckl every 10s with no +dirty check and no pruning (measured on VM 150: 198 MB rewritten every ~10s, +26.9 MiB/s, 32.3 TB in 18 days). The runtime's disable switch is a no-op through +`langgraph dev` in the pinned langgraph-cli (see docs/adr/0013), so the fix is a +size-capped tmpfs on /app/.langgraph_api: the writes land in RAM, not the NVMe. + +This test fails on the old config (no tmpfs floor) and passes once the tmpfs is +in place. It is a static config check, no container is started. The tmpfs +assertion is the load-bearing guard; the env-var assertion records the +forward-compatible flag that will take over if upstream honors it. +""" + +from __future__ import annotations + +import re +from pathlib import Path + +import yaml + +_REPO_ROOT = Path(__file__).resolve().parents[5] +_COMPOSE = _REPO_ROOT / "docker-compose.yml" + +_DISABLE_KEY = "LANGGRAPH_DISABLE_FILE_PERSISTENCE" +_API_DIR = "/app/.langgraph_api" + + +def _langgraph_service() -> dict: + compose = yaml.safe_load(_COMPOSE.read_text(encoding="utf-8")) + services = compose["services"] + assert "langgraph" in services, "compose has no langgraph service" + return services["langgraph"] + + +def _env_default_truthy(value: str) -> bool: + """True if a compose env value resolves to a truthy default. + + Accepts a bare ``true`` or a ``${VAR:-true}`` interpolation whose default + is truthy. + """ + value = value.strip() + if value.lower() == "true": + return True + m = re.fullmatch(r"\$\{[^:}]+:-([^}]*)\}", value) + return bool(m) and m.group(1).strip().lower() == "true" + + +def test_compose_mounts_size_capped_tmpfs_floor() -> None: + """THE fix: /app/.langgraph_api is a size-capped tmpfs so the runtime's + unavoidable 10s re-pickling hits RAM (capped), never the host NVMe.""" + volumes = _langgraph_service().get("volumes", []) + tmpfs_mounts = [ + v + for v in volumes + if isinstance(v, dict) + and v.get("type") == "tmpfs" + and v.get("target") == _API_DIR + ] + assert tmpfs_mounts, ( + f"compose langgraph service must back {_API_DIR} with a tmpfs so writes " + "never reach the host disk (ADR-0013)." + ) + size = tmpfs_mounts[0].get("tmpfs", {}).get("size") + assert isinstance(size, int) and size > 0, ( + f"the {_API_DIR} tmpfs must be size-capped so a runaway blob fails loud " + f"with ENOSPC in RAM instead of wearing the disk; got size={size!r}." + ) + + +def test_compose_sets_forward_compat_disable_flag() -> None: + """Records the forward-compatible disable flag. It is a no-op through + `langgraph dev` today (ADR-0013) but must stay set so it takes over if a + future upstream bump honors it.""" + env = _langgraph_service().get("environment", []) + # compose environment is a list of "KEY=VALUE" strings here. + pairs = dict(item.split("=", 1) for item in env if "=" in item) + assert _DISABLE_KEY in pairs, ( + f"compose langgraph service must set {_DISABLE_KEY} (ADR-0013)." + ) + assert _env_default_truthy(pairs[_DISABLE_KEY]), ( + f"{_DISABLE_KEY} must default to true; got {pairs[_DISABLE_KEY]!r} " + "(ADR-0013)." + ) From e3c5ddcd0f7e46ca12c93a24a347a91f2762bc09 Mon Sep 17 00:00:00 2001 From: Alex Ivantsov Date: Sun, 27 Sep 2026 11:41:54 -0400 Subject: [PATCH 2/2] style(tests): format the langgraph persistence config test --- .../unit/ops/test_langgraph_persistence_config.py | 11 +++-------- 1 file changed, 3 insertions(+), 8 deletions(-) diff --git a/packages/decepticon/tests/unit/ops/test_langgraph_persistence_config.py b/packages/decepticon/tests/unit/ops/test_langgraph_persistence_config.py index e6b72e6a2..3b910c18e 100644 --- a/packages/decepticon/tests/unit/ops/test_langgraph_persistence_config.py +++ b/packages/decepticon/tests/unit/ops/test_langgraph_persistence_config.py @@ -54,9 +54,7 @@ def test_compose_mounts_size_capped_tmpfs_floor() -> None: tmpfs_mounts = [ v for v in volumes - if isinstance(v, dict) - and v.get("type") == "tmpfs" - and v.get("target") == _API_DIR + if isinstance(v, dict) and v.get("type") == "tmpfs" and v.get("target") == _API_DIR ] assert tmpfs_mounts, ( f"compose langgraph service must back {_API_DIR} with a tmpfs so writes " @@ -76,10 +74,7 @@ def test_compose_sets_forward_compat_disable_flag() -> None: env = _langgraph_service().get("environment", []) # compose environment is a list of "KEY=VALUE" strings here. pairs = dict(item.split("=", 1) for item in env if "=" in item) - assert _DISABLE_KEY in pairs, ( - f"compose langgraph service must set {_DISABLE_KEY} (ADR-0013)." - ) + assert _DISABLE_KEY in pairs, f"compose langgraph service must set {_DISABLE_KEY} (ADR-0013)." assert _env_default_truthy(pairs[_DISABLE_KEY]), ( - f"{_DISABLE_KEY} must default to true; got {pairs[_DISABLE_KEY]!r} " - "(ADR-0013)." + f"{_DISABLE_KEY} must default to true; got {pairs[_DISABLE_KEY]!r} (ADR-0013)." )