diff --git a/docs/STATUS.md b/docs/STATUS.md index e855f05f..3df64898 100644 --- a/docs/STATUS.md +++ b/docs/STATUS.md @@ -66,6 +66,9 @@ What a reader should not have to dig for: episodes ingested by crashed attempts, a code patch between conditions, no trusted execution receipt, and execution from an uncommitted tree. All six, and the directory selection rule, are in `results/official-007-graphiti/provenance/README.md` beside the executed source. + 🔁 **Added the same day, and larger than all six: Graphiti searched a different corpus.** It + was offered 120 to 143 sessions per condition, without the ~4,900 document haystack every other + arm searched, so its row is not comparable to theirs and must not be read as a ranking. 3. **Claude Mem ran with a synchronous first-search guard**, which preregistration 074 requires to be labelled on any result. Its row says so in its integration field. 4. **Claude Mem was never sent a review invitation.** Nothing in this repository or the project diff --git a/results/official-007-graphiti/provenance/README.md b/results/official-007-graphiti/provenance/README.md index 27309785..b47eb33f 100644 --- a/results/official-007-graphiti/provenance/README.md +++ b/results/official-007-graphiti/provenance/README.md @@ -67,6 +67,26 @@ day. 2026-08-29 history rewrite and would republish content removed since. What executed is preserved here instead (next section). +## The corpus was not the haystack (added 2026-09-26, the same day) + +**This was missing from the first version of this page, and it matters more than any deviation +above.** Graphiti was offered only the condition's own sessions, without the shared ~4,900 document +haystack every other arm searched: + +| condition | sessions offered to Graphiti | offered to mempalace and cognee | +|---|---:|---:| +| present | 132 | 4,900 | +| absent | 120 | 4,888 | +| superseded | 143 | 4,911 | +| contradictory | 142 | 4,911 | +| adjacent | 132 | 4,900 | + +Read from `ingest[*].sessions_offered` in each `environment.json`. Preregistration 084 does not state +a corpus size, so this is not a deviation from its text, but it is a different retrieval problem: +with no distractors, finding the governing session is far easier. The joined comparison against +official-003 is therefore not like for like, and the Graphiti row must not be read as a ranking +against arms that searched the full haystack. + ## Source `source/` holds the working-tree copies of the files that define the arm and its execution: diff --git a/scripts/build_leaderboard.py b/scripts/build_leaderboard.py index 8f61fdf2..05c7b4ff 100644 --- a/scripts/build_leaderboard.py +++ b/scripts/build_leaderboard.py @@ -94,7 +94,14 @@ ), # Same reasoning: the run deviated from preregistration 084 in six stated ways, and a reader of # the row alone should not have to open results/official-007-graphiti/provenance/ to learn it. - "graphiti": ("Graphiti MCP server, run deviated from its preregistration", None, "Graphiti"), + "graphiti": ( + ( + "Graphiti MCP server, searched about 140 sessions instead of the 4,900 document " + "haystack, and deviated from its preregistration; not comparable" + ), + None, + "Graphiti", + ), } # ⛔ PRODUCT_ARMS is the list of arms that are MEASURED, not the arms that are hoped for. `mem0`, diff --git a/site/data/leaderboard.js b/site/data/leaderboard.js index 4e833968..3ff43097 100644 --- a/site/data/leaderboard.js +++ b/site/data/leaderboard.js @@ -308,7 +308,7 @@ window.AMB_LEADERBOARD = { }, { "name": "Graphiti", - "type": "Graphiti MCP server, run deviated from its preregistration", + "type": "Graphiti MCP server, searched about 140 sessions instead of the 4,900 document haystack, and deviated from its preregistration; not comparable", "success": 0.6059, "delta": 0.0228, "ci": [ diff --git a/site/index.html b/site/index.html index 0ecb4a15..ec7bf6d3 100644 --- a/site/index.html +++ b/site/index.html @@ -143,7 +143,7 @@

The arms

supermemory
official Claude Code lifecycle hooks, own run
product
cognee
MCP server, own run
product
Claude Mem
lifecycle hooks and MCP search, first-search guard, own run
product
-
Graphiti
MCP server, own run, deviated from its preregistration
product
+
Graphiti
MCP server, own run on a smaller corpus, not comparable
product

bare is the floor. placebo separates memory from diff --git a/site/method.html b/site/method.html index 995bfcde..f2067168 100644 --- a/site/method.html +++ b/site/method.html @@ -165,8 +165,9 @@

The arms

run on the same model and task grid "the product was never tried." Each run is joined to official-003 on task, seed and condition, so it is compared on the cells both admitted. Claude Mem ran with a - first-search guard and is labelled as such; Graphiti's run deviated from its - preregistration in six stated ways, listed with its artifacts + first-search guard and is labelled as such; Graphiti searched about 140 sessions + instead of the 4,900 document haystack and deviated from its preregistration in six + stated ways, so its row is not comparable to the others recall_prefetch