Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions docs/STATUS.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,9 @@ What a reader should not have to dig for:
episodes ingested by crashed attempts, a code patch between conditions, no trusted execution
receipt, and execution from an uncommitted tree. All six, and the directory selection rule,
are in `results/official-007-graphiti/provenance/README.md` beside the executed source.
🔁 **Added the same day, and larger than all six: Graphiti searched a different corpus.** It
was offered 120 to 143 sessions per condition, without the ~4,900 document haystack every other
arm searched, so its row is not comparable to theirs and must not be read as a ranking.
3. **Claude Mem ran with a synchronous first-search guard**, which preregistration 074 requires
to be labelled on any result. Its row says so in its integration field.
4. **Claude Mem was never sent a review invitation.** Nothing in this repository or the project
Expand Down
20 changes: 20 additions & 0 deletions results/official-007-graphiti/provenance/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,26 @@ day.
2026-08-29 history rewrite and would republish content removed since. What executed is
preserved here instead (next section).

## The corpus was not the haystack (added 2026-09-26, the same day)

**This was missing from the first version of this page, and it matters more than any deviation
above.** Graphiti was offered only the condition's own sessions, without the shared ~4,900 document
haystack every other arm searched:

| condition | sessions offered to Graphiti | offered to mempalace and cognee |
|---|---:|---:|
| present | 132 | 4,900 |
| absent | 120 | 4,888 |
| superseded | 143 | 4,911 |
| contradictory | 142 | 4,911 |
| adjacent | 132 | 4,900 |

Read from `ingest[*].sessions_offered` in each `environment.json`. Preregistration 084 does not state
a corpus size, so this is not a deviation from its text, but it is a different retrieval problem:
with no distractors, finding the governing session is far easier. The joined comparison against
official-003 is therefore not like for like, and the Graphiti row must not be read as a ranking
against arms that searched the full haystack.

## Source

`source/` holds the working-tree copies of the files that define the arm and its execution:
Expand Down
9 changes: 8 additions & 1 deletion scripts/build_leaderboard.py
Original file line number Diff line number Diff line change
Expand Up @@ -94,7 +94,14 @@
),
# Same reasoning: the run deviated from preregistration 084 in six stated ways, and a reader of
# the row alone should not have to open results/official-007-graphiti/provenance/ to learn it.
"graphiti": ("Graphiti MCP server, run deviated from its preregistration", None, "Graphiti"),
"graphiti": (
(
"Graphiti MCP server, searched about 140 sessions instead of the 4,900 document "
"haystack, and deviated from its preregistration; not comparable"
),
None,
"Graphiti",
),
}

# ⛔ PRODUCT_ARMS is the list of arms that are MEASURED, not the arms that are hoped for. `mem0`,
Expand Down
2 changes: 1 addition & 1 deletion site/data/leaderboard.js
Original file line number Diff line number Diff line change
Expand Up @@ -308,7 +308,7 @@ window.AMB_LEADERBOARD = {
},
{
"name": "Graphiti",
"type": "Graphiti MCP server, run deviated from its preregistration",
"type": "Graphiti MCP server, searched about 140 sessions instead of the 4,900 document haystack, and deviated from its preregistration; not comparable",
"success": 0.6059,
"delta": 0.0228,
"ci": [
Expand Down
2 changes: 1 addition & 1 deletion site/index.html
Original file line number Diff line number Diff line change
Expand Up @@ -143,7 +143,7 @@ <h2>The arms</h2>
<div class="arm-cell"><div class="arm-name">supermemory</div><div class="arm-type">official Claude Code lifecycle hooks, own run</div><span class="badge">product</span></div>
<div class="arm-cell"><div class="arm-name">cognee</div><div class="arm-type">MCP server, own run</div><span class="badge">product</span></div>
<div class="arm-cell"><div class="arm-name">Claude Mem</div><div class="arm-type">lifecycle hooks and MCP search, first-search guard, own run</div><span class="badge">product</span></div>
<div class="arm-cell"><div class="arm-name">Graphiti</div><div class="arm-type">MCP server, own run, deviated from its preregistration</div><span class="badge">product</span></div>
<div class="arm-cell"><div class="arm-name">Graphiti</div><div class="arm-type">MCP server, own run on a smaller corpus, not comparable</div><span class="badge">product</span></div>
</div>

<p class="prose mt-m"><code>bare</code> is the floor. <code>placebo</code> separates memory from
Expand Down
5 changes: 3 additions & 2 deletions site/method.html
Original file line number Diff line number Diff line change
Expand Up @@ -165,8 +165,9 @@ <h2>The arms</h2>
run on the same model and task grid</td>
<td>"the product was never tried." Each run is joined to official-003 on task, seed
and condition, so it is compared on the cells both admitted. Claude Mem ran with a
first-search guard and is labelled as such; Graphiti's run deviated from its
preregistration in six stated ways, listed with its artifacts</td>
first-search guard and is labelled as such; Graphiti searched about 140 sessions
instead of the 4,900 document haystack and deviated from its preregistration in six
stated ways, so its row is not comparable to the others</td>
</tr>
<tr>
<td><span class="m">recall_prefetch</span></td>
Expand Down
Loading