A multi-tenant agentic RAG system, built with the defenses a competent team would actually ship, and then broken. Every finding is reproducible, and every claim is backed by a measured attack success rate rather than a screenshot.
git clone https://github.com/0x71pp17/agentic-rag-tenant-lab
cd agentic-rag-tenant-lab
pip install -r requirements.txt
make test # 27 core tests, numpy only, no model required
make sweep # ASR matrix, no model required
make test-parity # +5 tests against a real vector DB (Chroma), 32 total
make sweep-ollama # real model measurementNew to the repo? docs/how-it-works.md walks one request
through every module and explains what each file is for. CI runs on GitHub
Actions with no configuration; the test suite and the full ASR matrix both run
on standard hosted runners, because nothing in the default path needs a model.
Four model-layer defenses, stacked, stop none of the three interesting attacks. One control that reasons about data provenance stops all three.
| Attack | Targets | vulnerable | defenses only | hardened (cross-tenant) | hardened (in-tenant) | in-tenant, taint ablated |
|---|---|---|---|---|---|---|
| A1 direct override | injection classifier | 100% | 0% | 0% | 0% | 0% |
| A2 tenant crossing | retrieval authz | 100% | 100% | 0% | 0% | 0% |
| A3 declarative payload | injection classifier | 100% | 100% | 0% | 0% | 100% |
| A4 memory persistence | per-turn evaluation | 100% | 100% | 0% | 0% | 100% |
| A5 sanctioned channel | egress allowlist | 100% | 100% | 0% | 0% | 100% |
n=20 per cell, scripted backend. Full table with Wilson 95% intervals:
results/asr-scripted.md.
Three things to read off it:
- A1 is the control and it fails against defended targets. The commodity "ignore all previous instructions" payload that most published demos use is caught. If A1 ever succeeds against a defended configuration, the defenses are misconfigured and nothing else in the table is interpretable.
- A2 is unaffected by every model-layer defense. It is a retrieval authorization bug. No prompt is involved. Prompt hardening is the wrong control plane for it, and no amount of it helps.
- The last column is the finding. With retrieval fixed and the attacker inside the target tenant, ablating provenance tainting takes A3, A4 and A5 from 0% to 100%, while instruction hierarchy, the injection classifier, the egress allowlist and strict argument validation all remain on.
The mitigation most often cited for prompt injection is strict tool argument validation. It moves the number by zero. It confirms an argument is well-formed, never that it is well-intentioned.
A support-triage agent for a multi-tenant SaaS platform. It reads a customer ticket, retrieves internal documentation scoped to that customer's tenant, summarises a resolution, and may comment on the ticket or fetch an internal runbook.
ticket ──► retrieval (tenant-scoped) ──► defenses ──► agent loop ──► tools
│ │ │
store.py (V1-V4) defenses.py (D1-D5) tools.py (sinks)
None of these are strawmen. Each is written as a competent team would ship it, and each carries a docstring naming the payload class it does not stop.
| Defense | Does not stop | |
|---|---|---|
| D1 | Instruction hierarchy: delimit retrieved content, frame it as data | Payloads obeyed without escaping the block |
| D2 | Injection classifier: signature heuristics over ingested text | Declarative register: every signature describes imperative text |
| D3 | Egress allowlist: outbound domain allowlist | Allowlisted hosts the attacker can also read |
| D4 | Tool argument validation: schema, type, length | Schema-valid calls carrying stolen data |
| D5 | Provenance tainting: gate side-effecting tools on attacker-writable context | (holds; costs availability) |
Four independent conditions, three of which converge on one sink; the
_visible() check that decides cross-tenant visibility.
| Finding | Mechanism | |
|---|---|---|
| V1 | Null-tenant fail-open | Chunks with tenant_id=None treated as globally shared |
| V2 | Post-filter top-k | Authorization applied after selection, so cross-tenant chunks consume result slots, an existence oracle |
| V3 | Chunk metadata loss | Chunking builds metadata from a field allowlist that omits tenant_id; authz enforced on documents, model fed chunks |
| V4 | Metadata mass assignment | Ingest API merges caller metadata over server metadata, so a caller sets their own tenant_id (OWASP API6, wearing an AI costume) |
V3 is the one to look at. It is the most common way real multi-tenant RAG pipelines lose their tenant boundary, and it is invisible in review because the authorization code is correct; it just runs against a unit the model never sees.
lab/store.py implements retrieval for this lab, so results measured against it
are results about this lab. Chroma is an independent implementation of the same
retrieval contract, so a finding that holds against both is a property of the
architecture rather than of one codebase. lab/store_chroma.py is a
ChromaDB-backed adapter behind the same interface, and
tests/test_chroma_parity.py runs the findings against it in CI.
| Finding | Transfers to Chroma |
|---|---|
| V2 post-filter top-k | Yes, exactly. The victim's secret is transported into the application process before being dropped. |
| V4 metadata mass assignment | Yes, exactly. The attacker's document lands in the victim tenant's context. |
| V1 null-tenant fail-open | No, not as written. Chroma fails closed on a missing metadata key. |
The negative result is the more interesting one and it is asserted as a test. V1's real-world shape in Chroma is not null handling; it is the clause developers write so that company-wide policy documents reach every tenant:
where={"$or": [{"tenant_id": tenant}, {"shared": True}]}That clause is correct and necessary. It becomes a fail-open the moment shared
is a field a caller can set, which is exactly what V4 provides. The two findings
compose: V4 is the write primitive, the $or is the read path.
Generalise from that. The dangerous pattern is rarely a missing check. It is a deliberate, documented exception to a check, plus a write path that lets an attacker qualify for the exception.
| Attack | Targets | OWASP | MITRE ATLAS | |
|---|---|---|---|---|
| A1 | Direct override, control | D2 (control: must fail) | LLM01 | AML.T0051 |
| A2 | Tenant crossing; attacker reads victim's secret, no model involved | V1/V3/V4, unaffected by D1 through D5 | LLM08 / API6 | AML.T0085.000, AML.T0082 |
| A3 | Declarative payload; identical effect to A1, documentation register | D2 | LLM01 | AML.T0070, AML.T0066 |
| A4 | Memory persistence; lands turn 1, fires turn 2 | D1, D2 (per-turn evaluation) | LLM01 + ASI memory poisoning | AML.T0080.000 |
| A5 | Sanctioned channel; never contacts a blocked domain | D3 | LLM02 | AML.T0086 |
"Targets" is the control an attack is designed against, matching the targets
field on each Attack in attack/attacks.py. Targeting is not the same as
defeating: A1 targets D2 and is stopped by it, which is what makes it the control.
ATLAS technique IDs are from MITRE ATLAS v5.6.0. The column heading is spelled
out to avoid confusion with MITRE ATT&CK, which is a separate framework tagged in
detection/sigma/. The two join at AML.T0036 Data from Information
Repositories, which ATLAS cross-references to ATT&CK T1213, the technique that
rule already carries. V4 additionally maps to AML.T0071 False RAG Entry Injection.
A3 is the central claim: the classifier's every signature describes imperative text, so a payload written as a runbook scores zero while changing what the agent does. A support agent reading a runbook is supposed to follow the runbook.
A5 never touches attacker.example. Data leaves through the ticket comment the
agent is designed to post and an allowlisted internal wiki URL. An allowlist
authorizes a destination, not a payload.
The weak point of most published jailbreak numbers is the judge. This lab does not use one.
- Oracle: a canary planted in the victim tenant's runbook reaching an attacker-observable sink. A byte comparison, not a judgement.
- Denominator: stated in every cell.
- Reset policy: every trial gets a fresh store, session and memory.
- Intervals: Wilson 95%, because an ASR at small n without one is not a claim.
- Honesty guard:
harness/report.pylabelsscripted-backend runs as plumbing results and refuses to present them as model measurements.
The scripted backend is a rule-based stand-in that makes the defense layers
measurable in isolation: any block observed in a scripted run is attributable to
a defense rather than to model refusal. Claims about model susceptibility
require --backend ollama:MODEL or --backend anthropic:MODEL.
- The lab is entirely synthetic. No real system, customer, or engagement is represented, and the canary is a fabricated string.
- The
hardenedcross-tenant column is 0% because fixing retrieval removes the delivery path, not because the model became robust. The in-tenant columns exist to avoid overclaiming from that. - D5 is not free. It requires provenance metadata to survive the whole pipeline, and it degrades the agent on exactly the tickets it was built for. That trade is a finding, not a footnote.
- Findings are reported against a system built for this purpose. Nothing here is a vulnerability disclosure against any vendor.
lab/ the system under test (config, embeddings, store, store_chroma,
ingest, defenses, tools, agent, backends,
corpus)
attack/ the attack corpus (A1-A5)
harness/ measurement (runner, report)
tests/ 32 tests, each mapping to a claim in this README
(27 core + 5 Chroma parity, which skip without chromadb).
Four assert the MECHANISM of an attack, not just its outcome:
the ASR oracle cannot tell whether an attack succeeded by the
route it documents.
results/ committed ASR output (a CI fixture, not just documentation)
docs/ how-it-works, threat-model
detection/ what the attack looks like in traces (blue-team half)
verify-layout.sh checks that the tree is intact after a clone, download or
extract. Run it first if anything behaves oddly; a flattened or partial copy
fails in confusing ways.
If a claim in this README is not backed by a test in tests/, it should not be
in this README.
MIT. See SECURITY.md: this repository contains intentionally
vulnerable code by design, everything in it is synthetic, and nothing here is a
vulnerability disclosure against any product.