Skip to content

Latest commit

 

History

History
247 lines (188 loc) · 20.5 KB

File metadata and controls

247 lines (188 loc) · 20.5 KB

ACC Value Proposition — Why ACC Instead of Another Agent Framework

Executive Summary

ACC (Agentic Cell Corpus) is the only open-source agent framework that:

  1. Enforces three-tier governance natively — constitutional rules (WASM, immutable), live-updatable setpoints (OPA bundle), and arbiter-signed adaptive rules generated from collective behaviour, with no external policy engine required in standalone mode.
  2. Runs autonomously at the edge — a single deploy_mode: edge switch reconfigures the entire stack for an offline-capable MicroShift node with NATS leaf-node buffering, eliminating the common "works in the cloud, broken in the field" deployment split.
  3. Treats agents as persistent biological cells — not stateless tool-calling functions — with stable identity, episodic memory, embedding-drift rogue detection, and a 5-level identity-preserving reprogramming ladder before termination is considered.

Comparison Matrix

Capability ACC LangChain Agents LlamaIndex Agents CrewAI AutoGen Haystack
Native governance tiers Cat-A (WASM immutable) + Cat-B (OPA live) + Cat-C (adaptive, signed) None None None None None
Edge / disconnected operation First-class deploy mode; NATS leaf node buffers signals offline Not supported Not supported Not supported Not supported Not supported
Persistent agent identity Stable agent ID across restarts; Ed25519-signed role updates Session-scoped Session-scoped Task-scoped Session-scoped Stateless
Episodic memory LanceDB/Milvus ICL episodes; arbiter promotes to Cat-C rules Plugin (Chroma, Pinecone) Plugin (LlamaIndex stores) Custom Plugin Custom
Rogue detection Embedding centroid drift + heartbeat absence scoring None None None None None
Cross-collective delegation ACC-9 bridge protocol; A-010 governance gate; JetStream queuing None None Role handoffs GroupChat None
Operator-native lifecycle Kubernetes operator (controller-runtime); CRD-driven; OLM-ready None None None None Helm chart only
Signed role updates Ed25519 arbiter countersign; cryptographic rejection of tampered updates None None None None None
LLM backend agnostic Ollama / Anthropic / vLLM / Llama Stack / any OpenAI-compat (Groq, Gemini, OpenRouter, HuggingFace, Together, Fireworks, 11+ providers); one env var switch Partial Partial OpenAI-first OpenAI-first Partial
NATS-native signaling JetStream at-least-once delivery; replay; dead-letter streams; 11 intra-collective signal types (TASK_PROGRESS, BACKPRESSURE, PLAN, KNOWLEDGE_SHARE, EVAL_OUTCOME, CENTROID_UPDATE, EPISODE_NOMINATE, …) HTTP polling HTTP polling Shared memory HTTP HTTP
Enterprise role library 30 pre-built roles across 9 business domains (sales, marketing, finance, HR, legal, IT, ops, etc.) + roles/TEMPLATE/ for custom roles None None None None None
OWASP LLM Top 10 guardrails LLM01–LLM08 (prompt injection, output validation, DoS shield, PII/PHI detection, excessive agency) with observe/enforce modes None None None None None
Built-in compliance framework EU AI Act risk classification, HIPAA PHI redaction, SOC2 control mapping, tamper-evident HMAC audit chain None None None None None

Deep Dives

Governance Tiers (Cat-A / Cat-B / Cat-C)

Most agent frameworks offer zero native governance. Rules, if any, are application-level Python code that an LLM response can silently bypass if the developer forgets to check.

ACC enforces governance at three distinct layers, each progressively softer:

Category A — Constitutional (immutable)

  • Implemented as a compiled WASM OPA binary (category_a.wasm) loaded into every agent process.
  • Runs in-process with sub-millisecond latency — no network call, no sidecar.
  • Immutable by design: the only way to change a Cat-A rule is to rebuild the WASM blob and roll new pods.
  • Examples: "Never exfiltrate PII", "Always respect rate limits", "Never publish to an external NATS bus without a Cat-B approval".

Category B — Live-Updatable Setpoints

  • OPA bundle server polled by agent OPA sidecars (configurable interval, default 30s).
  • Per-role overrides supported: the ingester can have a lower token_budget than the analyst.
  • Hot-reloaded without pod restart — the analyst can receive an updated rate_limit_rpm mid-task.
  • Examples: model token budgets, rate limits per role, domain allow-lists.

Category C — Adaptive, Arbiter-Signed

  • Generated by the arbiter from ICL episode patterns that recur across the collective.
  • Each Cat-C rule is signed with the arbiter's Ed25519 private key; agents verify before applying.
  • Arbiter enforces a minimum confidence threshold (spec.governance.categoryC.confidenceThreshold) before signing.
  • Examples: "When analyst confidence < 0.6 and task involves contract data, require synthesizer review", "After 3 consecutive ingestion failures from source X, quarantine source for 5 minutes".

Why this matters: A rogue LLM response cannot override a Cat-A rule from within the agent's own process. A misconfigured Cat-B bundle cannot be applied without OPA parsing it. A Cat-C rule cannot be installed without the arbiter's private key. This layered model makes governance a structural property of the runtime, not a convention in application code.


Edge-First Disconnected Operation

ACC is the only agent framework with a defined edge deployment profile. The edge profile is not "standalone with fewer features" — it is a distinct operational topology:

Edge Node (MicroShift / K3s)                  Datacenter Hub (OpenShift)
─────────────────────────────                  ──────────────────────────
Local fast path (works OFFLINE):               Hub NATS cluster:
  ingester → analyst → arbiter                   acc.sol-dc-01.task (hub tasks)
  μs latency, no hub needed                      acc.sol-edge-01.heartbeat (monitoring)
                                                  acc.bridge.*.delegate (delegated tasks)
NATS leaf node:
  local subjects stay local                     Bridge delegation (requires connectivity):
  bridge subjects flow to hub ──────────────►    hub analyst processes with 70B model
  JetStream queues pending tasks                  result returned via bridge subject
  drains automatically on reconnect              retry transparent to edge agent

What works offline:

  • All intra-collective tasks (ingester → analyst → arbiter pipeline)
  • Local Cat-A and cached Cat-B governance
  • Heartbeat accumulation in local JetStream (synced on reconnect)
  • LanceDB episodic memory (local NVMe)
  • Ollama inference (local 3B model)

What requires connectivity:

  • Cross-collective task delegation to the hub (ACC-9 bridge protocol)
  • ROLE_UPDATE hot-reload from hub arbiter
  • Cat-C rule sync from hub
  • Image pull for upgrades

Why other frameworks cannot do this: LangChain, CrewAI, and AutoGen treat agent-to-LLM communication as synchronous HTTP. An offline LLM endpoint means the agent is dead. ACC separates the signaling layer (NATS JetStream, offline-capable) from the inference layer (Ollama, running locally on the edge node), so the collective continues to function and accumulates work to sync back to the hub when connectivity is restored.


Role Infusion vs. System Prompts

Every agent framework lets you pass a system prompt to the LLM. ACC's role infusion is fundamentally different:

Dimension System prompt (other frameworks) Role infusion (ACC)
Persistence Lost on restart Loaded from ConfigMap/Redis/LanceDB; survives restart
Versioning None Semantic version (version: "1.2.0") tracked in heartbeat
Hot-reload Pod restart required NATS ROLE_UPDATE message; applied without restart
Signature verification N/A Ed25519 arbiter countersign; tampered updates rejected
Governance coupling None category_b_overrides carries per-role OPA setpoints
Operator rendering N/A Operator renders role definitions from AgentCollective CRD to ConfigMaps

A role definition is not just a string. It carries the agent's purpose, persona style, task-type allowlist, seed context, allowed actions, and OPA setpoint overrides — all in a versioned, signed document that the collective's arbiter can update at runtime without touching Kubernetes.


NATS over HTTP

Most agent frameworks use HTTP (REST or gRPC) for agent-to-agent and agent-to-orchestrator communication. ACC uses NATS JetStream:

Property HTTP polling NATS JetStream
Message delivery Best-effort (lost if agent is down during POST) At-least-once (JetStream persists until ACK)
Fan-out Requires load balancer or separate message bus Native subject wildcards (acc.*.task)
Replay Not possible JetStream stream replay by sequence or time
Backpressure Client must poll; polling interval trades latency vs. load Server-push; consumer ACK controls flow
Edge buffering Not possible Local JetStream stores messages during disconnect
Audit trail Application-level logging required JetStream stream is an immutable ordered log

For agent collectives that produce governance-relevant signals, the difference between "we think the message was delivered" (HTTP) and "we know the message was persisted and ACK'd" (JetStream) is the difference between an audit trail and a guess.


Operator-Native Lifecycle

ACC's Kubernetes operator is not a Helm chart wrapper. It is a reconciliation loop with 11 sub-reconcilers that understand the ACC domain:

  • PrerequisiteReconciler: detects KEDA, Gatekeeper, RHOAI, KServe, Prometheus — degrades gracefully when absent; suppresses irrelevant warnings in edge mode.
  • InfraReconciler: provisions NATS (single-node or 3-node JetStream cluster), Redis (standalone or Sentinel), OPA bundle server; renders NATS leaf-node config for edge.
  • CollectiveReconciler: creates agent Deployments per role; injects ConfigMaps with rendered acc-config.yaml and role definitions; creates KEDA ScaledObjects when available.
  • GovernanceReconciler: mounts WASM ConfigMap into each pod; syncs OPA bundle; optionally creates Gatekeeper ConstraintTemplates (skipped in edge mode).
  • ObservabilityReconciler: deploys OTel collector and PrometheusRules (skipped in edge mode).

This means an operator upgrade (OLM rolling) automatically re-renders all ConfigMaps, rolls agent Deployments in governance-respecting order (observer → analyst → synthesizer → ingester → arbiter), and handles NATS cluster config changes without manual intervention. No other agent framework has this.


When to Choose ACC

                    ┌──────────────────────────────────────┐
                    │ Do you need agents to run offline     │
                    │ at an edge location?                  │
                    └───────────────┬──────────────────────┘
                    YES             │             NO
                    ▼               │             ▼
              ┌────────────┐        │    ┌────────────────────┐
              │ Use ACC    │        │    │ Do you need native  │
              │ deploy_mode│        │    │ governance tiers?   │
              │ = edge     │        │    └────────┬───────────┘
              └────────────┘        │    YES      │     NO
                                    │    ▼        │     ▼
                                    │ ┌────────┐  │  ┌─────────────────┐
                                    │ │Use ACC │  │  │ Single-agent RAG│
                                    │ │rhoai or│  │  │ pipeline?       │
                                    │ │standal.│  │  │ → LangChain /   │
                                    │ └────────┘  │  │   LlamaIndex    │
                                    │             │  │                 │
                                    │             │  │ Multi-agent task│
                                    │             │  │ workflow?       │
                                    │             │  │ → CrewAI /      │
                                    │             │  │   AutoGen       │
                                    │             │  └─────────────────┘
                                    └─────────────┘

Use ACC when:

  • You are deploying to edge nodes, manufacturing floors, vehicles, or any environment with intermittent connectivity
  • Your domain requires auditable, cryptographically-verifiable governance decisions (regulated industries, defense, healthcare)
  • You need agents to accumulate knowledge over time and promote learned patterns to rules — not just look up a knowledge base
  • You are deploying on OpenShift and need an operator-managed lifecycle

Use LangChain / LlamaIndex when:

  • You are building a single-agent RAG pipeline with no governance requirements
  • You need the largest ecosystem of tool integrations
  • You need fast prototyping with minimal infrastructure

Use CrewAI / AutoGen when:

  • You need a quick multi-agent workflow with role-based task assignment
  • You don't need persistence, governance, or edge deployment
  • Your agents are stateless and disposable

Terminal UI — Built-In Collective Observability

ACC ships a Textual terminal dashboard (acc-tui) that no other agent framework provides as a first-class component. The TUI follows the same biological architecture as ACC itself — each screen corresponds to a layer of the cell model:

╔══ Soma (Dashboard) ══════════════════╗  ╔══ Nucleus (Role Infusion) ═══════════╗
║  Agent cards (drift sparkbar)         ║  ║  Role definition form                ║
║  Reprogramming ladder level           ║  ║  Purpose / Persona / Task types      ║
║  Governance panel (Cat-A/B/C events)  ║  ║  Seed context / Cat-B overrides      ║
║  Memory panel (ICL episodes)          ║  ║  ROLE_UPDATE → NATS → approval       ║
╚══════════════════════════════════════╝  ╚══════════════════════════════════════╝

╔══ Compliance ════════════════════════╗  ╔══ Performance ════════════════════════╗
║  OWASP LLM01–LLM08 guardrail status  ║  ║  LLM latency + token budgets          ║
║  Audit trail (HMAC chain)            ║  ║  Cat-B setpoints live view            ║
║  EU AI Act risk classification       ║  ║  Task queue depth + drain estimate    ║
║  Human oversight queue               ║  ║  Stress indicators + compliance score ║
╚══════════════════════════════════════╝  ╚══════════════════════════════════════╝

╔══ Comms (Signal Monitor) ════════════╗  ╔══ Ecosystem ══════════════════════════╗
║  All 11 NATS signal types live        ║  ║  Collective topology map              ║
║  TASK_PROGRESS, BACKPRESSURE, PLAN   ║  ║  Domain registry + centroids         ║
║  KNOWLEDGE_SHARE, EVAL_OUTCOME       ║  ║  ICL episode nominations             ║
║  CENTROID_UPDATE, EPISODE_NOMINATE   ║  ║  Cross-collective bridge status      ║
╚══════════════════════════════════════╝  ╚══════════════════════════════════════╝

The TUI connects to NATS as a read-only observer — no Redis or LanceDB access. It derives all display state from the full set of 11 ACC signal types (HEARTBEAT, TASK_COMPLETE, ALERT_ESCALATE, TASK_PROGRESS, QUEUE_STATUS, BACKPRESSURE, PLAN, KNOWLEDGE_SHARE, EVAL_OUTCOME, CENTROID_UPDATE, EPISODE_NOMINATE), making it deployable anywhere that can reach port 4222.

Multi-collective: A CollectiveTabStrip at the top of the TUI shows one tab per collective. Set ACC_COLLECTIVE_IDS=sol-01,sol-02 to monitor multiple collectives simultaneously. Switch tabs with Tab / Shift-Tab.

WebBridge: Set ACC_TUI_WEB_PORT=8765 to expose collective snapshots over HTTP (GET /api/snapshot), enabling CI pipelines and external dashboards to poll collective state without running a terminal.

Other agent frameworks offer third-party observability plugins (LangSmith, Weights & Biases). ACC's TUI is part of the core package, purpose-built for the governance model — it shows Cat-A trigger counts, reprogramming ladder escalation, role version drift, OWASP compliance scores, and the human oversight queue, none of which exist in any other agent framework's mental model.


Recently Shipped (v0.2.0)

  • ACC TUI Evolution (52 requirements): 6 biological screens (Soma / Nucleus / Compliance / Performance / Comms / Ecosystem), multi-collective CollectiveTabStrip, HTTP WebBridge (ACC_TUI_WEB_PORT), dynamic enterprise role loading.
  • Enterprise Role Library: 30 pre-built roles across 9 business domains (sales, marketing, product, customer success, finance, HR, legal, operations, IT/security). Each role ships with role.yaml, eval_rubric.yaml, and system_prompt.md. A roles/TEMPLATE/ directory enables custom role authoring.
  • Compliance Framework (25 requirements): OWASP LLM01–LLM08 guardrails, Cat-A CatAEvaluator with WASM/subprocess/passthrough modes, AuditBroker with HMAC tamper-evident chain (file and Kafka backends), HumanOversightQueue for EU AI Act Art. 14, compliance_health_score in StressIndicators.
  • LLM Independence (openai_compat backend): Single backend covering 11+ providers via the OpenAI Chat Completions API. Universal ACC_LLM_MODEL, ACC_LLM_BASE_URL, ACC_LLM_API_KEY_ENV env vars require no code changes to switch providers.
  • ACC-11 Domain-Aware Roles: domain_id, domain_receptors, and eval_rubric_hash on every role; PARACRINE signal receptor filtering at the agent membrane; DomainRegistry EMA centroid update.
  • ACC-10 Intra-Collective Protocol: 8 new signal types (TASK_PROGRESS, QUEUE_STATUS, BACKPRESSURE, PLAN, KNOWLEDGE_SHARE, EVAL_OUTCOME, CENTROID_UPDATE, EPISODE_NOMINATE); ProgressContext; Redis scratchpad for parallel plan steps; roles/ directory convention.

Roadmap Differentiators (in development)

The following phases are planned but not yet implemented. See docs/security-hardening.md for the full design.

  • NATS NKeys authentication (Phase 0c): Per-role cryptographic NKey authentication; per-subject publish/subscribe permissions including bridge subjects — closes the largest current attack surface (anonymous NATS).
  • Cilium L7 NetworkPolicy (Phase 1): L7-aware network policy enforcement; agent pods can only egress to NATS, Redis, their LLM backend, and (for edge) the hub's leaf node port 7422.
  • SPIFFE/SPIRE mTLS (Phase 2): Stable cryptographic workload identity for every agent pod; mTLS between NATS, Redis, and agents; SVIDs auto-rotate without pod restart. Agent ID moves from random UUID to a stable Deployment-label-derived value — heartbeats are auditable across restarts.
  • Tetragon kernel-level Cat-A (Phase 3): eBPF-powered kernel event stream feeds Cat-A governance decisions — execve, connect, file access events trigger governance responses that application code cannot bypass or falsify. Observe-only for 4 weeks before enforcement mode.
  • Real WASM Cat-A evaluation (Phase 3): Replace the _cat_a_allow = True placeholder in acc/governance.py with actual WASM OPA evaluation of the compiled category_a.wasm blob.