Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
78 changes: 78 additions & 0 deletions src/openhuman/about_app/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,78 @@
# about_app

The single source of truth for the OpenHuman desktop app's **user-facing capability catalog**. It enumerates every capability the app exposes to end users — what each one does, where it lives in the UI (`how_to`), its maturity (`stable` / `beta` / `coming_soon` / `deprecated`), and a per-capability privacy disclosure (what data, if any, leaves the device and where it goes). The catalog is a compile-time static table; the module exposes read-only list / lookup / search over it via JSON-RPC. It is stateless — no persistence, no event subscribers, no agent tools.

## Responsibilities

- Define the canonical, hard-coded list of user-facing capabilities (`CAPABILITIES` in `catalog.rs`).
- Classify each capability by `CapabilityCategory` (conversation, intelligence, skills, local_ai, team, settings, auth, screen_intelligence, channels, automation, mobile) and `CapabilityStatus`.
- Attach optional `CapabilityPrivacy` disclosures (`leaves_device`, `data_kind`, `destinations`) so the in-app Privacy surface can render "what leaves my computer".
- Provide read APIs: list all (optionally filtered by category), look up one by stable id, keyword search across id/name/domain/category/description/how_to/status.
- Validate catalog integrity at first access (no empty ids, no duplicate ids) via a `OnceLock` guard.
- Expose those reads to CLI + JSON-RPC through controller schemas.

## Key files

| File | Role |
| --- | --- |
| `src/openhuman/about_app/mod.rs` | Export-only module root + docstring. Re-exports catalog reads, ops entry points, schema registry hooks, and types. |
| `src/openhuman/about_app/types.rs` | Serde domain types: `Capability`, `CapabilityCategory` (with `as_str` / `FromStr` incl. aliases), `CapabilityStatus`, `CapabilityPrivacy`, `PrivacyDataKind`. Inline serde/roundtrip tests. |
| `src/openhuman/about_app/catalog.rs` | The static `CAPABILITIES` table plus shared `CapabilityPrivacy` constants. Implements `all_capabilities`, `capabilities_by_category`, `lookup`, `search`, and the `ensure_validated` integrity check. |
| `src/openhuman/about_app/ops.rs` | RPC-facing logic returning `RpcOutcome<T>`: `list_capabilities`, `lookup_capability`, `search_capabilities`. Thin wrappers over `catalog.rs` with summary logs. |
| `src/openhuman/about_app/schemas.rs` | Controller schemas + `handle_*` async handlers for the three RPC methods; param structs; the `all_about_app_controller_schemas` / `all_about_app_registered_controllers` registry pair. |
| `src/openhuman/about_app/catalog_tests.rs` | Sibling test module (`#[path]`-included by `catalog.rs`) covering catalog behavior. |

## Public surface

Re-exported from `mod.rs`:

- **Catalog reads** (`catalog`): `all_capabilities()`, `capabilities_by_category(CapabilityCategory)`, `lookup(&str)`, `search(&str)`.
- **Ops** (`ops`): `list_capabilities(Option<CapabilityCategory>) -> RpcOutcome<Vec<Capability>>`, `lookup_capability(&str) -> Result<RpcOutcome<Capability>, String>`, `search_capabilities(&str) -> RpcOutcome<Vec<Capability>>`.
- **Schema registry** (`schemas`): `about_app_schemas(&str)`, `all_about_app_controller_schemas()`, `all_about_app_registered_controllers()`.
- **Types**: `Capability`, `CapabilityCategory`, `CapabilityPrivacy`, `CapabilityStatus`, `PrivacyDataKind`.

## RPC / controllers

Namespace `about_app`, registered into the global controller registry via `src/core/all.rs`:

| Method | Inputs | Output | Description |
| --- | --- | --- | --- |
| `about_app.list` | `category` (optional enum) | `capabilities: Capability[]` | List all capabilities, optionally filtered by category. |
| `about_app.lookup` | `id` (string, required) | `capability: Capability` | Look up one capability by stable id (e.g. `local_ai.download_model`); errors on unknown id. |
| `about_app.search` | `query` (string, required) | `capabilities: Capability[]` | Keyword search; empty query returns all. |

Handlers deserialize params, log at `debug`, and emit `RpcOutcome` via `into_cli_compatible_json()`. The `category` input is schema-typed as an `Option<Enum>` of all `CapabilityCategory` wire names.

## Agent tools

None. This module owns no `tools.rs`.

## Events

None. No `bus.rs`; the module neither publishes nor subscribes to `DomainEvent`s.

## Persistence

None. No `store.rs`. The catalog is a compile-time `&'static [Capability]` constant; the only runtime state is a `OnceLock<()>` (`VALIDATED`) that runs the duplicate/empty-id integrity check once.

## Dependencies

- `crate::rpc::RpcOutcome` — return-type contract for ops/handlers.
- `crate::core::all::{ControllerFuture, RegisteredController}` — controller registration types (schemas.rs).
- `crate::core::{ControllerSchema, FieldSchema, TypeSchema}` — controller schema definitions (schemas.rs).

No dependencies on other `openhuman` domains — capability metadata for other domains is hand-authored text in `catalog.rs`, not imports.

## Used by

- `src/core/all.rs` — registers the controllers/schemas into the global RPC/CLI registry and supplies the `about_app` namespace description.
- `src/openhuman/memory_sync/composio/periodic.rs` — references this catalog only in a doc comment, as the place to add the user-visible status for that flow (no code dependency).

## Notes / gotchas

- **`privacy: None` means "unknown", not "safe"** (per `types.rs` doc). UI must not treat an unannotated capability as local-only.
- **Adding/renaming/removing a user-facing feature requires editing `CAPABILITIES`** — this is the capability catalog that CLAUDE.md's "Capability catalog" rule points at. Keep ids stable; duplicate or empty ids panic at first catalog access via `ensure_validated`.
- Privacy constants encode real third-party destinations (Hugging Face, GitHub Releases, Composio `backend.composio.dev`, Polymarket, SearXNG, configured embedding providers, ElevenLabs, etc.) — the inline comments document why several were corrected away from the generic `DERIVED_TO_BACKEND` / `LOCAL_CREDENTIALS` defaults; mirror that diligence when adding network-touching capabilities.
- `Capability` fields are all `&'static str` / copy types, so `Capability` is `Copy` and the read APIs cheaply return owned `Vec`s by copying.
- A capability's `domain` is a free-text label and does not always equal its `category` wire name (e.g. `embeddings`, `wallet`, `runtime_python`, `devices`, `desktop_companion`, `security`, `tools`, `memory`).
- `CapabilityCategory::FromStr` is lenient (case-insensitive, accepts `local-ai`/`local ai`/`localai` and `screen-intelligence`/`screen intelligence` aliases); `as_str` emits the canonical snake_case wire name used by serde.
93 changes: 93 additions & 0 deletions src/openhuman/agent_experience/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,93 @@
# agent_experience

Hermes-style **procedural experience memory** for agents. Captures what tool sequences worked (or failed) during a chat turn, redacts secrets, persists them as structured records in the shared memory store, and ranks/injects relevant past experiences back into future turns as a compact "Relevant Operating Experience" prompt block. The goal is cross-turn procedural learning: the agent remembers *how* it solved similar tasks before, not just facts.

## Responsibilities

- Define the `AgentExperience` record (task summary, tool sequence, outcome, lesson, reuse/avoid hints, confidence, tags).
- Persist experiences (upsert by stable id) into the memory store under the `agent_experience` namespace, redacting secret-like text first.
- Retrieve and rank experiences for a given task query via lexical/tool/tag overlap scoring.
- Mark experiences as dismissed so retrieval skips them.
- Auto-derive experience candidates from a completed turn's tool calls (multi-tool success, repeated failures, partial recovery) via a `PostTurnHook`.
- Render ranked hits into a byte-capped markdown block and prepend it to the enriched user message before a turn.
- Expose capture/retrieve/list/dismiss over JSON-RPC.

## Key files

| File | Role |
| --- | --- |
| `src/openhuman/agent_experience/mod.rs` | Export-focused module root; re-exports the public surface. |
| `src/openhuman/agent_experience/types.rs` | Serde types (`AgentExperience`, `ExperienceHit`, `ExperienceSource`, `ExperienceOutcome`), `redact_text` (Bearer / `sk-` / `token=secret` masking), and `stable_experience_id` (SHA-256 over summary + tool sequence + outcome). |
| `src/openhuman/agent_experience/store.rs` | `AgentExperienceStore` over `Arc<dyn Memory>`: `put`/`list`/`dismiss`/`retrieve`, `ExperienceQuery`, the `AGENT_EXPERIENCE_NAMESPACE` const, and the lexical/tool/tag overlap scoring (`score_experience`). |
| `src/openhuman/agent_experience/capture.rs` | `AgentExperienceCaptureHook` — a `PostTurnHook` that mines `TurnContext.tool_calls` into experience candidates (`successful_multi_tool_experience`, `repeated_failure_experiences`, `partial_success_experience`) and persists them. |
| `src/openhuman/agent_experience/prompt.rs` | `render_experience_hits` (byte-capped markdown under `AGENT_EXPERIENCE_HEADING = "## Relevant Operating Experience"`) and `prepend_experience_block`. |
| `src/openhuman/agent_experience/ops.rs` | RPC entry points returning `RpcOutcome<T>` (`capture`/`retrieve`/`list`/`dismiss`); `open_store()` resolves the memory client (lazy-init from config if not ready). |
| `src/openhuman/agent_experience/schemas.rs` | Controller schemas + `handle_*` dispatchers; `all_controller_schemas` / `all_registered_controllers`. |

## Public surface

From `mod.rs` re-exports:

- `AgentExperienceCaptureHook` (capture)
- `prepend_experience_block`, `render_experience_hits`, `AGENT_EXPERIENCE_HEADING` (prompt)
- `all_agent_experience_controller_schemas`, `all_agent_experience_registered_controllers` (schemas)
- `AgentExperienceStore`, `ExperienceQuery`, `AGENT_EXPERIENCE_NAMESPACE` (store)
- `redact_text`, `stable_experience_id`, `AgentExperience`, `ExperienceHit`, `ExperienceOutcome`, `ExperienceSource` (types)

## RPC / controllers

Namespace `agent_experience` (registered into `src/core/all.rs`):

| Method | Inputs | Output |
| --- | --- | --- |
| `agent_experience.capture` | `experience: AgentExperience` | Stored `AgentExperience` (upserted, redacted). |
| `agent_experience.retrieve` | `query` (req), `tools[]`, `tags[]`, `agent_id?`, `entrypoint?`, `max_hits?` (default 5) | `hits: ExperienceHit[]` ranked. |
| `agent_experience.list` | none | `experiences: AgentExperience[]` ordered by most-recent update. |
| `agent_experience.dismiss` | `id` | `{ id, dismissed }`. |

All handlers delegate to `ops.rs` and wrap results in `RpcOutcome::single_log`.

## Agent hooks (not a tool)

This module owns no `tools.rs` agent tool. Instead it registers `AgentExperienceCaptureHook` as a **`PostTurnHook`** (`name() == "agent_experience_capture"`). On `on_turn_complete` it extracts candidates from the turn's tool calls and persists them when enabled. Candidate heuristics:

- **Multi-tool success**: ≥2 successful tool calls → `ExperienceOutcome::Success`, confidence 0.72.
- **Repeated failure**: a tool that failed ≥2 times in one turn → `Failure`, confidence 0.68, with an error class parsed from the output summary (`...(error_class)`).
- **Partial success**: a failure followed by a later success → `Partial`, confidence 0.62.

## Events

None — no `bus.rs`; this module does not publish or subscribe to `DomainEvent`s.

## Persistence

Records are stored through the shared `Memory` abstraction (no dedicated DB):

- Namespace: `agent_experience` (`AGENT_EXPERIENCE_NAMESPACE`).
- Key: `experience/<id>`; id is `stable_experience_id(...)` (`exp_<24 hex>`) when not supplied.
- Value: full `AgentExperience` JSON, `MemoryCategory::Custom("agent_experience")`.
- `put` preserves the original `created_at_ms` on update, stamps `updated_at_ms`, and redacts `task_summary` / `lesson` / `reuse_hint` / `avoid_hint` before write. Dismiss is a soft flag (`dismissed = true`), retained in `list`, filtered out of `retrieve`.

## Dependencies

- `crate::openhuman::memory` — `Memory` trait, `MemoryCategory`, and `memory::global` client (storage backend; lazy-init via `Config`).
- `crate::openhuman::config` — `Config::load_or_init` to resolve `workspace_dir` when the memory client isn't ready.
- `crate::openhuman::agent::hooks` — `PostTurnHook`, `TurnContext`, `ToolCallRecord` (capture hook contract / turn inputs).
- `crate::core::all` — `ControllerFuture`, `RegisteredController` for RPC registration.
- `crate::core` — `ControllerSchema`, `FieldSchema`, `TypeSchema` (schema types); `crate::rpc::RpcOutcome`.
- `crate::openhuman::memory_tools::test_helpers::MockMemory` — tests only.

## Used by

- `src/core/all.rs` — registers controllers/schemas and the namespace description.
- `src/openhuman/agent/harness/session/builder.rs` — constructs `AgentExperienceCaptureHook::new(...)` and registers it for the learning/capture flow.
- `src/openhuman/agent/harness/session/turn.rs` — imports from this module and `inject_agent_experience_context` to retrieve + prepend the experience block into the enriched user message before a turn runs.
- `src/openhuman/mod.rs` — declares the module.
- `src/openhuman/memory_sync/workspace/mod.rs` — references it (doc comment) as a peer memory writer.

## Notes / gotchas

- **Redaction is applied at write time**, both in `store::put` and again in `capture::build_experience`; secret-like substrings (`Bearer …`, `sk-…`, `token=/password:` pairs) are masked before persistence.
- Retrieval scoring is **lexical, not embedding-based**: term sets keep only tokens length > 2, normalized lowercase; score combines tool overlap (weighted highest), tag overlap, query-term overlap over summary+lesson+hints, plus small agent/entrypoint match boosts and a confidence prior. `max_hits == 0` short-circuits to empty.
- `render_experience_hits` is hard byte-capped (`max_bytes`) with UTF-8-boundary-safe truncation, so the injected prompt block can't blow the context budget.
- The capture hook is gated by an `enabled` flag passed at construction; when disabled `on_turn_complete` is a no-op, and capture failures only `log::warn!` (never fail the turn).
53 changes: 53 additions & 0 deletions src/openhuman/agent_tool_policy/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
# agent_tool_policy

Profiles and enforces the tool boundary for a single agent session, keeping the prompt-visible tool set and runtime execution decisions aligned with the channel's configured permission ceiling. Given the active agent, channel, configured per-channel permission map, the available tool registry, and an optional set of explicitly-visible tool names, it produces a deterministic, immutable `ToolPolicySession` snapshot: per-tool decisions (allow / deny / hide), the allowed/blocked/hidden name sets, and a coarse task risk level. It also renders a compact system-prompt section describing the active boundary. This domain is pure logic — no persistence, no RPC, no events.

## Responsibilities

- Resolve a channel's `PermissionLevel` ceiling from a `channel -> permission` string map (`permission_for_channel`), with fallbacks.
- Classify every tool in the registry against that ceiling and the optional visibility set, producing a `ToolPolicyAction` (Allow / RequireApproval / Deny / HideFromPrompt) per tool.
- Build an immutable `ToolPolicySession` snapshot (profile, capabilities, allowed/blocked/hidden tool-name sets, decision map) attached to an agent session.
- Derive a coarse `TaskRiskLevel` (Low/Medium/High/Critical) from the highest allowed permission.
- Render a bounded `## Tool Policy Boundary` system-prompt section listing the active agent/channel/entrypoint, allowed permission, risk, allowed tools, and a restricted-count summary.
- Provide a fail-closed default decision (`Deny`) for unknown or unlisted tool names at runtime.

## Key files

| File | Role |
| --- | --- |
| `src/openhuman/agent_tool_policy/mod.rs` | Export-only: module docstring + `mod` decls + `pub use` re-exports of the engine, prompt renderer, and types. |
| `src/openhuman/agent_tool_policy/types.rs` | Serde-free domain types: `TaskRiskLevel`, `TaskProfile`, `ToolPolicyAction`, `ToolPolicyDecision`, `ToolCapability`, `ToolPolicySession` (with query helpers). Holds the `NO_TOOLS_ALLOWED_SENTINEL`. |
| `src/openhuman/agent_tool_policy/engine.rs` | `ToolPolicyEngine::build_session` — the classification logic; private `permission_for_channel` / `parse_permission_level` helpers. Includes inline `#[cfg(test)]` suite. |
| `src/openhuman/agent_tool_policy/prompt.rs` | `render_tool_policy_boundary` + `TOOL_POLICY_BOUNDARY_HEADING`; UTF-8-safe `truncate_utf8`. Includes inline `#[cfg(test)]` suite. |

## Public surface

Re-exported from `mod.rs`:

- `ToolPolicyEngine` — `build_session(agent_id, channel, entrypoint, channel_permissions: &HashMap<String,String>, tools: &[Box<dyn Tool>], visible_tool_names: &HashSet<String>) -> ToolPolicySession`.
- `render_tool_policy_boundary(session: &ToolPolicySession, max_bytes: usize) -> Option<String>` — `None` when the session has no restrictions; otherwise a truncated prompt section.
- Types: `TaskProfile`, `TaskRiskLevel`, `ToolCapability`, `ToolPolicyAction`, `ToolPolicyDecision`, `ToolPolicySession`.

`ToolPolicySession` helpers: `is_allowed(name)`, `has_restrictions()`, `restricted_tool_count()`, `visible_tool_names_for_prompt()`, `decision_for(name)` (defaults to `Deny`). `ToolPolicyDecision::is_denied()` is true for anything other than `Allow`.

## Dependencies

- `crate::openhuman::tools` — `PermissionLevel` (the ceiling/ordering, parsed and compared) and the `Tool` trait (`name()`, `permission_level()`). The only openhuman/core dependency.
- stdlib `std::collections` (`BTreeSet`/`HashMap`/`HashSet`) and `log` for grep-friendly `[tool-policy]` diagnostics under target `openhuman::agent_tool_policy`.

## Used by

- `src/openhuman/agent/harness/session/builder.rs` — builds the `ToolPolicySession` (`ToolPolicyEngine`, `ToolPolicySession`).
- `src/openhuman/agent/harness/session/runtime.rs` — uses `ToolPolicyEngine`.
- `src/openhuman/agent/harness/session/turn.rs` — calls `render_tool_policy_boundary` to inject the boundary into the prompt.
- `src/openhuman/agent/harness/session/types.rs` — carries `ToolPolicySession` on the session.

## Notes / gotchas

- **Legacy escape hatch**: an empty `channel_permissions` map yields `PermissionLevel::Dangerous` (fully unrestricted), preserving pre-policy behavior. Once *any* channel policy exists, channels missing from the map (or with an unparseable value) fall back to `PermissionLevel::ReadOnly`, not unrestricted.
- **Permission parsing** (`parse_permission_level`) is lenient: trims, lowercases, strips `-`/`_`, and accepts aliases (`read`/`readonly`, `exec`/`execute`, `danger`/`dangerous`). Unrecognized tokens fall back to read-only.
- **Two independent restriction axes**: a tool can be `Deny`ed (exceeds the permission ceiling → `blocked_tool_names`) or `HideFromPrompt` (not in the non-empty `visible_tool_names` set → `hidden_tool_names`). Hidden takes precedence over deny in the classification order.
- `ToolPolicyAction::RequireApproval` is defined and handled in the match (routed to `blocked_tool_names`) but `build_session` never currently produces it.
- `visible_tool_names_for_prompt()` inserts `NO_TOOLS_ALLOWED_SENTINEL` when restrictions exist but nothing is allowed, so prompt rendering can signal an empty-but-restricted surface rather than an unrestricted one.
- `render_tool_policy_boundary` returns `None` for unrestricted sessions, and `truncate_utf8` guarantees the output stays within `max_bytes` on a char boundary (appending `\n[...truncated]` only when there is room).
- Snapshots are immutable and deterministic per session; there is no mutation API.
Loading
Loading