V7 adds the first complete agent engineering evaluation system on top of the existing Research → Outcome → Calibration loop. It answers: what failed, why, where in the trajectory, which change caused the regression, and which model/strategy/skill configuration works better.
The two loops are independent but linkable:
TRACE → DATASET → EVALUATOR → EXPERIMENT → REGRESSION → IMPROVEMENT
RESEARCH → OUTCOME → CALIBRATION (existing, preserved)
Engineering metrics and investment outcomes can later be studied together (association only — never causation claims, spec §51).
Folio Agent Kernel
│
▼
Pi Runtime
┌──────────────────┴──────────────────┐
▼ ▼
Folio AgentEvent LangSmith Trace
│ │
│ ▼
│ Dataset
│ │
│ ▼
│ Evaluators (deterministic + judges)
│ │
│ ▼
└──────────────┬────────────────── Experiments
▼
Evaluation Result
Engineering / Research / Runtime Quality
Layers (spec §3):
- Agent Engineering — tool selection, arguments, trajectory, task completion, groundedness, failure recovery, latency, cost, stability.
- Financial Research — completeness, dimensions, evidence coverage, freshness, unsupported claims, bull/bear balance, provenance.
- Investment Outcome — existing system untouched (Opinion → Outcome → Performance → Calibration); V7 only records linkage data.
-
Pi Runtime spawns with repeated
--extensionflags (verified against Pi CLI source:--extension, -e <path>can be used multiple times). -
Extensions are modeled as
PiExtensionConfig[]:finagent(bundled) andlangsmith(bundled, MIT license). -
The LangSmith extension (
@langchain/langsmith-pi-extension, npm, MIT) reads config from env vars (highest precedence) or~/.pi/langsmith.json/<cwd>/.pi/langsmith.json:Variable Meaning TRACE_TO_LANGSMITHtrue/1/yes/onenablesLANGSMITH_PI_API_KEYAPI key (falls back to LANGSMITH_API_KEY)LANGSMITH_PI_ENDPOINTcustom/self-hosted API URL LANGSMITH_PI_PROJECTproject name (default pi-coding-agent)LANGSMITH_PI_METADATAJSON merged into root trace metadata LANGSMITH_PI_RUNS_ENDPOINTSreplica destinations -
Folio injects these into the Pi spawn env (main process only; API key comes from
safeStorage, never from renderer/env files). Tracing is OFF by default; enabling it restarts the Pi process (same as credential changes). -
Each trace root gets
metadata.thread_id= Pi session id. Because the Pi process is persistent, per-run metadata (folioRunId, symbol, strategy) cannot ride env vars —TraceCorrelationServicereconstructs the mapping by querying LangSmith for traces in the run's window scoped to the thread id, and persistsfolioRunId ↔ traceIdin the evaluation store (spec §52–55).
EvaluationCase (id, category, difficulty, input, expected behavior),
EvaluationDataset (versioned), EvaluationRun, EvaluationResultRecord,
EvaluationMetric/EvaluationScore, EvaluationExperiment +
ExperimentSummary, EvaluationBaseline + RegressionResult,
TraceReference, EvaluationSettings, PrivacyLevel,
EvaluationFailureMode (formal taxonomy — no free-text failure reasons).
EvaluationBackend is the provider-neutral boundary. First implementation:
LangSmithEvaluationBackend (minimal REST: /projects, /runs/query,
/runs/{id}/feedback). Plus LocalEvaluationBackend and NoopEvaluationBackend
for offline mode. Observability never breaks the agent: every backend
failure degrades to a status/log entry (spec §87).
- Tracing off by default; production default privacy
standard;fullis explicit opt-in. minimal: no prompt/answer/tool payloads — names, status, durations only.standard: prompt/answer/args allowed after redaction; portfolio tool results reduced to schema summaries (never raw holdings/positions/cash/ account ids/broker metadata).full: complete trace; credentials/tokens/secrets are still redacted.EvaluationRedactorenforces redaction locally AND the Finagent Pi extension redacts portfolio tool outputs whenFINAGENT_PRIVACY_LEVEL<full, so raw portfolio data never enters the Pi conversation (and thus never a trace).- API key:
safeStorageviaCredentialStore(providerlangsmith); renderer can set/delete/query state, never read the raw key. It never appears in logs, support bundles, or traces.
Priority: deterministic first, trajectory second, LLM judges last. Registered per metric with a versioned rubric (judge rubrics must version so history stays comparable — spec §81).
- Deterministic: task completion, tool coverage (required recall), tool precision (penalizes unnecessary calls), argument validity, tool error rate, max tool calls, evidence presence, provenance presence, freshness compliance, partial-failure honesty, latency, failure recovery.
- Judges (v1, spec §34–38): groundedness, research completeness, financial
reasoning, decision usefulness. Judge model is configured separately from the
agent under test (§80). Structured output only; parse failure →
judge_errorscore null, never crashes the experiment (§107).
folio-agent-benchmark-v1: 50–100 hand-authored high-quality cases (quality > quantity), categories: market, research, tool selection, tool arguments, grounded research, strategy/skill, provider failure, portfolio, compare, long-tail, adversarial.- Versioned:
folio-agent-v1≠folio-agent-v1.1; changing a case requires a version bump so historical experiments stay comparable. - Difficulty tags: golden, difficult, long_tail, tool_failure, regression, adversarial. Fixed real bugs become regression cases (highest gate weight).
- Real user traces may seed cases only after privacy cleanup.
bun run eval:smoke (10–20 high-value regression/golden cases, PR gate) and
bun run eval:full (entire benchmark). Modes: fixture (deterministic,
CI-safe) and live (real providers). Every experiment records metadata:
gitSha, folio/runtime/Pi versions, model, provider, thinking level, strategy,
skill versions, capability registry version, provider config, timestamps.
Cost control: maxCases, sampling, concurrency, timeout (spec §79).
EvaluationBaseline stores dataset version, commit, metrics, thresholds.
Critical metrics (task completion, tool accuracy, groundedness, failure
recovery) cannot regress past maxDelta without failing the gate. Small-sample
results are shown with sample counts; 0.91 vs 0.92 is never called a
meaningful improvement.
- PR:
eval:smokewith fixtures + deterministic evaluators (no market-data flakiness). - Main/nightly (manual): full benchmark.
- Release (manual): model/strategy experiments.
- Expensive cloud eval never runs on every commit.
- Settings → Agent / Evaluation: connection status, tracing toggle, project, privacy level, endpoint, test connection, open LangSmith; API key password field (configured/updatedAt only).
- Evaluation Center (internal/advanced): benchmark summary, experiment list, model comparison (only real metrics), failure-mode view (filterable), case detail (prompt, expected/actual, tool timeline, scores, failures, trace link), human review (👍/👎 + note).
wrong_tool, missing_tool, wrong_args, tool_loop, duplicate_tool, ignored_tool_result, provider_failure, no_evidence, unsupported_claim, premature_answer, context_miss, strategy_miss, timeout, runtime_error, judge_error, resource_unavailable.
EvaluationRun ↔ ResearchReport ↔ ResearchOpinion ↔ ResearchOutcome links are recorded (no causality claims). Future work: correlate engineering metrics with realized outcomes.
No Agent Runtime rewrite, no Outcome Engine rewrite, no LangSmith clone, no full trace viewer, no annotation platform, no default upload of production user data, no auto-online-eval, no auto skill/prompt modification, no self-modifying production behavior.
# In the repo (dev environment only — never part of app startup, spec §9):
pi install npm:@langchain/langsmith-pi-extension
# or rely on the bundled wrapper at .pi/extensions/langsmith/index.ts.
# Configure via electron Settings → Evaluation, or env for the CLI:
export LANGSMITH_PI_API_KEY=lsv2_… # stored in safeStorage when set via UI
export TRACE_TO_LANGSMITH=true
export LANGSMITH_PI_PROJECT=folio-agentLangfuse is a second, Folio-owned observability backend. It does not replace LangSmith's Pi extension. Folio exports the full run tree (input, retrieval/tool spans, synthesis, final report) and writes evaluation scores onto the same trace.
Enable from Settings → Evaluation, or via env for the CLI:
export LANGFUSE_TRACING=true
export LANGFUSE_PUBLIC_KEY=pk-lf-…
export LANGFUSE_SECRET_KEY=sk-lf-…
# optional self-hosted:
export LANGFUSE_HOST=https://cloud.langfuse.comSmoke a Deep Research export (uses the production ResearchRunner; writes to
Langfuse Cloud when keys are set, otherwise a local mock ingestion API):
bun run eval:langfuseTrace shape, score names, and failure isolation are documented in
docs/langfuse-tracing.md
(中文).