Tài liệu này là định hướng dài hạn. Nó KHÔNG mô tả trạng thái hiện tại — phần "Đã hoàn thành" nằm trong
ARCHITECTURE.md(mục Roadmap còn lại + changelog).Mỗi mục được gắn nhãn trạng thái để tránh làm lại thứ đã có:
- ✅ Đã có — đã tồn tại trong code; KHÔNG nằm trong phạm vi tương lai (chỉ ghi để tham chiếu).
- 🟡 Một phần — nền tảng đã có; roadmap chỉ gồm phần còn thiếu (ghi rõ delta).
- 🔵 Mới — chưa có; là công việc tương lai thực sự.
Nguyên tắc xuyên suốt (kế thừa từ ARCHITECTURE.md): local-first, thuần TypeScript, evolve không rewrite, verify 3 lớp (CORE + M1 + BUILD + tests), không phá hành vi cũ, không over-engineer, quyết định lớn ghi lại bằng ADR.
Ba lỗi phát hiện khi chạy autonomous thật (todo Flutter):
- ✅ Bug 2 — verify PASS giả (v0.0.76): task sinh code thiếu verify → PASS không chạy gì. Sửa: Planner gán verify mặc định theo stack; stack lạ → required criterion để FAIL.
- ✅ Bug 3 — autonomous mù thư mục / re-scaffold (v0.0.77):
summarizeWorkspace(từ ProjectIndex) + luật anti-re-scaffold trongDEFAULT_SYSTEM+ inject workspace overview vào prompt autonomous. Không còn re-runflutter createkhi project đã có. - ✅ Bug 1 — git checkpoint vỡ: ĐÃ sửa từ v0.0.46 (
shellQuotesingle-quote platform-agnostic + test regression). Log của user là từ build cũ; không cần code lại.
| # | Hạng mục | Trạng thái | Ghi chú (đã có gì trong code) |
|---|---|---|---|
| 1 | Refactor Agent Core (tách module) | 🟡 Một phần | Điều tra: toolLoopEngine/agentRuntime đã DI sạch; nợ thật ở chatViewProvider. ✅ bước 1 (v0.0.57) SessionCwdManager; ✅ bước 2 (v0.0.58) CommandExecutor; ✅ bước 3 (v0.0.60) WorkspaceMutator (mutation thuần; approval giữ ở provider); ✅ nối ModelRouter (v0.0.59). 🔶 Hợp nhất 2 đường agent (ADR-002, strangler 6 bước): ✅ bước 0+1 (v0.0.69) ADR + chat chạy trên cùng LLMProvider qua port ToolCaller. ✅ bước 2 (v0.0.70) ExecutionObserver seam + chat nối vào. ✅ bước 3 (v0.0.71) completionGuard (anti-fab/verify) thành module thuần. ✅ bước 4a (v0.0.72) toolLoopKernel thuần. ✅ bước 4b (v0.0.73) chat gọi kernel. ✅ bước 4c (v0.0.74) autonomous gọi kernel — HAI đường chung loop body. ✅ bước 5 (v0.0.75) AutonomyGate + setting agentAutonomy (mặc định edits). ✅ bước 6 (v0.0.78) bỏ hẳn chat-only mode — interactive luôn là agent. ADR-002 HOÀN TẤT (một lõi loop chung + client/sink/policy chung + autonomy policy + một interactive mode). Còn (rủi ro cao): xẻ webview render, Memory abstraction; "gọi autonomous như agent" (bước riêng sau) |
| 2 | Context Engine 2.0 | 🟡 Một phần | DefaultContextEngine: selection, ranking, compression, token budget, dependency-aware đã có |
| 3 | Project Intelligence (index bền) | 🟡 Một phần | ✅ bước 1 (v0.0.61): buildProjectIndex + WorkspaceIndexer (thuần). ✅ bước 2 (v0.0.62): nối ProjectIndex vào DefaultContextEngine — fast path tái dùng depGraph + symbol candidates. ✅ bước 3 (v0.0.63): ProjectIndexCache (cache + invalidation) — rebuild lazy khi TASK_COMPLETED mark stale; engine nhận getter đọc bản tươi mỗi task; best-effort giữ index tốt gần nhất khi build lỗi. Còn: call graph (hoãn Tree-sitter) |
| 4 | Tree-sitter + Static Analysis | 🔵 Mới (HOÃN) | ADR-001 hoãn Tree-sitter; dependencyGraph (regex) đã có |
| 5 | Project Knowledge Graph | 🟡 Một phần | ✅ bước 1 (v0.0.64): buildKnowledgeGraph(ProjectIndex) → KnowledgeGraph thuần — File/Symbol nodes + declares/imports edges + query API. ✅ bước 2 (v0.0.65): Context Engine truy vấn graph — coDeclarationPeers kéo file cùng khai báo symbol trùng tên (không import, không keyword) với SCORE_RELATED=45; graph cache theo WeakMap<ProjectIndex>. Còn: call edges (hoãn Tree-sitter/ADR-001); có thể nối blastRadius/#8 |
| 6 | Architecture Planning Phase | 🟡 Một phần | ✅ bước 1 (v0.0.68): kiểu ProjectBlueprint máy đọc chung + serialize/parse. Còn: LLM Architecture Planner sinh blueprint (project mới) → dẫn xuất DAG; ghi .agent/*.json |
| 7 | Existing Project Analysis Phase | 🟡 Một phần | ✅ bước 1 (v0.0.68): deriveBlueprintFromIndex(ProjectIndex) thuần — suy blueprint hiện trạng (components theo dir + responsibility từ symbol + entryPoints), không LLM. Còn: nối runtime/chat làm pha Project Understanding trước khi sửa |
| 8 | Impact Analysis | 🟡 Một phần | blastRadius.ts (reverse-dep) đã có ở autonomous. ✅ bước 1 (v0.0.66): analyzeImpact → ImpactReport thuần (files/tests/interfaces/entryPoints). ✅ bước 2 (v0.0.67): pre-edit trên đường autonomous — event IMPACT_ANALYZED richer, runtime gọi provider TRƯỚC execute (no I/O, dùng index cache), log ở service; giữ blastRadius cũ. Còn: nối đường chat; call-level (hoãn Tree-sitter) |
| 9 | Architecture Guardian | 🟡 Một phần | src/architecture/ + invariant tests đã có; chưa có vòng generate→guard→repair tự động |
| 10 | Machine-readable Blueprint | 🟡 Một phần | ProjectState/architecture.config.json máy đọc; chưa tách project/constraints/decisions |
| 11 | Tách Memory khỏi history | 🟡 Một phần | EventLog (episodic) + ProjectStore (project) đã có; thiếu các loại memory chuyên biệt |
| 12 | Decision Memory | 🔵 Mới | — |
| 13 | Failure Memory | 🟡 Một phần | selfDiagnostics đọc task.history phát hiện lặp; chưa có store bài học bền vững |
| 14–19 | Engineering Knowledge Engine (GitHub RAG) | 🔵 Mới | — |
| 20 | Multi-model Advisor | 🔵 Mới | Chỉ có provider Ollama local |
| 21 | Model Router | 🟡 Một phần | ✅ nối vào runtime (v0.0.59, trung tính): AutonomousController.resolveExecutionModel + event MODEL_SELECTED. Còn: route Planner, per-task routing, multi-tier (premium/cheap) — #20/#22 |
| 22 | Escalation System | 🔵 Mới | — |
| 23 | AI Council / Debate | 🔵 Mới | — |
| 24 | Verification Loop | 🟡 Một phần | LayeredVerifier (autonomous) + verify-reminder (chat) đã có; thiếu runtime-check khép kín |
| 25 | Self-review | 🔵 Mới | — |
| 26 | Automatic Failure Escalation | 🟡 Một phần | Repair bounded theo maxAttempts đã có; chưa escalate ra advisor/knowledge |
| 27 | Token/Credit Budget Manager | 🔵 Mới | — |
| 28 | Context Budgeting (phân bổ động) | 🟡 Một phần | Budget tổng đã có (getContextLength - overhead); chưa phân bổ theo hạng mục |
| 29 | Cost/Quality Telemetry | 🟡 Một phần | ✅ slice 1 (v0.0.54) TTFT/prefix-cache-hit; ✅ slice 2 (v0.0.55) telemetryAggregator; ✅ slice 3 (v0.0.56) dashboard section (per-model + quality + errors). Còn: escalation event + quality-per-model (phụ thuộc #21/#22) |
| 30 | Task Classification | 🟡 Một phần | TaskType đã tồn tại + router dùng nó; chưa có bộ phân loại intent tự động |
| 31 | Confidence Estimation | 🔵 Mới | — |
| 32 | Adaptive Context | 🟡 Một phần | Dependency expansion đã có; chưa mở/thu working set động trong lúc chạy |
| 33 | Long-running Project Memory | 🔵 Mới | Resume theo project đã có; chưa có "chắt lọc knowledge qua nhiều session" |
| 34 | Prompt/KV cache efficiency (cảm hứng LMCache) | 🔵 Mới | Prefix-stable ordering + context reuse; đo prefix-cache hit/TTFT (gộp với #29) |
Tách rõ trách nhiệm để không dồn vào toolLoopEngine/agentRuntime:
Agent Core
├── Intent
├── Context
├── Planning
├── Execution
├── Verification
├── Memory
└── Model Routing
Mục tiêu: biến prototype thành architecture mở rộng được. Ràng buộc: giữ interface hiện có
(ExecutionEngine, Verifier, ContextEngine, ModelRouter) làm ranh giới; refactor tăng dần,
mỗi bước verify 3 lớp.
Đã có (DefaultContextEngine): context selection, relevance ranking, compression, token
budgeting, task-specific context, dependency-aware retrieval (forward + reverse).
Delta còn thiếu (phạm vi tương lai):
- context prioritization tinh vi (không chỉ score đơn) — cân giữa named/keyword/dependency/recency;
- context freshness (biết context nào cũ/đã đổi trên đĩa để làm mới);
- truy vấn theo Knowledge Graph (#5) thay vì chỉ text search + import graph.
Kim chỉ nam vẫn là: Task → Context Engine → Relevant Working Set → LLM (KHÔNG "toàn bộ project → LLM").
Cho agent hiểu project một lần, thay vì khám phá lại mỗi task:
Project Intelligence
├── Project Tree
├── Symbol Index
├── AST
├── Dependency Graph (✅ dependencyGraph.ts — tái dùng)
├── Call Graph
├── Architecture Map
├── Entry Points
├── Tests
└── Configuration
Ghi chú: dependencyGraph.ts đã có làm điểm khởi đầu; phần index bền (persist + invalidation) là mới.
Source → Tree-sitter → AST → Symbol extraction → Dependency analysis → Project Graph
Chỉ triển khai khi có bằng chứng runtime rằng RegexCodeIntelligence là bottleneck (decision gate
trong ADR-001). Giữ abstraction CodeIntelligence để thay TreeSitterCodeIntelligence mà không đụng
ContextEngine. KHÔNG coi Tree-sitter là toàn bộ Project Intelligence.
Biểu diễn quan hệ giữa: Class, Function, Interface, File, Module, Feature, Dependency, Test, Route,
Database, API. Ví dụ: LoginUseCase → AuthRepository → FirebaseAuthRepository → FirebaseAuthDataSource.
Context Engine (#2) truy vấn graph thay vì tìm text thuần.
Đã có (v0.0.64, bước 1): context/knowledgeGraph.ts — buildKnowledgeGraph(ProjectIndex) dựng graph
thuần: File/Symbol nodes + declares (File→Symbol) / imports (File→File) edges + query API
(symbolsIn, filesDeclaring, importsOf/importersOf, relatedFiles). Là lớp view thống nhất trên
ProjectIndex.
Giới hạn (ADR-001): chỉ những edge regex chứng minh được (declares/imports). Chuỗi ví dụ
LoginUseCase → AuthRepository chỉ biểu diễn ở mức FILE-IMPORT, không phải call-level. Call edges
(Class/Function-level) hoãn tới Tree-sitter.
Đã có (v0.0.65, bước 2): Context Engine truy vấn graph — coDeclarationPeers thêm ứng viên là file
cùng khai báo symbol trùng tên với seed (quan hệ import graph + keyword + symbol-index bỏ sót),
SCORE_RELATED=45, reason co-declares a symbol with <seed>; graph build 1 lần/index qua WeakMap.
Còn: call edges (Class/Function-level) — hoãn Tree-sitter; có thể nối blastRadius/#8.
Project mới KHÔNG code ngay:
User Requirement → Requirement Analysis → Architecture Planner → Project Blueprint → Implementation Plan → Coding
Architecture là artifact máy đọc, không phải prose trong prompt:
.agent/
├── project.json
├── architecture.json
└── implementation-plan.json
Đã có (v0.0.68, bước 1): kiểu ProjectBlueprint máy đọc chung (planning/blueprint.ts) —
{version, goal?, origin, techStack?, components{name,responsibility,files}, decisions{title,rationale}, entryPoints} + serialize/parseBlueprint (khớp .agent/architecture.json).
Còn: Architecture Planner dùng LLM sinh blueprint cho project mới → dẫn xuất Implementation Plan (DAG);
ghi các artifact .agent/*.json.
Project có sẵn:
User Request → Project Understanding → AST/Graph → Architecture Analysis → Impact Analysis → Modification Plan → Coding
Không để model mò codebase mù quáng (chat đã có prime/auto-cd làm bước sơ khởi).
Đã có (v0.0.68, bước 1): deriveBlueprintFromIndex(ProjectIndex) thuần — suy blueprint hiện trạng
(components gom theo top-level dir + responsibility tóm từ symbol, entryPoints từ index), KHÔNG cần LLM.
Đây là "Project Understanding" ở mức máy đọc. AST/Graph + Impact đã có (#3/#5/#8).
Còn: nối runtime/chat để chạy pha này TRƯỚC khi sửa; kết hợp với analyzeImpact (#8) thành Modification Plan.
Đã có: blastRadius.ts tính reverse-dep transitive + emit BLAST_RADIUS_COMPUTED (autonomous).
Đã có (v0.0.66, bước 1): context/impactAnalysis.ts — analyzeImpact(seeds, ProjectIndex) → ImpactReport
thuần, phân loại tập impacted thành 4 nhóm: impactedFiles (importer transitive), impactedTests
(isTestFile), impactedInterfaces (symbol kind==="interface" + tên), impactedEntryPoints
(∩ entryPoints). Tái dùng blastRadius + ProjectIndex (không I/O). Trả lời "ai phụ thuộc? test nào?
interface/platform nào?" ở mức import-graph.
Giới hạn (ADR-001): suy từ import edges + regex symbol kind — cờ review surface, không phải chứng minh
call-level.
Đã có (v0.0.67, bước 2): pre-edit trên đường AUTONOMOUS — agentRuntime gọi analyzeImpact provider
NGAY TRƯỚC khi execute mỗi task, emit event richer IMPACT_ANALYZED (files/tests/interfaces/entryPoints).
Provider inject từ controller dùng index cache (không I/O); runtime giữ thuần (không biết ProjectIndex).
Log tóm tắt ở autonomousService. Giữ BLAST_RADIUS_COMPUTED cũ (sau tool loop) — backward-compat.
Còn: nối đường CHAT (webview-coupled — để bước riêng); call-level (hoãn Tree-sitter/ADR-001).
Đã có: src/architecture/ + invariant tests enforce ranh giới tầng (Phase 3/17).
Delta: vòng tự kiểm sau khi agent sinh code:
Generate → Analyze → Architecture Guardian → Violation? ─ No → Continue / Yes → Repair
Biến agent từ code generator thành architecture-aware coding system.
Đã có: ProjectState + architecture.config.json máy đọc được.
Delta: tách schema project.json / architecture.json / constraints.json / decisions.json
để Agent Core truy vấn trực tiếp thay vì Markdown thuần.
Đã có: EventLog (episodic JSONL), ProjectStore (project memory + resume).
Delta: phân loại memory chuyên biệt:
Memory
├── Project Memory (🟡 ProjectStore)
├── Architecture Memory
├── Decision Memory (#12)
├── Task Memory
├── Error / Failure Memory (#13)
├── User Preference Memory
└── Episodic History (🟡 EventLog)
Lưu quyết định + lý do + ngày + phương án đã loại. Ví dụ: "Domain layer must not depend on Firebase — vì cross-platform". Sau nhiều tháng agent vẫn biết tại sao kiến trúc như vậy.
Đã có: selfDiagnostics đọc task.history phát hiện lặp/kẹt trong một run.
Delta: store bài học bền vững qua session: Approach → Failed → Why → Store lesson
("Không dùng cách A vì gây circular dependency") để lần sau không lặp lại.
Không chỉ "GitHub → vector DB", mà:
Engineering Knowledge Engine
├── Repository Discovery (#15)
├── Repository Ranking / Quality Scoring (#16)
├── Repository Analysis
├── License Analysis (#19)
├── Architecture Extraction
├── Pattern Extraction
├── Solution Extraction (#17)
├── Knowledge Deduplication (#18)
├── Provenance (#19)
├── Semantic Search
└── Knowledge Graph
Agent tự tìm repo liên quan task (Flutter offline-first, SQLite sync, DI, MCP...). KHÔNG crawl hàng loạt:
Search → Quality filter → Top repositories → Targeted analysis.
Không chỉ theo stars: Architecture, Code quality, Tests, Maintenance, Documentation, Activity, Dependencies, Community, License.
Không lưu toàn bộ source. Trích: Patterns, Architectures, Solutions, Trade-offs, Implementation techniques, Testing strategies, Failure modes. (Quan trọng hơn vectorize toàn repo.)
Tránh đưa lại thứ model đã biết. Phân loại: Known concept / Known pattern / New implementation / New evidence / New trade-off / New failure mode / Current API-version. Chỉ giữ thứ thực sự bổ sung giá trị.
Mỗi knowledge item: source, repository, commit/version, license, extraction date. Quan trọng nếu thương mại hóa. Tuân thủ nghiêm license (attribution, giới hạn tái tạo nguyên văn).
Cho phép gọi OpenAI/Claude/DeepSeek/Gemini nhưng KHÔNG làm coding worker chính:
Local LLM = Worker · Premium AI = Expert · Agent = Orchestrator.
Ràng buộc local-first: premium là tùy chọn, mặc định tắt; không gửi code/secret ra ngoài nếu
user không bật rõ ràng (xem safety trong ARCHITECTURE.md).
Đã có: modelRouter.ts chọn model theo task/intent + capability constraints (trong model local).
Delta: mở rộng thành multi-tier:
Task → Model Router → [ Local | Cheap API | Premium | Multi-model ]
Local chỉ gọi premium khi: confidence thấp, architecture change, security-sensitive, large refactor,
unknown framework, repeated failure, critical decision.
Local attempt → Failure ×2 → ESCALATE → Premium Advisor. Premium chỉ dùng khi thực sự đáng tiền.
Chỉ cho vấn đề quan trọng: GPT proposal → Claude critique → DeepSeek alternative → Judge → Final.
KHÔNG dùng cho task đơn giản (tăng chi phí + latency).
Đã có: LayeredVerifier fail-fast (autonomous) + verify-reminder đa toolchain (chat, v0.0.53).
Delta: vòng khép kín gồm cả runtime check:
Plan → Code → Build → Test → Static analysis → Architecture check → Runtime check → Fix → Verify.
Sau khi code: "review what I just changed" — dùng context độc lập hoặc model khác để tránh cùng một reasoning path tự xác nhận chính mình.
Đã có: repair bounded theo maxAttempts, self-diagnostics phát hiện lặp.
Delta: khi fail lặp → Analyze failure → Retrieve knowledge (#14) → Ask premium advisor (#22)
thay vì loop cùng một cách sửa.
Agent biết: Task budget, Context budget, Model budget, API budget.
Simple → 0 premium · Medium → cheap · Complex → premium · Critical → multi-model.
Đã có: budget tổng (getContextLength - overhead, fallback DEFAULT_CONTEXT_BUDGET).
Delta: phân bổ theo hạng mục, ví dụ với 16k:
Task 2k · Architecture 1k · Relevant code 7k · Memory 1k · Knowledge 2k · Safety/system 2k
Agent tự cân đối thay vì một ngân sách phẳng cho code.
Theo dõi: Task, Model used, Input/Output tokens, Latency, Success, Retries, Premium escalations,
Final quality. Dùng để Model Router (#21) tự cải thiện ("task X: local success 91%, khỏi cần premium").
EventLog là nền thu thập sẵn có.
Bổ sung (cảm hứng LMCache — observability là first-class): đo các chỉ số hiệu năng prompt/cache mà tầng HTTP-to-Ollama quan sát được. ✅ ĐÃ LÀM (slice 1, v0.0.54):
- TTFT (time-to-first-token) — đo client-side ở đường stream (
ChatStats.ttftMs); - prompt-eval time (
promptEvalMstừprompt_eval_duration); - prefix-cache hit (heuristic) —
derivePrefixCacheHit: prompt lớn mà prefill ~0ms hoặc prompt-eval throughput vượt ngưỡng ⇒ prefix reuse (trả "unknown" khi thiếu tín hiệu); - persist qua event
TELEMETRY_RECORDED(metadata-only) vàoEventLog; module thuầntelemetry/promptTelemetry.ts.
Đây là điều kiện đo lường cho #34 (không tối ưu được cái không đo được). Vẫn thuần local, đọc từ metadata Ollama, không cần truy cập engine internals.
✅ ĐÃ LÀM (slice 2, v0.0.55): telemetry/telemetryAggregator.ts — aggregateTelemetry(events)
gấp event đã persist thành report: per-model (calls, tokens, avg TTFT/tokensPerSec/promptEvalMs,
prefix-cache hits/misses/hitRate) + quality per-run (completed/failed/blocked, totalAttempts, retries,
repairs, successRate) + llmErrorsByCode. Thuần, không thêm tracking runtime.
✅ ĐÃ LÀM (slice 3, v0.0.56): section Telemetry trong Dashboard — DashboardPanel fold
event qua aggregateTelemetry và render per-model (tokens/TTFT/prefill/tok-s/prefix-cache hit rate)
- quality (success rate/retries/repairs/blocked) + LLM errors. Consumer thật đầu tiên của aggregator.
Còn lại của #29: (a) gán quality theo model (cần multi-model routing #21/#22 để mỗi attempt biết
model); (b) escalation event (phụ thuộc #22 — hiện escalations=0). Cả hai KHÔNG làm độc lập được;
chờ tầng Premium/Router.
Đã có: TaskType (implementation/refactor/debug/test/build/review/architecture/analysis/research)
- router dùng nó. Delta: bộ phân loại tự động từ yêu cầu user → chọn pipeline khác nhau cho Create/Modify/ Refactor/Debug/Research/Explain/Test/Review/Architecture.
Agent tự đánh giá confidence để quyết định: tiếp tục / tìm thêm context / retry / escalate.
Đã có: dependency expansion khi build context.
Delta: mở/thu working set động trong lúc chạy:
task → discover file → new dependency → retrieve new context → update working set.
Làm việc cùng project qua nhiều tuần/tháng mà không giữ toàn bộ lịch sử — chỉ giữ knowledge đã chắt lọc từ các session trước.
Bối cảnh & phạm vi. LMCache (Apache-2.0) là tầng quản lý KV cache cho serving engine (vLLM, NVIDIA Dynamo...) ở quy mô GPU/phân tán: lưu bền và tái dùng KV cache qua request/session/engine, dồn ra bậc lưu trữ (GPU→CPU→disk→remote) để giảm TTFT. Nội dung được diễn giải lại cho phù hợp giấy phép; nguồn: kho LMCache trên GitHub.
Vì sao KHÔNG bê nguyên vào CodeForge: LMCache thao tác bên trong vòng inference (KV cache của model). CodeForge nói chuyện với Ollama qua HTTP API — không sở hữu vòng inference, không truy cập được KV cache. Do đó cố tình LOẠI các tính năng cần engine internals: tiered KV offloading, PD disaggregation, KV transfer (NVLink/RDMA/NIXL), non-prefix reuse/CacheBlend. Chúng không có bề mặt để cắm vào kiến trúc HTTP + local-first, và kéo vào sẽ phản nguyên tắc zero-heavy-dep.
Cái LẤY được là nguyên lý "reuse thay vì recompute", áp ở tầng CodeForge kiểm soát (prompt + context), tận dụng prompt/prefix cache có sẵn của Ollama/llama.cpp:
(a) Prefix-stable prompt ordering — khả thi, tác động rõ. Ollama/llama.cpp tái dùng phần prefix đã tính NẾU đầu prompt giống byte-for-byte giữa các lượt. Hiện CodeForge chèn context động (workspace prime, cwd, listing) vào phần đầu → prefix đổi liên tục → phá cache. Ý tưởng: sắp lại thứ tự message để phần ổn định (system tĩnh) đứng trước cùng, phần biến thiên (context/cwd/turn) đứng sau → tăng prefix-cache hit → giảm TTFT.
- Delta cụ thể: đây là tinh chỉnh Context Engine (#2) về thứ tự, không đổi nội dung.
- Lưu ý mâu thuẫn cần cân nhắc: thứ tự hiện tại đặt workspace-context sau system nhưng trước user message; tối ưu prefix có thể đòi tách "phần bất biến toàn phiên" ra trước. Là thay đổi nhỏ nhưng chạm code thật → thuộc roadmap, không sửa vội; cần đo trước/sau bằng #29.
(b) Application-level context reuse — tùy chọn, thận trọng. KHÔNG phải KV cache, mà cache kết quả ở tầng ứng dụng: context assembled cho một task (đắt vì đọc/rank/nén file) có thể tái dùng nếu working set chưa đổi, dùng chung cơ chế freshness của #2/#32. "Reuse thay vì recompute" áp cho context building, không cho inference.
(c) Cache observability — gộp vào #29. Đo prefix-cache hit / TTFT (xem #29) làm điều kiện đo lường: chỉ tối ưu (a)/(b) khi có số liệu chứng minh cải thiện, đúng tinh thần "observability first" của LMCache.
Điều kiện kích hoạt: làm sau khi #29 (telemetry TTFT/prefix-hit) có dữ liệu — nếu không đo được prefix-cache hit thì (a) là tối ưu mù. Ưu tiên thấp hơn các mục lõi (P0/P1).
USER
│
▼
Intent Analyzer (#30, #31)
│
▼
Context Engine (#2 🟡, #28 🟡, #32 🟡)
│
┌─────────────┼──────────────┐
▼ ▼ ▼
Project Model Memory Knowledge Engine (#14–19 🔵)
(#3 🔵) (#11–13) │
┌─────┼─────┐ GitHub / Docs / Patterns
AST Graph Architecture
(#4) (#5) (#9 🟡, #10 🟡)
│
└─────────────┬──────────────┘
▼
Task Planner (#6 🔵, #7 🔵)
│
▼
Model Router (#21 🟡, #22 🔵)
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Ollama Cheap API Premium AI (#20 🔵, #23 🔵)
│ │ │
└─────────────┼─────────────┘
▼
Execution Engine (✅ toolLoopEngine — sẽ tách theo #1)
│
▼
Tools (✅ ToolRegistry)
│
▼
Verification (#24 🟡)
│
┌─────────────┴─────────────┐
▼ ▼
Success Failure
│ │
▼ ▼
Memory Escalation (#26 🟡, #22 🔵)
- P0 #1 + #2-delta + #3/#5 — nền Project Intelligence + Knowledge Graph là điều kiện cho hầu hết phần sau.
- P1 Planning (#6/#7/#8) và Governance (#9/#10) — tận dụng graph vừa dựng.
- P1 Memory (#11–13) — Decision/Failure Memory nâng chất lượng dài hạn.
- P1 Premium/Router (#20–23) — chỉ sau khi có Escalation triggers rõ ràng (#22, #26, #31).
- P2/P3 — Telemetry (#29) nên làm sớm-vừa để dữ liệu hóa quyết định của Router; #34 (prompt/KV cache efficiency) phụ thuộc #29 (cần đo prefix-cache hit/TTFT trước khi tối ưu thứ tự prompt).
- Knowledge Engine (#14–19) — lớn và độc lập; làm khi các tầng lõi đã ổn.