Skip to content

Latest commit

 

History

History
483 lines (399 loc) · 30.2 KB

File metadata and controls

483 lines (399 loc) · 30.2 KB

CodeForge — Roadmap tương lai

Tài liệu này là định hướng dài hạn. Nó KHÔNG mô tả trạng thái hiện tại — phần "Đã hoàn thành" nằm trong ARCHITECTURE.md (mục Roadmap còn lại + changelog).

Mỗi mục được gắn nhãn trạng thái để tránh làm lại thứ đã có:

  • ✅ Đã có — đã tồn tại trong code; KHÔNG nằm trong phạm vi tương lai (chỉ ghi để tham chiếu).
  • 🟡 Một phần — nền tảng đã có; roadmap chỉ gồm phần còn thiếu (ghi rõ delta).
  • 🔵 Mới — chưa có; là công việc tương lai thực sự.

Nguyên tắc xuyên suốt (kế thừa từ ARCHITECTURE.md): local-first, thuần TypeScript, evolve không rewrite, verify 3 lớp (CORE + M1 + BUILD + tests), không phá hành vi cũ, không over-engineer, quyết định lớn ghi lại bằng ADR.


Bug fixes từ runtime (log thực tế của user)

Ba lỗi phát hiện khi chạy autonomous thật (todo Flutter):

  • ✅ Bug 2 — verify PASS giả (v0.0.76): task sinh code thiếu verify → PASS không chạy gì. Sửa: Planner gán verify mặc định theo stack; stack lạ → required criterion để FAIL.
  • ✅ Bug 3 — autonomous mù thư mục / re-scaffold (v0.0.77): summarizeWorkspace (từ ProjectIndex) + luật anti-re-scaffold trong DEFAULT_SYSTEM + inject workspace overview vào prompt autonomous. Không còn re-run flutter create khi project đã có.
  • ✅ Bug 1 — git checkpoint vỡ: ĐÃ sửa từ v0.0.46 (shellQuote single-quote platform-agnostic + test regression). Log của user là từ build cũ; không cần code lại.

Bản đồ trạng thái nhanh (đối chiếu code hiện tại)

# Hạng mục Trạng thái Ghi chú (đã có gì trong code)
1 Refactor Agent Core (tách module) 🟡 Một phần Điều tra: toolLoopEngine/agentRuntime đã DI sạch; nợ thật ở chatViewProvider. ✅ bước 1 (v0.0.57) SessionCwdManager; ✅ bước 2 (v0.0.58) CommandExecutor; ✅ bước 3 (v0.0.60) WorkspaceMutator (mutation thuần; approval giữ ở provider); ✅ nối ModelRouter (v0.0.59). 🔶 Hợp nhất 2 đường agent (ADR-002, strangler 6 bước): ✅ bước 0+1 (v0.0.69) ADR + chat chạy trên cùng LLMProvider qua port ToolCaller. ✅ bước 2 (v0.0.70) ExecutionObserver seam + chat nối vào. ✅ bước 3 (v0.0.71) completionGuard (anti-fab/verify) thành module thuần. ✅ bước 4a (v0.0.72) toolLoopKernel thuần. ✅ bước 4b (v0.0.73) chat gọi kernel. ✅ bước 4c (v0.0.74) autonomous gọi kernel — HAI đường chung loop body. ✅ bước 5 (v0.0.75) AutonomyGate + setting agentAutonomy (mặc định edits). ✅ bước 6 (v0.0.78) bỏ hẳn chat-only mode — interactive luôn là agent. ADR-002 HOÀN TẤT (một lõi loop chung + client/sink/policy chung + autonomy policy + một interactive mode). Còn (rủi ro cao): xẻ webview render, Memory abstraction; "gọi autonomous như agent" (bước riêng sau)
2 Context Engine 2.0 🟡 Một phần DefaultContextEngine: selection, ranking, compression, token budget, dependency-aware đã có
3 Project Intelligence (index bền) 🟡 Một phần ✅ bước 1 (v0.0.61): buildProjectIndex + WorkspaceIndexer (thuần). ✅ bước 2 (v0.0.62): nối ProjectIndex vào DefaultContextEngine — fast path tái dùng depGraph + symbol candidates. ✅ bước 3 (v0.0.63): ProjectIndexCache (cache + invalidation) — rebuild lazy khi TASK_COMPLETED mark stale; engine nhận getter đọc bản tươi mỗi task; best-effort giữ index tốt gần nhất khi build lỗi. Còn: call graph (hoãn Tree-sitter)
4 Tree-sitter + Static Analysis 🔵 Mới (HOÃN) ADR-001 hoãn Tree-sitter; dependencyGraph (regex) đã có
5 Project Knowledge Graph 🟡 Một phần ✅ bước 1 (v0.0.64): buildKnowledgeGraph(ProjectIndex) → KnowledgeGraph thuần — File/Symbol nodes + declares/imports edges + query API. ✅ bước 2 (v0.0.65): Context Engine truy vấn graph — coDeclarationPeers kéo file cùng khai báo symbol trùng tên (không import, không keyword) với SCORE_RELATED=45; graph cache theo WeakMap<ProjectIndex>. Còn: call edges (hoãn Tree-sitter/ADR-001); có thể nối blastRadius/#8
6 Architecture Planning Phase 🟡 Một phần ✅ bước 1 (v0.0.68): kiểu ProjectBlueprint máy đọc chung + serialize/parse. Còn: LLM Architecture Planner sinh blueprint (project mới) → dẫn xuất DAG; ghi .agent/*.json
7 Existing Project Analysis Phase 🟡 Một phần ✅ bước 1 (v0.0.68): deriveBlueprintFromIndex(ProjectIndex) thuần — suy blueprint hiện trạng (components theo dir + responsibility từ symbol + entryPoints), không LLM. Còn: nối runtime/chat làm pha Project Understanding trước khi sửa
8 Impact Analysis 🟡 Một phần blastRadius.ts (reverse-dep) đã có ở autonomous. ✅ bước 1 (v0.0.66): analyzeImpact → ImpactReport thuần (files/tests/interfaces/entryPoints). ✅ bước 2 (v0.0.67): pre-edit trên đường autonomous — event IMPACT_ANALYZED richer, runtime gọi provider TRƯỚC execute (no I/O, dùng index cache), log ở service; giữ blastRadius cũ. Còn: nối đường chat; call-level (hoãn Tree-sitter)
9 Architecture Guardian 🟡 Một phần src/architecture/ + invariant tests đã có; chưa có vòng generate→guard→repair tự động
10 Machine-readable Blueprint 🟡 Một phần ProjectState/architecture.config.json máy đọc; chưa tách project/constraints/decisions
11 Tách Memory khỏi history 🟡 Một phần EventLog (episodic) + ProjectStore (project) đã có; thiếu các loại memory chuyên biệt
12 Decision Memory 🔵 Mới —
13 Failure Memory 🟡 Một phần selfDiagnostics đọc task.history phát hiện lặp; chưa có store bài học bền vững
14–19 Engineering Knowledge Engine (GitHub RAG) 🔵 Mới —
20 Multi-model Advisor 🔵 Mới Chỉ có provider Ollama local
21 Model Router 🟡 Một phần ✅ nối vào runtime (v0.0.59, trung tính): AutonomousController.resolveExecutionModel + event MODEL_SELECTED. Còn: route Planner, per-task routing, multi-tier (premium/cheap) — #20/#22
22 Escalation System 🔵 Mới —
23 AI Council / Debate 🔵 Mới —
24 Verification Loop 🟡 Một phần LayeredVerifier (autonomous) + verify-reminder (chat) đã có; thiếu runtime-check khép kín
25 Self-review 🔵 Mới —
26 Automatic Failure Escalation 🟡 Một phần Repair bounded theo maxAttempts đã có; chưa escalate ra advisor/knowledge
27 Token/Credit Budget Manager 🔵 Mới —
28 Context Budgeting (phân bổ động) 🟡 Một phần Budget tổng đã có (getContextLength - overhead); chưa phân bổ theo hạng mục
29 Cost/Quality Telemetry 🟡 Một phần ✅ slice 1 (v0.0.54) TTFT/prefix-cache-hit; ✅ slice 2 (v0.0.55) telemetryAggregator; ✅ slice 3 (v0.0.56) dashboard section (per-model + quality + errors). Còn: escalation event + quality-per-model (phụ thuộc #21/#22)
30 Task Classification 🟡 Một phần TaskType đã tồn tại + router dùng nó; chưa có bộ phân loại intent tự động
31 Confidence Estimation 🔵 Mới —
32 Adaptive Context 🟡 Một phần Dependency expansion đã có; chưa mở/thu working set động trong lúc chạy
33 Long-running Project Memory 🔵 Mới Resume theo project đã có; chưa có "chắt lọc knowledge qua nhiều session"
34 Prompt/KV cache efficiency (cảm hứng LMCache) 🔵 Mới Prefix-stable ordering + context reuse; đo prefix-cache hit/TTFT (gộp với #29)

🟥 P0 — Nền tảng kiến trúc

1. Refactor Agent Core / Architecture — 🔵 Mới

Tách rõ trách nhiệm để không dồn vào toolLoopEngine/agentRuntime:

Agent Core
├── Intent
├── Context
├── Planning
├── Execution
├── Verification
├── Memory
└── Model Routing

Mục tiêu: biến prototype thành architecture mở rộng được. Ràng buộc: giữ interface hiện có (ExecutionEngine, Verifier, ContextEngine, ModelRouter) làm ranh giới; refactor tăng dần, mỗi bước verify 3 lớp.

2. Context Intelligence / Context Engine 2.0 — 🟡 Một phần

Đã có (DefaultContextEngine): context selection, relevance ranking, compression, token budgeting, task-specific context, dependency-aware retrieval (forward + reverse).

Delta còn thiếu (phạm vi tương lai):

  • context prioritization tinh vi (không chỉ score đơn) — cân giữa named/keyword/dependency/recency;
  • context freshness (biết context nào cũ/đã đổi trên đĩa để làm mới);
  • truy vấn theo Knowledge Graph (#5) thay vì chỉ text search + import graph.

Kim chỉ nam vẫn là: Task → Context Engine → Relevant Working Set → LLM (KHÔNG "toàn bộ project → LLM").

3. Project Intelligence — 🔵 Mới

Cho agent hiểu project một lần, thay vì khám phá lại mỗi task:

Project Intelligence
├── Project Tree
├── Symbol Index
├── AST
├── Dependency Graph   (✅ dependencyGraph.ts — tái dùng)
├── Call Graph
├── Architecture Map
├── Entry Points
├── Tests
└── Configuration

Ghi chú: dependencyGraph.ts đã có làm điểm khởi đầu; phần index bền (persist + invalidation) là mới.

4. Tree-sitter + Static Analysis — 🔵 Mới (đang HOÃN theo ADR-001)

Source → Tree-sitter → AST → Symbol extraction → Dependency analysis → Project Graph

Chỉ triển khai khi có bằng chứng runtime rằng RegexCodeIntelligence là bottleneck (decision gate trong ADR-001). Giữ abstraction CodeIntelligence để thay TreeSitterCodeIntelligence mà không đụng ContextEngine. KHÔNG coi Tree-sitter là toàn bộ Project Intelligence.

5. Project Knowledge Graph — 🟡 Một phần

Biểu diễn quan hệ giữa: Class, Function, Interface, File, Module, Feature, Dependency, Test, Route, Database, API. Ví dụ: LoginUseCase → AuthRepository → FirebaseAuthRepository → FirebaseAuthDataSource. Context Engine (#2) truy vấn graph thay vì tìm text thuần.

Đã có (v0.0.64, bước 1): context/knowledgeGraph.ts — buildKnowledgeGraph(ProjectIndex) dựng graph thuần: File/Symbol nodes + declares (File→Symbol) / imports (File→File) edges + query API (symbolsIn, filesDeclaring, importsOf/importersOf, relatedFiles). Là lớp view thống nhất trên ProjectIndex. Giới hạn (ADR-001): chỉ những edge regex chứng minh được (declares/imports). Chuỗi ví dụ LoginUseCase → AuthRepository chỉ biểu diễn ở mức FILE-IMPORT, không phải call-level. Call edges (Class/Function-level) hoãn tới Tree-sitter. Đã có (v0.0.65, bước 2): Context Engine truy vấn graph — coDeclarationPeers thêm ứng viên là file cùng khai báo symbol trùng tên với seed (quan hệ import graph + keyword + symbol-index bỏ sót), SCORE_RELATED=45, reason co-declares a symbol with <seed>; graph build 1 lần/index qua WeakMap. Còn: call edges (Class/Function-level) — hoãn Tree-sitter; có thể nối blastRadius/#8.


🟧 P1 — Planning trước khi Coding

6. Architecture Planning Phase — 🟡 Một phần

Project mới KHÔNG code ngay:

User Requirement → Requirement Analysis → Architecture Planner → Project Blueprint → Implementation Plan → Coding

Architecture là artifact máy đọc, không phải prose trong prompt:

.agent/
├── project.json
├── architecture.json
└── implementation-plan.json

Đã có (v0.0.68, bước 1): kiểu ProjectBlueprint máy đọc chung (planning/blueprint.ts) — {version, goal?, origin, techStack?, components{name,responsibility,files}, decisions{title,rationale}, entryPoints} + serialize/parseBlueprint (khớp .agent/architecture.json). Còn: Architecture Planner dùng LLM sinh blueprint cho project mới → dẫn xuất Implementation Plan (DAG); ghi các artifact .agent/*.json.

7. Existing Project Analysis Phase — 🟡 Một phần

Project có sẵn:

User Request → Project Understanding → AST/Graph → Architecture Analysis → Impact Analysis → Modification Plan → Coding

Không để model mò codebase mù quáng (chat đã có prime/auto-cd làm bước sơ khởi). Đã có (v0.0.68, bước 1): deriveBlueprintFromIndex(ProjectIndex) thuần — suy blueprint hiện trạng (components gom theo top-level dir + responsibility tóm từ symbol, entryPoints từ index), KHÔNG cần LLM. Đây là "Project Understanding" ở mức máy đọc. AST/Graph + Impact đã có (#3/#5/#8). Còn: nối runtime/chat để chạy pha này TRƯỚC khi sửa; kết hợp với analyzeImpact (#8) thành Modification Plan.

8. Impact Analysis — 🟡 Một phần

Đã có: blastRadius.ts tính reverse-dep transitive + emit BLAST_RADIUS_COMPUTED (autonomous). Đã có (v0.0.66, bước 1): context/impactAnalysis.ts — analyzeImpact(seeds, ProjectIndex) → ImpactReport thuần, phân loại tập impacted thành 4 nhóm: impactedFiles (importer transitive), impactedTests (isTestFile), impactedInterfaces (symbol kind==="interface" + tên), impactedEntryPoints (∩ entryPoints). Tái dùng blastRadius + ProjectIndex (không I/O). Trả lời "ai phụ thuộc? test nào? interface/platform nào?" ở mức import-graph. Giới hạn (ADR-001): suy từ import edges + regex symbol kind — cờ review surface, không phải chứng minh call-level. Đã có (v0.0.67, bước 2): pre-edit trên đường AUTONOMOUS — agentRuntime gọi analyzeImpact provider NGAY TRƯỚC khi execute mỗi task, emit event richer IMPACT_ANALYZED (files/tests/interfaces/entryPoints). Provider inject từ controller dùng index cache (không I/O); runtime giữ thuần (không biết ProjectIndex). Log tóm tắt ở autonomousService. Giữ BLAST_RADIUS_COMPUTED cũ (sau tool loop) — backward-compat. Còn: nối đường CHAT (webview-coupled — để bước riêng); call-level (hoãn Tree-sitter/ADR-001).


🟨 P1 — Architecture Governance

9. Architecture Guardian — 🟡 Một phần

Đã có: src/architecture/ + invariant tests enforce ranh giới tầng (Phase 3/17). Delta: vòng tự kiểm sau khi agent sinh code:

Generate → Analyze → Architecture Guardian → Violation? ─ No → Continue / Yes → Repair

Biến agent từ code generator thành architecture-aware coding system.

10. Machine-readable Project Blueprint — 🟡 Một phần

Đã có: ProjectState + architecture.config.json máy đọc được. Delta: tách schema project.json / architecture.json / constraints.json / decisions.json để Agent Core truy vấn trực tiếp thay vì Markdown thuần.


🟩 P1 — Memory

11. Tách Memory khỏi Conversation History — 🟡 Một phần

Đã có: EventLog (episodic JSONL), ProjectStore (project memory + resume). Delta: phân loại memory chuyên biệt:

Memory
├── Project Memory        (🟡 ProjectStore)
├── Architecture Memory
├── Decision Memory        (#12)
├── Task Memory
├── Error / Failure Memory (#13)
├── User Preference Memory
└── Episodic History       (🟡 EventLog)

12. Decision Memory — 🔵 Mới

Lưu quyết định + lý do + ngày + phương án đã loại. Ví dụ: "Domain layer must not depend on Firebase — vì cross-platform". Sau nhiều tháng agent vẫn biết tại sao kiến trúc như vậy.

13. Failure Memory — 🟡 Một phần

Đã có: selfDiagnostics đọc task.history phát hiện lặp/kẹt trong một run. Delta: store bài học bền vững qua session: Approach → Failed → Why → Store lesson ("Không dùng cách A vì gây circular dependency") để lần sau không lặp lại.


🟦 P1 — External Engineering Knowledge

14. Engineering Knowledge Engine (GitHub RAG) — 🔵 Mới

Không chỉ "GitHub → vector DB", mà:

Engineering Knowledge Engine
├── Repository Discovery       (#15)
├── Repository Ranking / Quality Scoring (#16)
├── Repository Analysis
├── License Analysis           (#19)
├── Architecture Extraction
├── Pattern Extraction
├── Solution Extraction        (#17)
├── Knowledge Deduplication    (#18)
├── Provenance                 (#19)
├── Semantic Search
└── Knowledge Graph

15. GitHub Repository Discovery — 🔵 Mới

Agent tự tìm repo liên quan task (Flutter offline-first, SQLite sync, DI, MCP...). KHÔNG crawl hàng loạt: Search → Quality filter → Top repositories → Targeted analysis.

16. Repository Quality Scoring — 🔵 Mới

Không chỉ theo stars: Architecture, Code quality, Tests, Maintenance, Documentation, Activity, Dependencies, Community, License.

17. Knowledge Extraction — 🔵 Mới

Không lưu toàn bộ source. Trích: Patterns, Architectures, Solutions, Trade-offs, Implementation techniques, Testing strategies, Failure modes. (Quan trọng hơn vectorize toàn repo.)

18. Knowledge Deduplication / Novelty Detection — 🔵 Mới

Tránh đưa lại thứ model đã biết. Phân loại: Known concept / Known pattern / New implementation / New evidence / New trade-off / New failure mode / Current API-version. Chỉ giữ thứ thực sự bổ sung giá trị.

19. Provenance + License Tracking — 🔵 Mới

Mỗi knowledge item: source, repository, commit/version, license, extraction date. Quan trọng nếu thương mại hóa. Tuân thủ nghiêm license (attribution, giới hạn tái tạo nguyên văn).


🟪 P1 — Premium AI Advisory Layer

20. Multi-model Advisor — 🔵 Mới

Cho phép gọi OpenAI/Claude/DeepSeek/Gemini nhưng KHÔNG làm coding worker chính: Local LLM = Worker · Premium AI = Expert · Agent = Orchestrator. Ràng buộc local-first: premium là tùy chọn, mặc định tắt; không gửi code/secret ra ngoài nếu user không bật rõ ràng (xem safety trong ARCHITECTURE.md).

21. Model Router — 🟡 Một phần

Đã có: modelRouter.ts chọn model theo task/intent + capability constraints (trong model local). Delta: mở rộng thành multi-tier:

Task → Model Router → [ Local | Cheap API | Premium | Multi-model ]

22. Escalation System — 🔵 Mới

Local chỉ gọi premium khi: confidence thấp, architecture change, security-sensitive, large refactor, unknown framework, repeated failure, critical decision. Local attempt → Failure ×2 → ESCALATE → Premium Advisor. Premium chỉ dùng khi thực sự đáng tiền.

23. AI Council / Multi-model Debate — 🔵 Mới

Chỉ cho vấn đề quan trọng: GPT proposal → Claude critique → DeepSeek alternative → Judge → Final. KHÔNG dùng cho task đơn giản (tăng chi phí + latency).


🟫 P2 — Verification & Self-improvement

24. Verification Loop — 🟡 Một phần

Đã có: LayeredVerifier fail-fast (autonomous) + verify-reminder đa toolchain (chat, v0.0.53). Delta: vòng khép kín gồm cả runtime check: Plan → Code → Build → Test → Static analysis → Architecture check → Runtime check → Fix → Verify.

25. Self-review — 🔵 Mới

Sau khi code: "review what I just changed" — dùng context độc lập hoặc model khác để tránh cùng một reasoning path tự xác nhận chính mình.

26. Automatic Failure Escalation — 🟡 Một phần

Đã có: repair bounded theo maxAttempts, self-diagnostics phát hiện lặp. Delta: khi fail lặp → Analyze failure → Retrieve knowledge (#14) → Ask premium advisor (#22) thay vì loop cùng một cách sửa.


🟥 P2 — Cost & Performance Intelligence

27. Token / Credit Budget Manager — 🔵 Mới

Agent biết: Task budget, Context budget, Model budget, API budget. Simple → 0 premium · Medium → cheap · Complex → premium · Critical → multi-model.

28. Context Budgeting (phân bổ động) — 🟡 Một phần

Đã có: budget tổng (getContextLength - overhead, fallback DEFAULT_CONTEXT_BUDGET). Delta: phân bổ theo hạng mục, ví dụ với 16k:

Task 2k · Architecture 1k · Relevant code 7k · Memory 1k · Knowledge 2k · Safety/system 2k

Agent tự cân đối thay vì một ngân sách phẳng cho code.

29. Cost/Quality Telemetry — 🔵 Mới

Theo dõi: Task, Model used, Input/Output tokens, Latency, Success, Retries, Premium escalations, Final quality. Dùng để Model Router (#21) tự cải thiện ("task X: local success 91%, khỏi cần premium"). EventLog là nền thu thập sẵn có.

Bổ sung (cảm hứng LMCache — observability là first-class): đo các chỉ số hiệu năng prompt/cache mà tầng HTTP-to-Ollama quan sát được. ✅ ĐÃ LÀM (slice 1, v0.0.54):

  • TTFT (time-to-first-token) — đo client-side ở đường stream (ChatStats.ttftMs);
  • prompt-eval time (promptEvalMs từ prompt_eval_duration);
  • prefix-cache hit (heuristic) — derivePrefixCacheHit: prompt lớn mà prefill ~0ms hoặc prompt-eval throughput vượt ngưỡng ⇒ prefix reuse (trả "unknown" khi thiếu tín hiệu);
  • persist qua event TELEMETRY_RECORDED (metadata-only) vào EventLog; module thuần telemetry/promptTelemetry.ts.

Đây là điều kiện đo lường cho #34 (không tối ưu được cái không đo được). Vẫn thuần local, đọc từ metadata Ollama, không cần truy cập engine internals.

✅ ĐÃ LÀM (slice 2, v0.0.55): telemetry/telemetryAggregator.ts — aggregateTelemetry(events) gấp event đã persist thành report: per-model (calls, tokens, avg TTFT/tokensPerSec/promptEvalMs, prefix-cache hits/misses/hitRate) + quality per-run (completed/failed/blocked, totalAttempts, retries, repairs, successRate) + llmErrorsByCode. Thuần, không thêm tracking runtime.

✅ ĐÃ LÀM (slice 3, v0.0.56): section Telemetry trong Dashboard — DashboardPanel fold event qua aggregateTelemetry và render per-model (tokens/TTFT/prefill/tok-s/prefix-cache hit rate)

  • quality (success rate/retries/repairs/blocked) + LLM errors. Consumer thật đầu tiên của aggregator.

Còn lại của #29: (a) gán quality theo model (cần multi-model routing #21/#22 để mỗi attempt biết model); (b) escalation event (phụ thuộc #22 — hiện escalations=0). Cả hai KHÔNG làm độc lập được; chờ tầng Premium/Router.


🟦 P3 — Advanced Agent Intelligence

30. Task Classification — 🟡 Một phần

Đã có: TaskType (implementation/refactor/debug/test/build/review/architecture/analysis/research)

  • router dùng nó. Delta: bộ phân loại tự động từ yêu cầu user → chọn pipeline khác nhau cho Create/Modify/ Refactor/Debug/Research/Explain/Test/Review/Architecture.

31. Confidence Estimation — 🔵 Mới

Agent tự đánh giá confidence để quyết định: tiếp tục / tìm thêm context / retry / escalate.

32. Adaptive Context — 🟡 Một phần

Đã có: dependency expansion khi build context. Delta: mở/thu working set động trong lúc chạy: task → discover file → new dependency → retrieve new context → update working set.

33. Long-running Project Memory — 🔵 Mới

Làm việc cùng project qua nhiều tuần/tháng mà không giữ toàn bộ lịch sử — chỉ giữ knowledge đã chắt lọc từ các session trước.


🟥 P2/P3 — Hiệu năng prompt & cache

34. Prompt/KV cache efficiency (cảm hứng LMCache) — 🔵 Mới

Bối cảnh & phạm vi. LMCache (Apache-2.0) là tầng quản lý KV cache cho serving engine (vLLM, NVIDIA Dynamo...) ở quy mô GPU/phân tán: lưu bền và tái dùng KV cache qua request/session/engine, dồn ra bậc lưu trữ (GPU→CPU→disk→remote) để giảm TTFT. Nội dung được diễn giải lại cho phù hợp giấy phép; nguồn: kho LMCache trên GitHub.

Vì sao KHÔNG bê nguyên vào CodeForge: LMCache thao tác bên trong vòng inference (KV cache của model). CodeForge nói chuyện với Ollama qua HTTP API — không sở hữu vòng inference, không truy cập được KV cache. Do đó cố tình LOẠI các tính năng cần engine internals: tiered KV offloading, PD disaggregation, KV transfer (NVLink/RDMA/NIXL), non-prefix reuse/CacheBlend. Chúng không có bề mặt để cắm vào kiến trúc HTTP + local-first, và kéo vào sẽ phản nguyên tắc zero-heavy-dep.

Cái LẤY được là nguyên lý "reuse thay vì recompute", áp ở tầng CodeForge kiểm soát (prompt + context), tận dụng prompt/prefix cache có sẵn của Ollama/llama.cpp:

(a) Prefix-stable prompt ordering — khả thi, tác động rõ. Ollama/llama.cpp tái dùng phần prefix đã tính NẾU đầu prompt giống byte-for-byte giữa các lượt. Hiện CodeForge chèn context động (workspace prime, cwd, listing) vào phần đầu → prefix đổi liên tục → phá cache. Ý tưởng: sắp lại thứ tự message để phần ổn định (system tĩnh) đứng trước cùng, phần biến thiên (context/cwd/turn) đứng sau → tăng prefix-cache hit → giảm TTFT.

  • Delta cụ thể: đây là tinh chỉnh Context Engine (#2) về thứ tự, không đổi nội dung.
  • Lưu ý mâu thuẫn cần cân nhắc: thứ tự hiện tại đặt workspace-context sau system nhưng trước user message; tối ưu prefix có thể đòi tách "phần bất biến toàn phiên" ra trước. Là thay đổi nhỏ nhưng chạm code thật → thuộc roadmap, không sửa vội; cần đo trước/sau bằng #29.

(b) Application-level context reuse — tùy chọn, thận trọng. KHÔNG phải KV cache, mà cache kết quả ở tầng ứng dụng: context assembled cho một task (đắt vì đọc/rank/nén file) có thể tái dùng nếu working set chưa đổi, dùng chung cơ chế freshness của #2/#32. "Reuse thay vì recompute" áp cho context building, không cho inference.

(c) Cache observability — gộp vào #29. Đo prefix-cache hit / TTFT (xem #29) làm điều kiện đo lường: chỉ tối ưu (a)/(b) khi có số liệu chứng minh cải thiện, đúng tinh thần "observability first" của LMCache.

Điều kiện kích hoạt: làm sau khi #29 (telemetry TTFT/prefix-hit) có dữ liệu — nếu không đo được prefix-cache hit thì (a) là tối ưu mù. Ưu tiên thấp hơn các mục lõi (P0/P1).


Kiến trúc tổng thể mục tiêu

                         USER
                           │
                           ▼
                    Intent Analyzer          (#30, #31)
                           │
                           ▼
                    Context Engine           (#2 🟡, #28 🟡, #32 🟡)
                           │
             ┌─────────────┼──────────────┐
             ▼             ▼              ▼
       Project Model     Memory       Knowledge Engine   (#14–19 🔵)
        (#3 🔵)          (#11–13)          │
       ┌─────┼─────┐                    GitHub / Docs / Patterns
      AST  Graph Architecture
     (#4)  (#5)  (#9 🟡, #10 🟡)
             │
             └─────────────┬──────────────┘
                           ▼
                      Task Planner           (#6 🔵, #7 🔵)
                           │
                           ▼
                      Model Router           (#21 🟡, #22 🔵)
                           │
             ┌─────────────┼─────────────┐
             ▼             ▼             ▼
          Ollama       Cheap API     Premium AI          (#20 🔵, #23 🔵)
             │             │             │
             └─────────────┼─────────────┘
                           ▼
                    Execution Engine         (✅ toolLoopEngine — sẽ tách theo #1)
                           │
                           ▼
                         Tools               (✅ ToolRegistry)
                           │
                           ▼
                      Verification            (#24 🟡)
                           │
             ┌─────────────┴─────────────┐
             ▼                           ▼
          Success                      Failure
             │                           │
             ▼                           ▼
          Memory                    Escalation           (#26 🟡, #22 🔵)

Thứ tự đề xuất (không ràng buộc)

  1. P0 #1 + #2-delta + #3/#5 — nền Project Intelligence + Knowledge Graph là điều kiện cho hầu hết phần sau.
  2. P1 Planning (#6/#7/#8) và Governance (#9/#10) — tận dụng graph vừa dựng.
  3. P1 Memory (#11–13) — Decision/Failure Memory nâng chất lượng dài hạn.
  4. P1 Premium/Router (#20–23) — chỉ sau khi có Escalation triggers rõ ràng (#22, #26, #31).
  5. P2/P3 — Telemetry (#29) nên làm sớm-vừa để dữ liệu hóa quyết định của Router; #34 (prompt/KV cache efficiency) phụ thuộc #29 (cần đo prefix-cache hit/TTFT trước khi tối ưu thứ tự prompt).
  6. Knowledge Engine (#14–19) — lớn và độc lập; làm khi các tầng lõi đã ổn.