Repository navigation
Conversation
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- README: 🇰🇷 한국 주식 리서치 section with architecture + sequence diagrams (mermaid), tool table, DART setup, market-specific notes - CLAUDE.md: trap list, key contracts, KR adaptation roadmap (Phase 1–4) Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Add Korean stock data pipe (Phase 1) DART-backed financial data for KOSPI/KOSDAQ tickers: - src/data/fetchers/dart-corp-codes.ts — fetch + unzip + parse DART corpCode.xml (fflate + fast-xml-parser, zero native deps) - src/data/ticker-registry.ts — memory+disk cache (7-day TTL) mapping 6-digit ticker / corp name → corp_code; stale-on-error fallback + in-flight dedup so DART downtime doesn't break analysis - src/tools/finance-kr/api.ts — DART HTTP wrapper; validates body `status`, strips crtfc_key from cached/logged URLs - src/tools/finance-kr/sub-tools/get-business-report.ts — first KR tool, fnlttSinglAcntAll.json (사업/반기/분기보고서, K-IFRS) - src/tools/finance-kr/get-financials-kr.ts — LLM router meta-tool; natural-language KR query → DART sub-tool(s) Wiring: - registry.ts gates registration on DART_API_KEY (ignores `your-`) - cli.ts summarizeToolResult handles get_financials_kr - prompts.ts routes 6-digit tickers to the KR tool typecheck + 60/60 tests pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Document DART_API_KEY in env.example Phase 1 gated get_financials_kr on DART_API_KEY but the example env file never listed it. Add it next to the other market-data keys. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Add Korean disclosure tools (Phase 2)
Add three DART-backed disclosure tools, the Korean equivalents of the US
filings/13F/insider tools:
- get_filings_kr — DART 공시검색 (/list.json), date/type filters
- get_large_holders_kr — 대량보유 상황보고 5%룰 (/majorstock.json)
- get_insider_trades_kr — 임원·주요주주 소유보고 (/elestock.json)
Each is a simple standalone tool taking a 6-digit ticker (no inner LLM
router); the main agent resolves company name → ticker and routes by ticker
pattern, per the CLAUDE.md "Decided" note. Ticker → corp_code via the existing
resolveTicker registry; results returned as { data } through formatToolResult.
majorstock/elestock have no server-side date filter, so the full history is
sorted by rcept_dt desc and capped client-side. DART status 013 (no data) is
treated as an empty result, not an error.
Wires registration (DART_API_KEY gate), the 6-digit routing rule in the system
prompt, and UI status lines (Found N filings/holders/reports). Shared helpers
in finance-kr/utils.ts with colocated tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Make get_filings_kr ordering explicit and guard list parsing
Address review nits on the filings tool:
- Send sort=date&sort_mth=desc instead of relying on DART's implicit default
ordering, so the "most recent first" promise holds regardless of pblntf_ty.
- Guard data.list with Array.isArray, matching get_large_holders_kr and
get_insider_trades_kr.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ings (Phase 3) (#5) * Add Korea-specific tools — short balance, foreign ownership, NPS holdings (Phase 3) Phase 3 of Korean stock support. Unlike Phases 1–2 (all clean DART JSON+key), none of these three live in a clean official keyed API, so each uses a different source: - get_foreign_ownership_kr — Naver mobile JSON (m.stock.naver.com trend), keyless, registered unconditionally. - get_short_balance_kr — KRX Data Marketplace login scrape (bld MDCSTAT30502). KRX now returns "LOGOUT" to anonymous requests, so a member session is required; login flow ported from pykrx (krx-session.ts) with a KRX_COOKIE override for social-login accounts. KRX keys on ISIN, not ticker → new krx-instrument-registry (bld MDCSTAT01901). Gated by KRX_ID/KRX_PW or KRX_COOKIE. - get_nps_holdings — data.go.kr odcloud dataset 3070507 (year-end, by stock name). Gated by DATA_GO_KR_SERVICE_KEY. All three live-verified (삼성 외국인 48.27%, 공매도 89/97 non-zero, NPS 1,200종목 삼성 지분율 7.26%). typecheck + 84 tests pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: README sourcing, ISO date normalization, more tests - README: split the 3 Korea-specific tools out of the "DART 기반" section into their own table with per-tool source/key; fix OPEN_DART_KEY → DART_API_KEY. - Normalize tool dates to ISO YYYY-MM-DD via new toIsoDate (KRX YYYY/MM/DD and Naver YYYYMMDD were inconsistent across tools). - Tests: toIsoDate, defaultRange, sessionFromCookie (caught + fixed an empty-key parse bug on leading-space segments), isLogoutBody. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…e 4) (#6) * Adapt skills for Korean market — DCF market-branch + 물적분할 skill (Phase 4) Phase 4 (skill adaptation) of Korean stock support. Pure markdown — skills auto-discovered by src/skills/registry.ts, no registry/prompt/cli.ts change. - dcf-valuation skill now branches on market (Step 0): 6-digit ticker → KR path using get_financials_kr, ~3% 국고채 risk-free, ~22% K-IFRS corporate tax, ~2% terminal growth, KRW. New sector-wacc-kr.md. 거래세/배당세 surface as an investor-level "세후 실현수익률" caveat, not in the intrinsic-value math. Kept a single skill (branches internally) so write-memo's dcf-valuation invocation keeps working. - New skill kr-spinoff-analysis: 물적분할/인적분할 events from the parent shareholder's POV (dilution / holding-co discount / double-counting) via get_filings_kr material disclosures. - registry.test.ts guard asserting all four builtin skills are discovered. bun test: 99 pass; typecheck clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Fix KR DCF skill: correct 2026 거래세 rate + Step 6 sensitivity override Review follow-ups on PR #6: - 증권거래세 was stated as ~0.18% "as of 2026" (the 2024 transitional rate). Because 금융투자소득세 was abolished, the rate reverted up for 2026: KOSPI ~0.20% all-in (0.05% 거래세 + 0.15% 농특세), KOSDAQ 0.20%. - Add the KR terminal-growth override (1.5/2.0/2.5) to Step 6 itself, where the sensitivity matrix is built — previously only stated in Step 4 and Step 8. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Translate KR skill content to Korean KR-facing skill prose → Korean for Korean users. Identifiers preserved: frontmatter name (dcf-valuation, kr-spinoff-analysis), tool names (get_financials_kr …), and the sector-wacc-kr.md link stay as-is. - kr-spinoff/SKILL.md: full Korean (100% KR-only skill) - dcf/sector-wacc-kr.md: full Korean - dcf/SKILL.md: 🇰🇷 KR override block *content* → Korean; US path and shared scaffolding stay English (the skill serves US stocks too) bun test (registry) 4 pass; typecheck clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Translate DCF skill fully to Korean (US path + shared scaffolding) The agent's whole flow targets Korean users, so translate the DCF skill end-to-end — Step 0, US path, shared mechanics (Steps 2/5/7), output format — not just the KR override blocks. Also translate the US sector-wacc.md companion. Machine-facing literals kept in English by design: frontmatter name, tool names (get_financials …), JSON field names (free_cash_flow …), file links, and the example query strings passed to the US get_financials tool (English-oriented router). Description keeps English trigger terms alongside Korean so English DCF queries still match. bun test (registry) 4 pass; typecheck clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Translate README fully to Korean Whole README → Korean for Korean users. Headings translated and the table-of-contents anchors updated to match. Code blocks, commands, URLs, env var names, badges, image tags, and the mermaid diagrams (already Korean) left intact. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * README: replace top architecture PNG with mermaid + add usage examples - Convert the static architecture screenshot at the top of the README to a maintainable mermaid flowchart (User → Agent Loop {LLM/Tools/ Scratchpad} → APIs, plus Evaluation Layer). - Add a "💡 사용 예시" section with concrete, agent-optimized example queries grouped by capability (financials, DCF, DART filings/ownership, Korea-specific data, spin-off analysis, cross-market, memo/sentiment), with key-requirement notes. TOC updated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Restructure README for a cleaner reading flow The structure had grown messy: the legal disclaimer sat above the overview, setup was split across three sections (Prerequisites / Install / Run) with Bun install duplicated, env keys were scattered (and missing the Korean ones), and the cross-market example was duplicated. - Reorder: title → architecture → badges → Overview → 🇰🇷 → Usage → Quick Start → Evaluate/Debug/WhatsApp → Contribute → License → Disclaimer (moved to the bottom). - Merge Prerequisites + Install + Run into one "🚀 빠른 시작" with a single consolidated .env block that now includes the Korean keys (DART / KRX / data.go.kr), web search, and X. - Fix env var bug: OPEN_DART_KEY → DART_API_KEY (matches the code). - De-duplicate the cross-market example (folded into Usage Examples). - TOC updated to the new order. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Update README.md * README: move architecture diagram into Overview, tidy TOC On top of the manual edits (tagline split, WhatsApp/Contribute sections removed): - Move the architecture mermaid from the very top into 개요 as an "### 아키텍처" subsection — the diagram no longer precedes what the project even is. - Clean the TOC: drop the dead links to the removed WhatsApp/Contribute sections, and group 평가/디버깅 under one "🔧 더 알아보기" section. TOC is now 7 entries that all resolve. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * README: use original architecture image instead of mermaid The Korean additions don't change the high-level architecture (they only add tools under Tools/Subagents and data sources under APIs/Other), so the original diagram is still accurate. Swap the redrawn mermaid back to the polished source image; the 🇰🇷 "동작 방식" mermaid still covers the Korea-specific routing detail. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Update README.md * README: remove the Korea-section mermaid diagrams Drop the 🇰🇷 "동작 방식" flowchart and the "종목 코드 해결 흐름" sequence diagram; keep the explanatory text (routing, TickerRegistry caching). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * README: add measured Dexter vs claude.ai comparison section Real head-to-head (2026-05) on 5 identical questions: Dexter (GPT-5.5 + finance harness) vs claude.ai (Opus 4.8 + web search). Both frontier models, so the section isolates the harness — primary-source tools (DART/KRX/Naver) + KR skills — as the differentiator. Honest: notes where claude.ai did better (Samsung P&L) and Dexter's stale hardcoded risk-free rate. ChatGPT-free results dropped as not meaningful. TOC updated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ook (#7) * KR financials: normalized summary-first + cross-signal research playbook Make Korean stock research produce a differentiated, data-grounded answer instead of a generic chatbot-style summary. - get_financials_kr: return a normalized per-period summary (revenue, operating profit, net income, EPS, assets/liabilities/equity, cash flow, capex + margins, ROE, FCF, YoY) in KRW instead of an 80-100KB raw DART dump. Raw line items are persisted to a pretty-printed file (rawLineItemsFile) for drill-down. Fixes the model failing to surface earnings: the old single-line payload blew the size cap and read_file could not page it, so revenue/OP/NI never reached the answer. Match by stable account_id then exact account_nm (IS+CIS for single-statement filers); no substring matching (avoids 순영업이익 -> 영업이익 contamination). - System prompt: add a Korean research playbook -- gather broadly for open-ended questions, synthesize foreign flows / shorts / holders / filings into ONE thesis, treat 대량보유 as a valuation modifier. Scope the broad sweep so narrow / skill-driven queries (DCF, 물적분할) gather only what they need. - registry: guard web_search providers against `your-` placeholder keys (a copied env.example was registering a web_search tool that 401'd on every call). - DCF / kr-spinoff skills: read the new normalized summary fields. - README: document the normalized summary + cross-signal synthesis; correct the now-stale benchmark row (삼성 손익 보류 -> reliable extraction). - Add normalize-financials unit tests and scripts/kr-eval.ts (a probe runner). Verified: typecheck + 113 tests; live before/after on 삼성전자, SK하이닉스, KB금융 (bank/holdco fallback), 네이버 DCF, LG화학 물적분할. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: bound raw-file cache, harden prune, DRY DART gate Follow-up to the summary-first change, applying code-review findings. - get-business-report: prune kr-financials-*.json dumps to the newest RAW_FILE_KEEP (was unbounded across sessions). Harden the cleanup — sanitize corp_code in the path (only externally-sourced component), never delete the file just written (keepName guard), and tolerate files vanishing under concurrent calls (per-file + per-unlink try/catch). The agent's call_*.txt persists are out of scope by prefix filter. - prompts: gate the KR research section with checkApiKeyExists('DART_API_KEY') instead of an inline process.env check — DRY, also honors .env, and treats a whitespace-only key as absent. - kr-eval: maxIterations 12 -> 10 so probe runs match app behavior. - normalize: document ROE basis (total net income / total equity, incl. NCI). - tests: prune safety invariants (preserves call_*.txt + the current file, bounded count, empty-dir no-op) and quarterly YoY-null when prior is absent. typecheck + 116 tests pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: correct stale CLAUDE.md note on final-answer flow CLAUDE.md claimed the final answer is "a separate LLM call with no tools bound". The agent loop has no such pass — tools stay bound on every call and the answer is just the text from the turn where the model stops emitting tool calls (handleDirectResponse). Clarify that answer quality is governed entirely by the system prompt. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ing (#8) * KR eval harness: fixed question bank + record/replay + LLM-judge scoring Extend the single-query seed (scripts/kr-eval.ts, unchanged) into a fixed-question + scorer harness for the Korean-stock agent path. scripts/kr-eval/ questions.ts — 12-question bank + rubric metadata (expected/required tools, judge dimensions per question) replay.ts — record (dump done.toolCalls → fixture) + replay (fixture-backed tool stubs via the Agent transformTools hook); match policy + _replay_miss sentinel. Records at the tool-output layer (free, since done.toolCalls already carries {tool,args,result}) rather than the fetch layer (no KRX cookie/ZIP/login reconstruction). judge.ts — per-dimension LLM judges via callLlm + Zod {score,comment} (earnings_yoy / cross_signal / governance / grounding), grounded in the Korean playbook in src/agent/prompts.ts scorer.ts — pure deterministic tool-firing gate + combine + aggregate runner.ts — entrypoint: mode (live/record/replay), env-gated skip, replay dummy-env shim, run agent, score, report report.ts — console report *.test.ts — scorer + replay unit tests (16 tests, no network/LLM) src change (only production touch): optional transformTools hook on AgentConfig, applied in Agent.create() (default no-op) — the eval seam replay uses to swap live KR tools for fixture stubs while the agent LLM still routes freely. package.json: kr-eval / kr-eval:record / kr-eval:replay scripts. Verified: bun run typecheck (src) + standalone tsc (scripts) clean; full suite 132 pass / 0 fail; kr-eval 16 pass / 0 fail; offline runner smoke test. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * KR eval: classify replay misses as inconclusive + add grounding to earnings Qs First live run surfaced two things on samsung-fundamentals: - replay diverged because the live agent took a tool path (read_file) the record run didn't, so the call had no fixture (replayMiss). A miss means the data wasn't fully pinned, so the score isn't comparable to a baseline. - the recorded get_financials_kr data has prior:null (no prior-year figures), yet the live answer scored earnings_yoy=1.0 — i.e. the YoY was likely ungrounded. That question had no grounding dimension to catch it. Changes: - scorer: a result with replayMisses is marked `inconclusive` (new status), excluded from pass/fail counts, dimension stats, and the exit code — so a non-faithful replay no longer reads as a false regression. report.ts renders INCONC + prompts to re-record the fixture. - questions: add `grounding` to samsung-fundamentals and cross-signal-alteogen so the judge cross-checks that YoY/figures are actually supported by tool data. Separately noted for follow-up: get_financials_kr returning prior:null for quarterly line items (DART frmtrm not mapped?) — a possible KR-financials bug, outside this harness. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * KR financials: read quarterly prior-year from frmtrm_q_amount (fixes null YoY) get_financials_kr returned prior:null for every quarterly/semiannual P&L line, so revenue/operating-profit/net-income YoY was silently uncomputable — and the agent would fill the gap from memory (ungrounded YoY). Surfaced by the KR eval harness on samsung-fundamentals. Root cause: for 분기/반기 reports DART leaves frmtrm_amount empty on flow statements (IS/CIS/CF) and carries the prior-year SAME period in frmtrm_q_amount (전년 동기 누적). The annual report and the quarterly balance sheet have no _q field. toMetric only read frmtrm_amount. Fix: prefer frmtrm_q_amount, fall back to frmtrm_amount — correct for all cases (annual → frmtrm_amount; quarterly flows → frmtrm_q_amount; quarterly BS → frmtrm_amount = 전년말). Verified against the real Samsung 2026 Q1 payload: revenue 133.9조 vs 79.1조 (+69.2%), OP +756%, NI +474% YoY. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * KR eval: first full fixture baseline + rubric calibration (12/12) Recorded all 12 questions live (KR_EVAL_MODE=record) → committed fixtures as the replay/regression baseline (156K total; small because we record tool outputs, not raw HTTP — no corp-code ZIPs). The first full run scored 9/12 and surfaced a rubric mis-calibration (mine, not the agent's): the cross_signal dimension (which requires weaving 수급+실적+지배구조 into one thesis) was applied to single-signal questions that only gather one signal, and governance (deep 지주사/승계 valuation) to a factual ownership-change lookup. The genuine synthesis/governance questions scored well (samsung-synthesis cross_signal 0.86, alteogen 0.94, large-holders 0.92, spinoff 1.00). Recalibrated dimensions to match question intent: - foreign-flow-hynix: cross_signal → grounding - institutional-vs-foreign: cross_signal → grounding - insider-trades-celltrion: governance → grounding cross_signal/governance now apply only to questions that actually call for synthesis/analysis. Live re-run of the three: PASS (grounding 0.88/1.00/0.82). Also confirms the frmtrm_q_amount fix (a42fcc4): samsung-fundamentals now passes earnings_yoy=1.00 AND grounding=1.00 (previously the YoY was ungrounded). Note: fixtures hold a point-in-time market snapshot; re-record to refresh. * KR eval: address PR #8 review (digest budget, replay match, cleanups) - judge.toolDigest: split the budget EVENLY across tool calls instead of greedy first-fit, so a later tool's data is never dropped wholesale — fixes biased grounding/cross_signal scoring on multi-tool answers. Add judge.test.ts. - replay.resolve: before falling back to the first recorded call, match on the entity key (ticker) so a multi-ticker tool can't return another ticker's data when only secondary args differ. - scorer: drop the unused compositeScore field (computed, never surfaced). - runner: record mode now exits 0 regardless of rubric scores (it's a capture action); pass/fail exit code applies to live/replay only. Comment the repeat>1 tool-firing caveat. - fixtures: drop meta.recordedAt so re-recording only diffs when data changes. - tsconfig.scripts.json + fold a scripts typecheck into `typecheck` so CI (bun run typecheck) now covers scripts/, which the project tsconfig excludes. typecheck (src + scripts) clean; bun test 137 pass / 0 fail. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * KR eval: relax governance threshold to 0.65 for synthesis questions Across 6 replay runs, governance on the two synthesis questions (samsung-synthesis, cross-signal-alteogen) clustered 0.70–0.85 — distinct from dedicated governance questions (large-holders 0.92+, spinoff 0.98). In a synthesis answer governance is one strand, not the focus, so the judge consistently scores it lower. The default 0.70 left two zero-margin cells (governance==0.70). Per-question threshold 0.65 adds a buffer without lowering the bar where governance IS the focus. Insurance, not a fix-for-failure: in 6 runs governance never dipped below 0.70, so both thresholds would have passed every run; this removes the boundary coin-flip. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* get_market_data_kr: 현재가·시총·PER/PBR·목표주가 컨센서스 (Naver keyless) The get_market_data equivalent for 6-digit KR tickers (get_market_data does not resolve them). A single keyless Naver /integration call returns: - price + daily change, 52-week range, session OHLC - market cap (조/억 파싱) and derived shares outstanding (= 시총 ÷ 현재가) - PER/PBR/EPS/BPS + forward 추정PER/EPS, dividend yield - analyst consensus: 목표주가 + mean recommendation + implied upside - same-industry peers (ticker, price, market cap) Fills the biggest KR gap: DCF/valuation previously inferred current price and share count via web_search. The DCF skill now sources both from this tool. - parseNaverMetric: strips 배/원/% suffixes (parseKrxNumber broke on them) - parseKoreanMarketCapToKRW: 조/억 → KRW - registered unconditionally (keyless, like get_foreign_ownership_kr) - wired into prompt routing + KR research section, DCF skill, write-memo KR branch - refreshed kr-eval fixtures so synthesis questions (#4/#14) capture the new tool - unit tests + live-verified across KOSPI / KOSDAQ / holding names Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * get_market_data_kr: review fixes — not-found guard, 4xx message, parser cleanup - fetchNaverIntegration: an unknown 6-digit code returns 409 (not 404); surface it as a clear "check the 6-digit ticker" error instead of a raw HTTP status the model can't act on - hasNoMarketData guard: return _error instead of an all-null snapshot when a 200 payload carries no usable data (name+price+marketCap all null). Cannot false-positive on a valid ticker, whose 200 always includes stockName - parseNaverMetric: drop unused 주 from the strip class — 주 only appears in Naver label keys (52주, 주당배당금), never in parsed values — + regression test - tests: hasNoMarketData false-positive/negative coverage, parseNaverMetric edges Verified the foreign-ownership sibling does NOT need the same 4xx branch: its /trend endpoint returns 200 + [] (not 409) for an unknown ticker, so an empty list is the correct response (indistinguishable from a valid no-data ticker). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* DCF 순부채 부호 수정: KR 이자부채(totalDebt)·단기금융상품 정규화 + EV→Equity 부호 규칙 DCF가 주식가치 = EV − 순부채에서 KR 경로의 순부채 항으로 부채총계 (totalLiabilities)를 "순부채 근사"로 차감했다. 부채총계는 매입채무·선수금 등 영업성 부채까지 포함하므로 이자부채가 아니며, 삼성처럼 순현금 기업에선 부호가 뒤집혀 적정가가 과소평가됐다(삼성 ~112k). 이번 PR(#11)이 DCF에 실제 가격을 공급하면서 더 도드라진 정확성 버그. 수정: - normalize-financials.ts: KR 정규화 요약에 balanceSheet.totalDebt(이자부채 합계)와 shortTermInvestments(단기금융상품) 추가. 새 sumMetrics 헬퍼 + DEBT_SUM_SPEC로 여러 차입금 라인을 합산(차입금 라벨이 없으면 null → 모델이 raw 파일 폴백). - skills/dcf/SKILL.md: line 42 KR override를 totalDebt/shortTermInvestments 기반으로 재작성, Step 5에 EV→Equity 브리지와 순현금 부호 규칙(Net Debt<0이면 EV에 가산)을 양 경로(US/KR) 모두 명시, Step 7에 순부채 부호·규모 sanity check 추가. - US 경로는 데이터(total_debt/cash_and_equivalents/current_investments)가 이미 모델에 전달되므로 마크다운만 명확화. account_id 검증: 실제 DART CFS 페이로드(삼성전자 005930, LG화학 051910, 알테오젠 196170)로 라벨/id 검증. 후보 5개 중 4개가 실데이터와 달라 교체(예: 삼성 단기차입금은 account_id '-표준계정코드 미사용-' → account_nm으로 매칭). LG화학이 차입금을 동일 라벨 두 행으로 보고해, 승인된 플랜의 "컴포넌트당 1행" 대신 매칭되는 모든 행을 합산하는 set 기반 방식으로 구현(안 그러면 ~12조 누락). 검증: bun run typecheck ✅ · bun test 187 pass ✅. 삼성 순현금 100.6조(=차입금 25.2조 − 현금 57.9조 − 단기금융상품 68.0조) 확인, 적정가 ~112k → ~144k. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * 리뷰 반영: 순부채 문서/주석 정밀화 (동작 변경 없음) 코드 리뷰 후속 다듬기 (4건, 모두 doc/comment — 동작·테스트 영향 없음): - SKILL.md 폴백 지침: totalDebt가 null일 때 raw 합산을 account_id가 아니라 account_nm(라벨) 우선으로 안내(삼성 단기차입금이 '-표준계정코드 미사용-'이라 라벨이 더 안정적이라는 이번 PR의 교훈과 일치). - SKILL.md Net Debt 식: shortTermInvestments/cashAndEquivalents가 null이면 0으로 본다고 명시(US 경로의 "없으면 0"과 대칭). - SKILL.md Step 5: LLM이 평문으로 읽는 텍스트라 HTML 엔티티 제거. - DEBT_SUM_SPEC 주석: 실측 검증 라벨(삼성·LG화학·알테오젠)과 best-effort 추정 라벨 (유동성사채/단기사채/신주인수권부사채 등)을 구분하고, "차입금" 소계+세부 동시보고 시 이중계상 가능성 caveat을 명시. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * README: 에이전트 진화 반영 + 챗봇 차별점 강화 - get_market_data_kr를 Korea-specific 툴 표에 추가(키 불필요 Naver; 그간 누락 — KR DCF 현재가/주식수 공급원) - DCF 자동 분기 설명에 "순부채=이자부채(부채총계 아님)−현금·단기금융상품, 순현금 가산" 규칙 반영 - "Dexter가 만든 차이"에 '정확한 밸류에이션'(이자부채 순부채) 추가, '1차 출처'에 키리스(외국인·현재가) 우위 명시, '감사·재현 가능'에 KR record/replay 하네스 근거 추가 - 개요에 '1차 출처 + 감사 가능' 차별점 불릿(비교 섹션 링크) - 한국 주식 평가 하네스(kr-eval record/replay) 섹션 추가 - env 블록을 env.example과 일치: LLM에 MOONSHOT/DEEPSEEK, 웹검색에 Perplexity/LangSearch 폴백 체인 - 비교표 네이버 DCF 행이 순부채 보정 이전 측정값임을 각주로 명시 - get_market_data_kr 사용 예시(현재가·목표주가 컨센서스) 추가 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: CLAUDE.md·AGENTS.md 드리프트 정정 (코드 기준, 동작 변경 없음) 코드와 어긋난 서술을 실제 구현에 맞게 정정: AGENTS.md (upstream 기여자 문서 드리프트): - Ink/React·`cli.tsx`·`src/hooks` → pi-tui·`cli.ts`·`components`(.ts) (src/hooks 미존재) - 가공의 툴명 financial_search/financial_metrics → 실제 툴(get_financials/get_market_data/ read_filings/stock_screener) + KR 툴 9종 전부 - web_search를 Exa→Perplexity→Tavily→LangSearch 폴백 체인으로 - "최종답변=별도 LLM 콜(no tools bound)" → 단일 패스(handleDirectResponse, 툴 항상 바인딩; 별도 마무리 패스 없음) - LLM 프로바이더/env에 Moonshot·DeepSeek + KR 데이터 키 추가, KR eval 하네스 명시 CLAUDE.md: - "FINANCIAL_DATASETS_API_KEY 없으면 get_* 툴 부재" 정정 → US 파이낸스 툴은 무조건 등록되고 호출 시점에 401(부재 아님). 실제 게이트(DART/검색/KRX/X) 명시 - TUI 트랩에 components(.ts)·src/hooks 미존재 추가 - Stale docs 섹션 갱신(AGENTS 정정 완료, README엔 Ink 없음) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ift canary) (#13) * read_filings_kr: governance 섹션 카테고리 추가 (지배구조·최대주주·계열회사·대주주 거래) 코리아 디스카운트 1차 변수(소유구조·특수관계자)를 inference가 아닌 DART 사업보고서 본문 verbatim으로 끌어오는 6번째 narrative 카테고리. - dart-document.ts: SectionCategory에 'governance' + sectionMatches 케이스. canonical 정기보고서 서식 VI 이사회 등 회사의 기관 / VII 주주에 관한 사항 / IX 계열회사 등 / X 대주주 등과의 거래내용을 번호(VI/VII/IX/X) + title-regex 이중 매칭 — 기존 risks 케이스와 동일 전략이라 보고서별 번호 drift에 견고. selectSections는 headerless·titleless 잡음 조각([] 라벨)을 필터. - read-filings-kr.ts: zod enum·플래너 프롬프트·툴 설명 배선 + 캐시 키 _dsd_sections → _dsd_sections_v2 범프(구 캐시가 새 카테고리를 조용히 누락하지 않도록 강제 재추출). - registry.ts: compactDescription에 지배구조·최대주주·계열회사 명시. 특수관계자와의 거래 주석은 본문 III. 재무 주석에 존재하고 매칭도 되나, 표 중심이라 dsdToPlainText가 표를 버려 prose가 거의 없어 자동 제외됨 — 정밀 거래내역은 get_financials_kr 영역. 본문 narrative(X. 대주주 등과의 거래내용 등) 까지 커버. 실검증: 005930 FY2025 사업보고서에서 VI/VII/IX/X 7블록·약 9.6KB 실제 prose 추출 확인. KR 스위트 green (governance 테스트 5건 신규), typecheck 클린. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * KR DART 견고성: 동시성 캡 + 020 쿼터 서킷브레이커 + 플레이북 keyless 티어링 (a) DART 호출 견고성 — dart-throttle.ts 신규: - 공유 카운팅 세마포어(기본 4, DART_MAX_CONCURRENCY로 조정)로 모든 DART 호출의 동시성 캡. getBusinessReport(연도 Promise.all) × get_financials_kr (sub-tool Promise.all) × executor(도구 병렬)의 곱셈식 fan-out이 무료 키 20,000/일 쿼터를 태우는 것을 차단. acquire는 단일 사용 releaser를 반환해 이중 해제로도 카운트가 음수가 되거나 cap을 넘기지 못함(구조적 안전). - 020 사용한도초과 서킷브레이커: 첫 020에 latch → 나머지 fan-out 버스트는 네트워크 없이 명확한 한국어 메시지로 즉시 실패, 5분 쿨다운 후 half-open 재시도. 지수 backoff은 일일 쿼터엔 무의미하므로 채택하지 않음. - api.ts get/getBinary가 슬롯+latch를 사용. assertDartOk가 020을 별도 분류해 latch(013 등 다른 상태는 기존 메시지·동작 그대로, latch 안 함). getBinary는 status XML/사용한도 본문으로 020 감지(콘텐츠타입 게이트가 먼저라 정상 ZIP은 오탐 불가). (b) 플레이북 keyless 티어링 — prompts.ts: - buildKoreanResearchSection을 순수 함수(hasDartKey: boolean)로 만들고, DART 키 없을 때 '' 대신 keyless 티어 반환. 항상 등록되는 Naver get_market_data_kr·get_foreign_ownership_kr만 참조(unbound DART 도구는 언급 안 함)해 수급·밸류에이션 추론이 keyless로도 살아남게 함. - DART 게이트를 env.ts hasDartKey()로 단일화 — registry(도구 등록)와 prompt(플레이북 티어)가 동일 판정을 써서 "플레이북이 등록 안 된 DART 도구를 가리키는" 불일치를 원천 차단(registry는 pure process.env 의미 유지 → registry.test 그대로 통과). 테스트: dart-throttle(세마포어 cap·gate 차단·단일사용 releaser·throw시 해제· 쿨다운 경계·half-open 재트립·sibling-trip 즉시실패), api(assertDartOk 020/013/ 000 분류), prompts(티어링) — 신규 16건. KR 스위트 green, typecheck 클린. 라이브: 5년치 fan-out이 cap 정확히 관측하며 전량 정상 반환, in-flight 0 정리. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * KR keyless 도구 partial-drift canary (Naver schema 변경 조용한 null 탐지) Naver 모바일 JSON은 계약이 없어 필드 하나만 rename되면 그 값이 조용히 null이 되고, 전부 null일 때만 발동하는 hasNoMarketData 가드를 통과한다. "일부는 있고 일부만 누락"(partial drift)을 _dataQualityWarning으로 surface해 모델이 그 값을 사실로 보고하지 않게 한다. (둘째 keyless 소스가 없어 fallback이 아닌 detection; 캐시 staleness 신호는 범위 밖.) - utils.ts: nullFields(스냅샷) + deadColumns(시계열, 전 row null) 헬퍼. - get-market-data-kr.ts marketDataQualityWarning: · critical(현재가·시총) 누락 → 경고 · EQUITY(PER 또는 EPS 있음)인데 PBR·BPS 둘 다 없음 → 경고. 펀드(ETN/ETF/REIT)· 상장당일 IPO는 실적·장부가가 본래 없어 PER/EPS 게이트로 조용히 통과(오탐 방지). · ETF/ETN 스키마(nav/fundPay/etfBaseIdx)는 시총·멀티플이 본래 없으므로 fund로 감지해 equity canary 자체를 건너뜀 — get_market_data는 equity 도구. · dealTrendInfos.closePrice가 rename되면 전일 lastClosePrice fallback이 stale 값으로 가려 canary가 못 보는 최악 케이스를 별도 감지(현재가 미갱신 경고). · PER/EPS/배당/추정치는 loss-maker·무배당·미커버에서 정상 null이라 drift 신호 에서 제외. - get-foreign-ownership-kr.ts: 6개 수급 컬럼(외국인·기관·개인 순매수·지분율·종가· 거래량)이 전 구간 null이면 경고. 정상 0은 0으로 파싱되므로 오탐 없음. 1~2행 창(신규상장)에선 "전 row null"이 자명히 참이라 ≥3행에서만 판정. - prompts.ts: _dataQualityWarning 처리 지침 — KR 대체 소스가 없으므로 omit 우선, 교차확인은 다른 도구가 같은 값을 줄 때만. 검증: 단위 테스트 신규(헬퍼·canary·ETF·book-drift·stale-close·foreign drift), KR 스위트 126/0, typecheck 클린. 라이브 오탐 0 — 대형주·KOSDAQ·우선주·지주사· 은행·SPAC·ETF·REIT 전부 clean. 적대적 리뷰 2렌즈 반영(closePrice 오탐누락·ETF 오탐·컬럼 커버리지·단일행·프롬프트 문구). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * README 전면 갱신: KR 견고성 레이어 + governance 본문 + 사실 정정 이번 브랜치의 변경(governance 섹션·DART 동시성/서킷브레이커·keyless 티어링· drift canary)을 README에 반영하고 낡은 사실을 바로잡음. - 🇰🇷 섹션 재구성: 도구 / 견고성·신뢰성(신규) / 종목코드 해결 / 고유처리. read_filings_kr에 지배구조·최대주주·특수관계자·계열사·대주주거래 추가. - 신규 "견고성·신뢰성": DART 쿼터 보호(동시성 캡 + 020 서킷브레이커), keyless 플레이북 티어링, partial-drift canary(_dataQualityWarning), record/replay 결정적 재현. - 사실 정정: DART 무료 한도 10,000 → 20,000건/일, clone URL을 포크 (dexter-with-korea)로, env에 DART_MAX_CONCURRENCY 추가. - 비교/한계 섹션에 견고성·컨센서스 깊이 한계 반영. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * README 재작성: 상속 구조 탈피, problem→tools→durability 아크로 업스트림 dexter의 구성(개요→사용예시→비교→빠른시작)을 따르지 않고 이 포크의 정체성(한국 1차 출처 + 견고성 레이어)에 맞춰 재구성: - 한국 1차 출처 에이전트로 정체성 재정의 + 단일 예시로 "무엇을 하나" 즉시 제시 - "왜 한국 주식은 따로 만들어야 했나" — 소스별 현실(DART 쿼터·KRX 로그인· Naver 무계약·NPS 연말스냅샷) 표로 견고성 레이어의 필요를 동기화 - 견고성을 독립 섹션으로 전면화(쿼터 보호·drift canary·keyless 티어링· record/replay) - 한국 시장 튜닝 / 빠른시작 / 예시 / 챗봇 대비 / 정직한 한계 / 개발·평가로 정리 내용 사실은 직전 갱신과 동일하게 코드 근거 유지. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* 한국 주식 리서치 품질개선: 종목 해석·사업부문 데이터·밸류에이션 스킬·eval 정확도 게이트 감사(46개 검증 갭) 기반 3-마일스톤 통합 업그레이드. M1 — 답변 품질 레이어(시스템 프롬프트=유일한 최종답 결정자): - 상대가치(기존 peers 활용)·QoQ+forward 모멘텀·실적의 질·밸류업(PBR<1 re-rating)·환율 strand·공매도 regime 조건화·종목 식별 가드 추가(prompts.ts 양 tier) - DCF: WACC를 bottom-up CAPM(실시간 국고채 Rf+β)로, 컨센서스 목표주가 대조, sector-wacc-kr는 sanity band화 M2 — 신뢰가능한 라우팅 + 분석 깊이: - resolve-kr.ts: 6자리 OR 회사명 통합 해석(DART 레지스트리 → keyless Naver 자동완성 fallback). 8개 도구에 배선해 모델의 코드 추측에 의한 '사일런트 다른 회사' 오류 제거 - get_segments_kr 신규: DART 문서의 사업부문별 요약 재무현황 표 복구(parseDsdTables). 삼성/현대차 라이브 검증 - 신규 스킬 4종(auto-discovered): kr-relative-valuation·kr-sotp-holdco·kr-earnings-quality·kr-shareholder-return. DCF Step0가 경기민감주/지주사를 라우팅 M3 — eval 안전망: - 결정론적 정확도 게이트: checkRequiredPhrases + checkNumericAnchors(조/억·% 파싱, 허용오차). buildQuestionResult hard-fail. samsung-fundamentals를 fixture 수치(133.9조/57.2조)에 앵커 → replay PASS 검증 - 질문뱅크 14→22(세그먼트·이름해석·지주사·밸류업·상대가치·은행·모멘텀·우선주), judge 차원 3종 추가 - ci.yml에 kr-eval-replay 잡(run-eval 라벨·시크릿 게이트, 키 없으면 graceful skip) 검증: typecheck 통과, 241 테스트 통과(2회 안정), 이름해석·세그먼트·numeric 게이트 라이브 검증. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * get_segments_kr: summarizeToolResult 케이스 추가 (TUI 상태줄) 신규 KR 도구가 cli.ts의 summarizeToolResult에 케이스가 없어 generic 'Received N fields'로 빠지던 것을 'Found N segment tables'로 수정 (CLAUDE.md 컨벤션). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
신뢰 핵심인 resolveKrSecurity의 fallback 순서(6자리 패스스루 → DART 레지스트리 이름 → keyless Naver → null)를 resolveTicker/resolveCorpName/fetchNaverAutocomplete 모킹으로 결정론적으로 커버. registry throw 시 keyless fallback, 비상장(stock_code 없음) 시 Naver 경로까지 검증. mock.restore()로 타 파일 누수 방지. 248 테스트 통과. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
현재 main(업스트림 동기화: subagents·web_fetch·ask_user_question·Fable 5 등) 위에 한국 주식 리서치 포크 전체를 통합. 충돌(agent.ts·prompts.ts·types.ts)은 양쪽 기능을 모두 보존하도록 union 해결: - AgentConfig: toolAllowlist/systemPromptOverride/agentLabel(업스트림) + transformTools(KR eval) - Agent.create: toolAllowlist 필터 + transformTools + CLI-only 필터 모두 적용 - 시스템 프롬프트: spawn_subagent 가이드 + KR _dataQualityWarning 가이드 공존 검증: bun install 후 typecheck 통과, 270 테스트 통과(업스트림+KR). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- desktop/: Electron+React GUI — 챗 / 업무(엑셀 변환) / 기록(history) / 설정(키 관리) - safeStorage 암호화 키 → sqlite → spawn 시 env 주입 (코어 무변경) - 대화·변환 아카이빙(sqlite), three.js 로고, 사이드바 접기, flat 모노톤 버튼 - src/sidecar/: stdio JSON-Lines 헤드리스 사이드카, Agent.run() 구동 (run/cancel/reset/convert) - core: text_delta 스트리밍 이벤트 추가; DEXTER_DIR override (userData 리다이렉트) - README 일반 사용자용으로 재작성; 개발자 내용은 DEVELOPMENT.md로 분리 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…-stock ranking, new KR skills (segments·relative-valuation·SOTP·shareholder-return) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
which bun is a Homebrew symlink; cpSync copied the link, so resources/bin/bun pointed at the build machine's Cellar path and the packaged app could not spawn the sidecar on any other machine. Resolve via realpathSync + dereference copy (remove stale dest first to avoid same-path error) + chmod 755. Also correct the header note: KRX short-balance is pure HTTP (no Chromium); only the browser tool needs it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- update gate: main fetches remote update.json (raw GitHub), compares vs app version. status required → full-screen lock + 업데이트 버튼; optional → dismissible banner. fail-open on network error so a working install is never bricked offline. - help: leads with clickable showcase prompts (→ prefill a new chat via seed), then API-key guides; flat divider sections (Conductor compact style). - settings flattened to the same compact row style (no boxed cards). - sidebar: English labels (History/Chat/Work) + Chat/Work glyph icons. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Builds on each OS's own runner so prepare-core stages the matching bun binary. Tag push (v*) → draft GitHub Release with dmg/exe; manual run → workflow artifacts. Unsigned by default (CSC off); add signing secrets later. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- build/icon.png: 1024px three.js render matching the in-app logo; wired to electron-builder mac/win/linux icon (force-added; root .gitignore excludes build/). - whole top of the window is now draggable (.titlebar-drag strip), not just the sidebar; collapse button / history search / update banner lifted above it + no-drag. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…tifact names - electron-builder artifactName fixed to Dexter-mac-arm64.dmg / Dexter-windows-x64.exe so releases/latest/download/* links are permanent across version bumps. - README top: shields badge buttons → direct download (non-technical users never need to find the Releases tab). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…n label Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… dep - SidecarManager tracks active run/convert ids and emits an error for each on unexpected exit/spawn-error, so chat/work UI recovers instead of spinning. - remove unused 'xlsx' (SheetJS) dependency — export uses exceljs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Desktop app (Electron + Bun sidecar) — chat / work(엑셀 변환) / update gate, + merge to 1.0.0
…Win more-info→run)
…amaged')
Unsigned arm64 app + quarantine → '"Dexter" is damaged and can't be opened' with
no Open option. Fix: (1) afterPack ad-hoc codesign; (2) prepare-core strips the 84
node_modules/.bin symlinks that made codesign --verify fail ('invalid destination
for symbolic link in bundle'), which had voided the signature. Now Gatekeeper shows
the normal 'unidentified developer' → right-click → Open. Verified: codesign --verify
--deep --strict passes, bundled sidecar still runs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
DCF·밸류에이션이 web_search 추론에 기대던 입력(무위험금리·환율)을 공식 한국은행 ECOS OpenAPI로 대체한다. - get_macro_rate_kr 도구 (+ src/data/fetchers/ecos.ts): 국고채 1/3/5/10년, 회사채 AA-(Kd 앵커), 원/달러 환율, 한국은행 기준금리. ECOS_API_KEY 게이팅. - DCF 스킬(Rf)·시스템 프롬프트(환율 양 tier)를 get_macro_rate_kr로 라우팅하고 도구 미등록 시에만 web_search 폴백 (소프트 유도, 하드 강제 없음). - 데스크톱: 설정탭 ECOS 입력칸(DATA_SOURCES)·도움말 항목·좌측 패널 ECOS 상태점. - 라이브 검증: 7/7 시리즈 정상, 환율/Rf 자연 라우팅 확인 (앱 end-to-end). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(kr): get_beta_kr 실측 회귀 베타 + ERP 4.87% 핀 DCF cost-of-equity의 두 soft 입력(β·ERP)을 출처있는 값으로 교체. 베타: web_search 추론 → 실측 회귀. - naver-price-history.ts: Naver 차트 API(keyless)로 종목·KOSPI/KOSDAQ 지수 일봉 - compute-beta-kr.ts: 2년 주간수익률 OLS 회귀 + Blume 보정(0.67·raw+0.33), 순수함수 - get-beta-kr.ts: 상장시장 자동판별, R²·관측치 품질 게이트(reliable 플래그+경고) - registry 무조건 등록(keyless), cli 요약 케이스, 콜로케이트 테스트 ERP: 무출처 5~7% 레인지 → Damodaran 한국 Total ERP 4.87%(2026.1) 단일 핀. 지배구조/코리아디스카운트 가산은 sector-wacc-kr.md 조정인자로 분리(이중계상 금지). SKILL.md Step8: 모든 WACC 입력에 항목·값·출처·as-of/방법 구조표 강제. 폴백 사용 시 출처칸에 명시 — 신뢰성을 출력에 드러내 감사 가능하게. typecheck + 290 tests green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(kr): get_beta_kr 신뢰성 임계값 제거 — 사실만 노출 builder 주관이 박힌 두 곳을 제거: - R²≥0.05 / MIN_RELIABLE_OBS(52/200/24) 컷오프 삭제. 출처없는 임의 경계라 reliable 판정 자체를 빼고 rSquared·observations·window를 사실로만 반환. "얼마나 믿을지"는 SKILL.md 소프트 가이던스로 에이전트가 맥락 판단(no-hard-forcing). - window에 requestedFrom vs coveredFrom 노출 — 신규상장(측정창 미달)을 임계값 없이 사실로 드러냄. - 2y weekly 디폴트를 Bloomberg 표준 관행으로 출처 귀속(설명·스키마·method 문자열). 값 자체는 못 없애도 "내 취향"이 아니라 "표준"임을 명시 + override 가능. marketToIndex export + 단위테스트 추가. typecheck + 292 tests green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(kr): DCF β 측정법을 계산 전 사용자에게 질문(소프트) 마지막 남은 builder 주관(2y weekly 디폴트)을 사용자 손으로 넘김. SKILL.md Step3 β 단계에 소프트 가이던스 추가 — 계산 전 ask_user_question으로 β 측정법(2y주간 권장 / 5y월간 / 1y일간)을 한 번 확인하고, 고른 값을 get_beta_kr의 years·frequency로 전달. 판단-입력 쿼리(DCF)에 한해서만 묻고, 이미 명시됐거나 비대화형(헤드리스· 서브에이전트, ask_user_question이 자동 폴백)이면 표준 2y weekly로 진행 — 맹목 질문 금지(no-hard-forcing). 코드 변경 없음(툴은 이미 파라미터 노출). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(desktop): ask_user_question 데스크탑 렌더링 (mid-turn 질문 패널) 지금까지 ask_user_question은 CLI에서만 떴다 — 사이드카가 requestUserInput을 넘기지 않고 프로토콜에 질문 메시지 타입도 없어, 데스크탑에선 조용히 'proceed with defaults'로 폴백했다. 이제 데스크탑에서도 에이전트가 턴 중간에 사용자에게 선택지를 띄운다. Core: - protocol.ts: question(sidecar→shell) / answer(shell→sidecar) 메시지쌍 추가 - user-input-bridge.ts: questionId로 pending Promise를 correlation, cancel/error/ reset 시 declined로 settle해 hang 방지. 순수 로직 + 단위테스트 - sidecar/index.ts: Agent.create에 requestUserInput 주입, answer/cancel/reset 라우팅 - ask_user_question 'CLI only' 라벨을 'CLI + desktop'으로 정정 Desktop: - shared/sidecar.ts: 프로토콜 미러 + Question/UserAnswers 타입 - ipc.ts + preload + DexterApi: chat:answer 채널 - QuestionPrompt.tsx: 옵션(단일/다중) + '직접 입력'(Other) 인라인 패널 - ChatView: question 메시지 처리 → 패널 렌더 → 답 회신, done/error/cancel 시 정리 typecheck(core+desktop node/web) + 281 tests green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(kr): 신뢰성 감사 P0 4건 — 우선주 분모·무형 capex·분기 라벨·장중가
2026-07-02 신뢰성 감사에서 확인된, 잘못된 수치가 자신있게 출력되는 P0 4건 수정:
1. DCF 우선주 분모 (~10%+ 과대): get_market_data_kr에 preferredListings 추가
(KRX 코드 규칙 끝자리 5/7/9 프로브 + 발행사 이름 검증, 라이브 검증: 삼성전자우,
현대차우/2우B/3우B). DCF Step 5에 'Equity Value − 우선주 시총 → 보통주 수로
분할' 지시 추가.
2. capex 무형자산 누락 (FCF 과대): capex를 CAPEX_SUM_SPEC(유형+무형자산 취득
합산)으로 교체 — 주파수이용권·자본화 개발비가 FCF에 반영된다. 처분 라인은
exact-match로 제외.
3. 분기 손익 라벨 정반대: DART 분기/반기 IS의 thstrm_amount는 3개월 단독이고
누적은 thstrm_add_amount인데 basis가 '누적'으로 단언했다 (실측: 005930 2025
Q3 매출 thstrm=86.1조(3M) vs add=239.8조(9M)). IS 메트릭에 ytdCurrent/
ytdPrior/ytdDisplay를 노출하고 BASIS_NOTES를 실제 의미대로 재작성.
4. 장중 전일가 혼합 (급락일 7%+ 오차): /basic(무캐시 실시간 quote) 병행 조회로
price/change/asOf/marketStatus/tradingStatus를 라이브 소스에서 취하고, 하락일
부호를 direction 코드로 강제. 거래정지·실시간 조회 실패는 _dataQualityWarning
으로 명시. 발행주식수 도출은 시총과 동일 시점 가격 쌍으로 계산.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(kr): 주관성 P1 3건 제거 + kr-eval CI 기본 경로 승격
제로 주관성 원칙(모든 입력은 출처 있는 값, 보수성은 민감도/오버라이드로) 위반
잔존분 정리:
1. sector-wacc-kr 조정인자: 무출처 ±0.5~2% WACC 직가산 지시를 '리스크 플래그'
체계로 전환 — 플래그는 도구 근거와 함께 서술, 수치 영향은 Step 6 민감도
±1% 열로, 수치 가산은 사용자 오버라이드로만.
2. DCF FCF 성장률: 10~20% 재량 haircut(같은 질문에 다른 답)을 15% 고정
기본값으로 핀하고 가정표 필수 행으로 노출; haircut 10/15/20% 성장률 민감도
행 추가; 5% 감쇠·15% 상한도 가정표 표기 의무화.
3. kr-sotp-holdco: '불확실하면 보수적으로' baked conservatism 제거 → 비상장
가치 low(장부가)/base(멀티플) 병기; '40~60% 할인이 흔하다' 무출처 앵커 제거
→ 동일 방법으로 계산한 동종 지주사 할인율 비교(출처 표기)로 대체; NAV 식의
'+순현금 −순부채' 이중계상 모호성을 단일 순액 항으로 수정; SOTP 표에
출처·기준일 열 의무화 + web_search 보완 지분율 잠정치 표기.
4. kr-eval CI: run-eval 라벨 전용이던 replay 잡을 main push + 주간 cron으로
승격(PR은 라벨 옵트인 유지 — fork는 secret 접근 불가). KR_EVAL_REQUIRE_RUN
신설: all-skip(키/fixture 미비) 실행이 초록으로 통과하지 않고 실패하도록.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(kr): 잔여 P1 7건 — NCI·이자분류·병합키·세그먼트단위·메모통화·salvage·browser직렬화
감사 잔여 P1 일괄 수정:
1. NCI 차감: balanceSheet.nonControllingInterests 노출(BS 전용 매칭 — IS/CIS의
동명 귀속 행 배제, 실측: 삼성 10.5조·LG화학 14.7조) + DCF Step 5 브리지를
'EV − Net Debt − NCI'로 수정(장부가 근사 캐비앳, 상장 자회사 시가 교차).
2. FCFF 이자분류: cashFlow.interestPaid + interestPaidClassification 노출
(IAS 7 표준 id로 operating/financing 판별, 이름만 매칭 시 null) + DCF에
'operating이면 FCFF = FCF + 세후이자 가산' 지시(이자 이중 벌점 제거).
3. get_financials_kr 병합 키: tool_ticker → tool_ticker(args) — 동일 종목
연간+분기/CFS+OFS 복수 호출의 silent drop 제거, 중복 키 #n 가드.
스테일 설명문 수정('Phase 2/3 미구현' 허위 안내, '이름 불가' 문구 →
실재하는 KR 도구 안내·이름 해석 지원 명시).
4. get_segments_kr 단위: '보통 백만원' 빌더 추정 제거 → 표 안 '단위:' 표기를
extractUnit으로 추출해 SegmentTable.unit으로 반환, 미검출 시 null +
'단위 확인 전 금액 크기 서술 금지' note (천원 filer 1000배 오차 차단).
5. write-memo 통화·가정표: 템플릿 $ 하드코딩 → {{currency}} 슬롯($/₩),
Step 8 채팅 형식 통화 파라미터화; 'Valuation Assumptions & Sources' 섹션
신설 — DCF 가정표(β·ERP·Rf·as-of)가 메모 산출물까지 전달되도록.
스테일 'KR 세그먼트 도구 없음' 문구 수정(get_segments_kr 반영).
6. maxIterations salvage: 한도 도달 시 수집 데이터를 버리고 사과문만 내던
동작 → scratchpad 도구 결과로 no-tools 최종 합성 1회(부분 답변 + 미확보
데이터 목록), 실패 시에만 기존 문구 폴백.
7. browser concurrencySafe: true → false — 모듈 전역 page/currentRefs 공유
상태에서 병렬 호출 시 타 페이지 내용 오귀속 race 제거.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(kr): P2 표기·라벨·가드 일괄 — 기준시점·β 품질 플래그·신선도 고지·검색 수치 규칙
기계적·저위험 P2 일괄 (구조 변경 성격의 eval 인프라·compaction 큐 보존은 별도):
- get_market_data_kr: valuation.basis 신설 — PER/EPS/PBR/BPS/배당의 기준시점
(Naver valueDesc: TTM 분기·배당 연도)을 노출해 낡은 지표가 '현재 지표'처럼
읽히지 않게 함 (라이브 확인: per 2026.03. / dividend 2025.12.).
- get_beta_kr: (a) 거래정지(거래량 0) 행을 회귀에서 제외 + 제외 수 보고
(평탄 0% 수익률의 β 하향 편향 제거), (b) 시장 판별 실패 시 KOSPI 폴백을
indexFallback 플래그+경고로 노출, (c) 소표본(n<60d/24w/12m) 시 R² 과대
경고 — 셋 다 _dataQualityWarning으로 모델에 전달.
- get_nps_holdings: asOf(연말 스냅샷 basis + snapshotFetchedAt) 노출, stale
fallback 시 경고 전달, 고정 UDDI 한계를 설명문에 명시.
- get_short_balance_kr: T+2~3 공표 지연 + 공매도 금지 regime 고지.
- get_business_report_kr: 연도 기본값을 보고서 유형별 공시 일정 기준으로
(분기·반기가 1년 묵은 기본값을 받던 문제 제거).
- web_search: 'Historical stock prices'가 When to Use에 오배치된 것 수정,
검색 유래 수치의 (출처, 날짜) 표기 규칙 추가; Perplexity answer가 LLM 생성
텍스트임을 _answerNote로 명시 + 결과 date 필드 보존; x_search에 '트윗은
사실 소스 아님' 캐비앗.
- prompts(키리스 VALUE-UP): '검색 미탐 = 환원 약속 부재' 단정 지시 제거 →
'미탐 ≠ 부재' + 한계 명시로 수정.
- channels(CLI): 수치에 as-of 날짜 병기 행동 지침 추가 (KR 전용이던 규칙의
일반화).
- KR 스킬 4종: earnings-quality(전환비율 1.0 = 항등 기준선 라벨 + 적자 가드 +
등급은 신호 집계로), relative-valuation(밴드 보수/낙관 = peer 분포 도출,
웹 선정 peer 표기, basis 병기), shareholder-return(보도 vs 공시 구분,
배당성향 분모 controllingNetIncome, 금융주 순현금 점검 제외),
spinoff(기사 게재일 ≠ 결의일).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(kr): 리뷰 반영 — 프로브 latch·priceSource 통일·스냅샷 라벨·salvage 캡·β gap
/review 8앵글×38후보 → 12건 검증(CONFIRMED 9)의 수정:
1. 우선주 프로브 latch: 일시 오류(네트워크·5xx)까지 영구 miss로 캐시하던 것을
확정 결과(409/404·타사 이름)만 latch하도록 분리. 일시 실패는 failedProbes로
반환되어 '목록 불완전 — 우선주 부재로 단정 금지' 경고가 붙는다.
(장수 프로세스에서 P0 재발 경로 차단)
2. priceSource 통일: 가격 소스('live'|'snapshotClose')를 명시 필드로 노출하고
asOf·신선도 경고·드리프트 canary를 전부 이 필드 기준으로 — /basic의
closePrice와 localTradedAt이 따로 실패해도 전일가가 라이브 타임스탬프를
달거나(억제된 경고), 라이브 가격에 거짓 stale 경고가 붙지 않는다.
3. 스냅샷 시점 혼합: 발행주식수를 동일 /integration 스냅샷 쌍(시총÷스냅샷
종가)으로 도출(캐시된 시총÷라이브가의 장중 변동 오염 제거 — 라이브 검증:
현대차 204.8M ≈ 실제 상장주식수), valuation.snapshotAsOf 신설(캐시 서빙 시
cachedAt)로 시총·OHLC·멀티플의 기준 시점을 정직하게 라벨.
4. salvage: compaction용 <analysis>/<summary> 형식을 강제하던 프리앰블/트레일러를
답변 전용으로 교체 + 방어적 스트리핑(원시 XML이 최종 답변에 노출 차단),
scratchpad 입력을 150K chars로 캡(head+tail — persist가 compaction 트리거를
억제해 무제한이던 경로).
5. β: 재개 직후 gap 수익률(멀티주 캐치업 vs 지수 1주) 드롭 + excludedGapReturns
보고 — 정지 행 제거가 만들던 2차 왜곡 제거.
6. K-코드 우선주 한계를 모델 표면에 명시(도구 설명 + DCF Step 5: 빈 배열 ≠
우선주 없음 확정, 정황 시 이름 조회/한계 명시).
7. CI: kr-eval-replay에 job-level concurrency(cancel-in-progress) + fixture
재기록 유지보수 규칙 명시. (fixture 재기록 자체는 keyed 환경 후속 작업)
경미: CLAUDE.md agent-loop 계약(no-tools salvage 예외) 갱신, defaultYear를
공시 마감+5일 기준 일단위로(5/20·8/20·11/20), 죽은 getNpsHoldings export 제거,
cache readCache가 cachedAt 노출.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(kr): get_equity_investments_kr — 타법인 출자현황으로 지주사 지분율 근거 확보
holdco SOTP에서 에이전트가 자회사 지분율을 사전지식으로 지어내던 근본 원인 수정:
read_filings_kr는 서술만 발췌(정형 표 스킵)하고, 5%룰(get_large_holders_kr)은
변동 시에만 보고라 수년째 안 움직인 지분(LG유플러스·생활건강)은 아예 안 떠서
지분율의 권위 소스를 읽는 도구가 없었다.
- DART otrCprInvstmntSttus 신규 도구: 기말 지분율·정확한 주식수·장부가·피출자사
재무 스냅샷(상장+비상장 전체), 연도 자동 폴백, 합계행 제거, 장부가 내림차순 캡
- kr-sotp-holdco Step 1/3을 이 도구 1차 소스로 재배선 + 미확인 지분율 단정 금지 가드
- holdco-sotp-lg 기대 도구 갱신 (grounding 0.35→0.82, replay PASS tools=100%)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(kr-eval): fixture 재기록 — 신규 8문항 최초 캡처 + 새 출력 필드 반영
keyed 환경에서 record 실행. 기존 2문항(foreign-flow·institutional)은 name 필드 등
새 도구 출력 반영, 신규 6문항(holdco·value-up·relative-value·bank·preferred·
name-resolution)은 최초 fixture — 영구 SKIP이던 문항들이 main CI에서 처음 채점된다.
holdco는 get_equity_investments_kr 경로로 기록 (replay PASS, grounding 0.82).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The eval gate was promoted to the main-push path but the repo has no OPENAI_API_KEY secret, so the first post-merge run failed: agent + judge (both gpt-5.5) errored on the missing key → Passed 0/20 → red main. Keep main green with zero token cost by default: run kr-eval-replay only on a `run-eval`-labeled PR, not on every main push. Also drop the now-purposeless weekly cron (its sole job was the eval sweep). Re-enable the always-on gate by restoring the push trigger + cron once OPENAI_API_KEY is set as a repo secret. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
README: 출처 목록에 한국은행 ECOS 추가, "판단을 주입하지 않습니다"(zero- subjectivity — ECOS Rf·실측 β·타법인출자 지분율) 원칙 명시, 기준시점(as-of)· 우선주 구분을 신뢰 항목으로 승격, 예시 질문을 실제 eval 문항(LG 지주사 NAV· 우선주 괴리율·밸류업·은행 PBR/ROE)으로 교체, 챗봇 비교표에 지배구조·지분율 행. DEVELOPMENT: 도구 표에 get_equity_investments_kr(타법인출자현황)·get_beta_kr· get_macro_rate_kr 누락분 추가. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
오픈 에이전트 루프(구조 해자) → 깊이의 차이 → 작동 방식 → 함정을 아는 도구 → 데이터 출처 → 무주관 신뢰 원칙 순으로 차별점을 계단식 구성. - 단일 자체완결 HTML (빌드·의존성 없음) - Pretendard + Newsreader(이탤릭 세리프) + Geist Mono, near-black 웜톤 미니멀 - taste-skill 디자인 원칙: 단일 웜골드 액센트, 1px 라인·여백 중심, 스크롤 리브만 Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
main push 시 landing/ 디렉터리를 Pages 아티팩트로 올려 루트 URL로 서빙. Settings → Pages → Source = "GitHub Actions" 선택 필요. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
get_consensus_kr(Naver keyless 증권사 실적 컨센서스, DCF 성장 anchor) + get_price_history_kr(일별 시계열·기간수익률·최대낙폭·KOSPI/KOSDAQ) 신설, 한국어 쿼리 유사도 감지 수정, Anthropic 대화 이력 incremental caching. 이후 8각도 셀프리뷰(23후보 검증)로 TZ 윈도우·장중 캐시·컨센서스 오단정·DCF 모순 수정 + 클린업 클러스터(중복 통합·microcompact·tier-aware 라우팅) 반영. 363 tests. 잔여: kr-eval replay 픽스처 재녹화(키 있는 환경에서 KR_EVAL_MODE=record).
…insider ownership 1.0.0(c5cb794) 이후 업스트림 18커밋 반영. 주요 도입분: - 권한 엔진 + bash 툴(allow/ask/deny 룰, 프로세스그룹 tree-kill). CLI 전용 - GPT-5.6 계열 + OpenAI Responses API 라우팅, Ollama Cloud 프로바이더 - insider/beneficial ownership (get_market_data 내부로 흡수) - 루브릭 기반 eval, x-search 레이트리밋 재시도, 포매터/타이머 누수 수정 충돌 2건 해소: - src/agent/types.ts: requestToolApproval 확장 시그니처(upstream, bash 승인용 command/decision) + desktop sidecar 주석(ours) - src/model/llm.test.ts: add/add. 우리 wire-payload 테스트 유지 + upstream GPT-5.6 Responses API 라우팅 테스트 병합 머지 후 수정: - read-filings-kr.test.ts: OpenAI fastModel 변경(gpt-5.4-mini → gpt-5.6-luna) 반영. KR 내부 요약 모델이 luna 로 내려가며 단가도 낮아짐 KR 영향 점검: *_kr 툴 17개 등록 유지, summarizeToolResult KR 케이스 14개 유지, CLI_ONLY_TOOLS 가 ['ask_user_question','bash'] 로 병합되어 desktop 채널엔 bash 미바인딩, 권한 룰은 .dexter/settings.json 의 permissions 키라 기존 설정과 충돌 없음. 검증: typecheck 클린, bun test 605 pass / 0 fail Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
upstream 은 gpt-5.5 를 모델 카탈로그에서 제거하고 DEPRECATED_MODEL_UPGRADES 로 settings.json 을 gpt-5.6-sol 로 덮어쓴다(loadConfig 가 즉시 저장). 사용자가 의도해 고른 모델이 말없이 바뀌므로 우리 fork 에서는 직전 세대까지 유지한다. - utils/model.ts: openai 카탈로그 맨 뒤에 gpt-5.5 복원(맨 뒤라 provider 기본값은 gpt-5.6-sol 유지) - utils/config.ts: DEPRECATED_MODEL_UPGRADES 에서 gpt-5.5 제거. 더 오래된 gpt-5.4/5.2 는 upstream 대로 자동 업그레이드 - utils/config.test.ts 신규: 마이그레이션 가드. config.ts 가 모듈 로드 시점에 SETTINGS_FILE 을 고정해 전체 스위트에서 import 순서 경합이 나므로 케이스마다 서브프로세스로 격리 - utils/model.test.ts: 카탈로그 기대값 갱신 참고: 모델 호출 자체는 upstream 에서도 정상이었다(resolveProvider 가 openai 로 fallback, useResponsesApi 는 gpt-5.6- 접두사에만 적용). 막혀 있던 것은 '선택 경로'다. 검증: typecheck 클린, bun test 609 pass / 0 fail. 가드 테스트는 upstream 동작으로 되돌리는 mutation 에서 정상 실패 확인. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
tool-executor 는 requestToolApproval 이 없을 때 'deny' 로 fail-closed 한다. 그런데 desktop 사이드카(src/sidecar/index.ts)는 Agent.create 에 channel 을 넘기지 않아 isCli 로 평가되고, CLI_ONLY_TOOLS 필터가 돌지 않는다. 결과적으로 bash 가 LLM 에 바인딩된 채 호출마다 조용히 거부돼 이터레이션만 소모한다. write_file/edit_file 은 머지 이전부터 같은 경로로 desktop 에서 항상 거부되고 있었다. 채널이 아니라 '승인 핸들러 유무'로 게이팅해 두 경우를 함께 해결한다. - permissions/engine.ts: APPROVAL_GATED_TOOLS 를 LEGACY_APPROVAL_TOOLS 옆에서 파생해 export (목록 드리프트 방지) - agent.ts: requestToolApproval 이 없으면 해당 툴을 제외 - agent.test.ts 신규: 바인딩 가드 ask_user_question 은 승인이 아니라 requestUserInput 으로 동작하므로 제외 대상이 아니다 — desktop 에서 계속 바인딩된다(테스트로 고정). 서브에이전트는 SUBAGENT_DISALLOWED_TOOLS/READ_ONLY_TOOLS 로 이미 제외돼 영향 없음. 검증: typecheck 클린, bun test 614 pass / 0 fail. 가드는 수정을 제거하는 mutation 에서 정상 실패 확인. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
업스트림 v1.0.3 머지 리뷰에서 나온 결함들. 각 항목은 회귀 테스트로 고정했고, 수정을 되돌리는 mutation 에서 정상 실패하는 것을 확인했다. 승인 게이트 우회 (src/permissions/): - command-parser: command 워드 없는 세그먼트를 버리던 필터 제거. `> file` 이나 `PATH=/tmp/evil` 만 있는 세그먼트는 효과가 실재하는데 매처와 built-in floor 양쪽에서 사라져 `ls && > file`, `PATH=/tmp/evil; ls` 가 Bash(ls:*) 하나로 ALLOW 됐다. 이제 보존되어 각각 ASK / DENY. - command-parser/rules: 리다이렉트 대상을 redirectTargets 로 보존해 secret floor 가 검사. `echo x > .env` 가 DENY. - rules.proposeRule: 플래그에 따라 파괴적이 되는 read-only 명령(find/sort/tree)에 bare 와일드카드를 제안하지 않음. `find . -name x` 승인이 `Bash(find:*)` 를 만들어 `find -delete`, `find -exec rm` 까지 덮던 문제. Bash(git status:*) 처럼 서브커맨드로 스코프된 제안은 유지. - engine: env 접두사, 경로 지정 command 워드를 auto-allow 대상에서 제외 (GIT_EXTERNAL_DIFF=./evil git diff, ./ls 가 규칙을 타던 문제). - read-only.safeUnless: --output=FILE, -oFILE 형태를 인식(기존엔 동어반복 조건이라 누락). tree 를 ALWAYS 에서 safeUnless(['-o']) 로. - rules.parseRule: Bash(:*) 를 tool-wide 로 승격하지 않고 malformed 처리. 데스크탑(주력 사용처): - providers: openai 항목이 머지로 카탈로그에서 사라진 gpt-5.4 를 계속 추천해 선택 시 404. 5.6 3종 + 5.5 로 동기화하고 defaultModel 을 gpt-5.6-sol 로. ollama-cloud 프로바이더 추가(업스트림 신규). - db: MODEL_ID_UPGRADES 에 gpt-5.4/5.2 자가치유 추가. 데스크탑은 .dexter/ settings.json 을 읽지 않아 코어의 마이그레이션이 닿지 않는다. - providers.test.ts 신규: 손으로 동기화하는 구조라 코어 카탈로그와의 드리프트를 CI 에서 잡도록 가드. 프롬프트 정직성: - buildCompactToolDescriptions/buildSystemPrompt 가 실제 바인딩된 툴만 안내. 데스크탑에선 write_file/edit_file/bash 가 언바인딩인데 프롬프트는 계속 광고해서, 모델이 메모를 저장하겠다고 한 뒤 실패하는 흐름이 나왔다. - Rule Management 문구도 파일 쓰기 가능 여부에 따라 분기. 기타: - cli: bash 결과 요약 케이스 추가(기존엔 generic fallback 으로 표시). - chat-log/subagent-group: ApprovalDecision 유니온 확장 + approvalLabel 공유. allow-always 가 성공 색상으로 Denied 로 렌더될 수 있던 잠재 버그. - llm.test: gpt-5.5 가 Responses API 로 새지 않는지 음성 어서션 추가. 검증: typecheck 클린, bun test 632 pass / 0 fail Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
merge: upstream v1.0.3 동기화 — bash 툴·권한 엔진, GPT-5.6, Ollama Cloud
ipc.ts 가 getSetting('modelId', 'gpt-5.5') 로 기본 모델을 하드코딩하고 있었다.
직전 커밋에서 providers.ts 의 defaultModel 만 gpt-5.6-sol 로 동기화하는 바람에
설정 화면이 보여주는 기본값과 실제 실행 모델이 어긋났다. 저장된 설정이 있는
기존 사용자는 영향이 없지만, 신규 설치는 설정엔 gpt-5.6-sol 이 뜨고 실제로는
gpt-5.5 로 돌아간다.
- ipc.ts: resolveRun() 이 provider 의 defaultModel 에서 폴백을 도출. chat:send
와 work:convert 두 곳 모두 사용.
- providers.test.ts: ipc.ts 에 모델 id 리터럴이 남으면 실패하는 가드 추가.
직전 가드는 providers.ts 만 봐서 이 불일치를 놓쳤다.
검증: typecheck 클린, bun test 633 pass / 0 fail. 가드는 하드코딩을 되돌리는
mutation 에서 정상 실패 확인.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
데스크탑의 History 한 행은 질문 하나와 답변 하나이고, 다른 걸 물으려면 새 대화를 연다. 그런데 사이드카는 프로세스 수명 내내 InMemoryChatHistory 를 유지하고 매 run 에 넘겨서, 의도한 모델과 엔진이 어긋나 있었다. 리셋은 새 대화/옛 대화 열기를 누를 때만 일어나 같은 창에서 계속 물으면 턴이 그대로 쌓였다. 그 상태의 실제 대가: - getRecentTurnsAsMessages 가 최근 10턴을 주입하고 그중 3턴은 답변 원문이다. 이 히스토리는 messages 배열에 실려 에이전트 루프가 도는 내내(최대 10 이터레이션) 반복해서 나간다. 연속된 무관한 질문이 서로 오염됐다. - saveAnswer 가 답변마다 요약용 LLM 호출을 한 번씩 더 냈다. 그 요약은 4턴째부터나 읽히는 값이라 단발 사용에서는 순수 낭비였다. - sidecar/index.ts: InMemoryChatHistory 제거. agent.run(req.query) 로 호출하고 reset 은 대기 중인 질문 해제만 담당. - single-shot.test.ts 신규: 사이드카가 히스토리를 안 물고 가는 것과, CLI 는 반대로 멀티턴을 유지하는 것을 함께 고정. CLI 는 그대로다 — controllers/agent-runner.ts 가 세션 내내 하나의 InMemoryChatHistory 를 넘기는 동작은 유지된다(테스트로 고정). 프로세스를 재시작하는 방식은 택하지 않았다. 콜드스타트를 실측하니 약 410~740ms 로 질문마다 붙는 순수 손해인데, 히스토리를 안 넘기면 같은 격리를 0ms 에 얻는다. 검증: typecheck 클린, bun test 637 pass / 0 fail. 가드는 히스토리를 되돌리는 mutation 에서 정상 실패 확인. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
UI 강제: History 한 행의 제목은 첫 질문에서만 뽑는다(ChatView.persist). 그래서 같은 행에 이어서 물으면 두 번째 이후 질문은 목록 어디에도 드러나지 않아 나중에 되찾을 수 없다. 답변이 끝나면 입력창 대신 '새 질문하기' 를 보여 다음 질문이 새 행으로 가도록 한다. 직전 커밋에서 사이드카가 단발이 됐으므로 엔진과 화면이 일치한다. - ChatView: answered 상태(정착된 assistant 턴 존재)면 composer 를 교체. 오류·중단도 non-pending assistant 텍스트로 남으므로 함께 커버된다 — 재시도 역시 새 대화가 맞다. - App: onNewChat 으로 기존 newChat 배선. CI 사각지대: 루트 tsconfig 는 include 가 src/**/* 라 desktop/ 을 전혀 보지 않는다. 즉 main/preload/renderer 가 깨져도 CI 는 초록이었다. 실제로 직전 커밋에서 추가한 desktop/src/main/providers.test.ts 가 desktop 자체 타입체크를 깨고 있었고 (bun:test 모듈 부재 + 루트 파일 cross-project 참조) 아무도 잡지 못했다. - tsconfig.node.json: src/**/*.test.ts 를 exclude. 이 테스트는 bun 런타임 전용이라 Electron tsc 프로젝트가 볼 대상이 아니다. - ci.yml: desktop-typecheck 잡 신설(npm ci 대신 --ignore-scripts 로 electron-rebuild 생략 — 타입체크에 네이티브 빌드는 불필요). 검증: 루트 typecheck 클린, bun test 637 pass / 0 fail, desktop npm run typecheck (node+web) 클린 — CI 와 동일한 --ignore-scripts 설치 경로로도 재확인. Electron 앱 자체를 띄워 눈으로 확인하지는 못했다(헤드리스 환경). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
사이드카가 channel 을 선언하지 않아 getChannelProfile 이 CLI_PROFILE 로 폴백했다. 그래서 주력 제품인 데스크탑이 모델에게 이렇게 지시받고 있었다: "Your output is displayed on a command line interface. Keep responses short", "Do not use markdown headers", "Tickers not names: AAPL not Apple Inc.". 정작 ChatView 는 react-markdown + remark-gfm 로 렌더링하고 독자는 터미널이 아니라 앱 창을 보는 투자자다. 답변 길이·구조·종목 표기가 근거 없이 깎이고 한국 종목이 005930 으로 나오던 원인. 업스트림 파일은 수정하지 않고 추가만 했다: - channels.ts: DESKTOP_PROFILE 신설 + CHANNEL_PROFILES 에 등록. CLI_PROFILE 과 WHATSAPP_PROFILE 은 한 글자도 건드리지 않았다. 정확성 규칙(as-of 병기, 검증보다 정확성, API 내부 노출 금지)은 형식 규칙이 아니므로 데스크탑에도 유지. 단발 세션이므로 '이전 대화 참조 금지, 후속 질문에 미루지 말 것'을 명시. - sidecar/index.ts: channel: 'desktop' 선언. ask_user_question 게이팅 수정 (위 변경의 전제): 채널만 바꾸면 CLI_ONLY_TOOLS 필터가 걸려 ask_user_question 이 함께 죽는다. 이 툴이 필요로 하는 건 '대화형 사용자'이고 그건 채널 문자열이 아니라 호출자가 requestUserInput 을 배선했는지의 문제다. 데스크탑은 배선하고 QuestionPrompt 로 직접 렌더링한다. - CLI_ONLY_TOOLS 는 bash 만 남김(셸은 터미널 사용자의 것). - HANDLER_GATED_TOOLS 신설: 핸들러가 없으면 제외. requestToolApproval 에 이미 적용한 능력 기반 게이팅과 같은 방식. 채널별 실측(프롬프트·바인딩): desktop : CLI 문구 없음, markdown 지시 있음, ask_user_question O, bash X cli : 기존 그대로, ask_user_question O, bash O whatsapp : 기존 그대로, 둘 다 X eval : 기존 그대로(cli 폴백), 둘 다 X 검증: 루트 typecheck 클린, bun test 647 pass / 0 fail, desktop typecheck 클린. 가드는 channel 제거 / ask_user_question 을 CLI_ONLY 로 되돌리는 mutation 에서 각각 정상 실패 확인. Electron 앱을 띄운 육안 확인은 못 했다(헤드리스). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1) 못 하는 일을 약속하지 않게 cron/heartbeat 실행기는 gateway.ts 에만 있고(데스크탑은 gateway 를 띄우지 않는다), cron 은 설명부터 결과를 WhatsApp 으로 전달한다고 되어 있다. browser 는 prepare-core.mjs 가 Playwright Chromium 을 의도적으로 스테이징하지 않는다. 셋 다 바인딩돼 있어서 '매일 아침 정리해드릴게요' 같은 응답 후 아무 일도 일어나지 않았다. - AgentConfig.unsupportedTools 신설. 채널 문자열로 추론하지 않고 호스트가 자기 한계를 직접 선언한다 — 같은 채널이라도 빌드마다 다를 수 있는 성질이라. - sidecar 가 ['cron','heartbeat','browser'] 선언. CLI/게이트웨이는 그대로 유지 (실측: desktop 에서만 빠지고 cli 는 3개 모두 바인딩). - 프롬프트 광고도 함께 사라진다(이미 바인딩 기준으로 필터링 중). 2) 복호화 불가 키가 초록불로 보이던 문제 ipc.statusFor 가 exists 를 DB 행 존재만으로 판정했다. safeStorage 는 OS 키체인에 묶여 있어 재서명된 빌드에서 행은 남고 복호화만 실패할 수 있고, sidecar 는 그런 키를 조용히 주입에서 건너뛴다. 결과적으로 설정은 전부 초록인데 모든 실행이 'API_KEY not found' 로 실패하고 원인을 알 방법이 없었다. previewLast4 가 이미 null 을 반환하고 있었는데 그 값을 버리고 있었다. 빈 문자열도 사용 불가로 처리. 3) DB 초기화 실패 시 창도 에러도 없이 종료되던 문제 initDb() 가 app.whenReady 안에서 guard 없이 호출돼, 네이티브 모듈 ABI 불일치나 DB 손상 시 createWindow() 전에 던지면서 앱이 아무것도 안 뜬 채 죽었다. try/catch + showErrorBox 로 원인과 userData 경로를 안내하고 종료(대화 기록 보존 안내 포함). 검증: 루트 typecheck 클린, bun test 650 pass / 0 fail, desktop typecheck 클린. 가드는 unsupportedTools 를 제거하는 mutation 에서 정상 실패 확인. Electron 앱 육안 확인은 못 했다(헤드리스) — 특히 2·3 은 실패 경로라 재현 자체가 어렵고, 코드 경로 수준까지만 검증된 상태다. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
KRX_COOKIE 제거: 소셜·네이버 로그인 계정을 위한 우회로였는데, DevTools 에서 Cookie 헤더를 복사해 붙여넣는 건 일반 회원 가입보다 오히려 어렵고 만료되면 조용히 실패한다. 경로를 없애고 data.krx.co.kr 일반 회원 가입을 안내한다. - registry: 게이팅을 ID/PW 만 보도록. krx-session: getCookieOverride 와 getKrxSession 의 쿠키 우선 분기 제거(sessionFromCookie 는 로그인 플로우 자체가 쓰므로 유지). krx-api: 에러 문구에서 쿠키 언급 제거. - env.example / AGENTS.md / DEVELOPMENT.md / HelpView 문구 정리. - registry.test.ts: ID/PW 만 인정하고 쿠키는 무시하는 것을 고정. DART 도구 개수 정정: '5개 도구' 로 안내하고 있었으나 실측 7개(get_financials_kr, get_filings_kr, get_large_holders_kr, get_equity_investments_kr, get_insider_trades_kr, read_filings_kr, get_segments_kr). data-sources.ts, HelpView, DEVELOPMENT.md 수정. Work 변환 취소: handleConvert 가 activeRuns 에 등록하지 않아 cancel 이 아무것도 매칭하지 못했다. 구조화 출력 호출이 멈추면 '변환 중…' 에서 앱 강제종료 외에 빠져나올 방법이 없었다. AbortController 등록 + callLlm 에 signal 전달, work:cancel IPC 와 취소 버튼 추가. 임의 타임아웃은 넣지 않았다 — 큰 시산표는 정상적으로도 오래 걸릴 수 있어 사용자가 직접 끊는 편이 맞다. History 삭제 확인: ✕ 가 항목을 여는 행 바로 옆에 있는데 확인창도 되돌리기도 없었다. 기록은 로컬 DB 에만 있고 내보내기 경로가 없어 오조작이 곧 영구 손실이다. confirm 추가. 에러 메시지는 영문 원문 유지 — 문제 해결에는 원문이 낫다는 판단. 검증: 루트 typecheck 클린, bun test 653 pass / 0 fail, desktop typecheck 클린. KRX 게이팅은 실측(쿠키만 있을 때 false, ID/PW 일 때 true, placeholder false) + 쿠키 경로를 되살리는 mutation 에서 가드 정상 실패 확인. Electron 앱 육안 확인은 못 했다(헤드리스). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
데스크탑이 주력 제품인데 bun test 는 UI 를 전혀 건드리지 않아, 변경을 실제 창에서 클릭해보기 전까지는 미검증 상태다. 그 과정을 매번 다시 알아내지 않도록 스킬로 묶었다. - driver.mjs: Playwright _electron REPL. launch/ss/click/click-text/set-input/ wait/poll-for/dialog/eval/text 명령. 파이프로 스크립트 실행도 지원. - SKILL.md: 셋업·명령표·실제로 도는 예시·함정·트러블슈팅. 안전: launch 가 매번 임시 --user-data-dir 을 만들고 quit 에서 지운다. 실제 설치본의 대화 기록과 암호화된 키는 읽지도 쓰지도 않는다(그래서 앱이 매번 빈 상태로 뜬다 — 데이터 손실이 아니라 격리 프로필). 실제 프로필은 DEXTER_REAL_PROFILE=1 을 명시해야만. 이번에 실제로 걸린 함정들을 문서와 코드에 박아뒀다: - npm install 이 electron 바이너리를 안 받아둘 수 있다(이전 --ignore-scripts 설치의 잔재) → node node_modules/electron/install.js. launch 가 이걸 감지해 안내한다. - playwright-core 는 루트 의존성이라 desktop/ 에 없다. - React 입력에 .value 를 대입하면 아무 일도 안 난다 — 네이티브 setter + 버블링 input 이벤트가 필요. 이걸 몰라 '취소 버튼이 안 뜬다'로 오진했었다(제품은 정상이었다). - 파이프 입력은 readline 이 버퍼된 줄을 한꺼번에 내보내고 async 핸들러를 기다리지 않아, 큐로 직렬화하지 않으면 모든 명령이 launch 전에 실행된다. - 사이드카 콜드스타트가 0.4~0.7s 라 변환 중 취소 버튼 같은 단명 상태는 poll-for 로. 검증: SKILL.md 의 예시 스크립트를 그대로 실행해 통과 확인(답변 후 composer 전환, History 삭제 확인창 dismiss 후 행 유지). 스크린샷으로 육안 확인. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1) sessionFromCookie 제거
KRX 쿠키 경로를 걷어내면서 파서만 남겨두고 '로그인 플로우가 쓴다'는 주석을 달았는데
사실이 아니었다. 로그인은 absorbCookies 를 쓰고 sessionFromCookie 는 프로덕션에서
한 번도 호출되지 않는다 — 자기 자신을 검증하는 테스트만이 이 export 를 살려두고
있었다. 함수를 지우고, 실제 프로덕션 코드인 cookieHeader(krx-api.ts:45 에서 사용)를
직접 검증하도록 테스트를 다시 썼다.
2) '새 질문하기' 가 아무 일도 하지 않는 상태
ChatView 의 메시지 로드가 useEffect([conversation?.id]) 에만 걸려 있어, chatId 가
이미 null 이면 null → null 은 변화가 아니라 effect 가 돌지 않는다. persist 는
done 에서만 호출되므로 취소(cancel)나 사이드카 레벨 오류로 끝난 턴은 저장되지 않고
chatId 가 null 로 남는다. 그 상태에서 answered 는 true 라 버튼은 떠 있는데 눌러도
화면이 그대로였다.
리뷰 때는 DB 저장 실패 경로만 보고 드물다고 판단했으나, 실제로 앱에서 재현해 보니
'중단' 버튼 한 번이면 걸리는 흔한 경로였다.
수정 전: 새 질문하기 클릭 → {입력창: false, 여전히잠김: true}
수정 후: 새 질문하기 클릭 → {입력창: true, 여전히잠김: false, 화면: chat-empty}
startNewChat() 이 부모에게 넘기기 전에 자기 상태를 먼저 비운다 — 부모의 id 가
실제로 바뀌었는지에 의존하지 않는다.
검증: 루트 typecheck 클린, bun test 653 pass / 0 fail, desktop typecheck 클린.
2번은 run-desktop 스킬로 수정 전/후를 같은 시나리오로 대조 확인.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1) 설정 화면에 '.env로 키 복사' CLI 에서 같은 키를 쓰려면 사람이 다시 입력해야 했다. 저장된 키를 KEY=value 줄로 클립보드에 복사한다. 복호화·조립·클립보드 쓰기를 전부 메인 프로세스에서 한다. 렌더러는 오늘까지 exists 와 last4 만 받아왔고(statusFor) 평문 키는 한 번도 건너간 적이 없다 — 그 속성을 깨지 않으려고 IPC 반환값도 env 변수 이름과 개수뿐이다. 복호화 불가 키는 건너뛰고 몇 개가 제외됐는지 알린다(다시 입력하라는 신호). 2) 볼드가 별표로 새는 문제 실제 답변에서 `**영업이익률 71.5%**와` 가 별표째 렌더됐다. micromark 로 확인: **71.5%**와 → bold 아님 (별표 노출) **100%**입니다 → bold 아님 **지속성**입니다 → bold 정상 **71.5%와** → bold 정상 CommonMark 는 구두점으로 끝난 강조 뒤에 글자가 바로 오면 닫지 못한다. 한국어는 퍼센트 뒤에 조사가 거의 항상 붙고, 데스크탑 프로파일이 '숫자를 볼드로' 지시하고 있어서 자주 터진다. 조사를 볼드 안에 넣도록 프로파일에 규칙을 추가했다. 검증: 루트 typecheck 클린, bun test 654 pass / 0 fail, desktop typecheck 클린. 내보내기는 앱을 띄워 버튼 클릭 → 토스트 → 클립보드 내용(KEY=value)까지 확인. 채널 프로파일은 실제 키로 질의해 마크다운 표·헤더·한글 종목명 렌더를 육안 확인 (기록은 키만 복사한 임시 프로필로 실행해 실제 DB 는 건드리지 않았다). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
앞 커밋에서 채널 프로파일에 '조사를 볼드 안에 넣어라' 규칙을 넣었는데, 실제 앱에서 같은 질문을 다시 던져 보니 모델이 그대로 `**...42.8%**로,` 를 냈다. 포맷 불변식을 모델 습관에 맡긴 판단이 틀렸다 — 규칙을 되돌리고 렌더러에서 결정적으로 처리한다. CommonMark 는 구두점으로 끝난 강조 뒤에 글자가 바로 오면 닫지 못한다. 한국어는 퍼센트 뒤에 조사가 붙으므로 데스크탑 답변에서 상시 발생한다. - renderer/markdown.ts: normalizeKoreanBold — 닫는 ** 앞이 구두점이고 뒤가 한글일 때만 조사를 볼드 안으로 옮긴다. 이미 파싱되는 형태(글자로 끝남 / 뒤가 공백·구두점) 는 손대지 않는다. - ChatView: 최종 답변과 추론 단계 두 렌더 지점 모두에 적용. - markdown.test.ts: 실제 답변에서 관찰된 문자열로 고정 + 무변경 케이스 검증. - channels.ts: 실패한 프롬프트 규칙 제거(매 요청 토큰만 쓰고 효과 없음). - tsconfig.web.json: 테스트 파일 제외(bun 런타임 전용). 검증(실제 키, 같은 질문): 수정 전: 별표 노출 2개, `**2026년 1분기 영업이익률은 42.8%**로,` 가 그대로 표시 수정 후: 별표 노출 0개, strong 요소 4개 — 볼드 정상 렌더 (스크린샷 확인) 루트 typecheck 클린, bun test 657 pass / 0 fail, desktop typecheck 클린. 실행은 키만 복사한 임시 프로필로 해 실제 DB 는 건드리지 않았다(대화 15개 유지). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
fix(desktop): 실행 모델 폴백을 카탈로그에서 도출 — 설정/실제 모델 불일치
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Author
|
Opened against the wrong repo (this is a fork-internal release PR) — apologies for the noise. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
v1.2.0 릴리스 준비
v1.1.2(6/29) 이후 main에 쌓인 변경을 릴리스합니다.
포함 내용
이 PR의 변경
desktop/package.json1.1.2 → 1.2.0update.jsonlatest 1.2.0 + 사용자용 노트머지 후
v1.2.0태그를 push하면 release.yml이 mac/win 인스톨러를 빌드해 드래프트 릴리스를 만듭니다.