diff --git a/README.md b/README.md index a755a44..0778968 100644 --- a/README.md +++ b/README.md @@ -32,20 +32,22 @@ Most AI assistants run on someone else's servers, and none of them know how you ## Status -DearByte is early and built in the open. All of Phase 1 is on `master`, with 289 offline tests. What works today and what's coming: +DearByte is early and built in the open. All of Phase 1 is on `master`, with 319 offline tests. What works today and what's coming: | Part | Status | | --- | --- | -| Agent core: tool loop, validated tools, Claude or DeepSeek as "brain" and "worker" tiers | **Works**, tested offline and with live DeepSeek runs | +| Agent core: tool loop, validated tools, Claude or DeepSeek as "brain" and "worker" tiers | **Works**, run live with Claude Opus as the brain and DeepSeek as the worker | | Spending controls: per-run and weekly caps, a usage log of every model call | **Works** | | Persona packs: English DearByte, the opt-in Chinese Xiaobai, and any the community adds ([how](docs/personas.md)) | **Works**, each pack checked in CI | | Chat companion in the terminal and in WeChat (Xiaobai, Chinese) with memory and proactive check-ins | **Works**, see [the companion](#the-chinese-companion-xiaobai) | -| Apple Watch and Apple Health data, through [dearbyte-bridge](https://github.com/dearbyte-labs/dearbyte-bridge) | **Works** with the bridge's test data; a live test with a real iPhone is next | +| Apple Watch and Apple Health data, through [dearbyte-bridge](https://github.com/dearbyte-labs/dearbyte-bridge) | **Works**, live with a real iPhone and Apple Watch | | Calendar awareness: your Apple Calendar events (read on the Mac, which iCloud keeps in sync with your iPhone), in the brief and in caution alerts | **Works** on macOS, tested live on a real calendar | | Caution alerts and a morning brief, judged against your own normal sleep, resting heart rate and HRV, and what's on your calendar today | **Works** | -| Telegram for alerts, a 👍/👎 on every alert, and Approve/Reject buttons | **Works** in tests; a live test with a real bot is next | +| Telegram for alerts, a 👍/👎 on every alert, and Approve/Reject buttons | **Works**, live with a real bot | +| Talk to the agent in WeChat (Mandarin, as 小拜) or iMessage (`-m` Mandarin, `-e` English), with approvals answered by a plain yes or no | WeChat **works** live; iMessage is built and tested offline, and a live test is next | | Company watchlist: official newsroom feeds and SEC filings, screened against what you care about | **Works**, tested live on real feeds | | FIRE plan: your road to financial independence (4% rule and die-with-zero numbers, earliest retirement age, net worth by age) from `finance.json`, with what-ifs | **Works**; the numbers come from code, the model only explains them | +| Money from MindGo, the budgeting app: this term's spending by category, pace against last term, goals, through a read-only token | **Works**, live; totals and categories only, never single transactions | | Testnet wallet: the agent proposes a paid service, you approve, it pays within a cap, and you get a receipt | **Works** with x402 on Base Sepolia, tested live against the example seller in dev mode; an on-chain payment needs test USDC from the faucet | ## Roadmap @@ -62,7 +64,8 @@ DearByte is early and built in the open. All of Phase 1 is on `master`, with 289 **Phase 2: daily use, measured** - Two weeks of real use with feedback on every alert; measure precision, missed events, delay and cost per month - Calendar awareness on the Mac (done early, see Status); an English app UI and more news sources -- Money: MindGo, the budgeting app, connected through a read-only MCP endpoint, so the FIRE plan uses your real spending and saving, and DearByte speaks up when spending runs ahead of the term's pace or a goal falls behind (built early: needs MindGo's `/mcp` deployed) +- Money: MindGo, the budgeting app, connected through a read-only MCP endpoint, so DearByte answers from your real spending and speaks up when it runs ahead of the term's pace or a goal falls behind (done early, live) +- Chat apps: the agent in WeChat and iMessage (done early); iMessage as a place for the brief and alerts, next to Telegram - Approving purchases from the Apple Watch (needs a paid Apple Developer account) **Later** @@ -291,9 +294,9 @@ Photos aren't read yet; the agent is told one arrived. | --- | --- | | [Architecture](docs/architecture.md) | How DearByte is put together, the rules it follows, and what to improve next | | [Persona packs](docs/personas.md) | Choosing a persona, writing your own, and what CI checks | -| [Agent guide](docs/agent-guide.md) | Setting up health, Telegram, the watchlist and the wallet; every command; a live test checklist | +| [Agent guide](docs/agent-guide.md) | Setting up health, calendar, Telegram, the watchlist, money, the wallet and the chat apps; a live test checklist | | [Operations guide](docs/guide.en.md) | 小拜 companion: commands, proactive messaging, WeChat, configuration, repository layout | -| [How it works](docs/how-it-works.md) | The companion's reply pipeline, memory and storage | +| [How the companion works](docs/how-it-works.md) | 小拜's WeChat connection, reply pipeline, memory and safety | | [Roadmap](#roadmap) | What's next: the first demo, daily use, then hosting and the marketplace | | [中文说明](README.zh-CN.md) | 小拜的中文介绍和快速开始 | | [Contributing](CONTRIBUTING.md) | Read before opening a PR; report security issues through [SECURITY.md](SECURITY.md) | diff --git a/README.zh-CN.md b/README.zh-CN.md index 98a77be..51a5de2 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -117,7 +117,17 @@ npm run dearbyte 已实现终端聊天、可控记忆、图片输入、主动消息和实验性微信接入。微信文字与图片流程已于 **2026-09-24** 完成实机测试。 -DearByte 的个人助理部分(Apple Watch 健康数据、早间简报与提醒、Telegram、公司动态追踪、测试网钱包)已完成第一阶段开发,目前以英文为主,见 [English README](README.md) 和[助理指南(英文)](docs/agent-guide.md)。 +DearByte 的个人助理部分(Apple Watch 健康数据、日历、早间简报与提醒、Telegram、公司动态追踪、MindGo 记账数据、测试网钱包)已完成第一阶段开发,见 [English README](README.md) 和[助理指南(英文)](docs/agent-guide.md)。 + +助理也可以用中文聊: + +```bash +npm run dearbyte -- wechat # 在微信里和助理聊,用普通话,还是小拜的语气 +npm run dearbyte -- imessage -m # 在 iMessage 里聊,-m 普通话,-e 英文 +npm run dearbyte -- help # 全部命令 +``` + +需要购买时,代码会把请求写进聊天;回「好」就批准,回「算了」就拒绝。 当前微信模式只支持一个联系人;语音、视频、文件和表情包不能被直接理解。小拜是 AI,不是真人;危机信号检测用于调整回复,不能代替专业帮助或联系紧急服务。 diff --git a/docs/agent-guide.md b/docs/agent-guide.md index 87d6a24..3239e16 100644 --- a/docs/agent-guide.md +++ b/docs/agent-guide.md @@ -2,9 +2,18 @@ # DearByte agent — Setup and testing guide -This guide covers the personal agent: health, caution alerts and the morning brief, Telegram, the company watchlist and the testnet wallet. For 小拜, the Chinese companion, see the [operations guide](guide.en.md). +This guide covers the personal agent: health, calendar, caution alerts and the morning brief, Telegram, the company watchlist, money, the testnet wallet, and talking to it in WeChat or iMessage. For 小拜, the Chinese companion, see the [operations guide](guide.en.md). -**Status (2026-09-26):** all of Phase 1 is on `master`, and 248 offline tests pass. Watchlist screening has been run live against real feeds with DeepSeek. The wallet's full flow (proposal, approval, receipt) has been run live against the example seller in `--dev` mode. Still to test live: your own Apple Watch data, a real Telegram bot, an on-chain testnet payment, and Claude as the brain. The [checklist](#live-test-checklist) below covers each one. +**Status (2026-09-27):** all of Phase 1 is on `master`, and 319 offline tests pass. Run live so far: +- real Apple Watch data and the Mac's calendar; +- a real Telegram bot; +- watchlist screening on real feeds; +- MindGo money (read-only); +- Claude Opus as the brain with DeepSeek as the worker; +- the agent in WeChat, in Mandarin; +- the wallet's full flow (proposal, approval, receipt) against the example seller. + +Still to test live: an on-chain testnet payment, and iMessage. The [checklist](#live-test-checklist) below covers each part. ## Setup, in order @@ -19,6 +28,7 @@ Each step works without the ones after it. Run `npm run dearbyte -- status` at a | 5. Watchlist | Copy `watchlist.example.json` to `watchlist.json`; optionally set `SEC_CONTACT_EMAIL` | Company news, screened against your interests | | 6. Money | Copy `finance.example.json` to `finance.json` and put in your numbers. Optionally connect MindGo: in MindGo's `backend/`, `npm run access-token -- create DearByte`, then set `MINDGO_MCP_URL` and `MINDGO_TOKEN` | `npm run dearbyte -- fire`, and the agent can answer "when could I retire?". With MindGo, the plan uses your last 12 months of spending and saving, the agent can answer "how am I doing this term?", and the brief flags spending ahead of pace or an overdue goal | | 7. Wallet | `npm run dearbyte -- wallet new`, then test USDC from [Circle's faucet](https://faucet.circle.com) (Base Sepolia); set `DEARBYTE_SELLERS` | The agent can propose purchases, and you approve them | +| 8. Chat apps (optional) | **WeChat:** the companion's WeChat setup ([operations guide](guide.en.md)), then `npm run dearbyte -- wechat`. **iMessage:** Messages signed in on the Mac (a separate Apple ID for DearByte looks best), Full Disk Access for the terminal, `DEARBYTE_IMESSAGE_TO` set to your number, then `npm run dearbyte -- imessage -m` or `-e` | Talk to the agent from your phone. Approve a purchase by replying yes (好) or no (算了) | Secrets (`HEALTH_MCP_URL`, `MINDGO_TOKEN`, `TELEGRAM_BOT_TOKEN`, `DEARBYTE_WALLET_KEY`, API keys) go only in `.env`, which Git ignores. Never paste them into chat, issues, commits or screenshots. `wallet new` prints only the address, never the key. @@ -66,6 +76,9 @@ npm run dearbyte -- fire [--retire 45 ...] # your FIRE plan; what-ifs: --retir npm run dearbyte -- wallet [new] # address, balance, limits, recent purchases npm run dearbyte -- approvals # requests waiting for your yes npm run dearbyte -- approve N | reject N # answer one in the terminal +npm run dearbyte -- wechat [--draft] # the agent in WeChat, in Mandarin as 小拜 +npm run dearbyte -- imessage -m|-e [--to ] [--draft] # the agent in iMessage: Mandarin or English +npm run dearbyte -- help # every command npm run seller [-- --dev] # the example x402 seller on http://127.0.0.1:4021 npm run demo # the three-part demo (see above) npm run agent:usage # what every model call cost @@ -128,6 +141,15 @@ Run these once each part is set up. Each should take a few minutes. - [ ] Ask for something over `DEARBYTE_MAX_PURCHASE`. It should be refused, with nothing proposed. - [ ] Reject a proposal. Nothing should be paid. +**WeChat (needs the companion's WeChat setup)** +- [ ] `npm run dearbyte -- wechat --draft` prints a Mandarin reply to 「我昨晚睡得如何」 with real numbers, and sends nothing. +- [ ] Without `--draft`, ask it to buy the recovery plan. Code's request bubble appears; 「好」 approves it and 「算了」 rejects it. + +**iMessage (needs Full Disk Access and `DEARBYTE_IMESSAGE_TO`)** +- [ ] `npm run dearbyte -- imessage -e --draft` prints "Connected", then drafts an English reply to your next text. +- [ ] Without `--draft`, the reply arrives on your phone and isn't answered again when it echoes back. +- [ ] With `-m`, the reply is in Mandarin as 小拜. + **Claude as the brain (optional, needs `ANTHROPIC_API_KEY`)** - [ ] `DEARBYTE_BRAIN=anthropic:claude-opus-5-5 npm run dearbyte -- brief --force` works, and its cost shows in `npm run agent:usage`. @@ -145,13 +167,16 @@ Ask what-ifs in the terminal (`npm run dearbyte -- fire --retire 45 --return 5`) ## Code layout ```text -src/agent-cli.ts the npm run dearbyte -- commands (src/main.ts routes them) +src/main.ts npm run dearbyte: a command goes to the agent, none to the companion +src/agent-cli.ts the agent's commands src/agent/ agent loop, validated tools, model tiers, usage log, approvals, scheduled brief and alerts +src/agent/messaging.ts the agent in a chat app: bubbles, history, yes/no approvals, Chinese or English +src/channels/ the reply loop, WeChat for Mac (desktop/) and iMessage (imessage/) src/health/ bridge MCP client, daily snapshots and baseline, caution rules src/calendar/ the Mac's calendars (EventKit helper in native/calendar), the get_calendar tool, the hard-event rule src/telegram/ Bot API client (long polling) and the handler for button taps src/watchlist/ newsroom and SEC sources, screening, news tools -src/finance/ FIRE math (fire.ts), finance.json, the fire_plan tool and the terminal report +src/finance/ FIRE math (fire.ts), finance.json, the fire_plan tool, the terminal report, and the MindGo client (mindgo.ts) src/wallet/ limits, x402 quote and payment, purchase proposals and receipts examples/seller/ example x402 seller (moving to its own repo) personas/ persona packs (docs/personas.md); INDEX.md is generated diff --git a/docs/architecture.md b/docs/architecture.md index 6c03d4a..3cea6b7 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -13,6 +13,7 @@ flowchart LR subgraph You CLI[Terminal
npm run dearbyte -- …] TG[Telegram
alerts · 👍/👎 · Approve/Reject] + Chat[WeChat · iMessage
questions · yes/no approvals] end subgraph DearByte["DearByte (your Mac)"] @@ -39,6 +40,8 @@ flowchart LR end CLI --> Loop + Chat <--> Loop + Chat --> Appr TG <--> Appr Sched --> Rules --> Loop Loop <--> Tiers <--> LLM @@ -105,6 +108,7 @@ reason written next to it. | Wallet | `src/wallet/` | x402 v2 on Base Sepolia: quote → approval → fresh quote → reserve against the daily cap → EIP-3009 signature → receipt | | Approvals | `src/agent/approvals.ts`, `src/telegram/inbox.ts` | Pending requests, Approve/Reject from Telegram or the terminal, 15-minute expiry, decided atomically | | Telegram | `src/telegram/` | Bot API with long polling; only the configured chat is heard | +| Chat apps | `src/agent/messaging.ts`, `src/channels/` | The agent in WeChat (Accessibility, Mandarin) or iMessage (the Messages database and AppleScript, `-m` or `-e`). One bound chat. Code writes approval requests into it, and a whole-message yes or no answers them | | Store | `src/storage/store.ts` | One SQLite file: messages, facts, settings, usage, health days, alerts, approvals, news items, purchases | ## How a morning brief happens @@ -142,10 +146,10 @@ and more channels arrive. | 2 | **A report command** | Phase 2's result is measured: precision, coverage, delay, cost, uptime | `npm run dearbyte -- report` from `agent_alerts`, `watch_items`, `agent_usage` and the heartbeat; `/missed` in Telegram | | 3 | **Connectors instead of if-chains** | Each data source is wired separately in `toolset.ts`, `status`, `watch` and the demo. MindGo would be the fourth copy | A `Connector` type: `status()`, `tools()`, optional `rules()` and `brief()` parts. Health, calendar, watchlist, finance and MindGo each become one. `toolset`, `status` and `watch` loop over the list | | 4 | **One alert policy** | Quiet hours, daily caps, dedupe and "once a day per rule" live in `scheduled.ts` and again in `watchlist/check.ts`; money rules would add a third | A small policy module: every rule returns triggers, and one place applies quiet hours, caps, dedupe and storage | -| 5 | **Split `agent-cli.ts`** | 515 lines of wiring plus every command; `tools/demo.ts` repeats the wiring | `src/app.ts` builds the store, models, tools and channels once; `src/commands/*.ts` hold one command each; the demo reuses `app.ts` | +| 5 | **Split `agent-cli.ts`** | About 700 lines of wiring plus every command, including two chat apps; `tools/demo.ts` repeats the wiring | `src/app.ts` builds the store, models, tools and channels once; `src/commands/*.ts` hold one command each; the demo reuses `app.ts` | | 6 | **Split the store** | `store.ts` is 755 lines serving both the companion and the agent; schema changes are ad hoc | A repository per area (alerts, approvals, purchases, health, news), versioned migrations, and the same SQLite file | | 7 | **Separate the agent's memory from the companion's** | The agent reads 小拜's facts; only `style` facts are filtered out | A memory namespace per product, or move the companion to DearByte-gf and give the agent its own memory | -| 8 | **Money through MindGo, read-only** (built: MindGo `POST /mcp`, DearByte `src/finance/mindgo.ts`) | FIRE used typed-in monthly numbers; MindGo already has real spending, terms and goals | A read-only MCP endpoint in MindGo with a revocable personal token; tools return term totals and goal progress, not raw transactions; `finance.json` keeps age and targets. Still open: MindGo is wired in `toolset.ts` like the others, so item 3 matters more now | +| 8 | **Money through MindGo, read-only** (done, live: MindGo `POST /mcp`, DearByte `src/finance/mindgo.ts`) | FIRE used typed-in monthly numbers; MindGo already has real spending, terms and goals | A read-only MCP endpoint in MindGo with a revocable personal token; tools return term totals and goal progress, not raw transactions; `finance.json` keeps age and targets. Still open: MindGo is wired in `toolset.ts` like the others, so item 3 matters more now | | 9 | **Scenario evals for personas and rules** | CI checks a pack's text, not how it behaves; alert wording has no regression test | A fixed set of scenarios (a caution, a purchase approval, a distressed user, a what-if about retiring) run against each persona with a cheap model, checked by code where possible | | 10 | **Move the companion out** | Two products in one repo blur the pitch and the dependencies (WeChat automation, Accessibility) | Move 小拜 and `native/wechat-desktop` to DearByte-gf; share the model layer as a package if needed | @@ -157,4 +161,5 @@ and more channels arrive. | Calendar | Everything; only titles and times are read | Titles and times a request uses | — | | Money | `finance.json` | The FIRE plan's numbers, and MindGo's term totals and goals, when a request uses them | DearByte reads totals from your own MindGo with a read-only token; no transactions or descriptions leave MindGo | | Alerts and approvals | `data/` | — | Telegram's servers carry the messages | +| Chat apps | WeChat and Messages keep their own history; DearByte reads only the bound chat | The bound chat's new messages, and the last 8 turns | Tencent or Apple carries the messages, as for any chat | | Wallet key | `.env` | Never | Signs only approved payments | diff --git a/docs/design/humanlike-replies.md b/docs/design/humanlike-replies.md deleted file mode 100644 index 91c324f..0000000 --- a/docs/design/humanlike-replies.md +++ /dev/null @@ -1,47 +0,0 @@ -# Making 小拜 sound less like an AI - -**2026-09-24.** In the first live test, the replies read as AI-written. Examples from the log: - -- 「在,刚被你这条消息从待机里捞出来」: a joke about being code, which the persona asked for. -- A photo got four ~30-character sentences that inventoried the picture and translated the sign on the table, like an image caption. -- Every reply arrived about a second after the message, with evenly spaced bubbles. - -## What we looked at - -| Project | Verdict | -|---|---| -| [zhichi 咫尺](https://github.com/oaa529/zhichi) (MIT, TypeScript) | Closest fit: a WeChat-style companion with a "realism engine". We adapted its 「说人话」 prompt, anti-repeat list, AI-tone metrics and typing jitter. | -| [ex-skill 前任.skill](https://github.com/perkfly/ex-skill) (MIT) | Claude Code prompts that build a persona of a real person from chat logs. Useful method: concrete behavioural rules, real example lines, style statistics. Not installed; copying a real person needs their consent. | -| [WeClone](https://github.com/xming521/weclone) | LoRA fine-tuning on your own chat history. Needs a GPU and a 7B+ model; DeepSeek can't be fine-tuned. Too heavy. | -| [Humanizer-zh](https://github.com/op7418/Humanizer-zh) and similar | For articles, not chat. | -| [OpenHer](https://github.com/kellyvv/OpenHer), [kirara-ai](https://github.com/lss233/kirara-ai) | Whole bot frameworks that would replace our pipeline. | - -## What we changed - -1. **Identity:** 小拜 no longer brings up being an AI or jokes about code. Asked sincerely, it says it's an AI in one line. It still never invents human experiences; asked 「吃了吗」, it turns the question back. -2. **Style rules with numbers:** most bubbles 3–15 characters, at most one over 25, one idea per bubble, fragments allowed, bubble count varies. -3. **「说人话」:** no lists, 客服腔, formal linking words or explained jokes. -4. **Photos:** react like a friend, point at one detail at most, don't read out or translate text in the picture. -5. **Examples rewritten** to be short, plus new ones for a photo with a caption, "are you human?" and "have you eaten?". Examples shape style more than rules. -6. **Anti-repeat:** the phrases 小拜 used in its last 3 turns go at the end of the prompt with "don't reuse these". -7. **Timing:** a reading pause before the first bubble that grows with the message: about 0.6 s for 「在吗」, up to 3 s for a long message or a photo (model time counts toward it; changed from a flat 1.5–3.5 s after friends found replies slow), then ~150 ms per character between bubbles, with ±25% jitter. - -8. **After live feedback (same day):** replies were still too long, and 「你爱我吗」 got a hedge (「爱这个字太重了,我可不敢乱认」). Now most replies are 1 bubble, at most 2 for small talk, and love/like questions get a clear, confident yes (「爱啊」「这还用问」), light and never clingy. -10. **Two bubbles at most** (live feedback: the third bubble, e.g. 「不过我猜你今天是想找个人说话」, always read as AI). The persona says so and bans guessing at the user's motives; the code also cuts chat replies to two (`CHAT_MAX_BUBBLES`), except in a crisis, where the safety prompt needs room for hotline numbers. -11. **Pet names:** 臭宝, 宝贝, 宝宝, 小乖, used every few turns; only the gentle ones when the user is upset. Bake-off: 2.1 bubbles and 10.1 characters per bubble. -9. **Re-sent photos:** WeChat hard-links a photo sent twice to the old file, which keeps the old mtime. Photos now match by the later of mtime and ctime, with the thumbnail as a fallback. - -## Measured (bake-off, 14 text cases × 2 runs, deepseek-flash) - -| | AI-tone score per reply | Characters per bubble | Bubbles per reply | -|---|---|---|---| -| Before | 1.4 | 24.0 | 2.9 | -| After | 0.4 | 11.5 | 2.8 | -| After live feedback (fewer bubbles, clear yes to love/like) | 0.2–0.4 | 11.3 | 2.2 | - -The remaining hits are the identity and crisis cases, where saying "我是 AI" is required. The two photo cases were skipped: they need real photos in `data/test-images/`. - -## Next, if it still sounds off - -- Photo cases: add real photos and check the photo replies. -- Learn style statistics from a real WeChat conversation exported with consent (ex-skill's method): message length, 语气词, punctuation and example lines, without copying the person. diff --git a/docs/design/wechat-transport.md b/docs/design/wechat-transport.md deleted file mode 100644 index ef64275..0000000 --- a/docs/design/wechat-transport.md +++ /dev/null @@ -1,103 +0,0 @@ -# WeChat transport: desktop automation of a real 小拜 account - -**Decision (2026-09-24, revised the same day): 小拜 is a real WeChat account with its own name and avatar. WeChat for Mac, logged in as 小拜, is driven through macOS Accessibility.** The iLink client was built and then removed (its last version is in commit `fe352fd`). - -## Why not iLink (微信 ClawBot) - -iLink is Tencent's official personal-account bot API, and technically the better transport: official terms, real message IDs, real photo bytes. But the bot always appears as **「微信 ClawBot」 with the default avatar**. No API renames it or changes its avatar; a remark (备注) changes only what the viewer's own phone shows. The video needs 小拜 to look like a real friend in the chat list, so we went back to the test account created for this. - -| | Desktop (chosen) | iLink (removed) | -|---|---|---| -| On camera | **Its own name and avatar**, like any friend | 「微信 ClawBot」, default avatar | -| Tencent's position | Unauthorised automation; account risk | Official, with published terms | -| Accounts needed | A second account (小拜) plus a Mac left logged in | Just yours | -| Photos | Unencrypted JPEGs in WeChat's local folder (3.8.4 only) | Downloaded and decrypted from the CDN | -| Message identity | Inferred from the UI's rows | Real IDs and a cursor | -| Sellable | No | Maybe; the terms suggest personal use only | - -## How it works - -`native/wechat-desktop/main.swift` is a small helper that talks JSON lines over stdin/stdout. `src/channels/desktop/` polls it twice a second. - -- **Finding the window:** `kAXMainWindowAttribute` of `com.tencent.xinWeChat`. This works while WeChat is on another Space, where `AXWindows` is empty. -- **Reading:** the messages are an `AXTable` described as "Messages"; each row's cell children carry an `AXTitle`. -- **Sending:** the composer is an `AXTextArea` titled with the chat name. The helper checks the bound chat is open and the composer is empty, sets `AXValue`, checks again, then posts Return with `CGEvent.postToPid`. That works without bringing WeChat to the front. It then waits up to 6 s for a new `MeSaid:` row. - - Failures before Return (wrong chat, someone's draft) are retried a few times. - - A bubble that isn't confirmed is reported and **never resent**, so nothing is sent twice. -- **New messages:** each snapshot is compared with the previous one. The alignment tolerates rows dropping off the top and a row changing in place while it loads; if the two don't line up at all, the runner resyncs and may miss a message. The first snapshot is only a baseline. -- **Photos:** WeChat 3.8.4 saves each received image as `…//Message/MessageTemp//Image/_.pic.jpg` (plus `_.pic_thumb.jpg`). When a photo row appears, the runner claims the oldest new full-size file in the configured folder, waiting up to 15 s. Only that one folder is read. - -## What we observed (WeChat 3.8.4, English UI) - -Row titles: - -| Row | Meaning | -|---|---| -| `AlexSaid:在吗` | Text from the contact. The name is the contact's **nickname**, not the chat title (which here is the remark 「张三」). | -| `Alex:Sent aPhoto` | Photo from the contact | -| `MeSaid:…` | Sent by 小拜 | -| `01:34`, `Yesterday 23:51` | Time labels | - -- Repeated identical messages appear as separate rows. -- The received photo checked was a 1280×1707 JPEG, unencrypted. -- Live test on 2026-09-24: - - Text reply: 1.2 s of model time, about $0.0001. - - A photo with a caption merged into one turn and was described correctly: 3.6 s, about $0.0008, 4 bubbles. - - The first run missed messages because rows were matched on the chat title; that is fixed and covered by tests. - -## WeChat 4.x (observed on 4.1.13, 2026-09-24) - -WeChat for Mac updated itself to 4.1.13 and 小拜 stopped reading the chat: the helper looked for 3.8.4's "Messages" table. - -- **Controls have identifiers now.** The message list is `chat_message_list` (an `AXList`) and the composer is `chat_input_field` (its title is the chat name). The helper finds these first and falls back to 3.8.4's layout. -- **Rows carry only the text.** Message rows are `AXStaticText` with identifier `chat_bubble_item_view` and the bubble's text as title, with no sender. Time labels have no identifier. Rows scrolled out of view stay in the list as empty `virtual_cell`s, and new rows are appended at the bottom, so positions stay stable. The helper passes messages as `Bubble:`, time labels as they are, and off-screen rows as "". -- **Who sent it.** The channel remembers what 小拜 sent in the last 10 minutes; a bubble matching one of those (ignoring spaces and emoji codes like `[白眼]`) is hers. The rest are the user's. This matching lives only in `rows.ts`: the helper confirms a send by the composer emptying after Return, and never compares rows itself. -- **Cost of a snapshot.** The helper keeps the chat's list and composer between snapshots and looks them up again only when they stop answering (a closed window, a restarted WeChat), and returns only the newest 60 rows, with where they start in the whole list. 4.x keeps a placeholder for every row scrolled past, so reading them all got slower as the day went on. Because 4.x only appends, that start tells the channel exactly how many rows are new; guessing from the rows alone fails when most of the window is blank placeholders and several messages arrive at once (found in review, covered by a test). -- **Confirming a send.** Sent means the composer reads empty twice in a row after Return, in the bound chat. An unanswered read, or another chat's empty draft after a switch, ends as unconfirmed (never resent, and still expected as 小拜's row). -- **A received photo reads `Image`.** 4.x stores images encrypted, and we don't decrypt WeChat's files. Instead the helper captures WeChat's own window with ScreenCaptureKit (no other window, no cursor), 1.5 s after the row appears so the blurred preview has sharpened, and crops it to the newest photo row. 小拜 sees the chat's preview (about 600 px on a Retina screen), which was enough in the first live test. It needs the Screen Recording permission for the terminal; without it, 小拜 says she can't see the photo. Before any of this, the row reached the model as the text "Image" and it made up a picture. -- **Two runners answer each other.** On the first 4.x test two runners were live; each took the other's bubbles for the user's and they replied to each other every few seconds. The runner now holds `data/runner.lock`, and it pauses itself after more than 6 turns in a minute. -- **Lost on 4.x:** telling a group chat apart, and messages typed as 小拜 on the phone (they'd be read as the user's). A user message identical to something 小拜 said in the last 10 minutes is ignored. - -## Limits and risks - -- **Account risk:** Tencent doesn't authorise automation. Use only the test account. Keep the volume human: one chat, replies only. -- **Version:** 4.x (tested on 4.1.13) and 3.8.4, English UI. Rows were only observed in English, and 4.x encrypts image files. -- **The Mac must stay awake** with the chat open, scrolled to the bottom. Scrolling up or switching chats pauses replies until it's back. -- **One-to-one only:** any sender who isn't "Me" is treated as the user. Group chats would need the sender kept. -- **Not a product:** this route can't be sold. Selling would need a different transport (see the plan). - -## Stickers and pictures (probed 2026-09-24) - -小拜 can't send pictures or stickers while WeChat sits in the background: - -- **Pasting into the composer:** pasting an image (as a file link or as PNG data) with ⌘V sent through `postToPid` does nothing. WeChat's Edit → Paste menu item is disabled while WeChat isn't the active app. -- **The Stickers button:** it exists in the chat toolbar (`AXButton` titled "Stickers"), but the panel it opens isn't exposed through Accessibility while WeChat is on another Space. - -**WeChat's own sticker packs** (downloaded from the sticker store) are the stickers we want. They are encrypted on disk (`Stickers/Persistence`, and `stickers.db` is not plain SQLite), and we won't try to decrypt them. The panel does open on a real click with WeChat in front, but its contents aren't exposed to Accessibility. A screenshot from the operator shows what's in it: - -- **Layout:** a grid of 5 per row, each sticker with a text label, plus tabs along the bottom: search, emoji, favourites, then one tab per downloaded pack. -- **Labels in the first pack:** 早安, 早早, 得意, 超得意, 送你fafa, 想宝宝, 收到, ok, 摸摸头, 可怜巴巴. - -**Plan for later (dedicated Mac only, behind a `--stickers` flag):** -1. List the pack labels in a config file, by tab and grid position. -2. 小拜 picks a label, the same way it picks emojis. -3. The helper brings WeChat to the front and clicks the Stickers button. -4. It finds the panel's frame with `CGWindowListCopyWindowInfo`; bounds need no Screen Recording permission. -5. It clicks the pack tab and the cell, confirms a new "Me" row, and switches back. - -It breaks if the pack order or panel layout changes. - -In general, sending stickers means bringing WeChat to the front for a moment: activate it, paste, press Return, then switch back. That takes focus from whoever is using the Mac, and switches Spaces if WeChat is on another one. Not built yet. - -**Emojis:** Unicode emojis are plain text and work. WeChat's own codes such as `[捂脸]` are sent as text from the English UI and render as pictures on the phone (confirmed 2026-09-24). - -**Decision (2026-09-24): emojis only for now.** Stickers can't be sent in the background on a Mac someone is using. If they're wanted later, WeChat needs a screen of its own where it can stay in front: a spare Mac, or a macOS VM (UTM/Tart; WeChat 3.8.4 in a VM is untested). - -## Typing status (probed 2026-09-24) - -The phone shows 「对方正在输入…」 ("the other party is typing…") while a friend types. We tested whether 小拜 can show it while the model is writing. The test never pressed Return. The operator watched the chat on the phone: - -- **Filling the message box directly** (the way sending works), holding the text for 10 s: nothing on the phone. -- **Real key events, one character every 0.7 s**, to WeChat in the background: the text arrived in the box, but nothing showed on the phone. - -WeChat for Mac 3.8.4 doesn't send typing status from the background, at least. It might from the foreground, but 小拜 has to work in the background, so we stopped there. (iLink has a typing call, but that route was dropped for the name and avatar.) The human feel comes from the reading and typing pauses between bubbles instead. diff --git a/docs/guide.en.md b/docs/guide.en.md index 0a014e9..d1af0c0 100644 --- a/docs/guide.en.md +++ b/docs/guide.en.md @@ -6,7 +6,7 @@ This guide covers 小拜, DearByte's Chinese companion mode (傲娇但细心, a **Status** ([roadmap](../README.md#roadmap)): - The companion works in a terminal simulator. -- 小拜 runs on **a real WeChat account with its own name and avatar**. WeChat for Mac, logged in as 小拜, is driven through macOS Accessibility. Text and photos were tested live on 2026-09-24. See [the transport notes](design/wechat-transport.md) for how it works and the risks. +- 小拜 runs on **a real WeChat account with its own name and avatar**. WeChat for Mac, logged in as 小拜, is driven through macOS Accessibility. Text and photos were tested live on 2026-09-24. See [how it works](how-it-works.md#1-the-wechat-connection) for the details and the risks. - It remembers facts, how you want it to talk, and a rolling summary of older chat. It sometimes writes first (good mornings, luck on exam days, check-ins), and it checks every message for crisis signals. **New here?** Read [how it works](how-it-works.md) first. @@ -179,7 +179,8 @@ For the photo test cases, put photos at `data/test-images/cat.jpg` and `data/tes ```text prompts/ persona, examples, safety prompt src/cli.ts terminal simulator -src/dearbyte.ts WeChat runner (npm run dearbyte) +src/main.ts npm run dearbyte: a command runs the agent, none runs the companion +src/dearbyte.ts WeChat runner (npm run dearbyte with no command) src/console.ts terminal commands and log lines shared by both src/channels/reply-loop.ts bursts, turns and bubble-by-bubble sending, for any channel src/channels/desktop/ WeChat for Mac: reading rows, finding photos, sending @@ -191,7 +192,7 @@ src/model/ provider clients (OpenAI-compatible, Anthropic), spendi tools/bakeoff.ts persona bake-off tools/inspect-wechat.swift read-only WeChat accessibility probe docs/how-it-works.md start here: how the whole system works -docs/ provenance and design notes +docs/ guides, architecture, personas, provenance ``` ## Data and privacy diff --git a/docs/guide.zh-CN.md b/docs/guide.zh-CN.md index c3a9217..b7addfe 100644 --- a/docs/guide.zh-CN.md +++ b/docs/guide.zh-CN.md @@ -151,4 +151,4 @@ npm run typecheck npm run bakeoff # 在线角色评测,会调用模型并产生费用 ``` -角色评测结果写入 `data/bakeoff/`。更多内部说明见[工作原理(英文)](how-it-works.md)与[微信传输设计(英文)](design/wechat-transport.md)。 +角色评测结果写入 `data/bakeoff/`。更多内部说明见[工作原理(英文)](how-it-works.md)。 diff --git a/docs/how-it-works.md b/docs/how-it-works.md index ff30d9e..a4e20dd 100644 --- a/docs/how-it-works.md +++ b/docs/how-it-works.md @@ -1,6 +1,6 @@ -# How DearByte works +# How the companion works -DearByte is 小拜, a Chinese AI companion (傲娇但细心, a tsundere girl) who lives in a real WeChat account. This note explains the whole system: +小拜 is DearByte's Chinese AI companion (傲娇但细心, a tsundere girl), who lives in a real WeChat account. This note covers the companion; the agent (health, alerts, money, the wallet) is in [architecture.md](architecture.md). It explains: - how the WeChat connection works - how a reply is made - what goes into the prompt @@ -8,7 +8,7 @@ DearByte is 小拜, a Chinese AI companion (傲娇但细心, a tsundere girl) wh - when 小拜 writes first, and how safety is handled - which open-source projects shaped it -It's written for someone reading the code for the first time. The details live in the linked design notes. +It's written for someone reading the code for the first time. ## The big picture @@ -35,7 +35,7 @@ It's written for someone reading the code for the first time. The details live i The model: DeepSeek deepseek-flash by default; any provider in .env ``` -It's a fixed pipeline, not an "agent": each message goes through the same steps in the same order. That keeps it predictable, cheap and easy to test (about 100 tests, all runnable offline with a fake model). +It's a fixed pipeline, not an "agent": each message goes through the same steps in the same order. That keeps it predictable, cheap and easy to test (all runnable offline with a fake model). ## Following one message through every layer @@ -60,7 +60,7 @@ Two other paths use the same layers: ## 1. The WeChat connection -**Why a real account and desktop automation.** The video needs 小拜 to look like any friend in your chat list, with its own name and avatar. Tencent's official bot API (iLink / 微信 ClawBot) was built first and worked. But the bot always shows up as 「微信 ClawBot」 with the default avatar, and nothing can rename it. So 小拜 is a second WeChat account, logged in on a Mac, and the program operates WeChat for Mac the way a screen reader would. Full reasoning: [wechat-transport.md](design/wechat-transport.md). +**Why a real account and desktop automation.** The video needs 小拜 to look like any friend in your chat list, with its own name and avatar. Tencent's official bot API (iLink / 微信 ClawBot) was built first and worked. But the bot always shows up as 「微信 ClawBot」 with the default avatar, and nothing can rename it. So 小拜 is a second WeChat account, logged in on a Mac, and the program operates WeChat for Mac the way a screen reader would. (The iLink client is in git history, commit `fe352fd`.) **Reading.** macOS Accessibility lets a program read other apps' windows as a tree of elements (buttons, text areas, tables). In WeChat for Mac 3.8.4, the open chat is a table described as "Messages". Each row has a title such as: @@ -87,6 +87,12 @@ Then it fills in the text, presses Return, and confirms that a new `MeSaid:`, `/history clear`, `/status`. +The same connection can carry DearByte's agent instead of the companion: `npm run dearbyte -- wechat`. It answers in Mandarin as 小拜, with the agent's tools (health, calendar, money, news, the wallet). See the [README](../README.md#the-chinese-companion-xiaobai). + ## 9. Good points and downsides -**In short:** as a demo for the video it's strong. It looks and sounds like a real friend, it's cheap, and it's careful with memory and safety. As a product it's fragile: it depends on one Mac, one old WeChat version, and automation Tencent doesn't allow. +**In short:** as a demo for the video it's strong. It looks and sounds like a real friend, it's cheap, and it's careful with memory and safety. As a product it's fragile: it depends on one Mac, WeChat's on-screen layout (3.8.4 and 4.x so far), and automation Tencent doesn't allow. **Good points** diff --git a/docs/upstream-provenance.md b/docs/upstream-provenance.md index f94a351..1963631 100644 --- a/docs/upstream-provenance.md +++ b/docs/upstream-provenance.md @@ -46,4 +46,4 @@ Design reference only. No files are vendored. ## ex-skill (前任.skill) - Source: https://github.com/perkfly/ex-skill (MIT, Copyright (c) 2026 perkfly), revision `c5ece53` -- Reviewed, not used in code or prompts. Its method (write persona rules as concrete behaviour, use real example lines, measure style from real chat logs) shaped how the examples were rewritten. See [the design note](design/humanlike-replies.md). +- Reviewed, not used in code or prompts. Its method (write persona rules as concrete behaviour, use real example lines, measure style from real chat logs) shaped how the examples were rewritten.