Skip to content

A/B tests decide themselves: 50+ per arm, by reply rate, winner becomes the default - #128

Merged
ralyodio merged 5 commits into
mainfrom
feat/planner-ab-winner
Oct 7, 2026
Merged

ralyodio merged 5 commits into
mainfrom
feat/planner-ab-winner

Conversation

@ralyodio

@ralyodio ralyodio commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Third Hunter outreach-planner feature, fully automated. The planner's rule: test one variable at a time, with at least 50 recipients per variant, measure by reply rate (not opens), and carry the winner forward.

What changes

  • decideAbTest (domain, pure): every arm needs 50+ sends. The leader wins once a one-sided two-proportion z-test against the runner-up reaches 95%. If every arm reaches 200+ sends without that, the best arm wins anyway, and A keeps its place when nothing beat it. Opens never count (Apple Mail fetches every pixel).

  • promoteAbWinners runs hourly per workspace in the worker, before cadences:

    • the winning angle becomes the step's intent
    • the variants are cleared
    • the test is recorded in ab_results with every arm's numbers
    • an event lands in the live feed

    A later test on the same step only counts runs after the decision.

  • Bug fix: the per-variant report counted only state = 'replied', but the mailbox poll records replies as responded. Every reply the poll picked up was missing from A/B results. The query now lives in the pipeline (variantStats) and reads both.

  • Surfaces: GET /cadences/:id/variants now includes decided, plus a new GET /ab-results, og cadences ab <id> and MCP get_ab_results.

Migration 0054 adds one new table.

Tests: domain ab-testing.test.ts (5) and pipeline ab-winners.test.ts (2: promotion with poll-recorded replies, and a test still short of 50 left alone). The API variants, MCP, CLI and migration suites pass locally (73 tests).

🤖 Generated with Claude Code

ralyodio and others added 5 commits October 7, 2026 02:12
…retire bounced addresses

Hunter's outreach planner, run by the worker instead of remembered by a person:

- Every address is verified before its first message and again after 90 days
  (MX, then an SMTP RCPT probe where port 25 allows; a blocked port sends on
  MX alone). Verdicts are cached in email_verifications, 25 fresh checks per run.
- A bounce the delivery report pins to an address marks it invalid for good
  and stops the person's cadences. The sender query skips invalid addresses.
- A campaign whose bounce rate passes 2% over 50+ sends pauses itself,
  re-verifies every queued address, drops the failures and resumes on a fresh
  window once a complete pass finds nothing waiting (campaign_list_health).
- Accept-all addresses send only while the campaign bounces at or under 1%.
- GET /autogtm/campaigns/:id/list-health, og leads health, MCP get_list_health.

Migration 0053.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…udget holder

Hunter's planner reaches four people per target company, one at a time, a
month apart: budget holder, pain feeler (end user), blocker (legal, security,
IT, finance) and champion (manager). Now automatic:

- classifyPersona reads the job title deterministically (no model on the
  send path).
- Autopilot orders each company's queued cards by persona behind its best
  card, without changing the order between companies.
- A colleague emailed in the last 21 days, or on an active cadence, who has
  not answered holds the company; everyone else there waits, card intact.
  A reply releases it. Deliberately reverses the old "two mailboxes at one
  company both get written to in one run" test.
- Drafts get a per-role angle added to the step's guidance.
- GET /autogtm/campaigns/:id/accounts, og leads accounts, MCP
  get_campaign_accounts: who was reached, who is next, missing personas.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts:
#	apps/api/src/autogtm-docs.ts
#	apps/api/src/autogtm.ts
#	apps/cli/src/commands.ts
#	apps/mcp/src/tools.ts
#	packages/domain/src/index.ts
#	packages/pipeline/src/autopilot.ts
#	packages/pipeline/src/index.ts
…es the default

Hunter's planner: test one variable, 50+ recipients per variant, measure by
reply rate, carry winners forward. Now automatic, hourly per workspace:

- decideAbTest: every arm needs 50+ sends; the leader wins at a one-sided
  95% two-proportion z-test against the runner-up, or once every arm has
  200+ the best wins (A keeps its place on a tie). Opens are not an input.
- promoteAbWinners makes the winning angle the step intent, clears the
  variants and records the test in ab_results; later tests on the step
  count only runs after the decision.
- Fix: the variant report counted only `replied`, but the mailbox poll
  writes `responded`, so every polled reply was missing from A/B results.
  The query now lives in the pipeline (variantStats) and reads both.
- GET /cadences/:id/variants now includes `decided`; GET /ab-results;
  og cadences ab <id>; MCP get_ab_results.

Migration 0054.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts:
#	packages/domain/src/index.ts
#	packages/pipeline/src/index.ts
@ralyodio
ralyodio merged commit 29deff4 into main Oct 7, 2026
4 checks passed
@ralyodio ralyodio mentioned this pull request Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant