A Rust-based load testing tool for benchmarking Keycast's HTTP RPC endpoint (POST /api/nostr).
# Build the tool
cargo build --release -p keycast-loadtest
# Create test users (against running server)
./target/release/keycast-loadtest setup \
--url http://localhost:3000 \
--users 100 \
--output ./test-users.json
# Run load test
./target/release/keycast-loadtest run \
--url http://localhost:3000 \
--users-file ./test-users.json \
--concurrency 50 \
--duration 60 \
--seed 371 \
--scenario warm-cache \
--method get-public-key \
--output ./results.jsonCreates users via HTTP registration API with full OAuth flow (register + authorize + token exchange).
keycast-loadtest setup \
--url <server-url> \
--users <count> \
--output <json-file>What happens:
- Registers user with email/password
- Approves OAuth authorization (with PKCE)
- Exchanges code for access token (UCAN with
bunker_pubkey) - Saves credentials to JSON file
Important: Users are created against a specific SERVER_NSEC. If the server restarts with a different key, you need to recreate users.
Sends concurrent RPC requests and measures performance.
keycast-loadtest run \
--url <server-url> \
--users-file <json-file> \
--concurrency <num> \
--duration <seconds> \
--scenario <warm-cache|cold-start|mixed> \
--method <get-public-key|sign-event> \
--output <results.json>Scenarios:
| Scenario | Behavior | Use Case |
|---|---|---|
warm-cache |
Reuses the first user | Steady-state performance |
cold-start |
Each request uses different user | Cache miss impact |
mixed |
80% repeat / 20% new users | Realistic traffic |
Methods:
| Method | What it does | CPU Cost |
|---|---|---|
get-public-key |
Returns user's pubkey | Minimal |
sign-event |
Schnorr signature | ~1-3ms |
keycast-loadtest report --input ./results.jsonClient Request
│
▼
┌─────────────────────────────────────────────────────────┐
│ POST /api/nostr │
│ Authorization: Bearer <UCAN with bunker_pubkey> │
│ Body: {"method": "get_public_key", "params": []} │
└─────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ UCAN Token Verification │
│ - Verify signature against SERVER_NSEC pubkey │
│ - Extract: user_pubkey, redirect_origin, bunker_pubkey│
└─────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ Handler Cache Lookup (LRU) │
│ Key: bunker_pubkey (32 bytes) │
│ Capacity: 1,000,000 (configurable via HANDLER_CACHE_SIZE)
│ TTL: 1 hour idle timeout │
│ │
│ HIT ──► Return cached HttpRpcHandler │
│ MISS ──► Load from DB, cache, return │
└─────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ Execute RPC Method │
│ - Validate authorization (cached: expires_at, revoked)│
│ - Check permissions (cached: policy rules) │
│ - Perform operation (sign/encrypt/decrypt) │
└─────────────────────────────────────────────────────────┘
The HttpRpcHandler caches everything needed to process requests without DB hits:
- User's signing keys (decrypted)
- Authorization metadata (expires_at, revoked_at)
- Permission rules (from policy)
- Cache keys (bunker_pubkey, authorization_handle)
| Component | Typical Latency |
|---|---|
| Network (local) | <1ms |
| Network (Cloud Run) | ~200ms |
| Cache hit | <1ms |
| Cache miss (DB load) | 20-100ms |
| Schnorr signing | 1-3ms |
| NIP-44 encrypt/decrypt | <1ms |
After running tests, check metrics at /api/metrics:
curl http://localhost:3000/api/metrics | grep http_rpcKey metrics:
keycast_http_rpc_requests_total- Total requestskeycast_http_rpc_cache_hits_total- Cache hitskeycast_http_rpc_cache_misses_total- Cache misses (DB loads)keycast_http_rpc_success_total- Successful requestskeycast_http_rpc_auth_errors_total- Auth failures
The capacity command runs a predeclared sequence and writes a versioned,
aggregate-only evidence bundle. Copy capacity-plan.json.example, replace its
illustrative objectives and revision placeholders with values approved for the
specific run, and keep separate plans and outputs for Cloud Run and GKE.
keycast-loadtest capacity \
--plan ./capacity-plan.json \
--output ./capacity-evidenceEach plan must include ramp, spike, soak, recovery, rollout, and scale-down
exactly once in that order. Every phase declares its method, cache scenario,
concurrency, duration, client-side unintentional-error and p95 latency limits,
and whether the profile SLO applies. The plan seed deterministically
offsets the user-selection schedule and is recorded in every phase. Registration
is excluded because its generated identities are not seed-reproducible. The
command records ordinary phase or SLO misses and continues so later recovery can
still be measured. Absolute profile stop conditions are checked during each
phase and stop active work when breached. evidence.json is also written when a
phase errors, preserving all completed and partial evidence.
Intentional capacity shedding is distinct from caller-specific HTTP 429 rate
limiting and from ordinary dependency failures. It is recognized only as HTTP
503 with the machine-readable response code admission_rejected. A spike may
declare the following expectation once the target implements that contract:
"intentional_shedding": {
"status": 503,
"error_code": "admission_rejected",
"min_rate": 0.01
}When the field is null, the evidence records shedding as not-tested; the
rehearsal may pass its other checks but certification_ready remains false.
Unmarked 503 responses remain unintentional errors. A final platform capacity
profile must test and pass the declared shedding expectation.
Healthy-phase availability counts every failed request, including unexpected
HTTP 429 responses; phase error limits exclude intentional shedding.
Localhost targets run without an additional flag. A non-local test environment requires an exact host authorization so a stale plan cannot silently target a different service:
keycast-loadtest capacity \
--plan ./capacity-plan.json \
--allow-host keycast.staging.example \
--output ./capacity-evidenceThe capacity command rejects the production Keycast host even when passed via
--allow-host. Production exercises require a separately approved operational
procedure, monitoring, stop authority, and rollback plan.
evidence.json records the profile platform separately from the derived local
or authorized-non-production execution environment. It also records the plan,
revision identifiers, environment-parity statement, phase summaries, SLO and
threshold checks, the stop point, and an overall result. It does
not contain the credential file's tokens, generated passwords, request bodies,
or raw server errors. When available it preserves before/after server metric
snapshots, but their difference is meaningful only if that endpoint aggregates
all serving instances. The bundle remains rehearsal evidence: dependency
headroom, autoscaling timelines, notification receipts, recovery timing, and
rollback drill records must be gathered by the platform runbook before
certifying a platform.
Cloud Run routes requests away from high-CPU instances even if their concurrency limit isn't reached. Session affinity is broken when an instance hits max CPU—requests go to other instances automatically.
For load testing: Use multiple simulated users with different cookies (this tool does this) to distribute load across instances realistically.
-
Cloud SQL Connection Pool (db-g1-small: ~25 max connections)
- Each Cloud Run instance uses ~2-5 connections
- Under load, instances scale up → connection exhaustion
- Fix: Upgrade DB tier or use Cloud SQL Proxy pooler
-
Cold Instance Caches
- New Cloud Run instances have empty handler caches
- All requests hit DB until cache warms
- Mitigated by: session affinity, min instances
-
Network Latency
- ~200ms baseline to Cloud Run (geographic)
- Not reducible without edge deployment
Keycast Load Test Results
=========================
Target: http://localhost:3000/api/nostr
Scenario: WarmCache
Method: GetPublicKey
Concurrency: 50
Duration: 25.0s
Users: 20
Summary:
Requests: 21933
Throughput: 877.2 req/s
Success Rate: 100.0%
Latency (ms):
Min: 0.8
p50: 44.0
p95: 49.3
p99: 58.4
Max: 178.4
| Variable | Default | Description |
|---|---|---|
HANDLER_CACHE_SIZE |
1,000,000 | Max handlers in cache per instance |
Use the included script to profile CPU usage with flamegraph:
# Basic usage (requires sudo for dtrace on macOS)
sudo ./tools/loadtest/flamegraph.sh
# Custom parameters
sudo ./tools/loadtest/flamegraph.sh \
--users 100 \
--concurrency 100 \
--duration 60 \
--scenario warm-cache \
--method sign-event
# Reuse existing users (faster iteration)
sudo ./tools/loadtest/flamegraph.sh --skip-setup --duration 30Options:
| Option | Default | Description |
|---|---|---|
--users |
50 | Number of test users to create |
--concurrency |
50 | Concurrent requests |
--duration |
30 | Test duration (seconds) |
--scenario |
warm-cache | warm-cache, cold-start, mixed |
--method |
get-public-key | get-public-key, sign-event |
--output |
/tmp | Output directory |
--skip-setup |
false | Reuse existing users file |
Output:
keycast-flamegraph-TIMESTAMP.svg- Interactive flamegraph (open in browser)flamegraph-results-TIMESTAMP.json- Load test resultsflamegraph-users.json- Reusable test users
Prerequisites:
# Install flamegraph
cargo install flamegraph
# Build with debug symbols (done automatically by script)
CARGO_PROFILE_RELEASE_DEBUG=true cargo build --release --bin keycastIMPORTANT: When investigating performance issues, always check these FIRST before building synthetic benchmarks:
# Enable SQLx query logging
RUST_LOG=sqlx=debug ./target/release/keycast
# Or check PostgreSQL directly
psql -c "SELECT query, calls, mean_exec_time FROM pg_stat_statements ORDER BY mean_exec_time DESC LIMIT 10;"What to look for:
- Queries running on every request that should be cached
- INSERT/UPDATE statements in read-heavy paths
- Missing indexes (high
mean_exec_time) - N+1 query patterns
Real example: The TenantExtractor was doing INSERT ... ON CONFLICT on every request instead of caching the tenant lookup. This caused ~160x slowdown (1.8k → 623k req/s after fix).
curl http://localhost:3000/api/metrics | grep -E "cache|rpc|latency"Verify cache hit rates match expectations. 99%+ cache hits with high latency = bottleneck is elsewhere (not cache misses).
# Enable tokio-console (if configured)
RUSTFLAGS="--cfg tokio_unstable" cargo build --release
# Or use tracing to spot blocking operations
RUST_LOG=tokio=trace ./target/release/keycastLook for:
spawn_blockingoverload- Blocking operations on async threads
- Task starvation
Only after ruling out obvious issues:
sudo ./tools/loadtest/flamegraph.sh --duration 30| Pitfall | Symptom | Fix |
|---|---|---|
| DB query per request | High latency, low CPU | Add caching layer |
| Uncached extractor | Middleware runs on every request | Cache in global state |
| Blocking in async | Low throughput, thread starvation | Use spawn_blocking |
| Lock contention | High CPU, low throughput | Use lock-free structures (dashmap, concurrent caches) |
| Serialization overhead | High CPU in serde | Use zero-copy or pre-serialize |
# Run with debug logging
RUST_LOG=debug cargo run -p keycast-loadtest -- setup --url http://localhost:3000 --users 10
# Run tests
cargo test -p keycast-loadtest