Notice when a dependency has stopped answering the client your process uses, and restart instead of hanging.
Zero dependencies. No build step. One module.
npm install @profullstack/watchdogA web process can be wedged and look perfectly healthy. It has happened to us three times on two sites:
-
2026-09-07, Postgres. Every page hung forever. The container was up, the database was healthy, the platform reported the service Online. The process had run about 29 hours and its connection pool had lost every slot. A query queued for a connection that was never going to arrive, and the pool had no queue deadline, so the request never failed and never answered.
-
2026-09-13 and 2026-09-22, Redis. Same shape. The shared ioredis client is built with
maxRetriesPerRequest: nullbecause BullMQ requires it, and that turns a disconnect into a hang rather than a rejection. Code written as "a Redis blip must not take the site down" catches nothing, because there is no error to catch:try { const hit = await redis.get(key); // never returns } catch { // never runs }
Every time, /healthz answered in milliseconds, because a health endpoint that
touches no dependency is measuring the wrong thing. Restarting the datastore did
not help either: the wedge is on the client side of the socket, so only replacing
the process recovered it. Every time, a person had to notice first.
The probe goes through the same client the requests use. During the Postgres outage a freshly opened pool inside that same container ran a real query in 51ms. A watchdog on its own connection would have reported everything fine throughout.
The timeout is a race, not a driver option. The symptom is a promise that never settles, so awaiting the command alone hangs the watchdog in exactly the case it exists for.
Most services want the same two, so there is one call for it:
import { watchDependencies } from '@profullstack/watchdog';
const watchdogs = watchDependencies({
postgres: () => healthcheck(), // resolve truthy, or it counts as a failure
redis: () => connection.ping(), // must answer PONG
});
process.on('SIGTERM', async () => {
watchdogs.stop(); // FIRST, or a clean drain looks like a wedge
await drain();
});That reads DB_WATCHDOG_INTERVAL_MS, DB_WATCHDOG_TIMEOUT_MS,
DB_WATCHDOG_FAILURES and the REDIS_WATCHDOG_* equivalents, all with working
defaults, so a service needs no new variables. Redis is allowed one more failure
than the pool, because a healthy Redis is routinely unreachable for a minute or
two while it reloads its snapshot from disk, and a watchdog that trips during a
normal restart is worse than no watchdog at all.
Pass the probes rather than the clients: this package stays zero-dependency and
never needs to know whether you are on bun:sql, pg or ioredis.
import { startWatchdogs } from '@profullstack/watchdog';
const watchdogs = startWatchdogs([
{
subject: 'the database pool',
probe: async () => {
if (!(await healthcheck())) throw new Error('select 1 did not come back');
},
},
{
subject: 'redis',
// Redis often restarts by reading a snapshot before it accepts anything, so
// it earns a longer rope than the pool does.
failures: 4,
probe: async () => {
if ((await connection.ping()) !== 'PONG') throw new Error('PING did not come back');
},
},
]);
process.on('SIGTERM', async () => {
// First, before you close the clients: a clean drain must not look like a wedge.
watchdogs.stop();
await drain();
});startWatchdog(options) is the single one. startWatchdogs(specs, shared)
applies shared under each spec, so common timings live in one place and a spec
overrides what it needs.
| option | default | what it is |
|---|---|---|
probe |
required | runs a trivial command on the shared client; throw to report trouble |
subject |
'the dependency' |
names the client in the log and the give-up reason |
intervalMs |
30000 |
gap between probes |
timeoutMs |
10000 |
how long one probe may take before it counts as a failure |
failures |
3 |
consecutive failures before giving up |
onGiveUp |
process.exit(1) |
what to do when the client is declared gone |
log |
console |
needs error and warn; either may be missing |
The bar is consecutive failures, so a slow minute never costs a restart. The wedge this watches for does not recover on its own, so its count never clears.
check() runs one probe now and resolves true if the client answered. That is
what the tests drive, and it is useful behind an admin route.
It reads as drastic for a web server. It is the cheapest correct move: the failure is process-local state that no request can repair, a restart demonstrably clears it, and a platform replaces the container in about a minute.
Hanging forever is not the safer option. It is the outage.
Note that a platform health check usually gates a new deploy and is never re-run, so nothing else is going to restart a wedged-but-alive container.
The watchdog above runs on a timer inside your process. If the event loop itself is blocked, a synchronous loop or a regex gone exponential, that timer never fires either. Only something outside the process can see that.
Docker already is that something. A HEALTHCHECK fails, Docker counts the
failures, marks the container unhealthy, and then does nothing. A restart
policy only fires when the process exits, and a wedged process has not exited.
On our production box that was 183 health-checked containers with nothing acting
on the verdict.
watchdog-autoheal reads Docker's verdict and acts on it:
npm install -g @profullstack/watchdog
watchdog-autoheal status # healthy, unhealthy, and which containers have no healthcheck
watchdog-autoheal --dry-run # what one pass would restart
watchdog-autoheal # one pass: restart what Docker marked unhealthyRun it every minute from cron, under flock so passes never overlap:
* * * * * flock -n ~/.local/state/autoheal.lock watchdog-autoheal --notify 'mail -s autoheal you@example.com' >> ~/.local/state/autoheal.log 2>&1or loop it under a service manager with watchdog-autoheal watch --interval 30.
It does two things on purpose:
- It trusts Docker's streak.
unhealthyalready meansretriesconsecutive failures afterstart_period. Adding a streak of its own would only add minutes to every real wedge. - Restarts are budgeted. A container that is unhealthy because its database
is down, or because the image is broken, comes back unhealthy. After
--max-restarts(3) inside--windowminutes (60) it stops restarting that container and reports it once, through--notify, as needing a person.
Opt a container out with the label autoheal=false. status --json and
--json give machine-readable output.
The policy is a library too, so it can sit behind another runtime or an admin route:
import { healOnce, listDockerContainers, restartDockerContainer } from '@profullstack/watchdog/autoheal';
const { state, restarted } = await healOnce({
list: () => listDockerContainers(),
restart: (c) => restartDockerContainer(c),
state: previousState,
});A container with no healthcheck is invisible to it, and status lists those.
The healthcheck worth adding is one request to your own /healthz: if the event
loop is blocked, that request times out, which is exactly the signal.
MIT