Skip to content

IDBMirrorVFS: recover from a commit whose IndexedDB transaction aborts - #371

Merged
rhashimoto merged 5 commits into
rhashimoto:masterfrom
lalexdotcom:fix/idb-mirror-commit-abort
Oct 10, 2026
Merged

rhashimoto merged 5 commits into
rhashimoto:masterfrom
lalexdotcom:fix/idb-mirror-commit-abort

Conversation

@lalexdotcom

@lalexdotcom lalexdotcom commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

As offered in #363. This one turned out trickier than I expected 😅

What happens

#commitTx adds a transaction to the connection's view (#acceptTx, #setView) before its IndexedDB transaction commits. When that transaction aborts, on a quota error for instance, nothing undoes it:

  • With synchronous=full, SQLite gets SQLITE_IOERR, but txActive is kept. The connection's next commit stores the failed commit's rows along with its own.
  • With synchronous=normal, the commit has already returned SQLITE_OK when the abort arrives. The connection goes on from a view that was never stored, and its next commit writes only its own pages on top of the stored database. A new connection then fails integrity_check (Page 28: never used), or cannot read the database at all (database disk image is malformed) when the aborted transaction was larger than the page cache.
  • In locking_mode=exclusive, commits already queued behind the aborted one are stored the same way.

The change

When a commit's IndexedDB transaction aborts, the file's AbortController is aborted. File already declared that field, but it was never created. From then on:

  1. Nothing built on the aborted view is stored. #commitTx refuses the transaction with an error, so SQLite gets SQLITE_IOERR. A commit already queued behind the aborted one checks for the abort in its first IndexedDB request and aborts itself.
  2. The view is reloaded from IndexedDB at the next SHARED lock, immediately in the synchronous=full error path, and when #commitTx refuses a transaction, unless SQLite has a rollback journal for it. Loading moves out of jOpen into #loadFile so that all three can use it.
  3. A write transaction that began on the aborted view gets SQLITE_BUSY at RESERVED, so it starts over on the reloaded view.
  4. The journal of an aborted view is removed when the file is closed.

Why this way

Writes never fail; the commit is refused instead. Failing every call after the abort, as OPFSPermutedVFS does, is not enough here. When a batch-atomic write fails with an SQLITE_IOERR code, SQLite retries the commit with a rollback journal (pager.c). That journal stays in the VFS, and the next open plays it back over the stored database: with that approach, reopening after a synchronous=normal abort failed integrity_check every time. Here writes only reach txActive, so they can always succeed. The refusal comes at SQLITE_FCNTL_SYNC, outside the batch, and SQLite writes no journal.

The view is reloaded rather than the connection failed. Once nothing built on the aborted view can be stored, failing the connection until it is reopened would be safe too. But every application would have to reopen after a quota error. Reloading is safe at two points:

  • at SHARED, where SQLite checks its cache against the file's change counter;
  • after a commit fails, because SQLite then discards its cache, in exclusive locking mode too (pager.c). That is why synchronous=full recovers at once, even in exclusive mode, and why exclusive synchronous=normal recovers after the commit that is refused.

There is one exception to the second point. If the refused transaction had spilled to a rollback journal, SQLite rolls it back through that journal after the error. That writes pages of the aborted view over the reloaded one, and the next commit stored them: the aborted rows, or a corrupt database, in every run. So the view is not reloaded when the database has a journal in the VFS.

SQLITE_BUSY at RESERVED. With synchronous=normal, the abort usually becomes known only when the next write transaction takes RESERVED, because reading the tx store waits for the aborted IndexedDB transaction. By then SQLite has already validated its cache against the aborted view, so that transaction cannot go on. The check just above it, for a view that is out of date, already returns SQLITE_BUSY for this reason. SQLite then releases its lock, and with a busy timeout it retries from SHARED (btree.c), which reloads the view. With a busy timeout the application never sees the abort; without one it gets SQLITE_BUSY once. Refusing only at commit time gave SQLITE_IOERR in both cases.

A gate request only while another commit is pending. IDBTransaction.abort() throws once commit() has been called, so the aborted transaction cannot cancel the ones queued after it. Instead, a queued transaction issues a request first. That request runs only after the earlier transactions have finished, so its callback can see the abort and abort its own transaction. The gate is used only when another commit is still pending, which never happens with synchronous=full:

  • on every commit, it cost 15-30 % on synchronous=full commits in exclusive mode;
  • awaiting the previous commit instead made exclusive synchronous=normal commits 1.7 to 2 times slower on Chromium.

The journal is removed on close. In exclusive locking mode with synchronous=normal, a transaction larger than the page cache writes a journal before its commit is refused. Played back at the next open, that journal stores the aborted rows.

What it costs

  • With synchronous=normal, the aborted commit is lost. In exclusive locking mode, so are the commits that returned before the connection learned of the abort; otherwise none can, since the next write transaction waits for the aborted one at RESERVED. The stored database stays as it was before them.
  • In exclusive locking mode with synchronous=normal, the lock is never released, so the view is reloaded only when a commit is refused. That commit fails with SQLITE_IOERR, and until then reads still show the lost commit. If the refused transaction was larger than the page cache, commits fail until the database is reopened.
  • Nothing measurable changes while commits succeed. 300 single-row commits take as long as on master, within run-to-run variation: 9 interleaved runs, Chromium and Firefox, asyncify and jspi, full and normal, normal and exclusive locking mode.

Test

test/vfs_commit_abort.js, with a worker of its own, is wired into test/IDBMirrorVFS.test.js. The worker replaces IDBTransaction.prototype.commit so that, once armed, the next read-write transaction aborts, optionally after being kept alive for 300 ms with requests. The tests:

  • An aborted commit, for each synchronous setting and each locking mode. The aborted rows never reach the store, and the connection's next insert works. In exclusive normal, it comes after the one commit that fails, which must stay invisible. A new connection counts 201 rows and passes integrity_check.
  • A commit queued behind the aborted one (exclusive, normal) is not stored.
  • A journal written on the aborted view is not played back, neither on the reloaded view nor when the database is reopened.

Before closing, the worker lets pending commits finish. Closing first makes the commit's broadcast throw on the closed channel (InvalidStateError). That also happens on master, and this PR does not address it.

yarn web-test-runner test/IDBMirrorVFS.test.js

On master, all six tests fail on asyncify and jspi, for example:

  • Expected 204 to be 201. with synchronous=full;
  • the integrity_check result with synchronous=normal;
  • Expected 11 to be 500. for the queued commit;
  • Expected 3 to be 0. for the journal.

With this change, the file's 190 tests pass, 3 runs of 3, on Chromium and on Firefox (Playwright).

Each part of the change is needed by a test. Removing one at a time:

removed fails
SQLITE_BUSY at RESERVED normal / normal locking: the next insert gets SQLITE_IOERR
reload at SHARED normal / normal locking
reload in the full error path full / exclusive locking: the connection stays failed
reload when a commit is refused normal / exclusive locking: the connection stays failed
skipping that reload while a journal exists the journal test: the aborted rows are stored
refusal in #commitTx exclusive normal, the queued commit and the journal tests
gate request the queued commit test: the update is stored
journal removal on close the journal test: the aborted rows are stored

The whole suite on this branch: 6274 passed, 0 failed.

Checklist

  • I grant to recipients of this Project distribution a perpetual,
    non-exclusive, royalty-free, irrevocable copyright license to reproduce, prepare
    derivative works of, publicly display, sublicense, and distribute this
    Contribution and such derivative works.
  • I certify that I am legally entitled to grant this license, and that this
    Contribution contains no content requiring a license from any third party.

lalexdotcom and others added 3 commits October 3, 2026 16:24
The view already includes a transaction when its IndexedDB commit
aborts. A later commit then carried the failed rows (synchronous=full)
or was stored on top of a state that never was (synchronous=normal),
including commits already queued behind the aborted one.

Such a view is no longer published: the commits built on it are
refused or dropped, the view is reloaded from IndexedDB when SQLite
next validates its cache, and its journal is not played back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
In exclusive locking mode the lock is never released, so a connection
with synchronous=normal stayed failed until reopened after an abort.
SQLite discards its cache after the refused commit's error, so the view
can be reloaded there, except when SQLite rolls the transaction back
through a journal written on the aborted view.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lalexdotcom

Copy link
Copy Markdown
Contributor Author

I added a commit, 3367cb6, which removes one of the costs listed in the description: in exclusive locking mode with synchronous=normal, the connection no longer stays failed until it is reopened.

In that mode the lock is never released, so the view could not be reloaded at SHARED. But when #commitTx refuses a transaction built on the aborted view, SQLite gets SQLITE_IOERR and discards its cache, in exclusive mode too. The view is now reloaded right there, as the synchronous=full error path already does. Only that commit fails; the next one works.

There is one exception: a refused transaction that had spilled to a rollback journal. After the error, SQLite rolls it back through that journal, which writes pages of the aborted view over the reloaded one. The next commit then stored them: the aborted rows, or a corrupt database, in every run. So the reload is skipped while the database has a journal in the VFS, and that case still needs a reopen.

The exclusive normal test now expects the next insert to succeed without a reopen; it fails without this commit. The journal test fails if the reload ignores the journal. With the change, the file's 190 tests pass, 3 runs of 3 on Chromium and Firefox, and the whole suite passes (6274). I updated the description to match.

@rhashimoto rhashimoto left a comment •

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR!

The approach looks fine. It took me a while to figure out the new #commitTx() code so I suggested some reorganization.

The code excerpts below came through in a confusing order. My comments should make more sense if you read them in the file view.

Comment thread src/examples/IDBMirrorVFS.js
Comment thread src/examples/IDBMirrorVFS.js
Comment thread src/examples/IDBMirrorVFS.js Outdated
Comment thread src/examples/IDBMirrorVFS.js Outdated
Comment thread src/examples/IDBMirrorVFS.js Outdated
The commit fence is now awaited inside an async executor instead of
gating a nested write function. Also declare commitsPending with the
other File properties and note a faster way to load the database.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lalexdotcom

Copy link
Copy Markdown
Contributor Author

Thanks for the review! All three are in 874303e: commitsPending is in the property list, the TODO is in #loadFile, and #commitTx is reorganized as you suggested. The file's 190 tests pass, 3 runs of 3 on Chromium and Firefox, and the gate test still fails without the fence.

One thing to be aware of with the async executor: an exception inside it no longer reaches complete. It rejects the executor's own promise instead, which nothing awaits. I checked two cases with a standalone IndexedDB probe, with the same results on Chromium and Firefox:

  • If the transaction aborts while the fence is pending, onabort still rejects complete, but await idbX(fence) also leaves an unhandled rejection (AbortError). That one is only noise.
  • If the write body throws after a put(), it is no longer running inside a request's event handler, so IndexedDB does not abort the transaction. The requests already issued auto-commit and complete resolves. The previous version aborted in that case on the gated path, though not on the other one.

Nothing in that body should throw today, so I left it as you wrote it. If you'd rather close the gap, wrapping the body in a try/catch that calls idbTx.abort() would cover both paths. I'm happy to add it.

@lalexdotcom
lalexdotcom requested a review from rhashimoto October 9, 2026 08:15
@rhashimoto

Copy link
Copy Markdown
Owner

Nothing in that body should throw today, so I left it as you wrote it. If you'd rather close the gap, wrapping the body in a try/catch that calls idbTx.abort() would cover both paths. I'm happy to add it.

You're right, that's totally valid. I was tempted to weasel out of another round of changes, but yes, please wrap in a try/catch.

The write body now runs after an await, outside any request event
handler, so an exception there no longer aborts the IndexedDB
transaction: the requests already issued auto-commit and the commit
resolves. Abort explicitly so a failed write is never stored.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lalexdotcom

Copy link
Copy Markdown
Contributor Author

Done 😉

@rhashimoto

Copy link
Copy Markdown
Owner

Thanks! LGTM.

@rhashimoto
rhashimoto merged commit d09953e into rhashimoto:master Oct 10, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants