Skip to content

Define multipart/form-data - #1922

Open
annevk wants to merge 5 commits into
mainfrom
annevk/multipart/form-data
Open

annevk wants to merge 5 commits into
mainfrom
annevk/multipart/form-data

Conversation

@annevk

@annevk annevk commented Apr 14, 2026 •

Copy link
Copy Markdown
Member

Based on Andreu's work at https://github.com/andreubotella/multipart-form-data with mostly editorial changes and some corrections.

This defines a multipart/form-data serializer, which HTML's form submission will use as well, and a parser for formData(), replacing the rough approximation that was there before. In particular:

  • A FormData body is serialized when it is extracted and kept as a Blob. This gives the body a length, so Content-Length is set, and means a redirect sends the same bytes again, matching the boundary in the Content-Type header, even if the FormData object has since changed.
  • A boundary consists of 27 to 70 ASCII alphanumerics and hyphens and has at least 95 bits of entropy.
  • Names are newline normalized and filenames have lone surrogates replaced. In both, LF, CR, and " are escaped as %0A, %0D, and %22. The parser does not reverse this escaping.
  • The parser is strict: there cannot be a preamble, header lines have to end in CRLF, header names have to be tokens, and header values cannot contain 0x00 or be folded. Transport padding and the epilogue are ignored.
  • How a part's Content-Disposition header is parsed remains an XXX, as that ought to use the same parser as downloads, which is yet to be defined.
  • A file part's type is extracted from its Content-Type header as elsewhere in Fetch, defaulting to text/plain, and its lastModified is 0. Other parts are UTF-8 decoded.

Tests: web-platform-tests/wpt#59216

HTML PR: whatwg/html#12992

Fixes whatwg/html#6424 and fixes whatwg/html#7413.

(See WHATWG Working Mode: Changes for more details.)


Preview | Diff

Comment thread fetch.bs
Comment thread fetch.bs Outdated
Based on Andreu's work at https://github.com/andreubotella/multipart-form-data with mostly editorial changes and some corrections.

WPT coverage at fetch/api/response/response-form-data.html
Parsing:

* Look for CRLF -- boundary in the body loop rather than the bare boundary.
  As written a part whose value merely contained the boundary would remove
  more bytes than the body has. Gecko and WebKit have the same bug.
* Allow transport padding after the close delimiter, per RFC 2046. This
  works in all three implementations.
* Reject an empty boundary parameter, reachable through boundary="".
* Do not reverse the %0A, %0D and %22 escapes. No implementation does, and
  as 0x25 (%) is not escaped they are ambiguous anyway.
* Normalize a part's Content-Type the way the File constructor normalizes
  its type member, instead of only dropping non-ASCII values.
* Leave parsing the Content-Disposition value undefined and marked XXX. The
  strict matching this had is stricter than every implementation, and this
  ought to share a Content-Disposition parser with downloads.

Serializing:

* Set up the stream with byte reading support, as Response's body is a byte
  stream for every other BodyInit.
* Acquire the file reader once instead of on each pull, which would throw as
  the stream is already locked.
* Take the chunk off the list before recursing into the pull algorithm.
* Empty the chunks on cancelation so no further file is opened.
* Compute the length before creating the stream, which consumes the chunks.
@annevk
annevk force-pushed the annevk/multipart/form-data branch from 4af534c to 7c13fd7 Compare September 9, 2026 15:56
This reuses extract a MIME type for a part's Content-Type. A header's value is
now a byte sequence, as a part's can contain 0x00 (NUL).
@annevk
annevk marked this pull request as ready for review September 11, 2026 15:54
Comment thread fetch.bs Outdated
@zcorpan

zcorpan commented Sep 30, 2026

Copy link
Copy Markdown
Member

AI re-review at df46626:

  • fetch.bs:10285-10293: using extract a MIME type for a part's Content-Type matches no browser:

    • Content-Type: text stays text in all three; the spec gives "".
    • text/html;;x="y" stays raw in all three; the spec gives text/html;x=y.
    • text/html, text/plain gives text/html in Chrome and the raw string in Firefox and Safari TP; the spec gives text/plain.

    Is that intended?

  • fetch.bs:10256-10273: nit, body goes from a string to a byte sequence partway through. Maybe use a separate variable.

Possible test additions: a Content-Type with a non-ASCII parameter value (should give ""), and junk after padding on a non-final delimiter (--boundary x\r\n, should fail).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

How to handle lone surrogates in filenames in form submission Fully define multipart/form-data and allow for streaming

4 participants