Skip to content

Latest commit

 

History

History
294 lines (198 loc) · 121 KB

File metadata and controls

294 lines (198 loc) · 121 KB

xls-codec

GitHub npm npm version CI

Hand-written reader and writer for the legacy Excel Binary File Format (.xls, BIFF8) as specified by [MS-XLS], mapping a workbook's record stream onto the same document-schema.js spreadsheet model ooxml.js's xlsx support and odf.js's ods support target. Worker-isomorphic: the same code runs under Node and inside a Cloudflare Workers isolate.

A .xls file is not one format but two nested ones. The outer container is an [MS-CFB] compound file — the same "filesystem in a file" that carries .doc and .ppt — holding a stream named Workbook. Inside that stream is BIFF8: a flat sequence of records, each a two-byte type, a two-byte size, and that many bytes of data, organised into substreams delimited by BOF/EOF. This package leaves the outer layer to archive-codec's bounded CFB reader and writer and implements the inner one, from the record framing up to a ContentDocument and back.

Status

Under active development, with real, tested read and write support. Built and shipped:

  • Record framing (src/biff/records.ts, src/biff/record-writer.ts) — the three-component record structure of [MS-XLS] 2.1.4 in both directions, with the 8224-byte data ceiling enforced and every malformed or oversized stream thrown on rather than silently truncated or split into a Continue chain the writer does not implement.
  • Continuation-aware cursor and strings (src/biff/cursor.ts, src/biff/strings.ts, src/biff/string-writer.ts) — Continue records ([MS-XLS] 2.4.58) joined per the rules of the record being continued on read, including the case a naive reader gets wrong: an XLUnicodeRichExtendedString ([MS-XLS] 2.5.293) resuming after a boundary re-states its own fHighByte flag, which may differ from the flag the string started with. All three string shapes (XLUnicodeString, ShortXLUnicodeString, XLUnicodeRichExtendedString) are read and written, compressed (one byte per UTF-16 code unit) whenever every character allows it and uncompressed otherwise.
  • Workbook globals, read (src/workbook/globals.ts) and write (src/workbook/globals-writer.ts) — BoundSheet8 (sheet names, tab order, hidden state, type, and substream offsets), SST with its Continue chain on read, Format (custom number-format codes), Font, XF's fixed prefix plus its trailing CellXF/StyleXF fill/border payload in both directions (src/biff/xf-colors.ts's shared bit-layout packing/unpacking; see Cell decoration), Palette in both directions, the fifteen mandatory built-in Style records, and Date1904.
  • Worksheet substreams, read (src/workbook/sheet.ts) and write (src/workbook/sheet-writer.ts) — Dimensions, Row (height and hidden state), ColInfo (width and hidden state), MergeCells, the print-settings record group (see Print settings), and the cell-value family: Number, BoolErr, LabelSst, Blank, and now Formula (see Formula writing) on write (MulBlank, RK, MulRk, and Label stay read-only — see below).
  • Formula writing, same-sheet only (src/biff/ptg-writer.ts) — a ContentSheetCell.formula's own text, compiled back into a Formula record's Ptg token stream: literal operands, same-sheet cell/range references ($-qualified), every arithmetic/comparison/unary/percent operator, explicit parentheses, and a function call resolved by name against ptg-functions.ts's own Ftab table (PtgFunc when its fixed arity matches, PtgFuncVar otherwise) — see Formula writing for the exact scope boundary and why it stops at one sheet.
  • Cell comments (src/workbook/comment-writer.ts) — a ContentSheetCellComment's text and author written back out as the Note/Obj/Txo triple src/workbook/comments.ts already reads, one object id per commented cell; replies and createdAt have nowhere to land, matching the reader's own documented gap (legacy BIFF8 has no threading or per-comment timestamp at all).
  • Number-format classification and date serials (excel-number-format, src/serial.ts) — what turns a bare number into the schema's own percentage/currency/date/time/dateTime value kinds and back, honouring the workbook's own epoch flag (the writer always emits the 1900 system) and refusing the 1900 system's phantom leap day in both directions. The classification itself is a dependency, not local code: this package shares it with ooxml.js's xlsx support, since it is the identical mini-language in both formats (#848). A cell's own numberFormatCode is preserved verbatim on write when present; absent, it resolves to a representative code for its value kind (General for a plain number/string/boolean/error, 0% for a percentage, a [$USD]#,##0.00-shaped bracket format for a currency naming an ISO 4217 code — the one carrier the code survives the round trip through, the identical encoding ooxml.js's own xlsx writer states for the same schema field — a bare $ format for a currency with no code or a symbol that is not an ISO code, mm-dd-yy/h:mm:ss/m/d/yy h:mm for date/time/dateTime), and the workbook-wide Format/XF table is deduplicated across every sheet so two cells sharing one code share one entry.
  • Formula text recovery, including shared, array, and external-reference formulas (src/biff/ptg.ts, src/biff/ptg-functions.ts, src/workbook/sheet.ts's collectFormulaGroups, src/workbook/globals.ts's readSupBook) — a Formula record's compiled rgce token stream ([MS-XLS] 2.5.198's Ptg vocabulary) read and rebuilt into the infix text a spreadsheet application would show: literal operands (PtgInt/PtgNum/PtgStr/PtgBool/PtgErr/PtgMissArg), cell and range references including their 3D (cross-sheet) forms (PtgRef/PtgArea/PtgRef3d/PtgArea3d, $-qualified per their own relative/absolute flags), every arithmetic/comparison/unary/percent operator and explicit parentheses, function calls through both PtgFunc (fixed arity, resolved from a curated table of [MS-XLS]'s own Ftab grammar) and PtgFuncVar (variable arity, its own on-disk cparams), a shared formula (PtgExp joined against the ShrFmla record that follows its group's base cell, its relative PtgRefN/PtgAreaN tokens re-expanded for each referencing cell's own position, alongside any ordinary, non-relative token the same shared expression carries), an array formula (PtgExp joined against an Array record instead, its expression returned with no further wrapping — Excel's own {...} CSE bracing is formula-bar display syntax, never written into the formula itself, matching ooxml.js's own xlsx convention; a PtgArray array-constant literal's {...} is genuine syntax, not this bracing, and is resolved from its own PtgExtraArray trailer wherever one appears, array-entered or not), and a 3D reference's sheet name, resolved through EXTERNSHEET and SupBook for both a self-referencing workbook and a genuinely external one (its own file name and sheet name(s) recovered as far as SupBook's virtPath/rgst fields allow); a DDE/OLE/add-in/same-sheet/unused link, an unresolved sheet index, or an undecoded virtPath form has no real name to recover, and leaves the whole containing formula unresolved exactly like any other unsupported construct, rather than writing a fabricated placeholder into what would otherwise be real formula text (src/workbook/globals.ts's own sheetRanges) — see "Formula expressions" under Read-side gaps below for the exact boundary of what still resolves to nothing at all.
  • Schema mapping — readXlsContent/readXls (src/content.ts) as before, now also populating ContentSheetCell.formula wherever the Ptg reader above resolves it; writeXlsContent/writeXls (src/write.ts) the counterpart, taking a ContentDocument/DocumentTree of kind: 'spreadsheet' and producing genuine .xls bytes: a real BIFF8 Workbook stream (globals substream, one worksheet substream per sheet, BoundSheet8.lbPlyPos patched to each sheet's real byte offset once every substream's length is known) wrapped in a real [MS-CFB] compound file via archive-codec's writeCompoundFile.
  • Document metadata — title/subject/author/keywords/createdIso/modifiedIso read from a "\x05SummaryInformation" stream when one is present, and written back to one whenever the input's metadata carries anything that stream can hold (see Metadata).
  • Cell decoration — a cell's background fill (every named FillPattern value, solid and pattern alike) and per-side borders, read from and written to XF's trailing CellXF payload plus the workbook's own Palette record, in both directions and verified against real LibreOffice-produced BIFF8, not just this package's own round trip (see Cell decoration).
  • Print settings — every field of ContentSheetPrintSettings: page size and orientation, all four margins, gridline and row/column-header printing, page order, print scale or fit-to-page counts, manual page breaks, the print range, and the repeated header rows and columns — read from and written to the nine worksheet-substream records and the two built-in defined names that carry them, in both directions and verified against real LibreOffice-produced BIFF8 (see Print settings).
  • Cell alignment — a cell's own horizontal (left/center/right/justify) and vertical (top/middle) alignment, read from and written to XF's trailing CellXF/StyleXF payload's own leading word, in both directions and verified against real LibreOffice-produced BIFF8 (see Cell alignment).
  • Per-cell fonts — a cell's own font (ContentSheetCell.font: bold/italic/underline/strike, family, size, colour), read from the Font record its XF's own ifnt indexes into the workbook's font table and written back into an interned font table the writer builds from the cells it writes (biff/font.ts holds both directions of the one record layout). A cell's font is derived by diffing that entry against the table's own first font — the Normal style's, which is what "the format's default" concretely means for a given file — so only properties the cell genuinely differs in are stated, the same default-omission policy alignment and fill already apply.
  • Defined names — a workbook's own defined names, read from and written to the Lbl records ([MS-XLS] 2.4.150) of the globals substream (workbook/defined-names.ts): the name as spelled (a built-in under its _xlnm. spelling), the refersTo rebuilt from the record's compiled Ptg stream by the same parser a cell formula uses, and a sheet-local scope carried as scopeSheetIndex translated into the document's own sheets-array position. The two print built-ins stay where they already live, on print settings.
  • Data validation, read-only (#1098) — a sheet's own Dv records (src/workbook/data-validation.ts) into ContentSheet.dataValidations: type, comparison operator, one or two formulas (via the same Ptg reader Formula records already use), allow-blank/show-input/show-error flags, error style, and prompt/error titles and messages, resolved directly against [MS-XLS]'s own published Dv/DVal/DVParsedFormula/SqRefU field layouts rather than a producer-specific convention.
  • Conditional formatting, the CF12-era spellings, written (#1186) — every rule variant the schema models beyond cellIs writes as one CondFmt12 plus one CF12 record (workbook/conditional-format-write.ts): the text family as genuine ct 0x02 formula conditions in Excel's own generated shape (the search text as the formula's PtgStr operand, the identical spelling the reader's readCfTextFilterRule recovers it from), the operand-free family, top10, aboveAverage, timePeriod, and the three visual-scale rules (colorScale/dataBar/iconSet) through their CFGradient/CFDatabar/CFMultistate payloads with CFColor RGB triples and CFVO thresholds compiled through the same compileFormulaText cell formulas use. A cellIs rule still writes as a base CondFmt/CF pair; the two families share one sheet, base groups first. ipriority is minted (smallest unused positive integer) when a rule states no priority, since [MS-XLS] 2.4.43 requires it present and unique. What genuinely has no BIFF8 spelling throws a BiffWriteError naming it — see the Writer scope table below.
  • Charts, drawings, and images, read; images and non-chart embedded objects, now also written (#924, #1186) — a worksheet's own MS-ODRAW (Escher) drawing layer (src/drawing/escher.ts, src/drawing/shapes.ts, src/drawing/blips.ts), the workbook-wide Blip Store every picture shape's own pib property indexes into, and an embedded chart's own nested substream (src/workbook/chart.ts), joined into ContentSheet.images and ContentSheet.embeddedObjects by src/workbook/drawing.ts — see "Charts, drawings, and images, resolved for real" under Read-side gaps below. The write direction (src/drawing/escher-writer.ts, src/workbook/drawing-writer.ts) builds the identical Escher container tree back out — the workbook-wide Blip Store deduplicated by image bytes, a per-sheet shape tree, and the Obj records pairing each shape with what it holds — and a non-chart embedded object's own document rides inside a genuine [MS-XLS] Embedding Storage (src/workbook/embedded-object.ts); see Images and embedded objects, written.

Verified primarily by round trip (src/write.test.ts, plus a dedicated test/workers/write.test.ts proving the whole write path inside a real workerd isolate, not just Node): build a ContentDocument, write it, read it back through this package's own independently-pinned reader, and check the result. Every record's own byte layout is additionally cited to its [MS-XLS] section in the writer's source, matching the reader's own convention.

Writer scope

What writeXlsContent/writeXls cover: every ContentCellValue kind a real .xls can hold (number, percentage, currency, boolean, date, time, dateTime, string, error; an empty cell is written as a Blank record when it carries formatting and as nothing at all when it does not — see below), merged ranges (colSpan/rowSpan), row heights and hidden rows, column widths and hidden columns, multiple sheets, explicit and default number formats, a shared string table deduplicated across the whole workbook, every field of a sheet's own print settings, a cell's own alignment, a cell's own background fill and per-side borders (see Cell decoration), a same-sheet formula (see Formula writing), a cell's own comment text and author (see Cell comments, written), a sheet's own images and non-chart embedded objects (see Images and embedded objects, written), a sheet's own data validations (one Dval wrapper plus one Dv record per rule, every schema field through the same compileFormulaText formula compiler cell formulas use), and a sheet's own conditional formats in full: a cellIs rule as one CondFmt plus one CF record (its style's text colour and solid background encoded as a DXFN through the workbook's own palette), every other variant the schema models as one CondFmt12 plus one CF12 record (the CF12-era spelling, its thresholds, template parameters, and CFColor triples written as the exact inverse of the reader's own walk). What it deliberately does not:

Not written Why
Formula records for a 3D (cross-sheet or external-workbook) reference, an array-constant literal or CSE array formula, a defined name, or a call to a function outside Ftab's own vocabulary Same-sheet formulas write (see Formula writing); these four constructs each throw a BiffWriteError naming the construct instead, rather than emit a token stream the writer cannot prove round-trips. A 3D reference needs a SupBook/ExternSheet pair this writer only ever mints today for the two built-in print-settings names (see Print settings), not for an arbitrary formula; an array constant/CSE formula needs a PtgExtraArray/Array-record trailer this writer does not build; a defined name has nowhere to resolve against, the same gap the Lbl row below describes; and Excel 2007+ added many worksheet functions BIFF8's own Ftab enumeration never named, resolved through a PtgNameX/add-in mechanism this writer does not implement.
Per-cell font The reader does not read one back: ContentSheetCell has no cell-level font field at all (ooxml.js's xlsx reader makes the identical font-scope choice), so writing a real value here would be unverifiable by round trip. Every XF this writer emits still references the same single font.
MulBlank/RK/MulRk Pure compaction optimisations over information a plain Blank/Number/LabelSst/BoolErr record already carries losslessly. (Blank itself is written, for a formatted empty cell — see the empty row below.)
An empty-kind cell carrying no formatting Written as nothing at all, which is what round-trips: content.ts's reader drops an unformatted blank cell it reads, and a merged range's empty anchor is independently reconstructed from MergeCells alone. A formatted one (a fill, a border, or a non-default alignment) is the opposite case — its formatting exists only in the XF a cell record names, so it gets a real Blank record ([MS-XLS] 2.4.20) and round-trips with that formatting intact.
A 'chart'-kind embedded object Writing one means embedding a genuine BIFF8 chart substream — the whole [MS-XLS] chart grammar its series data links drive — which is a chart engine of its own rather than a container for the flattened series table the schema carries; writeXlsContent throws BiffWriteError naming it rather than approximating a chart no BIFF8 record could actually redraw. A sheet's own images and every other embedded-object kind (wordprocessing/presentation/spreadsheet/drawing/formula) do write — see Images and embedded objects, written.
Defined names (Lbl) Read only for the two built-in ones a sheet's print range and repeated header bands live in (Print settings); a user-defined name has nowhere to land in document-schema.js's spreadsheet model, so there is no round trip to verify a writer for it against.
A data validation whose formula text compileFormulaText cannot encode Every other rule shape writes (see the lead paragraph); a formula naming a construct outside the same-sheet Ptg vocabulary the cell-formula writer supports throws the identical BiffWriteError that cell formula already does, rather than emit a token stream this writer cannot prove round-trips. A comparison-typed rule carrying no operator, or a two-operand operator (between/notBetween) carrying fewer or more than its two formulas, is a malformed model and throws naming it.
The conditional-format shapes the CF12 record vocabulary itself cannot state Every rule variant the schema models now writes (workbook/conditional-format-write.ts): a cellIs rule as one CondFmt plus one CF record — its operator, both formula operands through compileFormulaText, and its style's text colour and solid background as a DXFN resolved through the workbook's own palette, the write-side inverse of readCondFmtGroup/readCf/parseDxfStyle — and every other variant as one CondFmt12 plus one CF12 record, the CF12-era spelling (ct 0x02 formula condition carrying the search text for the text family, ct 0x05 filter templates for the operand-free/top10/aboveAverage/timePeriod family, CFGradient/CFDatabar/CFMultistate for the visual-scale rules), with an absent priority minted the smallest unused ipriority because [MS-XLS] 2.4.43 requires the field present and unique. What the record vocabulary genuinely diverges from the schema on throws a BiffWriteError naming it rather than being silently dropped or approximated: a timePeriod naming thisYear/lastYear/nextYear (ODF's own calcext extension values — no icfTemplate names a year-scoped period); an iconSetType outside [MS-XLS] 2.5.36's seventeen built-in sets (a producer-extensible vocabulary with no iIconSet byte); an icon-set threshold count disagreeing with the set the rule names (cStates is pinned to the set's own icon count); stopIfTrue on a colour-scale/data-bar/icon-set rule (2.4.43 pins fStopIfTrue to zero for the visual-scale types); an aboveAverage standard-deviation count above 2 (2.5.23's iParam table admits 0/1/2); and two rules declaring the same priority (contradictory ordering rather than something to renumber). A cellIs rule's own priority/stopIfTrue/source still have no base-CF field and are not written; the CFEx compatibility spelling — a legacy CF-plus-extension pair keeping a pre-2007 Excel able to evaluate a rule — is not written either, the CF12 spelling being what this package's own reader resolves either way.
A comment's replies or createdAt A cell's text and author write (see the Status list above); legacy BIFF8's Note/Obj/Txo triple has no threading and no per-comment timestamp at all, the same gap src/workbook/comments.ts's own read side already documents, so there is nothing for these two fields to round-trip against.
A page size no iPaperSize code names Written as [MS-XLS] 2.4.257's own custom-paper value rather than as a named paper it is not — the dimensions themselves are unwritable, since Setup addresses paper only by code. See Print settings.
Window1/Window2, CodePage, Index/DBCell, the legacy interface records (InterfaceHdr, WriteAccess, …) UI and interoperability bookkeeping [MS-XLS]'s own grammar names in the globals/worksheet substreams alongside the content-carrying records above, not data. Index/DBCell specifically is a pure cell-lookup performance optimisation (see [MS-XLS]'s own "Retrieval of Last-Calculated Cell Values Without Loading Cell Table") that this reader — and Excel's own reader — does not require to find a cell; real, well-established minimal BIFF8 writers (e.g. Python's xlwt) omit the same set and produce files Excel opens correctly. The calculation-state records (CalcCount, CalcRefMode, CalcIter, CalcDelta, CalcSaveRecalc) sat in this row until print settings needed them — see Print settings for why the writer emits them now.
Continue-chain splitting A record whose data would exceed the 8224-byte single-record ceiling ([MS-XLS] 2.1.4) — an extremely long shared string, an enormous shared string table, or thousands of merged ranges in one sheet — is refused with a thrown BiffWriteError rather than silently split across Continue records.

Column widths round-trip to the nearest pixel Excel's own integer-pixel-grid quantization allows (matching the read direction's own "honestly approximate" contract, units.ts), never narrower than requested. A .xls cell outside BIFF8's own grid (65536 rows, 256 columns) is refused rather than silently wrapped or truncated.

Read-side gaps

Each deliberate rather than overlooked:

  • Formula expressions, including shared, array, and external-reference formulas. A Formula record's compiled Ptg token stream (src/biff/ptg.ts) is walked and rebuilt into real formula text — literal operands, cell/range references ($-qualified, including 3D cross-sheet references), every arithmetic/comparison/unary/percent operator, explicit parentheses, and both fixed- and variable-arity function calls, resolved by name against [MS-XLS]'s own built-in function table (src/biff/ptg-functions.ts, covering the whole published table — [MS-XLS] 2.5.198.17 — cited to that table's own iftab index; PtgFunc's fixed argument count is a curated subset of it, since PtgFunc's own token carries no count and only a function [MS-XLS]'s grammar states a fixed, non-optional arity for is resolved through it, empirically confirmed against real LibreOffice-written BIFF8 rather than assumed from the grammar alone). Three constructs the read side used to leave ContentSheetCell.formula silently absent for are now resolved: a shared formula — a PtgExp token is joined (src/workbook/sheet.ts's collectFormulaGroups) against the ShrFmla record that follows its group's base Formula record, and the shared expression's own relative tokens (PtgRefN/PtgAreaN, [MS-XLS] 2.5.198.88/2.5.198.31) are re-expanded for each referencing cell's own row/column, including the format's own sheet-edge wraparound rule, while any ordinary (non-relative) token the same expression carries is resolved once and reused unchanged for every referencing cell; an array (CSE) formula — the same PtgExp join against an Array record instead, its expression returned exactly as parsed, with no formula-bar bracing added (Excel's own {...} CSE braces are display syntax the formula-bar shows for an array-entered cell, never characters written into the formula itself — matching ooxml.js's own xlsx convention for the identical construct); and a PtgArray array-constant literal ({1,2,3}-style, whether inside an array formula or an entirely ordinary one like =SUM({1,2,3})) — this bracing genuinely is real, retypeable formula syntax, unlike CSE bracing, and its values are read from the token's own PtgExtraArray trailer rather than left unresolved. A 3D reference's sheet name resolves through EXTERNSHEET and SupBook for both a self-referencing workbook and a genuinely external one — SupBook's virtPath (the external workbook's own path, decoded for the common relative-path forms: same-drive, startup, alternate-startup, and library-relative) and rgst (its sheet names) recover a real [Workbook.xlsx]Sheet1-style label wherever they can, and that label is trusted as real formula text only when both halves were genuinely recovered from the file. An absolute drive volume, a UNC share, or a transfer-protocol URL in virtPath; a virtPath whose final file-path segment carries a bracket character anywhere in it, not only one already shaped like the bracketed sheet-name form ([Book.xlsx]Sheet1) — a bracket-named file is legal per [MS-XLS]'s own grammar, and a bracket's position within that final segment does not settle whether it is genuinely that form: the grammar's own bracketed alternative opens with a literal [ before any directory at all, so a workbook sitting in a subdirectory ([sub + a directory separator + Book.xlsx]Sheet1) puts its closing bracket, non-leading, in the very same final segment this reader isolates once split on the separator — a real instance of the bracketed form just as much as the leading-bracket case ([Book.xlsx]Sheet1 with no directory at all) is; a bracket elsewhere in an otherwise plain, single-directory file name (My[Draft].xlsx) is the case this reader can actually rule the bracketed form out for, since with no directory separator to split on, that whole segment's own leading character settles it — but returning it as a plain file name would still collide with this reader's own [${fileName}] wrapping of the resolved label, producing a raw, unbalanced bracket inside what should read as one clean [Book.xlsx]-style pair — so the whole segment is declined regardless of where the bracket falls or whether this reader could rule the bracketed form out, trading a known false negative for never emitting a mangled label; a SupBook that is a DDE/OLE data source, an add-in, a same-sheet, or an unused link in the first place; or a sheet index [MS-XLS] itself marks unresolvable or workbook-level — none of these has a real name to recover, and each leaves the WHOLE containing formula unresolved, exactly like any other construct this reader cannot turn into real formula text, rather than writing a fabricated placeholder into what a spreadsheet application would otherwise treat as live formula content. What remains genuinely unresolved for the same reason, still leaving formula absent for that cell: a PtgTbl data table, a defined name (PtgName/PtgNameX, for the same reason Lbl is not read at all yet — see below), and a natural-language "Elf" reference.
  • Cell decoration and alignment resolved for real; font still not. XF's trailing CellXF payload's fill pattern/colours and per-side border style/colour are read and resolved through the Palette record (or the default colour table when one is absent) — see Cell decoration for the full mapping, including every named pattern beyond solid, and how this was verified against real LibreOffice-produced BIFF8. The same payload's leading word is resolved too — see Cell alignment for the full alc/alcV mapping. Font records are not read at all: ContentSheetCell has no cell-level font field, and ooxml.js's xlsx reader makes the identical scope choice.
  • Print settings, resolved for real. Every field of ContentSheetPrintSettings is read from the records that carry it, with Excel's own "Normal" preset standing in per field for whatever the file leaves unstated — see Print settings for the record map, the two of BIFF8's own conditional rules that decide how to read Setup, and the three things that genuinely do not come through.
  • Cell comments, resolved for real (#949). A legacy BIFF8 comment is split across three record kinds a single self-contained xlsx <comment> element never needs: a Note record ([MS-XLS] 2.4.179, wrapping a NoteSh structure) anchors the comment to a cell and names its author, but carries no text of its own; the text lives in a Txo record ([MS-XLS] 2.4.329), its characters carried across trailing Continue records; the two are joined through an Obj record ([MS-XLS] 2.4.181) whose FtCmo names an object id and type, the Note's own idObj field naming that same id, and a Txo always immediately following the Obj record whose shape it belongs to (src/workbook/comments.ts). Read in one independent pass over the whole worksheet substream rather than assuming stream order between a Note and the Obj+Txo pair it names — both are resolved regardless of which comes first. Legacy BIFF8 has no threading and no per-comment timestamp at all (unlike xlsx's own [MS-XLSX] threaded-comments extension), so ContentSheetCellComment.replies and .createdAt are never populated from an .xls source; rich per-character formatting within a comment (TxORuns) is skipped rather than modelled, the same scope limit this reader already applies to the shared string table's own rich runs, since ContentSheetCellComment.text is a plain string with nowhere to carry them.
  • Data validation, resolved for real (#1098). Dv ([MS-XLS] 2.4.95) is a fixed, published binary structure — a bit-packed flags DWORD (valType/errStyle/fAllowBlank/fShowInputMsg/fShowErrorMsg/typOperator), four XLUnicodeStrings (prompt/error titles and messages), one or two DVParsedFormula structures, and a trailing SqRefU range list — read directly against that spec (src/workbook/data-validation.ts), unlike calcext:condition's own LibreOffice-source-transcribed mini-language odf.js's equivalent reader needed. A DVParsedFormula's own rgce is the identical Ptg token grammar a cell's Formula record carries, so its formula text is recovered with the same parseFormulaText this reader already uses for ordinary cell formulas. DVal ([MS-XLS] 2.4.96), the record that precedes a sheet's own Dv collection, carries only input-window UI state (position, drop-down Obj reference) with no counterpart in ContentSheetDataValidationSchema and is not read for its own fields.
  • Charts, drawings, and images, resolved for real (#924). A worksheet's own MsoDrawing records ([MS-XLS] 2.1.7.20.3-adjacent, RECORD_MSODRAWING) carry MS-ODRAW (Escher) bytes that split across as many records as the drawing needs and simply concatenate, in stream order, into one continuous Escher record tree (src/drawing/escher.ts, reading [MS-ODRAW]'s own OfficeArtRecordHeader/container-vs-atom framing directly against its spec rather than from memory). Within that tree, src/drawing/shapes.ts walks the DgContainer/SpgrContainer/SpContainer hierarchy for each real top-level shape's own type, id, and cell anchor (ClientAnchorSheet, [MS-XLS] 2.5.163 — one Sp atom's recInstance and ClientAnchor's corner-cell-plus-fractional-offset pair, 1/1024ths of a cell's width and 1/256ths of its height), skipping the invisible "patriarch" group root every drawing wraps its real shapes in. A picture shape's own Opt property table names a pib — a 1-based index into the workbook-wide Blip Store (OfficeArtBstoreContainer, read once from the globals substream's own MsoDrawingGroup stream, src/drawing/blips.ts) — resolved to real image bytes for a PNG or JPEG blip (OfficeArtBlipPNG/OfficeArtBlipJPEG's own literal file bytes, past their rgbUid header); a DIB/EMF/WMF/PICT/TIFF blip is recognised but produces no image, since ContentImageBlockSchema.format has no lossless slot for any of them. Since neither an Escher shape nor a MsoDrawing record names what kind of object it actually holds, each shape is paired 1:1, in document order, with the worksheet substream's own Obj records (skipping any Note-type Obj src/workbook/comments.ts already accounts for) — the same positional correlation every other real BIFF8 reader relies on. An Obj naming a chart locates the chart's own nested BOF(dt=0x0020)/EOF substream — which sits INSIDE the worksheet substream, immediately after the shape that anchors it ([MS-XLS] "Chart Area") — via src/biff/substreams.ts's own stack-based substream nesting (a genuinely new capability this issue needed: splitSubstreams previously treated any BOF met mid-substream as ending it outright, which would have silently dropped every worksheet record written after an embedded chart). src/workbook/chart.ts reads that substream's Series/AI (BRAI)/SeriesText records — a series' own name and its category/value data links — resolving each link's value either from the chart's own on-disk SERIESDATA cache (Dimensions/SIIndex/Number/BoolErr/Blank/Label, present when the chart's data lives on another sheet or is genuinely external) or, for the common case of a chart plotting its own sheet's cells, by resolving the AI's PtgArea3d/PtgRef3d reference token directly against that sheet's own already-mapped cells — and lands as a ContentEmbeddedObject with objectKind: 'chart', its cached series/category table flattened into a small one-sheet spreadsheet ContentDocument, the identical shape ooxml.js's own xlsx/pptx chart reading already produces (#719): document-schema.js has no richer, chart-type-aware object model yet, so this reader targets the fallback that exists rather than inventing one. Every other recognised shape (an autoshape, a text box, a line, a group) lands as objectKind: 'drawing', a single-page ContentDocument holding one sized-and-positioned ContentShape with no text content modelled. Not modelled at all: a shape's own rotation, fill/line formatting, or text runs; a group's own nested transform (a nested nested group's shapes are still recovered, cell-anchored, but a group's own rotation/scale is not composed into its children's positions); and a chart's own type, axes, and legend, which document-schema.js's flattened-table representation has no field for regardless of what this reader could recover.
  • Conditional formatting, base rules resolved for real (#1102). A CondFmt record ([MS-XLS] 2.4.56) marks the start of one to three CF records ([MS-XLS] 2.4.42) sharing one cell-range list, read together as a group (src/workbook/conditional-format.ts) since a CF's own ranges live on its parent CondFmt, not on itself. Base BIFF8 (Excel 97) conditional formatting has exactly two rule shapes: a comparison ("Cell Value Is") condition, promoted to ContentSheetConditionalFormat's cellIs variant with its formula(s) recovered through the same parseFormulaText this reader already uses for cell formulas and Dv; and a formula condition, left unpromoted for the same "no closed-form structure without a general formula engine" reason data-validation.ts and ooxml.js's own xlsx cfRule reading already draw. A rule's resulting style — font colour and fill background — comes from its CF's own DXFN structure, resolved through the identical icv/Palette colour machinery a regular cell's fill and border already use (Cell decoration); a DXFN naming a user-defined number format (DXFNumUsr) degrades that one rule's style to absent rather than risk misreading the font/fill fields that follow it, since [MS-XLS]'s own prose for DXFNumUsr's length field does not settle whether it counts itself. Every richer rule type Excel 2007+ added has no representation in the base CF record at all; it rides a CF12 record instead, or, for a rule Excel keeps expressible as a legacy formula condition for pre-2007 readers, a CFEx extension record that attaches richer metadata to an ordinary CondFmt-owned CF without promoting it into a CF12 record at all.
  • Conditional formatting, CF12's colour scale/data bar/icon set rules resolved for real (#1104). CF12's own ct field picks one of six rule shapes ([MS-XLS] 2.4.43); this reader promotes the three that share a genuine array-of-thresholds building block already modelled in document-schema.js — colour scale, data bar, and icon set (src/workbook/conditional-format-12.ts) — each reading an array of CFVO threshold objects (the identical num/percent/max/min/formula/percentile vocabulary ContentSheetConditionalFormatValueSchema already carries for ooxml.js's xlsx reader and odf.js's ods reader) paired with CFColor values. A CFColor naming an indexed colour resolves through the same icv/Palette machinery as base CF's own DXFN, then Excel's own TintAndShade model applies the colour's own tint/shade Xnum (xf-colors.ts's applyTint); one naming a plain RGB triple resolves directly, tint included; one naming an automatic or theme colour degrades the whole rule to absent, since this package has no BIFF8 Theme reader to resolve a theme reference against. Icon set's own iIconSet byte maps to the identical ECMA-376 ST_IconSetType vocabulary (3Arrows, 3TrafficLights1, 5Quarters, and the rest of the seventeen built-in sets) ooxml.js's xlsx cfRule reading already carries as a plain string. A record longer than the 8224-byte single-record ceiling continues onto ContinueFrt12 rather than the plain Continue every other record in this package joins against; groupRecords (src/biff/substreams.ts) strips its own restated FrtRefHeader before joining it.
  • Conditional formatting, CF12's filter-dispatched template rules resolved for real (#1106). ct 0x05 ("filter") dispatches further through icfTemplate, alongside a 16-byte CFExTemplateParams block always present regardless of ct. Of that block's five variants, only two need real parsing — CFExFilterParams (top10's own fTop/fPercent/iParam) and CFExAveragesTemplateParams (the aboveAverage family's own standard-deviation count) — since CFExDefaultTemplateParams (duplicateValues/uniqueValues/the four blank/error conditions) is 16 reserved bytes and CFExDateTemplateParams's own dateOp field is a fixed restatement of icfTemplate for all ten date/time periods, so both dispatch directly off icfTemplate with no further byte reading. Unlike colour scale/data bar/icon set, [MS-XLS] does not force a ct 0x05 rule's own cbDxf to zero, so its DXFN12 can carry a genuine font/fill override; that structure resolves through the identical parseDxfStyle base CF's own DXFN already uses. A malformed iParam of zero degrades a top10 rule to absent rather than promoting a rank ContentSheetConditionalFormatSchema itself requires to be positive.
  • Conditional formatting, containsText/notContainsText/beginsWith/endsWith and CFEx's own legacy-CF extension resolved for real (#1100). icfTemplate 0x0008 ("Contains text") is the one filter-dispatched template that is NOT actually a ct 0x05 rule: CFExTextTemplateParams carries only ctp, naming which of the four text sub-types a rule is, and neither it nor ct 0x05's own CFFilter rgbCT has anywhere to carry the literal search text (confirmed against a second, independent transcription of both structures — kinkou/unxls's own cfextexttemplateparams/cffilter readers, which read CFFilter as an unconditional fixed six bytes with no variable trailer). Excel instead expresses these four rule kinds as a genuine ct 0x02 formula condition — either directly on a CF12 record, or (to stay evaluable by pre-2007 Excel, which silently skips an unrecognised CF12/CondFmt12 "future record" wholesale) via a CFEx record ([MS-XLS] 2.4.63) extending an ordinary legacy CF — with the literal text present as a string-constant (PtgStr) operand somewhere in the formula's own token stream, matching Excel's real, independently-confirmed generated formula for each sub-type (NOT(ISERROR(SEARCH("text",cell))) for containsText, and the ISERROR/LEFT+LEN/RIGHT+LEN equivalents for the other three — LibreOffice's own xecontent.cxx GetFixedFormula emits exactly these shapes for its xlsx compatibility formula). readCfTextFilterRule (src/workbook/conditional-format-12.ts) extracts that operand via ptg.ts's extractFirstStringLiteral, reached from both readCf12's own ct 0x02 branch and readCfEx (src/workbook/conditional-format-ex.ts). A CFEx record with fIsCF12 set (extending a genuine CF12 record rather than a legacy CF) stays unread: that CF12 carries no ranges of its own — a CondFmt12 normally supplies them — and this reader has no established link back to one for it. CFEx's own nID cross-references the nID field [MS-XLS] 2.1.7.20.6's own worksheet-substream grammar places on every CondFmt record (every CFEx on a sheet follows every CondFmt/CondFmt12 group, per that same grammar), so workbook/sheet.ts keeps a per-nID map of each group's own resolved ranges and raw CF operands as it walks the substream, for a later CFEx to resolve its own icf index against.
  • Not read at all: a CFEx record that extends a genuine CF12 record rather than a legacy CF (see above). Defined names (Lbl) are read only for the two built-in ones a sheet's print range and repeated header bands live in (Print settings); a user-defined name has nowhere to land in document-schema.js's spreadsheet model, so it is skipped.
  • RC4-encrypted and XOR-obfuscated workbooks, decrypted for real (#1108, #922). A FilePass record ([MS-XLS] 2.4.117) means every record after it is ciphertext under one of three schemes; this reader implements two of them, dispatching on the record's own wEncryptionType — the [MS-OFFCRYPTO] 2.3.6.1 "RC4 encryption header" and 2.3.7's own XOR obfuscation (Method 1) — readXlsContent/readXls take an optional password, verify it against the header's own fields (RC4's EncryptedVerifier/EncryptedVerifierHash, decrypted then compared; XOR obfuscation's key/verificationBytes, plain unencrypted checksums compared directly), then decrypt every subsequent record's data (src/workbook/encryption.ts) using the RC4/MD5 and XOR-obfuscation primitives archive-codec shares with doc-codec. XOR obfuscation's own per-record XorArrayIndex — (streamOffset + recordDataLength) % 16, recordDataLength always the record's own full declared size regardless of how much of a given span is actually being decrypted — is confirmed against a genuine Excel-generated fixture, not just the published spec text (see archive-codec's own account of why that spec text alone is not trustworthy here). [MS-XLS] 2.2.10's own exclusion list — BOF, FilePass, and four shared-workbook revision-tracking records this package otherwise never reads, plus BoundSheet8's own lbPlyPos field specifically — is honoured exactly for both schemes, re-derived independently in this package's own tests rather than trusted by construction. A missing or incorrect password throws rather than returning a garbled document; the newer "RC4 CryptoAPI encryption header" scheme remains explicitly out of scope (ppt-codec's own encryption, tracked separately on #1116).

This package is wired into documents.js's conversion registry (xlsToPdf/pdfToXls, convertDocument("xls", ...), and every same-variant spreadsheet bridge) — see that package's own README Fidelity table for exactly which pairs route and which don't. Per-cell font, defined names, and a foreign (non-ISO-code) currency symbol are the remaining read+write scope gaps that are genuinely permanent rather than pending (see the Writer scope table below); formula writing, cell comments, data validation writing, conditional-format writing in full (base cellIs and the CF12-era variants alike), and a sheet's own images and non-chart embedded objects are all covered, each within the boundary the sections and the Writer scope table describe. A 'chart'-kind embedded object is the one drawing-layer construct that still throws rather than writes — see Images and embedded objects, written.

Cell decoration

A cell's own background fill — solid, or a genuine two-colour pattern — and per-side borders are read from and written to XF's trailing CellXF/StyleXF payload ([MS-XLS] 2.4.353) and the workbook's own Palette record ([MS-XLS] 2.4.188), verified both by round trip and against a real, independent BIFF8 implementation — ExaDev/documents.js#815 as a scoped chunk of that issue's own broader tracking, not a claim of closing it outright. src/biff/xf-colors.ts is the one place the payload's border/fill bit layout is packed or unpacked, shared by workbook/globals.ts's read side and biff/xf-writer.ts's write side, so the two directions cannot silently disagree on what a given byte means.

Colour resolution. A fill or border colour is a 7-bit icv index into BIFF8's own colour table ([MS-XLS] "Icv"): 0-7 name eight fixed built-in colours (this package's writer never emits one of these, per the spec's own "SHOULD NOT be ≤ 0x07"; the reader still resolves them, for a real third-party file that does), 8-63 index into either the workbook's own Palette record when one is present or a fixed 56-entry default table when it is not. The writer scans every distinct decoration colour a workbook's cells use before writing anything: when every one already matches the default table exactly, no Palette record is written at all, keeping an undecorated-adjacent file as minimal as it always was; the moment even one colour falls outside that table, a real 56-entry Palette record is minted, with every distinct colour the workbook actually uses (not only the non-default ones) assigned its own dedicated slot, so the whole table is self-consistent rather than a mix of "the file's own entries" and "the implicit default".

Fill patterns beyond solid, resolved for real (#951). [MS-XLS]/[MS-XLSB]'s own FillPattern enumeration names nineteen values: no fill, solid, and seventeen genuine two-colour patterns — five grey shades (FLSMEDGRAY/FLSDKGRAY/FLSLTGRAY/FLSGRAY125/FLSGRAY0625) and the stripe/crosshatch families. ContentSheetCell.background is document-schema.js's own ContentCellFill: a discriminated 'solid'/'pattern' shape, 'pattern' naming a closed ContentCellPatternType vocabulary spanning both this format's own FillPattern and WordprocessingML's ST_Shd (see that schema's own top comment for the full citation). FLSSOLID resolves to a 'solid' fill of the pattern's own foreground colour (icvFore), which [MS-XLS] itself documents as the only colour a solid fill actually renders ("If this value is 1, then only icvFore is rendered") — icvBack carries no meaning for it and is never consulted. Every other named FillPattern value resolves to a real 'pattern' fill via xf-colors.ts's own FILL_PATTERN_TO_PATTERN_TYPE, an exact 1:1 mapping onto the identical ECMA-376 ST_PatternType token name (FLSMEDGRAY through FLSGRAY0625, values 0x02-0x12, onto mediumGray through gray0625 in that order) — carrying whichever of the pattern's own foreground (icvFore, drawn as the pattern's strokes) and background (icvBack, the colour its gaps show through) colours actually resolve to a fixed RGB value; either may be an "Automatic" icv this package cannot express and is left unstated, matching ContentCellFillSchema's own "a colour can defer instead of asserting" convention. A reserved FillPattern value beyond 0x12 still reads as no background at all, there being no pattern name to give it. Writing a 'pattern' fill states its own colours (automatic where the fill leaves one unstated) under the FillPattern value the same table's inverse names for it; a WordprocessingML-only pattern name (the percentN family, or a stripe/cross member ST_Shd names but ST_PatternType does not) throws BiffWriteError rather than writing the wrong pattern or silently dropping it.

Borders. Each of a cell's four sides carries its own [MS-XLS] BorderStyle line-style token and colour, mapped onto ContentBorder's widthPt/style pair the same way ooxml.js's own typed/xlsx/styles.ts maps xlsx's border tokens: four named weights — hair/thin/medium/thick, at 0.5/0.75/1.5/2.25pt, derived from Excel's own documented 96-DPI rendering and held once in document-schema.js's own border-weight module, which this package and ooxml.js both import rather than each keeping a copy of it — crossed with a pattern (solid/dashed/dotted/double); the dash-family tokens (dashDot, dashDotDot, and their medium/slant variants) collapse to 'dashed', the closest ContentStrokeStyle member, exactly as ooxml.js's equivalent table does for xlsx's own dash tokens. Diagonal borders (dgDiag/grbitDiag) are out of scope — ContentCellBordersSchema has no diagonal member — and are always read as absent, always written as none.

Verified against a real, independent BIFF8 implementation, not just this package's own reader/writer pair. A .xls built directly by LibreOffice (soffice --headless --convert-to xls, from a hand-authored .fods declaring real fo:background-color/fo:border* cell styles) is read correctly by this package's own reader — every fill and border colour matched exactly against what LibreOffice's own subsequent re-export of the same bytes independently confirms them to be, and every border's named weight and pattern resolved to the same [MS-XLS] BorderStyle token LibreOffice wrote. The named weight is where the two implementations stop agreeing on a number, and necessarily so: BIFF8 stores a name (thin, medium, thick), not a width, so LibreOffice's re-export states its own rendered widths (0.74/1.76/2.49pt) for the same tokens this package renders at 0.75/1.5/2.25pt. Both are honest readings of the same bytes. And a .xls this package writes — including one whose colours force a real Palette record, proving that path specifically — opens in LibreOffice with the correct fill and borders, confirmed by converting it back through soffice --headless --convert-to fods and inspecting the cell styles LibreOffice itself recovers from it.

The decorated-blank case was checked the same way, in both directions and against the same implementation: a .fods whose one styled cell has a fill and four borders but no value converts to a .xls in which LibreOffice writes a real Blank record, and this reader recovers that cell's fill colour and all four border colours exactly; a .xls this package writes for the equivalent empty cell converts back to a .fods in which LibreOffice recovers a valueless cell carrying the same fo:background-color and fo:border colours — at its own 0.74pt rendering of the thin weight, per the note above.

A genuine two-colour pattern fill was not checked against LibreOffice in either direction: ODF's own style:table-cell-properties has no attribute for a two-colour pattern fill for a hand-authored .fods to state one through, so there was no third-party bytes to compare against. Reading and writing every named FillPattern value is instead pinned against bytes hand-built from [MS-XLS]/[MS-XLSB]'s own enumeration (biff/xf-colors.test.ts) and against a whole-document round trip through this package's own reader and writer (write.test.ts).

Decoration on a cell with no value. A cell can be empty and still have something to show, and BIFF8 says so with a Blank record ([MS-XLS] 2.4.20) — a cell header naming an XF and nothing else, written precisely because that XF carries a fill or a border. Both directions honour it: a Blank or MulBlank whose XF resolves to real decoration is read as an empty-kind cell carrying that decoration rather than dropped, and an empty cell carrying a background or a border is written back as a Blank record pointing at an XF encoding it. An empty cell with no decoration is still written as nothing at all and still read as absent, which is what keeps ContentSheet's cell array sparse; "decoration this reader can express" is the same test the value-cell path applies, so a reserved FillPattern value with no pattern name, or an unresolvable colour on every field a cell states, leaves a blank cell dropped exactly as before.

A per-cell font remains out of scope in both directions — see Read-side gaps and Writer scope above. The payload's leading word, once similarly out of scope, is now Cell alignment below.

Cell alignment

A cell's own horizontal and vertical alignment are read from and written to the same XF trailing CellXF/StyleXF payload Cell decoration covers ([MS-XLS] 2.4.353), but a different pair of fields within it: alc and alcV, the leading word's own bits 0-2 and 4-6 ([MS-XLS] HorizAlign and VertAlign) — verified both by round trip and against a real, independent BIFF8 implementation, the identical bar Cell decoration and Print settings are held to. src/biff/xf-colors.ts is the one place alc/alcV are packed or unpacked in either direction (resolveHorizontalAlignment/horizAlignTokenFor and resolveVerticalAlignment/vertAlignTokenFor), so the two directions cannot silently disagree on what a given bit pattern means.

Only the members ContentSheetCell.alignment/verticalAlignment can express survive, matching ooxml.js's own xlsx alignment policy exactly. HorizAlign names eight members; only ALCLEFT/ALCCTR/ALCRIGHT/ALCJUST have an Alignment counterpart (left/center/right/justify). ALCGEN — general alignment — is the identical semantics to alignment being absent (the schema's own "numeric right, text left" value-kind default already means the same thing), so it round-trips to and from undefined rather than a literal member that would override the default it means to request. ALCFILL/ALCCONTCTR/ALCDIST are real members with no schema counterpart at all and are left unread, the same policy odf.js's own alignment reader applies to a construct its schema has no member for. VertAlign names five members; only ALCVTOP/ALCVCTR have a verticalAlignment counterpart (top/middle). ALCVBOT is the schema's own documented default for an absent verticalAlignment ("no value-kind default to fall back to, so its own absence means 'bottom' outright"), so it round-trips to and from undefined the same way ALCGEN does; ALCVJUST/ALCVDIST are left unread for the identical reason ALCFILL/ALCCONTCTR/ALCDIST are.

Interned into the cell-XF table alongside decoration, not a separate table. write.ts's own buildCellXfPlan already deduplicated cells sharing an identical (number format, decoration) pair into one XF record; the interning key is now a (number format, alignment, vertical alignment, decoration) tuple, so two cells sharing all four still share one record and a cell differing in only its alignment still mints its own. A cell with General formatting, no alignment, and no decoration still resolves to the workbook's own implicit GENERAL_CELL_XF_INDEX with no new XF record at all, exactly as before alignment was modelled.

Formatting on a cell with no value, widened. Cell decoration's own Blank-record rule (a decorated empty cell is worth a record; an undecorated one is written as nothing at all) now triggers on alignment too: written-cells.ts's cellCarriesFormatting (renamed from cellCarriesDecoration when this was added) treats a non-default alignment or verticalAlignment the same way it already treats a background or a border. An empty cell stating only alignment: 'center' therefore gets a real Blank record and round-trips with that alignment intact, rather than being silently dropped as if it had nothing to say.

Verified against a real, independent BIFF8 implementation, not just this package's own reader/writer pair. A .xls this package writes — one cell per Alignment member, one per verticalAlignment member, and one combining a non-default horizontal and vertical alignment on the same cell — opens in LibreOffice 26.2.5.2 with every alignment intact, confirmed by converting it back through soffice --headless --convert-to fods and inspecting the cell styles LibreOffice itself recovers: each left/center/right/justify cell reads back as LibreOffice's own fo:text-align="start"/"center"/"end"/"justify", each top/middle cell as style:vertical-align="top"/"middle", and the combined cell as both at once on the identical style. Going the other way, a .xls built directly by LibreOffice (soffice --headless --convert-to xls, from a hand-authored .fods declaring the same real fo:text-align/style:vertical-align cell styles) is read correctly by this package's own reader, every cell recovering the exact alignment LibreOffice's own style declared.

Formula writing

src/biff/ptg-writer.ts compiles a ContentSheetCell.formula's own text back into a Formula record's Ptg token stream — the write-side counterpart of src/biff/ptg.ts, and, alongside cell comments below, the first two of xls-codec's read+write scope gaps this package's own README used to list as Formula records/Note/Txo being read-only.

Scope: same-sheet formulas over Ftab's own function vocabulary, nothing more. A tokenizer and a small recursive-descent parser (Excel's own documented operator precedence, narrowed to the operators ptg.ts's reader reconstructs) turn the formula text into a tree, then a second pass walks that tree post-order to emit Ptg tokens exactly the shape ptg.ts expects to read back: literal operands (PtgInt/PtgNum/PtgStr/PtgBool/PtgErr), a same-sheet cell or range reference (PtgRef/PtgArea, each $-qualified per the text's own absolute/relative markers), every arithmetic/comparison/unary/percent operator, an explicit parenthesis (PtgParen, restated unconditionally so the round trip preserves exactly what the author typed rather than only what precedence strictly requires), and a function call resolved by name against ptg-functions.ts's own Ftab table — PtgFunc when the name's own fixed arity (where Ftab's grammar states one) matches the call's argument count, PtgFuncVar otherwise. Every Formula record this writer produces carries its own complete, independent rgce — never a PtgExp pointing at a shared or array formula group — which is entirely legal BIFF8 (shared-formula compression is an optimisation, not a requirement) and sidesteps needing to build a ShrFmla/Array record pair at all.

What throws instead of writing unreadable bytes, each because the construct needs infrastructure this writer does not have, or has nowhere to resolve against: a 3D (cross-sheet or external-workbook) reference — resolving one to an ixti needs a SupBook/ExternSheet pair, which globals-writer.ts today only ever mints for the two built-in print-settings names, not for an arbitrary formula; an array-constant literal ({1,2;3,4}) or a CSE array formula — both need a PtgExtraArray/Array-record trailer this writer does not build; a defined name — document-schema.js's spreadsheet model has nowhere a user-defined name lives, the identical gap the Lbl row of the Writer scope table already describes; and a function name outside Ftab's own vocabulary — Excel 2007+ added many worksheet functions BIFF8's Ftab enumeration never named (resolved instead through a PtgNameX/add-in mechanism this writer does not implement). Every one of these throws BiffWriteError naming the construct.

The cell's own cached value is written into the Formula record's FormulaValue field alongside rgce, exactly as a real producer does: a numeric/temporal value kind writes as a plain IEEE 754 double, boolean/error write through FormulaValue's own tagged shape, and string writes the tagged shape plus a following String record carrying the cached text — matching workbook/sheet.ts's own taggedFormulaValue read side field for field. A formula whose value resolves to empty is refused: the one BIFF8 encoding that could carry it (a tagged "blank" result) reads back through this package's own reader as an empty string, not an empty cell, so writing it would silently change what round-trips.

Verified by round trip (src/write.test.ts's own formula records suite) rather than against a third-party BIFF8 implementation — unlike Cell decoration/Cell alignment/Print settings, no LibreOffice-authored fixture exists yet to cross-check a written Formula record's bytes against.

Cell comments, written

src/workbook/comment-writer.ts writes a ContentSheetCellComment back out as the same Note/Obj/Txo triple src/workbook/comments.ts already reads (see that module's own top comment for the full citation of how the three record kinds join): one Note record per commented cell naming its own Obj record by object id, that Obj record's FtCmo+FtNts pair marking it a comment, and a Txo record plus one Continue record carrying the comment's own text and a minimal, unformatted TxORuns trailer ([MS-XLS]'s own "cbRuns MUST be >= 16 and a multiple of 8" rule needs at least one real run plus its terminating TxOLastRun, even for plain text with no rich formatting). Object ids are assigned sequentially per sheet, starting at 1 — unique within the substream because workbook/drawing-writer.ts's own picture and embedded-object Obj records continue numbering from one past this count, rather than starting over at 1 themselves (see Images and embedded objects, written), matching [MS-XLS] 2.5.92's own "id MUST be unique among all Obj records of the substream" rule.

Only text and author round-trip: legacy BIFF8's NoteSh carries no reply structure and no per-comment timestamp at all (unlike xlsx's own [MS-XLSX] threaded-comments extension), so replies and createdAt have nowhere to write to, the identical gap the reader's own Read-side gaps section documents. A comment with no recorded author writes stAuthor as an empty string rather than a non-empty placeholder — NoteSh's own field documents a length-1 minimum, but this package's reader only ever promotes a non-empty stAuthor to ContentSheetCellComment.author, so a placeholder would round-trip back as a fabricated author nobody wrote; an empty string is the one spelling that round-trips as "no author" through this reader specifically. Rich per-character formatting within a comment (TxORuns' own genuine run array) is never written, matching the reader's own identical scope limit.

Verified by round trip (src/write.test.ts's own cell comments suite): a comment's text and author on a valued cell, a comment anchored to an otherwise-empty cell, a comment with no author, an empty-text comment, and several comments on one sheet each keeping their own cell and text.

Images and embedded objects, written

src/drawing/escher-writer.ts and src/workbook/drawing-writer.ts are the write direction of src/drawing/escher.ts/src/drawing/shapes.ts/src/drawing/blips.ts and src/workbook/drawing.ts (see "Charts, drawings, and images, resolved for real" under Read-side gaps): one MS-ODRAW (Escher) container tree per sheet, the workbook-wide drawing group every sheet's Blip Store references share, and the Obj records pairing each Escher shape with what it actually holds.

buildDrawingWritePlan runs once, workbook-wide, before any sheet's own records are written — the same shape every other cross-sheet table in this writer takes (the number-format, colour, font, and shared-string plans in write.ts). A ContentSheetImage resolves into the workbook's own Blip Store (OfficeArtBstoreContainer): two images sharing identical base64 bytes dedupe onto one BSE entry, whose cRef counts the references, rather than minting a second copy of the same picture. rgbUid, the BSE's own hash of the blip's literal file bytes, is computed with a hand-written MD4 implementation (src/drawing/md4.ts, pinned against RFC 1320's own test vectors) — no platform crypto API exposes MD4, the digest [MS-ODRAW] itself specifies for this field. Only PNG and JPEG blips write (ContentImageBlockSchema.format's other members have no MSOBLIPTYPE this writer emits a Blip Store entry for); any other format throws BiffWriteError naming it. Shape ids allocate workbook-wide from 1024 — Excel's own convention for a file's first drawing group — the invisible per-sheet "patriarch" root taking one id ahead of that sheet's own real shapes, so the Escher stream's shape order and the worksheet substream's Obj records pair up 1:1 exactly as the reader's own positional correlation expects.

A ContentEmbeddedObject writes through the identical Picture Obj machinery a plain image uses, but names an Embedding Storage instead of a Blip Store index: its FtPictFmla sub-record ([MS-XLS] 2.5.150) carries a storage id, and the outer [MS-CFB] compound file gets a sibling MBD<8-hex-digit id>/Package stream beside the Workbook stream ([MS-XLS] 2.1.7's own Embedding Storage naming convention) — writeXlsContent collects these from buildDrawingWritePlan and adds them to writeCompoundFile's own stream list. What actually rides inside that Package stream (src/workbook/embedded-object.ts) is this package's own JSON serialisation of the embedded object's objectKind/document/source — the same honest boundary rtf-codec's own embedded-object module draws for the identical reason: xls-codec cannot depend on ooxml.js/odf.js (format codecs are peers, never one another's dependency), so a wordprocessing/presentation/spreadsheet/drawing/formula object's own ContentDocument cannot be re-serialised into a real docx/pptx/xlsx/odf/MathML byte stream here the way a genuine OLE server would. frame and the anchor quartet are deliberately excluded from that payload: the Escher anchor (ClientAnchorSheet) is the one placement authority the format itself carries, both directions deriving every placement field from it, and a second, redundant copy inside the payload could only ever disagree with it. A 'chart'-kind embedded object throws BiffWriteError by name rather than being approximated — see the Writer scope table.

A record whose Escher bytes exceed the 8224-byte single-record ceiling — a sheet's own MsoDrawing record, or (once a workbook's images push its Blip Store past that size) the globals substream's MsoDrawingGroup record — chains onto Continue records via writeRecordChain rather than being refused, the one deliberate exception to the Continue-chain-splitting row the Writer scope table states for every other record kind; reading a chained record back correctly required a corresponding fix on the read side, described below.

A latent read-side bug, found and fixed while building this. Verifying an image large enough to force a Continue chain surfaced a bug already on main, in code this feature did not touch: both workbook/drawing.ts's readSheetDrawing (a sheet's own MsoDrawing bytes) and content.ts's concatDrawingGroupBytes (the workbook-wide MsoDrawingGroup bytes) read only record.blocks[0] — the base record's own data — silently discarding every Continue-chained block beyond the first. Since a real picture's blip bytes routinely exceed one record's 8224-byte ceiling, this would have corrupted the Escher stream, and therefore the shape tree, for almost any real-world image, whether written by this package or by Excel itself. Both call sites now concatenate every block in the group, confirmed by a test that forces the chain and checks the image reads back byte-for-byte.

A genuine write-side bug, found and fixed the same way. The same large-image test also caught OfficeArtClientAnchorSheet's own corner-cell offsets (dxL/dyT/dxR/dyB) being written as 4-byte fields when the record's own declared recLen (18 bytes total) and the reader's readClientAnchor both treat them as 2-byte fields — every field after the first offset landed at the wrong byte position, so any shape's placement beyond a trivial single-shape sheet came back with a wildly wrong size and anchor. Fixed to match the reader's own 2-byte reads. A second, independent bug in the same commit — bytesFromBase64's own decode alphabet was missing /, one of the two non-alphanumeric characters standard base64 uses — meant any image whose base64 encoding happened to contain that character (which is most real images) silently dropped it, corrupting the written bytes. Both were caught by the same test, which decodes real base64 test fixtures large enough to statistically require a / and small enough to fit in one record, so it also stands as regression coverage for both defects independently of the Continue-chain case that surfaced them.

Verified by round trip (src/write.test.ts's own images and embedded objects suite), entirely within this package: a single image's format, bytes, and cell anchor; an image large enough to force the workbook-wide Blip Store onto a Continue chain, confirmed by inspecting the raw record stream rather than trusting the round trip alone; many images on one sheet, forcing that sheet's own MsoDrawing record onto its own Continue chain; a non-chart embedded object through its own Embedding Storage; and the 'chart' refusal. Not cross-checked against a real, independent BIFF8 implementation the way Cell decoration/Cell alignment/Print settings are — a genuinely useful follow-up, since every other decoration feature this package writes has that independent check and this one does not yet.

Print settings

Every field of document-schema.js's own ContentSheetPrintSettings is read from and written to the BIFF8 records that carry it, verified both by round trip and against a real, independent BIFF8 implementation — ExaDev/documents.js#815 as a scoped chunk of that issue's own broader tracking, not a claim of closing it outright. The two directions had to land together: before this, the reader returned Excel's fixed "Normal" preset for every sheet regardless of what the file said, so there was nothing to verify a writer's output against.

One sheet's print settings live in two substreams, not one. The page setup is in the sheet's own substream, as the optional records of [MS-XLS] 2.1.7.20.6's own GLOBALS and PAGESETUP productions. The print range and the repeated header bands are not there at all: BIFF8 keeps them in the workbook globals substream, as ordinary defined names ([MS-XLS] 2.4.150's Lbl record) carrying a built-in name index rather than a user-typed name, scoped to one sheet through the record's own itab. src/workbook/print-names.ts is both directions of exactly those two names; src/biff/print-setup.ts is the one place the Setup record's own bit layout and paper-size code table are packed or unpacked, shared by the read and write sides so the two cannot silently disagree.

Schema field Where BIFF8 keeps it
pageSize Setup's iPaperSize code ([MS-XLS] 2.4.257), transposed to landscape by its own fPortrait/fNoOrient flags
margins LeftMargin/RightMargin/TopMargin/BottomMargin ([MS-XLS] 2.4.151, 2.4.219, 2.4.328, 2.4.27), each a single Xnum of inches
gridlines PrintGrid's fPrintGrid ([MS-XLS] 2.4.202)
headers PrintRowCol's printRwCol ([MS-XLS] 2.4.203) — the row/column header chrome, not a repeated print band
pageOrder Setup's fLeftToRight
scalePercent / fitToPages Setup's iScale or its iFitWidth/iFitHeight pair, selected by WsBool's own fFitToPage bit ([MS-XLS] 2.4.351)
manualBreaks HorizontalPageBreaks/VerticalPageBreaks ([MS-XLS] 2.4.142, 2.4.343)
printRange The built-in Print_Area defined name (Lbl with built-in index 0x06), whose rgce is a PtgArea3d naming the range
repeatRows / repeatColumns The built-in Print_Titles name (index 0x07), whose rgce is a PtgMemFunc wrapping one or two PtgArea3d tokens joined by PtgUnion

Every record behind these is optional, and every field falls back independently. A sheet whose page setup was never touched carries no Setup and no margin records at all, so something has to stand in — Excel's own "Normal" preset (top/bottom 0.75in, left/right 0.7in, Letter paper, gridlines and headers not printed, pages down-then-over), the same constants ooxml.js falls back to for an xlsx carrying no pageMargins element, so the same untouched sheet reads identically from either format. The fallback is per field rather than wholesale: a sheet declaring a left margin and nothing else keeps its real left margin and takes the preset for the other three.

Two of BIFF8's own conditional rules are honoured rather than flattened. A Setup record whose fNoPls bit is set declares its own paper size and scale undefined ([MS-XLS] 2.4.257: "whether the iPaperSize, iScale, iRes, iVRes, iCopies, fNoOrient, and fPortrait data are undefined and ignored"), so neither is read from it — the page size falls back to the preset and no scalePercent is reported, rather than a paper code the file itself disowns being resolved into a confident page size. And WsBool's fFitToPage decides which of Setup's two mutually exclusive scaling fields is live: real producers write both regardless of which is active (confirmed against LibreOffice-written BIFF8, which carries iScale=100 alongside a real fit-to-page pair, and a real iScale alongside iFitWidth=iFitHeight=1), so reading both would report a scale and a page count that contradict each other.

A repeated band's axis is its shape, not a field. BIFF8 has no flag saying which axis a Print_Titles band repeats along: a repeated row band is written as an area spanning every column of the sheet ($1:$2, columns 0-255) and a repeated column band as one spanning every row ($A:$A, rows 0-65535). Both directions use that shape as the discriminant, and an area spanning both axes at once names the whole sheet — which is neither — so it is left unclassified rather than assigned to whichever branch happened to be tested first.

Three things do not come through, each for a reason in the format rather than an oversight:

  • A page size no iPaperSize code names loses its dimensions. Unlike xlsx's own pageSetup element, which can state an explicit paperWidth/paperHeight pair, Setup addresses paper only by code; its escape hatch for a size outside the table is a printer-defined custom size carried in a separate Pls record ([MS-XLS] 2.4.199), a printer driver's opaque DEVMODE blob rather than a pair of dimensions any reader could recover a size from. So the writer emits iPaperSize 0 — that section's own "custom printer paper sizes" — which is true, where substituting Letter would not be; the reading application then falls back to its own default paper (this package to the preset, LibreOffice to its locale's), and every other print setting on the sheet still comes through. Refusing the file outright was the first thing tried and is the wrong trade: a spreadsheet converted from a slide deck or a drawing carries that source's own canvas as its page size, which is almost never a named paper, and failing the conversion would lose the cells too. The code table this package does map is the office paper sizes of [MS-XLS]'s own 118-entry enumeration — the Letter/Legal/Tabloid/Executive/Statement family, A2 through A6, B4/B5 in both the JIS and ISO spellings, Folio and Quarto — each derived from the inches or millimetres that table itself states rather than from pre-converted points, and each matched within half a point so a page size picking up conversion drift on its way between codecs still resolves.
  • A "fit to as many pages as necessary" axis has no schema spelling. Setup documents iFitWidth/iFitHeight of 0 as "use as many pages as necessary to print the columns/rows in the sheet", and ContentSheetPrintSettings.fitToPages requires both counts to be positive. A fit-to-page sheet with an auto axis therefore reports no fitToPages at all rather than a fabricated 1, which would claim the sheet is pinned to a single page along an axis the file left free.
  • An explicit 100% scale reads back as no declared scale. Setup's iScale is a mandatory field of a mandatory record with no spelling for "this sheet declares no scale", so an untouched sheet still states 100 — and ContentSheetPrintSettings already means exactly that by carrying no scalePercent at all. The two spellings print identically, so the reader collapses them onto the absent one rather than putting a field carrying no actionable information on every sheet of every workbook it reads. A scale that is genuinely anything else is reported exactly as the file states it.

Two further deliberate narrowings, both because the schema models less than BIFF8 states: a Print_Area naming several disjoint areas (legal in BIFF8, and what Excel writes for a multi-area print selection) yields only the first, since printRange models one rectangle and merging several into their bounding box would claim cells print that do not; and a BIFF8 page break carries an extent along the perpendicular axis, which manualBreaks — an index with no extent — cannot express, so a partial break is carried as a full one and two breaks on the same row collapse into one.

Verified against a real, independent BIFF8 implementation, not just this package's own reader/writer pair. Three .xls files built directly by LibreOffice (soffice --headless --convert-to xls, from hand-authored .fods files declaring a real page layout, print range, header rows and columns, and manual breaks) read back through this package with every field matching what was authored: A4 landscape, four distinct margins, gridlines and headers on or off, overThenDown page order, an 80% scale and a 2x3 fit-to-page pair, a row break at row 10, a column break at column 3, a B2:D6 print range, two repeated header rows, and one repeated header column. Going the other way, a .xls this package writes for the same content opens in LibreOffice with every one of those fields intact, confirmed by converting it back through soffice --headless --convert-to fods and comparing LibreOffice's own re-export of our file against its re-export of its own, attribute by attribute. A fourth file, written with a page size no paper code names, opens in LibreOffice with its own default paper substituted and every other print setting and cell intact — the custom-paper behaviour above, checked rather than assumed. The Print_Area record this writer emits is byte-for-byte the one LibreOffice writes for the same range (src/workbook/print-names.test.ts asserts exactly that against the real bytes); the Print_Titles record differs by one byte, a trailing PtgParen display token this writer has no reason to emit.

Why this writer now emits the calculation-state records. [MS-XLS] 2.1.7.20.6's GLOBALS production makes CalcCount, CalcRefMode, CalcIter, CalcDelta and CalcSaveRecalc mandatory ahead of PrintRowCol, and this writer previously omitted them along with the rest of BIFF8's UI and interoperability bookkeeping. That turned out to matter: LibreOffice's own importer silently discards whichever page-settings record comes first in a worksheet substream, so with PrintRowCol in that slot, a .xls this package wrote with row and column headers enabled opened in LibreOffice with them off — while every other print setting in the same file came through correctly. Moving any other record into that slot fixes it, and the records the grammar already required there are the honest way to do it. Confirmed by writing the same workbook with and without them and re-reading each through soffice --convert-to fods.

Metadata

A .xls's title, author, and dates do not live in any BIFF8 record at all — they live in a "\x05SummaryInformation" stream, a genuinely different format ([MS-OLEPS] Property Set Streams, [MS-OSHARED] 2.3.3.2.2's own naming of the specific properties Office uses) that happens to sit beside Workbook in the same [MS-CFB] compound file. readXlsContent reads that stream when present (archive-codec's readSummaryInformation, since the property-set format itself is zero document-format knowledge, exactly as the [MS-CFB] container it sits inside is) and maps it onto document-schema.js's LayoutMetadata (archive-codec's own summaryInformationToLayoutMetadata — the mapping is format-agnostic, so it lives there rather than being copied in this package, alongside doc-codec's and ppt-codec's identical need for it); writeXlsContent does the inverse (src/metadata.ts's layoutMetadataToSummaryInformation, which validates createdIso/modifiedIso as real dates and throws a BiffWriteError naming the offending field before delegating to archive-codec's own mapping), including a "\x05SummaryInformation" stream in its writeCompoundFile call only when the input's metadata actually carries something that stream can hold — an input whose metadata is {}, or carries only fields the mapping below has no destination for, produces no stream at all, matching what an absent-metadata read already returns.

The mapping is not 1:1, and each gap is permanent rather than a remaining TODO:

Direction Fields covered Gap
SummaryInformation → LayoutMetadata title, subject, author, keywords, createdIso, lastSavedIso → modifiedIso comments and lastPrintedIso have no LayoutMetadata field to land in — no other codec in the family has a "last printed" or free-text "comments" concept, so these are read from the stream but never reach a ContentDocument.
LayoutMetadata → SummaryInformation the same six fields, in reverse creator, producer, and language have no SummaryInformation equivalent: producer is a PDF-only concept in this schema, and creator/language are not among the fields the stream this package writes covers.

Only the fixed SummaryInformation property set is read or written — the sibling "\x05DocumentSummaryInformation" stream (company, manager, and custom user-defined properties, [MS-OLEPS]'s two-property-set spelling) is not attempted at all, an explicit scope boundary archive-codec's own oleps support shares.

Getting started

Requires Node.js >=20 and pnpm 11.6.0.

pnpm install
pnpm build          # tsdown -> dist/ (ESM + CJS + .d.ts, one file set per src module)
pnpm typecheck      # tsc -p tsconfig.json && tsc -p tsconfig.node.json (dual tsconfig)
pnpm lint           # eslint . --fix --cache --max-warnings 0
pnpm test           # vitest run --project unit
pnpm test:watch     # vitest --project unit
pnpm test:workers   # vitest run --config vitest.workers.config.ts, inside a real Cloudflare Workers (workerd) isolate
pnpm test:smoke     # builds dist/, then loads the built ESM and CJS barrels and every advertised deep import

Usage

import { isXlsFile, readXls, readXlsContent } from "xls-codec";

const bytes = new Uint8Array(await file.arrayBuffer());

if (isXlsFile(bytes)) {
  const content = readXlsContent(bytes); // ContentDocument, kind: 'spreadsheet'
  const tree = readXls(bytes); // DocumentTree, the same read decomposed
}

readXlsContent mirrors ooxml.js's readXlsxContent deliberately, down to returning a ContentDocument rather than a bare { metadata, sheets } object, so a caller can hold either behind one type. The difference is the input: an .xls has no Package equivalent to decode first, so these take the file's raw bytes and select the Workbook stream themselves.

The record layer is exported in its own right, for a caller inspecting a workbook rather than converting it:

import { readRecords, readWorkbookStreams } from "xls-codec";

const { workbook } = readWorkbookStreams(bytes); // the raw BIFF8 record stream out of the compound file, plus the optional "\x05SummaryInformation" stream beside it
for (const rec of readRecords(workbook)) {
  console.log(rec.type.toString(16), rec.data.length);
}

Microsoft Works spreadsheets (.xlr)

Works 9's .xlr is BIFF8 in the same compound-file container, carrying the identical Workbook stream alongside a Works-specific WksSSWorkBook stream (SheetJS format notes). Because this package selects the Workbook stream by name and ignores every other stream in the container, an .xlr reads through exactly the same path with no special-casing; isXlsFile accepts one, and there is a test pinning that.

Architecture

Layered bottom-up, each layer testable against hand-built byte sequences taken from the spec's own field-layout tables:

  • src/biff/record-types.ts — the record type numbers, each cited to [MS-XLS] 2.3.1's own enumeration rather than copied from another implementation.
  • src/biff/records.ts — the record framing, and nothing above it. Deliberately does not merge Continue records: whether a continuation's bytes simply append or re-state a flag byte first is decided by the record being continued, so the blocks are reported as written.
  • src/biff/cursor.ts — a field cursor over one record's blocks that reads across a continuation boundary transparently while keeping the boundary observable, which is exactly what the string reader needs.
  • src/biff/strings.ts, src/biff/rk.ts, src/biff/errors.ts — the shared value encodings: the three string shapes, the RkNumber packed-numeric encoding, and the BErr error-value vocabulary.
  • src/biff/ptg.ts, src/biff/ptg-functions.ts — the Ptg compiled-formula token stream, walked as a postfix expression and rebuilt into infix formula text (an operand stack tagged with each entry's own operator precedence, so a child is parenthesised only when its precedence genuinely requires it), and the built-in worksheet-function name/fixed-arity table PtgFunc/PtgFuncVar resolve against.
  • src/biff/ptg-writer.ts — the write direction: a tokenizer and recursive-descent parser turning same-sheet formula text into a tree, then a post-order walk emitting Ptg tokens — see Formula writing for its exact scope.
  • src/workbook/comment-writer.ts — the write direction of src/workbook/comments.ts: a cell's own Note/Obj/Txo triple, text and author only — see Cell comments, written.
  • src/workbook/globals.ts, src/workbook/sheet.ts — the two substream readers, each walking the record sequence its ABNF in [MS-XLS] 2.1.7.20.3 / 2.1.7.20.5 defines; globals.ts also resolves a 3D reference's own ixti to a sheet scope through EXTERNSHEET and SupBook (a plain sheet range for a self-referencing workbook, a fully-formatted label — a real external workbook/sheet name or a diagnostic placeholder — otherwise), which sheet.ts threads into ptg.ts for a Formula record's own 3D references, and reads a Palette record and each XF's trailing fill/border payload for Cell decoration. sheet.ts's own collectFormulaGroups additionally joins a shared or array formula's PtgExp-bearing member cells against the ShrFmla/Array record that carries the group's real expression (see "Formula expressions" under Read-side gaps). Both readers contribute to Print settings, which BIFF8 splits between them: sheet.ts reads the page-setup record group, globals.ts the two built-in defined names carrying the print range and the repeated header bands.
  • src/biff/xf-colors.ts — the Icv colour table (both the eight fixed colours and the 56-entry default palette), the BorderStyle/FillPattern vocabularies, and the CellXF/StyleXF trailing payload's own border/fill bit-layout packing and unpacking, shared by globals.ts's read side and biff/xf-writer.ts's write side — see Cell decoration.
  • src/biff/print-setup.ts, src/workbook/print-names.ts — the two halves of Print settings. The first is the Setup record's own flag bit layout and iPaperSize code table, packed and unpacked in one place exactly as xf-colors.ts does for the XF payload; the second is both directions of the built-in Print_Area/Print_Titles defined names, which live in the globals substream rather than the sheet's own and so are read by globals.ts and written by globals-writer.ts.
  • excel-number-format, src/serial.ts — number-format classification and date-serial conversion, the two pieces of xlsx semantics BIFF8 shares because ECMA-376 inherited them from BIFF. The classifier itself is a dependency shared with ooxml.js, not a module in this package (#848) — classifyNumberFormat and BUILTIN_NUMBER_FORMATS still ride this package's own barrel (export * from "excel-number-format" in src/index.ts), so import { classifyNumberFormat } from "xls-codec" is unchanged.
  • src/drawing/escher.ts, src/drawing/shapes.ts, src/drawing/blips.ts — MS-ODRAW (Escher), read as its own record tree independent of BIFF8's own record framing: the shared OfficeArtRecordHeader container/atom reader, the per-sheet shape walk (type, id, cell anchor, a picture's own pib blip reference), and the workbook-wide Blip Store a pib resolves against.
  • src/workbook/chart.ts, src/workbook/drawing.ts — an embedded chart's own nested substream, read into a flattened series/category table; and the orchestration that pairs each Escher shape with the Obj record naming what it holds (a picture, a chart, or a generic shape) and produces ContentSheet.images/ContentSheet.embeddedObjects — see "Charts, drawings, and images, resolved for real" under Read-side gaps.
  • src/drawing/escher-writer.ts, src/drawing/md4.ts, src/workbook/drawing-writer.ts, src/workbook/embedded-object.ts — the write direction of the four modules above: the Escher container tree builder, a hand-written MD4 implementation for a Blip Store entry's own rgbUid hash, the workbook-wide drawing plan and Obj-record writers, and the Embedding Storage Package-stream wrapper an embedded object's own document rides inside — see Images and embedded objects, written.
  • src/content.ts — the mapping onto document-schema.js.
  • src/metadata.ts — wraps archive-codec's own SummaryInformationProperties <-> LayoutMetadata mapping with this package's createdIso/modifiedIso date validation, throwing BiffWriteError for a malformed one rather than letting an opaque RangeError escape the FILETIME conversion (see Metadata).

Deliberately not depended on

No third-party spreadsheet or compound-file library — not xlsx/SheetJS, exceljs, or cfb — enforced by an ESLint no-restricted-imports rule in this package's own config. The format is implemented against its published Open Specification, and the compound-file layer comes from archive-codec, a sibling in this workspace.

Conventions

  • Worker-isomorphic (see the family-wide convention): runtime src/ must not import node:*, a bare Node builtin, or use the Buffer global — enforced by a no-restricted-imports/no-restricted-globals ESLint rule and exercised in CI by running a test suite inside an actual workerd isolate (pnpm test:workers). Test files under src/**/*.test.ts and src/test-support/ are exempt and may use Node APIs for fixtures.
  • Only src/index.ts may be named index.* — a custom ESLint rule (local/no-non-barrel-index) rejects any other module using an index basename, since that would be a hidden entry point the exports map in package.json doesn't advertise.
  • Every record layout is cited to its own [MS-XLS] section, by URL, at the point it is read. A field offset with no citation is a field offset nobody can check.

Install

pnpm add xls-codec
# or
npm install xls-codec

Release and publishing

Release, CI, and commit-message conventions are all workspace-wide, not package-local — see the monorepo root README for the mechanism (topological per-package semantic-release via @exadev/semantic-release-workspace, OIDC trusted npm publishing, automatic sibling dependency-range rewriting) and its post-release republishing and attestation note on the restored GitHub Packages mirrors, npm aliases, and SBOM/provenance signing.

Contributing

Conventional Commits, enforced workspace-wide by commitlint through a root commit-msg hook. Work inside packages/xls-codec/; see CONTRIBUTING.md for the shared git hooks and history conventions.

References

License

MIT