docboss.dev

Benchmarks

DOCX and DOC benchmarks, measured on real documents.

docboss against docx2txt, python-docx, docx2python, mammoth, pandoc, antiword and catdoc on the test documents of LibreOffice, Apache POI and python-docx: speed, word recall against LibreOffice, parallel throughput, one large document, memory, damaged files and rendering fidelity. The tables where another engine wins are here too. Every number comes from the scripts in the docboss repository.

3,304 files/s

DOCX text over 80 real-world test files: about 4× docx2txt and 28× mammoth.

3,130 files/s

DOC text over 80 legacy .doc files: about 19× antiword and catdoc.

0.992 word recall

Against LibreOffice’s text export on DOCX, the highest of the six engines measured; 0.985 on DOC.

140/150 damaged files

Damaged DOCX files that still give text, where the other engines manage 1 to 8. No crash and no hang on 300 damaged files.

DOCX text extraction

80 sampled DOCX files, 76 handled by every engine · files per second, higher is better

docboss3,304
docx2txt757
python-docx725
docx2python127
mammoth117
pandoc7.4

Best of 3 per file after a warm-up pass, aggregated over the files every engine handled. Markdown and HTML follow the same pattern: docboss 3,324 and 3,355 files/s, mammoth 109 and 111, pandoc 7.6 and 7.5.

DOC text extraction

80 sampled DOC files, 66 handled by every engine · files per second

docboss3,130
catdoc161
antiword158

antiword refuses 12 of the 78 files the others open.

Parallel, DOCX

300 files, 12 workers, each engine's best route · files per second

docboss13,839
docx2txt1,262
python-docx1,082
mammoth316
docx2python248
pandoc41

docboss through extract_texts, one call on a Rust thread pool (5.1× its sequential speed); the others through processes, pandoc through threads.

Parallel, DOC

180 files, 12 workers · files per second

docboss5,685
catdoc490
antiword392

The DOC sample holds fewer, larger files, so the speedup is lower: 2.5× for docboss.

Word recall against LibreOffice, DOCX
EngineMeanWeighted
docboss0.9920.997
docx2txt0.9740.991
docx2python0.9740.993
mammoth0.9560.992
pandoc0.8940.971
python-docx0.8480.978

Share of the reference's words, with multiplicity, found in each engine's output; 57 files with reference text. python-docx's adapter reads body paragraphs and table cells, so headers, footers, footnotes, comments and text boxes are left out.

One large document, DOCX text, best of 3
EngineTime
docboss85 ms
docx2txt808 ms
python-docx1,490 ms
mammoth5,917 ms
pandoc8,044 ms
docx2python8,780 ms

A generated document of 3,000 sections (1,188 pages in docboss's layout), each a heading, a paragraph with bold and italic runs, a bullet list with a nested numbered list and a four-row table. docboss's output is the longest, 1.64 million characters against about 0.98 million for most others, because it keeps list labels and table cell separators.

The same document as DOC
EngineTime
catdoc54 ms
docboss105 ms
antiword2,552 ms

catdoc wins, about 2× faster than docboss: it streams the text out of the piece table without building a document model, while docboss parses formatting, styles, 9,001 lists and tables into its model first.

Peak memory, DOCX
Engine100 filesLargest (2.96 MB)
docx2txt30.1 MB24.6 MB
docboss33.2 MB34.7 MB
docx2python55.0 MB30.3 MB
python-docx58.8 MB41.6 MB
mammoth82.2 MB32.7 MB
pandoc29.5 MB (child 134.8)29.4 MB (child 139.9)

Peak RSS of a fresh process that imports the engine and extracts the text; the interpreter's own baseline is 22 to 35 MB. For command-line engines the child process's peak is given too. docx2txt peaks slightly lower than docboss.

Peak memory, DOC
Engine100 filesLargest (6.24 MB)
catdoc23.3 MB (child 2.0)24.0 MB (child 2.9)
antiword23.8 MB (child 2.5)23.7 MB (child 2.9)
docboss40.4 MB63.5 MB

antiword and catdoc win: they stream text in about 2 to 3 MB of their own, while docboss holds the whole document model and the file's streams. 12 MB of the largest file's peak are its metafile and bitmap pictures, decoded into the model to be drawn.

150 damaged DOCX files
EngineTextErrorCrashHang
docboss1401000
docx2txt814200
docx2python414600
pandoc414600
python-docx214800
mammoth114900

Byte flips, truncations, packages rebuilt with a damaged word/document.xml, and packages with a zeroed central directory. docboss reads a ZIP from its local headers when the central directory is damaged, keeps what inflated before a corrupt deflate stream, and balances broken XML.

150 damaged DOC files
EngineTextErrorCrashHang
docboss142800
catdoc1004901
antiword717630

Byte flips, truncations and a 512-byte sector overwritten or zeroed, each file in its own process with a 30 s timeout. Of the 8 docboss refuses, 3 are encrypted originals and 5 have lost their FIB. antiword crashed with a segmentation fault 3 times.

Rendering fidelity against LibreOffice
FormatFilesSSIM meanSSIM medianSSIM p10MAD medianPage counts equal
DOCX590.9520.9850.8790.5556 of 60
DOC290.8980.9660.6341.8327 of 30

LibreOffice converts each document to PDF, which is rasterized, and docboss renders the same pages from the document itself, both at scale 1.5; each pair is scored with windowed SSIM and mean absolute pixel difference, up to 5 pages per file. Mostly white pages score high whatever their content, so the p10 column is the honest one: a tenth of the DOC files score under 0.64.

Method

How the numbers were taken.

Corpora. The real-world test documents of LibreOffice, Apache POI and python-docx, fetched at pinned revisions by fetch_corpus.sh: 1,697 DOCX and 180 DOC files after removing duplicates. Each script takes a deterministic, evenly spaced sample.

Timing. Best of 3 per file after one warm-up pass, so the file cache and imports are hot, aggregated only over the files every engine handled, so each total compares the same workload.

Recall. LibreOffice converts each file to text once. An engine’s recall on a file is the share of the reference’s words, with multiplicity, found in its output. Files whose reference has no words are skipped, since LibreOffice’s text export leaves out text boxes. On DOC, the weighted recall of catdoc and antiword is dominated by one 74,813-word Russian novel that both decode to 0.1% of its words, so their mean recall is the fairer comparison.

LibreOffice converts the same samples at 16.1 (DOCX) and 14.8 (DOC) files/s as one batch, and takes a median 827 ms (DOCX) and 725 ms (DOC) per file one process at a time.

Versions. docboss 0.1.0 (release build of master), python-docx 1.2.0, docx2python 3.7.1, mammoth 1.13.0, docx2txt 0.9, pandoc 3.9 through pypandoc_binary 1.17, antiword 0.37, catdoc 0.95, LibreOffice 26.8.0.3. Apple M3 Pro, 12 cores, macOS 26.4, Python 3.12.10, one session on 2026-09-30.

Numbers depend on the machine; compare rows against each other. Every script writes its results-*.json with the machine, versions and date, and benchmarks/README.md has the commands to reproduce each table.

What is not measured

Where the tables stop.

  • Markdown and HTML are timed but not scored for quality against the other engines' output; recall scores plain text only.
  • Rendering is compared with LibreOffice only.
  • DOCX writing and the async remote reader are not timed.