Benchmarks
DOCX and DOC benchmarks, measured on real documents.
docboss against docx2txt, python-docx, docx2python, mammoth, pandoc, antiword and catdoc on the test documents of LibreOffice, Apache POI and python-docx: speed, word recall against LibreOffice, parallel throughput, one large document, memory, damaged files and rendering fidelity. The tables where another engine wins are here too. Every number comes from the scripts in the docboss repository.
DOCX text over 80 real-world test files: about 4× docx2txt and 28× mammoth.
DOC text over 80 legacy .doc files: about 19× antiword and catdoc.
Against LibreOffice’s text export on DOCX, the highest of the six engines measured; 0.985 on DOC.
Damaged DOCX files that still give text, where the other engines manage 1 to 8. No crash and no hang on 300 damaged files.
DOCX text extraction
80 sampled DOCX files, 76 handled by every engine · files per second, higher is better
Best of 3 per file after a warm-up pass, aggregated over the files every engine handled. Markdown and HTML follow the same pattern: docboss 3,324 and 3,355 files/s, mammoth 109 and 111, pandoc 7.6 and 7.5.
DOC text extraction
80 sampled DOC files, 66 handled by every engine · files per second
antiword refuses 12 of the 78 files the others open.
Parallel, DOCX
300 files, 12 workers, each engine's best route · files per second
docboss through extract_texts, one call on a Rust thread pool (5.1× its sequential speed); the others through processes, pandoc through threads.
Parallel, DOC
180 files, 12 workers · files per second
The DOC sample holds fewer, larger files, so the speedup is lower: 2.5× for docboss.
| Engine | Mean | Weighted |
|---|---|---|
| docboss | 0.992 | 0.997 |
| docx2txt | 0.974 | 0.991 |
| docx2python | 0.974 | 0.993 |
| mammoth | 0.956 | 0.992 |
| pandoc | 0.894 | 0.971 |
| python-docx | 0.848 | 0.978 |
Share of the reference's words, with multiplicity, found in each engine's output; 57 files with reference text. python-docx's adapter reads body paragraphs and table cells, so headers, footers, footnotes, comments and text boxes are left out.
| Engine | Time |
|---|---|
| docboss | 85 ms |
| docx2txt | 808 ms |
| python-docx | 1,490 ms |
| mammoth | 5,917 ms |
| pandoc | 8,044 ms |
| docx2python | 8,780 ms |
A generated document of 3,000 sections (1,188 pages in docboss's layout), each a heading, a paragraph with bold and italic runs, a bullet list with a nested numbered list and a four-row table. docboss's output is the longest, 1.64 million characters against about 0.98 million for most others, because it keeps list labels and table cell separators.
| Engine | Time |
|---|---|
| catdoc | 54 ms |
| docboss | 105 ms |
| antiword | 2,552 ms |
catdoc wins, about 2× faster than docboss: it streams the text out of the piece table without building a document model, while docboss parses formatting, styles, 9,001 lists and tables into its model first.
| Engine | 100 files | Largest (2.96 MB) |
|---|---|---|
| docx2txt | 30.1 MB | 24.6 MB |
| docboss | 33.2 MB | 34.7 MB |
| docx2python | 55.0 MB | 30.3 MB |
| python-docx | 58.8 MB | 41.6 MB |
| mammoth | 82.2 MB | 32.7 MB |
| pandoc | 29.5 MB (child 134.8) | 29.4 MB (child 139.9) |
Peak RSS of a fresh process that imports the engine and extracts the text; the interpreter's own baseline is 22 to 35 MB. For command-line engines the child process's peak is given too. docx2txt peaks slightly lower than docboss.
| Engine | 100 files | Largest (6.24 MB) |
|---|---|---|
| catdoc | 23.3 MB (child 2.0) | 24.0 MB (child 2.9) |
| antiword | 23.8 MB (child 2.5) | 23.7 MB (child 2.9) |
| docboss | 40.4 MB | 63.5 MB |
antiword and catdoc win: they stream text in about 2 to 3 MB of their own, while docboss holds the whole document model and the file's streams. 12 MB of the largest file's peak are its metafile and bitmap pictures, decoded into the model to be drawn.
| Engine | Text | Error | Crash | Hang |
|---|---|---|---|---|
| docboss | 140 | 10 | 0 | 0 |
| docx2txt | 8 | 142 | 0 | 0 |
| docx2python | 4 | 146 | 0 | 0 |
| pandoc | 4 | 146 | 0 | 0 |
| python-docx | 2 | 148 | 0 | 0 |
| mammoth | 1 | 149 | 0 | 0 |
Byte flips, truncations, packages rebuilt with a damaged word/document.xml, and packages with a zeroed central directory. docboss reads a ZIP from its local headers when the central directory is damaged, keeps what inflated before a corrupt deflate stream, and balances broken XML.
| Engine | Text | Error | Crash | Hang |
|---|---|---|---|---|
| docboss | 142 | 8 | 0 | 0 |
| catdoc | 100 | 49 | 0 | 1 |
| antiword | 71 | 76 | 3 | 0 |
Byte flips, truncations and a 512-byte sector overwritten or zeroed, each file in its own process with a 30 s timeout. Of the 8 docboss refuses, 3 are encrypted originals and 5 have lost their FIB. antiword crashed with a segmentation fault 3 times.
| Format | Files | SSIM mean | SSIM median | SSIM p10 | MAD median | Page counts equal |
|---|---|---|---|---|---|---|
| DOCX | 59 | 0.952 | 0.985 | 0.879 | 0.55 | 56 of 60 |
| DOC | 29 | 0.898 | 0.966 | 0.634 | 1.83 | 27 of 30 |
LibreOffice converts each document to PDF, which is rasterized, and docboss renders the same pages from the document itself, both at scale 1.5; each pair is scored with windowed SSIM and mean absolute pixel difference, up to 5 pages per file. Mostly white pages score high whatever their content, so the p10 column is the honest one: a tenth of the DOC files score under 0.64.
Method
How the numbers were taken.
Corpora. The real-world test documents of LibreOffice, Apache POI and python-docx, fetched at pinned revisions by fetch_corpus.sh: 1,697 DOCX and 180 DOC files after removing duplicates. Each script takes a deterministic, evenly spaced sample.
Timing. Best of 3 per file after one warm-up pass, so the file cache and imports are hot, aggregated only over the files every engine handled, so each total compares the same workload.
Recall. LibreOffice converts each file to text once. An engine’s recall on a file is the share of the reference’s words, with multiplicity, found in its output. Files whose reference has no words are skipped, since LibreOffice’s text export leaves out text boxes. On DOC, the weighted recall of catdoc and antiword is dominated by one 74,813-word Russian novel that both decode to 0.1% of its words, so their mean recall is the fairer comparison.
LibreOffice converts the same samples at 16.1 (DOCX) and 14.8 (DOC) files/s as one batch, and takes a median 827 ms (DOCX) and 725 ms (DOC) per file one process at a time.
Versions. docboss 0.1.0 (release build of master), python-docx 1.2.0, docx2python 3.7.1, mammoth 1.13.0, docx2txt 0.9, pandoc 3.9 through pypandoc_binary 1.17, antiword 0.37, catdoc 0.95, LibreOffice 26.8.0.3. Apple M3 Pro, 12 cores, macOS 26.4, Python 3.12.10, one session on 2026-09-30.
Numbers depend on the machine; compare rows against each other. Every script writes its results-*.json with the machine, versions and date, and benchmarks/README.md has the commands to reproduce each table.
What is not measured
Where the tables stop.
- Markdown and HTML are timed but not scored for quality against the other engines' output; recall scores plain text only.
- Rendering is compared with LibreOffice only.
- DOCX writing and the async remote reader are not timed.