Compare
docboss vs python-docx.
Both read Word files from Python. python-docx is the pure-Python library most people start with, and its object model lets you change a document and save it. docboss is written from scratch in Rust, reads DOCX and DOC, includes headers, footers, notes and comments in its text, renders pages, and is about 4.6 times faster on the reading path. All numbers below come from one benchmark run.
| docboss | python-docx | |
|---|---|---|
| Engine | Written from scratch in safe Rust; no C dependencies | Pure Python |
| Formats | .docx, .docm, .dotx, .dotm, .doc and .dot | .docx |
| Text extraction | 3,304 files/s | 725 files/s |
| Open and parse | 3,785 files/s | 1,406 files/s |
| Word recall (mean) | 0.992 | 0.848 |
| Text the benchmark reads | Body, tables, text boxes, headers, footers, footnotes and endnotes | Body paragraphs and table cells |
| Parallel, 12 workers | 13,839 files/s (extract_texts) | 1,082 files/s (processes) |
| One 1,188-page document | 85 ms | 1,490 ms |
| Peak memory, 100 files | 33.2 MB | 58.8 MB |
| 150 damaged DOCX files | Text from 140 | Text from 2 |
| Markdown, HTML, JSON | Yes | Not in the benchmark |
| Page rendering | PNG, PPM, BMP, JPEG | Not in the benchmark |
Every row comes from the docboss benchmark harness: python-docx 1.2.0, the same 80 sampled files, the same machine and session.
Migration
The reading path, side by side.
extract_text walks the whole document, so the paragraph and table loops go away, and list paragraphs keep their labels. For per-paragraph work, blocks() gives each paragraph’s kind, style, heading level and list label.
from docx import Document
doc = Document("report.docx")
for paragraph in doc.paragraphs:
print(paragraph.text)
for table in doc.tables:
for row in table.rows:
print("\t".join(cell.text for cell in row.cells))import docboss
doc = docboss.Document("report.docx") # or "legacy.doc"
print(doc.extract_text(headers_footers=True, comments=True))
for block in doc.blocks(): # heading, list_item, paragraph, table
print(block.kind, block.text)Where python-docx is ahead
What python-docx does that docboss does not.
- Changing an existing document and saving it. docboss does not edit a document in place; it writes fresh DOCX files from its model, from Markdown or from a Rust builder.
- A Python object model for building documents. docboss's builder is Rust only; from Python it composes DOCX from Markdown.
Where docboss is ahead
What docboss does that python-docx does not.
- Legacy .doc files through the same API, and password-protected DOCX and DOC.
- Headers, footers, footnotes, endnotes, comments and text boxes in the text, with the document's own list labels.
- Markdown, HTML, JSON and a block view, and pages rendered to images.
- Damaged files: text from 140 of 150 where python-docx returned it for 2.
- Remote files read with HTTP range requests.
Questions
- Is docboss a drop-in replacement for python-docx?
- No. The APIs differ: python-docx exposes paragraphs, runs and tables as objects you can change and save, while docboss reads a document into an immutable model and writes fresh DOCX files from it. Moving the reading path over is a few lines per call site.
- Is docboss faster than python-docx?
- On the benchmark's 80 real-world DOCX files, docboss extracts text at 3,304 files per second against python-docx's 725, about 4.6 times faster, and opens and parses at 3,785 against 1,406. On one generated 1,188-page document it takes 85 ms against 1,490 ms.
- Why is python-docx's recall lower?
- The benchmark's python-docx adapter reads body paragraphs and table cells, which is what its API exposes, so headers, footers, footnotes, comments and text boxes are left out. Its mean word recall against LibreOffice is 0.848; docboss's is 0.992.
- Does docboss read .doc files?
- Yes. Word 97-2003 files and the formatting of Word 6 and 95 files are read into the same model, through the same Document class. The benchmark measures python-docx on DOCX only.