docboss.dev

Compare

docboss vs python-docx.

Both read Word files from Python. python-docx is the pure-Python library most people start with, and its object model lets you change a document and save it. docboss is written from scratch in Rust, reads DOCX and DOC, includes headers, footers, notes and comments in its text, renders pages, and is about 4.6 times faster on the reading path. All numbers below come from one benchmark run.

At a glance
docbosspython-docx
EngineWritten from scratch in safe Rust; no C dependenciesPure Python
Formats.docx, .docm, .dotx, .dotm, .doc and .dot.docx
Text extraction3,304 files/s725 files/s
Open and parse3,785 files/s1,406 files/s
Word recall (mean)0.9920.848
Text the benchmark readsBody, tables, text boxes, headers, footers, footnotes and endnotesBody paragraphs and table cells
Parallel, 12 workers13,839 files/s (extract_texts)1,082 files/s (processes)
One 1,188-page document85 ms1,490 ms
Peak memory, 100 files33.2 MB58.8 MB
150 damaged DOCX filesText from 140Text from 2
Markdown, HTML, JSONYesNot in the benchmark
Page renderingPNG, PPM, BMP, JPEGNot in the benchmark

Every row comes from the docboss benchmark harness: python-docx 1.2.0, the same 80 sampled files, the same machine and session.

Migration

The reading path, side by side.

extract_text walks the whole document, so the paragraph and table loops go away, and list paragraphs keep their labels. For per-paragraph work, blocks() gives each paragraph’s kind, style, heading level and list label.

before.py · python-docx
from docx import Document

doc = Document("report.docx")
for paragraph in doc.paragraphs:
    print(paragraph.text)
for table in doc.tables:
    for row in table.rows:
        print("\t".join(cell.text for cell in row.cells))
after.py · docboss
import docboss

doc = docboss.Document("report.docx")   # or "legacy.doc"
print(doc.extract_text(headers_footers=True, comments=True))

for block in doc.blocks():              # heading, list_item, paragraph, table
    print(block.kind, block.text)

Where python-docx is ahead

What python-docx does that docboss does not.

  • Changing an existing document and saving it. docboss does not edit a document in place; it writes fresh DOCX files from its model, from Markdown or from a Rust builder.
  • A Python object model for building documents. docboss's builder is Rust only; from Python it composes DOCX from Markdown.

Where docboss is ahead

What docboss does that python-docx does not.

  • Legacy .doc files through the same API, and password-protected DOCX and DOC.
  • Headers, footers, footnotes, endnotes, comments and text boxes in the text, with the document's own list labels.
  • Markdown, HTML, JSON and a block view, and pages rendered to images.
  • Damaged files: text from 140 of 150 where python-docx returned it for 2.
  • Remote files read with HTTP range requests.

Questions

Is docboss a drop-in replacement for python-docx?
No. The APIs differ: python-docx exposes paragraphs, runs and tables as objects you can change and save, while docboss reads a document into an immutable model and writes fresh DOCX files from it. Moving the reading path over is a few lines per call site.
Is docboss faster than python-docx?
On the benchmark's 80 real-world DOCX files, docboss extracts text at 3,304 files per second against python-docx's 725, about 4.6 times faster, and opens and parses at 3,785 against 1,406. On one generated 1,188-page document it takes 85 ms against 1,490 ms.
Why is python-docx's recall lower?
The benchmark's python-docx adapter reads body paragraphs and table cells, which is what its API exposes, so headers, footers, footnotes, comments and text boxes are left out. Its mean word recall against LibreOffice is 0.848; docboss's is 0.992.
Does docboss read .doc files?
Yes. Word 97-2003 files and the formatting of Word 6 and 95 files are read into the same model, through the same Document class. The benchmark measures python-docx on DOCX only.