Compare
docboss vs mammoth and pandoc.
mammoth is a pure-Python DOCX converter, and pandoc a Haskell document converter, here driven through pypandoc_binary. All three turn a DOCX into Markdown or HTML. docboss does it about 30 times faster than mammoth and 440 times faster than pandoc on the benchmark corpus, and reads .doc files through the same call.
| docboss | mammoth | pandoc | |
|---|---|---|---|
| Engine | Safe Rust, no C dependencies | Pure Python | Haskell binary |
| Formats measured | DOCX and DOC | DOCX | DOCX |
| Markdown | 3,324 files/s | 109 files/s | 7.6 files/s |
| HTML | 3,355 files/s | 111 files/s | 7.5 files/s |
| Text | 3,304 files/s | 117 files/s | 7.4 files/s |
| Word recall (mean) | 0.992 | 0.956 | 0.894 |
| Parallel, 12 workers | 13,839 files/s | 316 files/s (processes) | 41 files/s (threads) |
| One 1,188-page document | 85 ms | 5,917 ms | 8,044 ms |
| Peak memory, 100 files | 33.2 MB | 82.2 MB | 29.5 MB, child 134.8 MB |
| 150 damaged DOCX files | Text from 140 | Text from 1 | Text from 4 |
mammoth 1.13.0 and pandoc 3.9, measured by the docboss benchmark harness in one session. Speed rows count only the files every engine handled.
Migration
Markdown and HTML, side by side.
docboss opens the file once and both conversions run on the same parsed model. Headings come from outline levels and heading styles, lists nest by numbering level, and footnotes become [^1] references. More on the Markdown output →
before.py · mammoth
import mammoth
with open("report.docx", "rb") as handle:
html = mammoth.convert_to_html(handle).value
with open("report.docx", "rb") as handle:
markdown = mammoth.convert_to_markdown(handle).valueafter.py · docboss
import docboss
doc = docboss.Document("report.docx") # or "legacy.doc"
html = doc.extract_html() # images embedded as data: URIs
markdown = doc.extract_markdown(images="reference", image_prefix="images/")Where they are ahead
Where the others win.
- pandoc's own process stays small on memory: 29.5 MB for the Python side against docboss's 33.2 MB, though its child process peaks at 134.8 MB.
Where docboss is ahead
What docboss adds.
- About 30 times mammoth's speed and 440 times pandoc's on Markdown and HTML.
- Legacy .doc files, encrypted DOCX and DOC, and remote files over HTTP range requests.
- Pages rendered to PNG, PPM, BMP or JPEG, and DOCX written back out.
- Damaged files: text from 140 of 150, where mammoth returned it for 1 and pandoc for 4.
Questions
- How much faster is docboss than mammoth?
- On the benchmark's 80 DOCX files, docboss converts to Markdown at 3,324 files per second against mammoth's 109 and to HTML at 3,355 against 111, about 30 times faster. pandoc does 7.6 and 7.5.
- What HTML does docboss write?
- Semantic HTML: docboss writes h1 to h6, p, strong, em, u, s, sup, sub, lists with start attributes, tables with colspan and rowspan from merged cells, links, images and a footnotes section, as a body fragment or with --standalone a whole document.
- Does docboss read .doc files?
- Yes, through the same calls. mammoth and pandoc were measured on DOCX only.