docboss.dev

Compare

docboss vs mammoth and pandoc.

mammoth is a pure-Python DOCX converter, and pandoc a Haskell document converter, here driven through pypandoc_binary. All three turn a DOCX into Markdown or HTML. docboss does it about 30 times faster than mammoth and 440 times faster than pandoc on the benchmark corpus, and reads .doc files through the same call.

At a glance
docbossmammothpandoc
EngineSafe Rust, no C dependenciesPure PythonHaskell binary
Formats measuredDOCX and DOCDOCXDOCX
Markdown3,324 files/s109 files/s7.6 files/s
HTML3,355 files/s111 files/s7.5 files/s
Text3,304 files/s117 files/s7.4 files/s
Word recall (mean)0.9920.9560.894
Parallel, 12 workers13,839 files/s316 files/s (processes)41 files/s (threads)
One 1,188-page document85 ms5,917 ms8,044 ms
Peak memory, 100 files33.2 MB82.2 MB29.5 MB, child 134.8 MB
150 damaged DOCX filesText from 140Text from 1Text from 4

mammoth 1.13.0 and pandoc 3.9, measured by the docboss benchmark harness in one session. Speed rows count only the files every engine handled.

Migration

Markdown and HTML, side by side.

docboss opens the file once and both conversions run on the same parsed model. Headings come from outline levels and heading styles, lists nest by numbering level, and footnotes become [^1] references. More on the Markdown output →

before.py · mammoth
import mammoth

with open("report.docx", "rb") as handle:
    html = mammoth.convert_to_html(handle).value
with open("report.docx", "rb") as handle:
    markdown = mammoth.convert_to_markdown(handle).value
after.py · docboss
import docboss

doc = docboss.Document("report.docx")   # or "legacy.doc"
html = doc.extract_html()               # images embedded as data: URIs
markdown = doc.extract_markdown(images="reference", image_prefix="images/")

Where they are ahead

Where the others win.

  • pandoc's own process stays small on memory: 29.5 MB for the Python side against docboss's 33.2 MB, though its child process peaks at 134.8 MB.

Where docboss is ahead

What docboss adds.

  • About 30 times mammoth's speed and 440 times pandoc's on Markdown and HTML.
  • Legacy .doc files, encrypted DOCX and DOC, and remote files over HTTP range requests.
  • Pages rendered to PNG, PPM, BMP or JPEG, and DOCX written back out.
  • Damaged files: text from 140 of 150, where mammoth returned it for 1 and pandoc for 4.

Questions

How much faster is docboss than mammoth?
On the benchmark's 80 DOCX files, docboss converts to Markdown at 3,324 files per second against mammoth's 109 and to HTML at 3,355 against 111, about 30 times faster. pandoc does 7.6 and 7.5.
What HTML does docboss write?
Semantic HTML: docboss writes h1 to h6, p, strong, em, u, s, sup, sub, lists with start attributes, tables with colspan and rowspan from merged cells, links, images and a footnotes section, as a body fragment or with --standalone a whole document.
Does docboss read .doc files?
Yes, through the same calls. mammoth and pandoc were measured on DOCX only.