docboss.dev

DOCX to Markdown

Word to Markdown, three ways.

docboss turns a .docx or .doc file into GitHub-flavored Markdown: headings from styles, bold, italic and strikethrough merged across runs, nested lists, pipe tables, footnotes and endnotes, links and images, in one pass. The same Rust engine answers from Python, from Rust and from the command line.

Command line

One command, one file.

cargo install docboss-cli puts the docboss binary on your path. md writes to stdout unless -o names a file, and takes an http(s):// URL as well as a path.

~/reports
$ docboss md report.docx > report.md
$ docboss md report.docx --images images -o report.md   # export images and link them
$ docboss md report.docx --embed-images                 # data: URIs instead
$ docboss md report.docx --title --page-breaks
$ docboss md legacy.doc                                 # Word 97-2003 too
$ docboss convert legacy.doc -o legacy.md

Output

What a real file gives.

The Markdown of libreoffice-rich.docx, a test fixture written by LibreOffice: a Heading 1 and a Heading 2, a bold run, a bulleted list, a table with a merged cell, a footnote and an endnote.

report.md
$ docboss md libreoffice-rich.docx
# Fixture Heading

Text with **bold words** and a note[^1] here.

Commented word.

Picture:

- Item one
- Item two

| A1 | B1 |
| --- | --- |
| Span |  |

## Endnote section

Ends here[^i].

[^1]: Footnote text.

[^i]: Endnote text.

Python

One method on the document.

pip install docboss installs an abi3 wheel for CPython 3.12 and later. extract_markdown also takes headers_footers, notes and comments.

convert.py
import docboss

doc = docboss.Document("report.docx")
markdown = doc.extract_markdown(images="reference", image_prefix="images/")
inline = doc.extract_markdown(images="embed")
bare = doc.extract_markdown(images="omit", title=True, page_breaks=True)

Rust

Two crates.

docboss-core opens the document and docboss-output writes it; write_markdown writes to any io::Write.

main.rs
use docboss_output::{to_markdown, MarkdownOptions};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let doc = docboss_core::open("report.docx")?;
    println!("{}", to_markdown(&doc, &MarkdownOptions::default()));
    Ok(())
}
and back
$ docboss create md notes.md -o notes.docx --size a4

What is preserved

What ends up in the file.

  • Headings from the paragraph's outline level or a Heading N or Title style.
  • Bold, italic and strikethrough spans merged across adjacent runs, with whitespace kept outside the markers.
  • A paragraph set entirely in a monospace font, or in a code style, as a fenced code block.
  • Lists nested by numbering level: bullets as -, numbered items as 1.
  • Tables as pipe tables, with merged cells, nested tables and multi-paragraph cells flattened with <br>.
  • Footnotes and endnotes as [^1] references with definitions at the end; hyperlinks and HYPERLINK fields as links.
  • Images as ![description](name), exported, embedded or omitted.
3,324 files/s

DOCX to Markdown over 80 real-world files; mammoth does 109, pandoc 7.6.

0 C dependencies

Its own ZIP, XML and compound-file readers; abi3 wheels for CPython 3.12+.

Questions

How do I convert a DOCX file to Markdown in Python?
pip install docboss, then docboss.Document("report.docx").extract_markdown() returns the document as GitHub-flavored Markdown. The CLI equivalent is docboss md report.docx.
Where do the headings come from?
From the paragraph's outline level or a Heading N or Title style. A document that sets its headings as bold Normal paragraphs gives bold paragraphs, not headings.
What happens to tables?
They become pipe tables. Merged cells, nested tables and multi-paragraph cells are flattened, with <br> between the paragraphs of a cell.
What happens to images?
Each image becomes ![description](name). --images DIR on the CLI, or images="reference" with image_prefix in Python, writes or links the files; --embed-images or images="embed" inlines them as data: URIs, and images="omit" leaves them out.
Does it work on old .doc files?
Yes. A Word 97-2003 file goes through the same model as a DOCX, so docboss md legacy.doc and Document("legacy.doc").extract_markdown() give the same kind of output.
Can it go the other way?
Yes. docboss create md notes.md -o notes.docx and docboss.md.to_docx() compose CommonMark and GFM into a DOCX: headings to Heading1 through Heading6, lists to numbering definitions, tables with a header row, block quotes to the Quote style, footnotes to footnotes.