Python
Fast Word document reading in Python.
pip install docboss puts a DOCX and DOC engine written in Rust behind a small Python API: text with list labels, Markdown, HTML, JSON and a block view, page rendering, images, DOCX writing, Markdown to DOCX, encrypted files and an asyncio reader for remote files. Heavy calls release the GIL. abi3 wheels for CPython 3.12 and later on Linux x86_64 and macOS arm64, MIT OR Apache-2.0.
_docboss.pyi, py.typed).Extract
Text, Markdown, HTML, JSON and blocks.
Document opens a path or data= bytes and detects the format from them. Documents are immutable and can be used from any thread. Images in Markdown are "reference", "embed" or "omit". Bad input, a wrong password and unsupported formats raise DocbossError.
import docboss
doc = docboss.Document("report.docx") # or Document(data=raw_bytes), password="..."
print(doc.format, doc.metadata.title) # "docx" or "doc"
text = doc.extract_text(headers_footers=True, notes=True, comments=False, list_labels=True)
md = doc.extract_markdown(images="reference", image_prefix="images/")
html = doc.extract_html(standalone=True)
model = doc.to_json(pretty=True)
for block in doc.blocks(): # paragraph, heading, list_item or table
print(block.kind, block.heading_level, block.list_label, block.text)Many files
One call, a Rust thread pool.
extract_texts reads every file on a Rust thread pool without the GIL and returns the texts in input order. With strict=True, the default, the first failure raises; with strict=False a failed file gives None.
texts = docboss.extract_texts(["report.docx", "legacy.doc"], threads=8)
lenient = docboss.extract_texts(["report.docx", "not-a-document.txt"], strict=False)
print(lenient[1]) # None: that file failed, the others still came backRender
Pages as PNG, PPM, BMP or JPEG bytes.
Pages index from 0, negative from the end. render_pages renders every page, or the indexes given in pages=, in parallel. Layout runs once per document and is cached, so rendering pages one by one costs no second layout.
from pathlib import Path
print(doc.page_count(), "pages")
Path("first.png").write_bytes(doc.render(0, scale=2.0))
Path("last.jpg").write_bytes(doc.render(-1, format="jpeg", jpeg_quality=85))
for index, png in enumerate(doc.render_pages(scale=1.5)): # in parallel
Path(f"page-{index + 1}.png").write_bytes(png)
for image in doc.images(): # name, file_name, content_type, data
Path(image.file_name).write_bytes(image.data)Write and report
DOCX out, and what the reader approximated.
to_docx writes a fresh, deterministic DOCX, also for a .doc input. docboss.md.to_docx composes CommonMark and GFM into DOCX bytes. diagnostics lists every item the lenient reader approximated or dropped.
docboss.Document("legacy.doc").save_docx("legacy.docx")
data = docboss.Document("report.docx").to_docx()
docx = docboss.md.to_docx("# Notes\n\n- one\n- two\n", images_dir="notes/")
for diagnostic in doc.diagnostics: # "approximated" or "dropped"
print(diagnostic.severity, diagnostic.location, diagnostic.message)Async and remote
asyncio, fetching only the bytes it needs.
AsyncDocument.open, from_bytes and open_url are coroutines. The extraction methods mirror Document’s and are awaitable, and await doc.document() returns a synchronous Document for rendering and DOCX writing.
import asyncio
import docboss
async def main() -> None:
doc = await docboss.AsyncDocument.open_url("https://example.com/report.docx")
print(doc.format, doc.part_names()[:3])
text = await doc.extract_text()
markdown = await doc.extract_markdown(images="omit")
styles = await doc.part("word/styles.xml")
print(doc.bytes_fetched, "bytes in", len(doc.requests), "requests")
asyncio.run(main())Speed
Why it is fast.
DOCX text over 80 real-world files, 4.6× python-docx.
Through extract_texts on 12 cores, 300 DOCX files.
DOCX to Markdown, about 30× mammoth.
The work happens in Rust with the GIL released, so documents extract and render in parallel from Python threads. Method and full tables →
Scope
What it does not do.
- Editing a document in place. docboss reads documents and writes fresh DOCX files from its model.
- Running macros.
- Writing encrypted DOCX; a decrypted document is written as a plain one.
- Password-protected Word 6 and 95 files are refused with an error.