Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Extracting text

docboss text report.docx
docboss text report.docx --headers-footers --comments
docboss text legacy.doc --no-labels -o legacy.txt

Plain text has one line per paragraph. List paragraphs start with the label the document's numbering gives them (1., a), •), computed from the list definitions, their start values and overrides, the way Word numbers them. Table rows become one line each with tab-separated cells. Footnote and endnote references become [1] and [i], and the notes' text follows the main text; --no-notes leaves both out. --headers-footers writes each section's headers before it and its footers after it, and --comments marks comment anchors [c1] and appends the comments with their authors.

What is left out, as ECMA-376 and [MS-DOC] define it: hidden text (vanish), deleted revisions (inserted ones are kept), and field instructions; a field contributes its cached result, so a table of contents or a PAGE field reads as the file shows it. Text boxes are part of the text, written after the paragraph that anchors them.

import docboss

doc = docboss.Document("report.docx")
text = doc.extract_text(headers_footers=True, notes=True, comments=False, list_labels=True)

Many files at once

import docboss

texts = docboss.extract_texts(["report.docx", "legacy.doc"], threads=8)
lenient = docboss.extract_texts(["report.docx", "not-a-document.txt"], strict=False)
print(lenient[1])   # None: that file failed, the others still came back

extract_texts reads every file on a Rust thread pool in one call, without the GIL, and returns the texts in input order. With strict=True (the default) the first failure raises DocbossError; with strict=False a failed file gives None. It is the fastest way to feed a pipeline.

Blocks

For retrieval pipelines, the block view keeps each paragraph's role:

docboss json report.docx --blocks
import docboss

for block in docboss.Document("report.docx").blocks():
    print(block.kind, block.heading_level, block.list_label, block.text)

kind is paragraph, heading, list_item or table; headings come from outline levels and heading styles, list items carry their level and label. Empty paragraphs are left out.