Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Extracting images

docboss images report.docx -o images
docboss images legacy.doc -o images

images writes every picture the document carries at its stored size and in its stored format (PNG, JPEG, GIF, BMP, EMF, WMF, TIFF), named after its part (image1.png). DOC pictures come from the Data stream and the OfficeArt BLIP store; DIB pictures are written as BMP, and identical DOC pictures are stored once.

from pathlib import Path

import docboss

for image in docboss.Document("report.docx").images():
    Path(image.file_name).write_bytes(image.data)
    print(image.name, image.content_type)

name is the part name inside the package (word/media/image1.png), file_name its last segment, and content_type is detected from the bytes. The Markdown and HTML writers reference or embed the same images (see Markdown, HTML and JSON).