Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Remote documents over HTTP

Every read command accepts an http(s):// URL:

docboss text https://example.com/report.docx
docboss info https://example.com/legacy.doc

The file is not downloaded whole. For a DOCX, docboss fetches the end of the file to find the ZIP central directory, then only the entries the reader needs: content types, relationships, document.xml, styles, numbering, notes and comments. Images are fetched only when a command needs them. For a DOC, it fetches the compound-file header, the FAT and directory sectors, and then only the sectors of the Word streams. Requests are coalesced and cached, so no byte is fetched twice; a server that ignores Range costs one full download. On a terminal, a line of coverage shows which parts of the file were fetched.

On Apache POI's saut_page.docx (2.96 MB, mostly images), extracting the text fetched 17.9 KB in 3 requests; a DOC reads its Data stream too, so it saves less.

Rust

docboss-aio with the http feature:

use docboss_aio::{AsyncDocument, ReadOptions};
use docboss_output::{to_text, TextOptions};

async fn run() -> docboss_aio::Result<()> {
let remote = AsyncDocument::open_url("https://example.com/report.docx").await?;
let doc = remote.read(&ReadOptions::default()).await?;
println!("{}", to_text(&doc, &TextOptions::default()));
println!("{} of {} bytes in {} requests", remote.bytes_fetched(), remote.len(), remote.requests().len());
Ok(()) }

ReadOptions { media: true, .. } fetches images too; otherwise AsyncDocument::load_media fetches them later. AsyncDocument::open and from_bytes give the same interface over local files and memory, and from_backend over any byte source that implements Backend.

Python

import asyncio

import docboss

async def main() -> None:
    doc = await docboss.AsyncDocument.open_url("https://example.com/report.docx")
    print(doc.format, doc.part_names()[:3])
    text = await doc.extract_text()
    markdown = await doc.extract_markdown(images="omit")
    styles = await doc.part("word/styles.xml")
    print(doc.bytes_fetched, "bytes in", len(doc.requests), "requests")

asyncio.run(main())

AsyncDocument.open(path), from_bytes(data) and open_url(url) are coroutines and take password=. The extraction methods mirror Document's and are awaitable: extract_text, extract_markdown, extract_html, to_json, blocks, images, and part(name) for one ZIP entry or stream. bytes_fetched and requests report what was read so far, and await doc.document() returns a synchronous Document for rendering, DOCX writing and the rest of its API; images are fetched when a call needs them.