Files

OCRmyPDF batch web interface

A small FastAPI application that lets people drop several PDFs or images onto a web page, OCRs them in the background, and hands back the results individually or as one zip archive.

This is separate from misc/_webservice.py, the Streamlit app, which handles one file at a time and exposes every OCRmyPDF option. This one trades option coverage for batch throughput.

Running it

With Docker (recommended — Tesseract, Ghostscript and friends are already installed):

docker compose -f misc/docker-compose.webui.yml up --build
# then open http://localhost:8000/

From a source checkout:

uv sync --extra webui
uv run python -m webui

Configuration

All settings come from environment variables; see the table in docs/docker.md. The two that matter most:

  • OCRMYPDF_WEBUI_WORKERS — how many files are OCR'd at the same time.
  • OCRMYPDF_WEBUI_OCR_JOBSocrmypdf --jobs for each of those files.

Their product should be about the number of cores available. Defaults split the machine automatically.

How it works

browser ──POST /api/batches──> FastAPI ──> ThreadPoolExecutor
   │                                          │
   │                                          └─> subprocess: python -m ocrmypdf
   └──GET /api/batches/{id} (poll ~1.5s)──> per-file status + log tail
  • A batch is one submission: N files plus one set of options. Each file becomes a job with its own status, so one bad file does not sink the batch.
  • OCRmyPDF runs out of process. A crash, hang, or runaway allocation stays in a child process that the server can time out and kill.
  • Batches are deleted from disk once their TTL lapses, whether or not they were downloaded. A background thread sweeps every 60 seconds and also removes directories left behind by a previous process.
  • State lives in memory, so the server runs as one process. Scale with OCRMYPDF_WEBUI_WORKERS, not with server workers.

Notes on input handling

User input never reaches a command line unchecked:

  • Options are parsed by a Pydantic model with extra="forbid"; modes and output types are enums, numbers are bounded, and languages must appear in tesseract --list-langs output for this container.
  • Uploaded filenames are used only as display text and as the filename of a download. Files on disk get generated names, and the positional arguments to ocrmypdf are preceded by --.
  • Uploads are streamed to disk and aborted past the size limit, so an oversized request is not buffered in memory.

There is deliberately no authentication. Put this behind a reverse proxy if it needs to be reachable from anywhere untrusted.

API

Method Path Purpose
GET /api/config Limits, installed languages, defaults
POST /api/batches Upload files (files) + options (options, JSON)
GET /api/batches/{id} Poll batch and per-file status
DELETE /api/batches/{id} Cancel and delete immediately
GET /api/batches/{id}/files/{n} Download one result
GET /api/batches/{id}/files/{n}/log ocrmypdf output for one file
GET /api/batches/{id}/download Zip of all successful results
GET /healthz Liveness probe

Interactive docs are at /api/docs.

Example:

curl -sS -X POST http://localhost:8000/api/batches \
  -F files=@scan1.pdf -F files=@scan2.pdf \
  -F 'options={"languages":["eng","fra"],"mode":"skip-text","deskew":true}'

Tests

tests/test_webui.py starts a real uvicorn server and runs real OCR. It skips itself unless the webui extra is installed:

uv sync --extra webui --group test
uv run pytest tests/test_webui.py

Licensing

OCRmyPDF uses Ghostscript, which is AGPLv3+. This subpackage is distributed under AGPLv3+ (rather than OCRmyPDF's MPL-2.0) to make it plain that SaaS deployments must comply with Ghostscript's terms.