OCRmyPDF batch web interface
A small FastAPI application that lets people drop several PDFs or images onto a web page, OCRs them in the background, and hands back the results individually or as one zip archive.
This is separate from misc/_webservice.py, the Streamlit app, which
handles one file at a time and exposes every OCRmyPDF option. This one
trades option coverage for batch throughput.
Running it
With Docker (recommended — Tesseract, Ghostscript and friends are already installed):
docker compose -f misc/docker-compose.webui.yml up --build
# then open http://localhost:8000/
From a source checkout:
uv sync --extra webui
uv run python -m webui
Configuration
All settings come from environment variables; see the table in
docs/docker.md. The two that matter most:
OCRMYPDF_WEBUI_WORKERS— how many files are OCR'd at the same time.OCRMYPDF_WEBUI_OCR_JOBS—ocrmypdf --jobsfor each of those files.
Their product should be about the number of cores available. Defaults split the machine automatically.
How it works
browser ──POST /api/batches──> FastAPI ──> ThreadPoolExecutor
│ │
│ └─> subprocess: python -m ocrmypdf
└──GET /api/batches/{id} (poll ~1.5s)──> per-file status + log tail
- A batch is one submission: N files plus one set of options. Each file becomes a job with its own status, so one bad file does not sink the batch.
- OCRmyPDF runs out of process. A crash, hang, or runaway allocation stays in a child process that the server can time out and kill.
- Batches are deleted from disk once their TTL lapses, whether or not they were downloaded. A background thread sweeps every 60 seconds and also removes directories left behind by a previous process.
- State lives in memory, so the server runs as one process. Scale
with
OCRMYPDF_WEBUI_WORKERS, not with server workers.
Notes on input handling
User input never reaches a command line unchecked:
- Options are parsed by a Pydantic model with
extra="forbid"; modes and output types are enums, numbers are bounded, and languages must appear intesseract --list-langsoutput for this container. - Uploaded filenames are used only as display text and as the
filenameof a download. Files on disk get generated names, and the positional arguments toocrmypdfare preceded by--. - Uploads are streamed to disk and aborted past the size limit, so an oversized request is not buffered in memory.
There is deliberately no authentication. Put this behind a reverse proxy if it needs to be reachable from anywhere untrusted.
API
| Method | Path | Purpose |
|---|---|---|
GET |
/api/config |
Limits, installed languages, defaults |
POST |
/api/batches |
Upload files (files) + options (options, JSON) |
GET |
/api/batches/{id} |
Poll batch and per-file status |
DELETE |
/api/batches/{id} |
Cancel and delete immediately |
GET |
/api/batches/{id}/files/{n} |
Download one result |
GET |
/api/batches/{id}/files/{n}/log |
ocrmypdf output for one file |
GET |
/api/batches/{id}/download |
Zip of all successful results |
GET |
/healthz |
Liveness probe |
Interactive docs are at /api/docs.
Example:
curl -sS -X POST http://localhost:8000/api/batches \
-F files=@scan1.pdf -F files=@scan2.pdf \
-F 'options={"languages":["eng","fra"],"mode":"skip-text","deskew":true}'
Tests
tests/test_webui.py starts a real uvicorn server and runs real OCR. It
skips itself unless the webui extra is installed:
uv sync --extra webui --group test
uv run pytest tests/test_webui.py
Licensing
OCRmyPDF uses Ghostscript, which is AGPLv3+. This subpackage is distributed under AGPLv3+ (rather than OCRmyPDF's MPL-2.0) to make it plain that SaaS deployments must comply with Ghostscript's terms.