Split the hOCR transformation code into three distinct layers:
1. ocr_element.py - Generic OcrElement dataclass that represents OCR
output structure from any source (hOCR, ALTO, custom engines).
Includes helper classes: BoundingBox, Baseline, FontInfo.
2. hocr_parser.py - HocrParser class that parses hOCR XML files into
OcrElement trees, extracting bbox, baseline, textangle, confidence,
font info, direction, and language.
3. pdf_renderer.py - PdfTextRenderer class that renders OcrElement
trees to PDF text layers, handling text positioning, baseline
rotation, LTR/RTL, and word break injection.
The existing HocrTransform class is preserved for backward compatibility,
now delegating to the new components internally.
This separation enables:
- Support for non-hOCR OCR output formats
- Independent improvements to text rendering
- Reuse of OcrElement for other purposes
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>