With garbage collection it reduces waste on the worst case file.
That's nice. 1 MB -> 105 MB -> 1.5 MB.
Indicates really problem is using PyPDF2 to watermark.
Currently hacked into --output-type pdf.
Previously page splitting occurred in a single process because it was
not believed to affect performance much. It turned out to be an expensive
operation.
It now scales better with large page sizes although this has a negative
effect on small files.
Overall time changes as follows:
7 page file, 9.02s -> 9.56s
731 page file, 213s -> 97s
WITH --tesseract-timeout 0 --output-type pdf --skip-text
i.e. you don't get a 2.2x speed gain when OCR is available.
Squashed a commit to ix test suite failure on --rotate-pages
Squashed a commit to remove debug code
Here we are manually scaling the pt width used for the BoundingBox and
the Text element when manually adding whitespace to account for
limitations of the PDF.js viewer. This fixes an initial regression
noticed when selecting text elements in Chrome and PDFium. The width
of the Text element and BoundBox had not been adjusted for the
additional whitespace so the highlighting was offset slightly.
This commit includes an optional work around for limitations of the
PDF.js viewer described in
https://github.com/jbarlow83/OCRmyPDF/issues/133. Here is explicitly
add an addition space to text elements before drawing them on the PDF
canvas when using the HOCR renderer. This option does not apply to
other pdf renderers in OCRmyPDF and is turned off by default.
Previously we threw an exception if the output name was a directory (only after doing OCR) and would trigger a PermissionError on trying to flip permission bits of /dev/null due to shutil.copyfile implementation. Instead of copying file use shutil.copyfileobj which should also respect umask etc.
If OCR is skipped due to --tesseract-timeout or similar, and the skip page is rotated with /Rotate, and the skip page was deskewed or had other image processing, then the skip page was created with the wrong dimensions causing the output page to be cropped.