With garbage collection it reduces waste on the worst case file.
That's nice. 1 MB -> 105 MB -> 1.5 MB.
Indicates really problem is using PyPDF2 to watermark.
Currently hacked into --output-type pdf.
Previously page splitting occurred in a single process because it was
not believed to affect performance much. It turned out to be an expensive
operation.
It now scales better with large page sizes although this has a negative
effect on small files.
Overall time changes as follows:
7 page file, 9.02s -> 9.56s
731 page file, 213s -> 97s
WITH --tesseract-timeout 0 --output-type pdf --skip-text
i.e. you don't get a 2.2x speed gain when OCR is available.
Squashed a commit to ix test suite failure on --rotate-pages
Squashed a commit to remove debug code
Homebrew removed python3 and python now defaults to version 3. Here we
use `brew upgrade python` to upgrade the pre-installed version of
python to python3.
Homebrew removed python3 and python now defaults to version 3. Here we
use `brew upgrade python` to upgrade the pre-installed version of
python to python3.
Here we are manually scaling the pt width used for the BoundingBox and
the Text element when manually adding whitespace to account for
limitations of the PDF.js viewer. This fixes an initial regression
noticed when selecting text elements in Chrome and PDFium. The width
of the Text element and BoundBox had not been adjusted for the
additional whitespace so the highlighting was offset slightly.
This commit includes an optional work around for limitations of the
PDF.js viewer described in
https://github.com/jbarlow83/OCRmyPDF/issues/133. Here is explicitly
add an addition space to text elements before drawing them on the PDF
canvas when using the HOCR renderer. This option does not apply to
other pdf renderers in OCRmyPDF and is turned off by default.