Compare commits

..
94 Commits
Author SHA1 Message Date
James R. Barlow 21fb6c82ca v13.6.2 release notes 2022-07-25 23:48:22 -07:00
James R. Barlow 27f7b9f255 Fix missing TypeAlias on <3.10 2022-07-25 16:32:55 -07:00
James R. Barlow 6f31a92ffb _windows: Avoid messy next()/StopIteration in generator 2022-07-23 15:40:05 -07:00
James R. Barlow da2276788c Fix _windows module typing 2022-07-23 15:32:14 -07:00
James R. Barlow dc6f1a266a Modernize type annotations 2022-07-23 00:39:24 -07:00
James R. Barlow 9c8ddd853d Typing adjustments 2022-07-23 00:07:50 -07:00
James R. Barlow 014d0302f2 Pre-commit: update 2022-07-22 23:49:51 -07:00
James R. Barlow 65568b3dbc Merge commit '05e2b6698dacce898abc356222124c7a1609f569' 2022-07-18 14:12:01 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
05e2b6698d build(deps): bump docker/setup-buildx-action from 1 to 2 (#994)
Bumps [docker/setup-buildx-action](https://github.com/docker/setup-buildx-action) from 1 to 2.
- [Release notes](https://github.com/docker/setup-buildx-action/releases)
- [Commits](https://github.com/docker/setup-buildx-action/compare/v1...v2)

---
updated-dependencies:
- dependency-name: docker/setup-buildx-action
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2022-07-18 14:10:45 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2f8e0f7d95 build(deps): bump actions/download-artifact from 2 to 3 (#995)
Bumps [actions/download-artifact](https://github.com/actions/download-artifact) from 2 to 3.
- [Release notes](https://github.com/actions/download-artifact/releases)
- [Commits](https://github.com/actions/download-artifact/compare/v2...v3)

---
updated-dependencies:
- dependency-name: actions/download-artifact
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2022-07-18 14:10:33 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
7e7553fc6b build(deps): bump actions/setup-python from 2 to 4 (#996)
Bumps [actions/setup-python](https://github.com/actions/setup-python) from 2 to 4.
- [Release notes](https://github.com/actions/setup-python/releases)
- [Commits](https://github.com/actions/setup-python/compare/v2...v4)

---
updated-dependencies:
- dependency-name: actions/setup-python
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2022-07-18 14:10:19 -07:00
James R. BarlowandH. Felix Wittmann 6b425aaebe Add shim for cancel_futures in older Pythons
Thanks @hfwittmann

Closes #993

Co-authored-by: H. Felix Wittmann <hfwittmann@users.noreply.github.com>
2022-07-17 16:02:45 -07:00
James R. Barlow 725af43bc3 docs: fix badges for debian 2022-07-12 14:56:00 -07:00
James R. Barlow 5c60309609 Merge remote-tracking branch 'origin/master' 2022-07-12 02:23:17 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
7d5cd55909 build(deps): bump codecov/codecov-action from 1 to 3 (#990)
Bumps [codecov/codecov-action](https://github.com/codecov/codecov-action) from 1 to 3.
- [Release notes](https://github.com/codecov/codecov-action/releases)
- [Changelog](https://github.com/codecov/codecov-action/blob/master/CHANGELOG.md)
- [Commits](https://github.com/codecov/codecov-action/compare/v1...v3)

---
updated-dependencies:
- dependency-name: codecov/codecov-action
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2022-07-12 02:23:02 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
ec4a06fad2 build(deps): bump docker/login-action from 1 to 2 (#988)
Bumps [docker/login-action](https://github.com/docker/login-action) from 1 to 2.
- [Release notes](https://github.com/docker/login-action/releases)
- [Commits](https://github.com/docker/login-action/compare/v1...v2)

---
updated-dependencies:
- dependency-name: docker/login-action
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2022-07-12 02:22:53 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
48c6e2318e build(deps): bump docker/setup-qemu-action from 1 to 2 (#987)
Bumps [docker/setup-qemu-action](https://github.com/docker/setup-qemu-action) from 1 to 2.
- [Release notes](https://github.com/docker/setup-qemu-action/releases)
- [Commits](https://github.com/docker/setup-qemu-action/compare/v1...v2)

---
updated-dependencies:
- dependency-name: docker/setup-qemu-action
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2022-07-12 02:22:38 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2eafa5e070 build(deps): bump actions/upload-artifact from 2 to 3 (#989)
Bumps [actions/upload-artifact](https://github.com/actions/upload-artifact) from 2 to 3.
- [Release notes](https://github.com/actions/upload-artifact/releases)
- [Commits](https://github.com/actions/upload-artifact/compare/v2...v3)

---
updated-dependencies:
- dependency-name: actions/upload-artifact
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2022-07-12 02:21:27 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
6aa04d7569 build(deps): bump actions/checkout from 2 to 3 (#991)
Bumps [actions/checkout](https://github.com/actions/checkout) from 2 to 3.
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/checkout/compare/v2...v3)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2022-07-12 02:21:12 -07:00
James R. Barlow b9bffa97ba Fix ABCMeta typing 2022-07-12 02:20:22 -07:00
James R. Barlow 777ba99ccc v13.6.1 release notes 2022-07-12 02:11:26 -07:00
James R. Barlow 001b3324f1 Require setuptools-scm 7.0.5 to ensure sdists work alright 2022-07-12 02:09:37 -07:00
James R. Barlow b1f2d257e2 Typing improvements 2022-07-09 02:24:19 -07:00
James R. Barlow 24a08e5170 hocrtransform: suppress deprecation warning from importing reportlab 2022-07-07 02:09:47 -07:00
James R. Barlow adf97fd82c pipeline: eliminate useless test 2022-07-07 02:08:14 -07:00
James R. Barlow a60ea72517 docs: improve remarks about lossy JBIG2 and lossy image transformations 2022-07-04 23:01:20 -07:00
James R. Barlow 59f967cdcd Activate GHA dependabot 2022-07-04 22:43:19 -07:00
James R. Barlow 6e439ee89e Modernize setuptools usage and setuptools_scm 2022-07-04 02:26:14 -07:00
James R. Barlow da38e1b035 docs: adding missing plugins 2022-07-04 02:16:20 -07:00
James R. Barlow 28c60c4f82 v13.6.0 release notes 2022-07-03 15:35:22 -07:00
James R. Barlow a5efc4af9b unpaper: replace input pnm with png
Unpaper or its underlying libraries don't seem to accept pnms with an
odd integer width. Although it's not clear if this is the issue at all.

In any case, keeping the image a PNG works around the issue. unpaper
only accepted PNM input in the past, which is why we send it PNM.
Since it now accepts PNG, we might as well use PNG.

Unpaper can write PNG as output too, but this added a few seconds to
the test suite was not committed.

Related issues:

https://github.com/ocrmypdf/OCRmyPDF/issues/887

https://github.com/ocrmypdf/OCRmyPDF/issues/665

https://github.com/unpaper/unpaper/issues/82
2022-07-03 15:32:16 -07:00
James R. Barlow 1141235c42 Merge remote-tracking branch 'origin/master' 2022-06-24 01:05:08 -07:00
2b6b7a4975 Update README to show nix package manager install option (#904)
ocrmypdf now works on M1 Mac (aarch64-darwin)

Co-authored-by: xave <xavieking@gmail.com>
2022-06-19 01:47:10 -07:00
James R. Barlow b062c9e8c0 pluginspec: hint that optimize may need to implement initialize
[ci skip]
2022-06-19 01:35:30 -07:00
James R. Barlow ed632ae366 docs: update batch to avoid suggesting Docker volumes
[ci skip]
2022-06-19 01:01:39 -07:00
James R. Barlow af742229e7 docs: fix sentence fragment in batch page
Closes #980

[ci skip]
2022-06-19 00:41:58 -07:00
Alexander JaustandGitHub e2d998245d Fix type in cookbook.rst (#978)
Add missing dash in warning about `--clean-final` and `--remove-background` commands.
2022-06-19 00:32:35 -07:00
James R. Barlow 8c58e95c3a Add new initialize hookspec to make suppress plugins easier 2022-06-19 00:31:42 -07:00
James R. Barlow 61600111d3 test_pdfinfo: refactor by extracting fixtures 2022-06-18 16:29:57 -07:00
James R. Barlow e4c45e3d3b Plugins should ideally not import from ocrmypdf._* 2022-06-18 16:03:55 -07:00
Julius BullingerandGitHub 7cabbb125f watcher: Add an option to archive processed originals (#951)
* watcher: Add an option to archive processed originals

This adds a feature from existing OCRmyPDF watchdog Docker containers like meyay/ocrmypdf-batch and unze/ocrmypdf-watchdog. With this option, the input directory can be kept clean from already processed files, without losing the originals.

* docs: Improve watcher.py's Docker parameters documentation
2022-06-17 15:17:03 -07:00
James R. Barlow d8753dc790 Fix Windows issue if no exception 2022-06-13 01:46:59 -07:00
James R. Barlow ef43d7e016 v13.5.0 release notes 2022-06-13 01:30:27 -07:00
James R. Barlow 17a5b8b43c Refactor reporting of optimization failures 2022-06-13 01:30:15 -07:00
James R. Barlow 13d11e76e5 optimize plugin: solve linearization and "is optimization enabled?" issues 2022-06-13 00:59:41 -07:00
James R. Barlow 61069660a2 Move optimization options to plugin 2022-06-12 02:42:16 -07:00
James R. Barlow 685a06c93d Move optimize into a builtin plugin 2022-06-12 02:23:13 -07:00
James R. Barlow 6cdf68363a docs: copyright year 2022-06-12 00:31:01 -07:00
James R. Barlow 522ff3c21a Remove major version pins 2022-06-12 00:31:01 -07:00
James R. Barlow 10245dc954 Refactor a helper to use the Python API eventually 2022-06-12 00:31:01 -07:00
James R. Barlow 3d4f80639d Remove test that is now always skipped 2022-06-12 00:31:01 -07:00
James R. Barlow db9a22c9dd api: call enable_ansi_support unconditionally
coloredlogs documentation says this is (now?) permitted.

Also import it from the right place.
2022-06-12 00:31:01 -07:00
James R. Barlow 31683530f8 api/cli: replace private variable access with method 2022-06-12 00:31:01 -07:00
James R. Barlow 0e550a1c6d graft: use pikepdf's random name API instead of our own
Minor change to PDF output.
2022-06-12 00:31:01 -07:00
James R. Barlow b17fb61389 Configure pylint in pyproject and delint 2022-06-12 00:30:44 -07:00
James R. Barlow d640c2ded3 Tidy some if condition -> EAFP 2022-06-11 00:21:21 -07:00
James R. Barlow a0ac448d52 Tidy some old-school %s strings 2022-06-11 00:15:19 -07:00
James R. Barlow e3ba13e365 docs: update installation notes
Add Snap. Remove Mageia because I have no idea if it still works. Drop a
lot of old versions and old notes. Change pip3->pip except for old py2/py3 images.
Try to make things less version sensitive.
2022-06-10 01:33:05 -07:00
James R. Barlow 0cd04abc4e snap: add logo 2022-06-09 16:59:20 -07:00
Alexander LangankeandJames R. Barlow ee81f3968f Add snapcraft.yaml 2022-06-08 01:22:18 -07:00
James R. Barlow 21cacad93b unpaper: super syntax 2022-06-02 17:05:59 -07:00
James R. Barlow 3589f4e7d1 unpaper: fix unsaved change 2022-06-02 02:22:48 -07:00
James R. Barlow 1cdc2591e5 v13.4.7 release notes 2022-06-02 02:00:16 -07:00
James R. Barlow e05f9575a8 Merge remote-tracking branch 'origin/master' 2022-06-02 01:57:39 -07:00
James R. Barlow 10c703e119 unpaper: use TemporaryDirectory(ignore_cleanup_errors=True) where available
Fixes #974 when used in conjunction with Python 3.10.

Reviewed other uses of TemporaryDirectory in ocrmypdf and decided it
was not worth fixing them since neither are exactly production code.
2022-06-02 01:56:41 -07:00
James R. Barlow 0ac15dd0b2 Suppress libxmp DeprecationWarning during test 2022-06-01 00:46:16 -07:00
Robert SchützandGitHub 808b24d59f ignore PermissionError when calling os.nice() (#973) 2022-05-28 18:07:54 -07:00
James R. Barlow c082526dea Test pypy-3.8 instead of 3.7 2022-05-26 13:52:56 -07:00
James R. Barlow 33cdabaf65 tests: account for test that expected pngquant for windows 2022-05-26 13:52:22 -07:00
James R. Barlow 94f8e36601 ci: don't install pngquant for windows anymore 2022-05-26 13:01:58 -07:00
James R. Barlow 865002c7be v13.4.6 release notes 2022-05-26 00:59:14 -07:00
James R. Barlow 5d0cc0a092 tests: Extract some test fixtures for better clarity 2022-05-26 00:57:31 -07:00
James R. Barlow 6c427f82ea Add test case for corrupt ICC profiles 2022-05-26 00:41:19 -07:00
James R. Barlow e7a44ba87a info: adjust ICC warning message 2022-05-26 00:21:31 -07:00
James R. Barlow c311768452 Merge branch 'corrupt-icc' of https://github.com/oscherler/OCRmyPDF into oscherler-corrupt-icc 2022-05-26 00:14:47 -07:00
James R. Barlow f53fedee63 pre-commit: autoupdate 2022-05-25 15:21:55 -07:00
James R. Barlow 87838127b0 info: replace introspection with explicit f-string 2022-05-25 13:16:40 -07:00
Olivier Scherler 4db4df5c72 Log a warning instead of failing on images with a corrupt ICC profile. 2022-05-25 12:37:21 +02:00
James R. Barlow 11125c5367 v13.4.5 release notes 2022-05-24 16:52:43 -07:00
James R. Barlow e648411067 Remove pdfminer.six upper version restriction
Needing to update this every time has become more inconvenient than dealing
with the occasional breakage in a new release.
2022-05-24 16:51:38 -07:00
Ben BeasleyandGitHub 11365575d7 Allow pdfminer.six 20220524 (#971) 2022-05-24 16:33:44 -07:00
James R. Barlow 845cb5c40c docs: clarify that --pages and --skip-text exclusions apply to image processing and OCR
Closes #950
2022-05-16 13:20:47 -07:00
James R. Barlow b699e158be Fix references to old repo at jbarlow83/OCRmyPDF 2022-05-16 12:48:10 -07:00
leonnicolasandGitHub 603da52026 Fix small typo in update api.py (#963)
[ci skip]
2022-05-14 21:45:23 -07:00
James R. Barlow 8d0765a5e0 v13.4.4 release notes 2022-05-14 14:53:03 -07:00
James R. Barlow 1ca327e13b Update pdfminer.six version 2022-05-14 14:51:59 -07:00
James R. Barlow f504fd1875 Change Docker base to 22.04 2022-04-28 22:46:30 -07:00
James R. Barlow cf7c20ca16 v13.4.3 release notes 2022-04-14 20:19:31 -07:00
James R. Barlow b00fe3dc5d pytest.skip() - remove kwarg entirely, to avoid breaking older pytest and not getting warns from newer pytest 2022-04-14 20:15:00 -07:00
James R. Barlow e6aa3a4299 tests: explain why CacheOcrEngine needs lock 2022-04-05 16:16:51 -07:00
James R. Barlow 24f1b57288 Merge branch 'master' of github.com:ocrmypdf/OCRmyPDF 2022-04-05 16:04:45 -07:00
James R. Barlow 43302d7e12 Fix pytest.warns() on older pytest
Thanks @QuLogic
2022-04-05 16:02:50 -07:00
Christopher BeschandGitHub fed0226761 Add Gentoo Language Installation Instructions (#936) 2022-04-05 00:26:24 -07:00
Joseph MorrisandGitHub 27e22b4f07 Update jbig2.rst (#911)
Adding Ubuntu package names for dependencies I needed to save others time. The tequired package for leptonica is particularly confusing since configure just says "Error! Leptonica not detected." and there are multiple leptonica packages
2022-04-05 00:25:28 -07:00
112 changed files with 1870 additions and 816 deletions
+1 -1
View File
@@ -1,7 +1,7 @@
# OCRmyPDF # OCRmyPDF
# #
FROM debian:bookworm-slim as base FROM ubuntu:22.04 as base
ENV LANG=C.UTF-8 ENV LANG=C.UTF-8
ENV TZ=UTC ENV TZ=UTC
+4 -1
View File
@@ -1 +1,4 @@
ref-names: $Format:%D$ node: $Format:%H$
node-date: $Format:%cI$
describe-name: $Format:%(describe:tags=true)$
ref-names: $Format:%D$
@@ -22,7 +22,7 @@ Run with verbosity or higher `-v1` to see more detailed logging. This informatio
**Example file** **Example file**
If your issue is a problem that affects only certain files, and we will require an input file (PDF or image) that demonstrates your issue. If your issue is a problem that affects only certain files, and we will require an input file (PDF or image) that demonstrates your issue.
Please provide an input file with no personal or confidential information. At your option you may [GPG-encrypt the file](https://github.com/jbarlow83/OCRmyPDF/wiki) for OCRmyPDF's author only. Please provide an input file with no personal or confidential information. At your option you may [GPG-encrypt the file](https://github.com/ocrmypdf/OCRmyPDF/wiki) for OCRmyPDF's author only.
Links to files hosted elsewhere are perfectly acceptable. You could also look in ``tests/resources`` and see if any of those files reproduce your issue. Links to files hosted elsewhere are perfectly acceptable. You could also look in ``tests/resources`` and see if any of those files reproduce your issue.
+1 -1
View File
@@ -19,7 +19,7 @@ A clear and concise description of any alternative solutions or features you've
**Example file** **Example file**
If your issue concerns how OCRmyPDF processes certain files, and please provide an example file that helps illustrate how OCRmyPDF's output could be improve. If your issue concerns how OCRmyPDF processes certain files, and please provide an example file that helps illustrate how OCRmyPDF's output could be improve.
Please provide an input file with no personal or confidential information. At your option you may [GPG-encrypt the file](https://github.com/jbarlow83/OCRmyPDF/wiki) for OCRmyPDF's author only. Please provide an input file with no personal or confidential information. At your option you may [GPG-encrypt the file](https://github.com/ocrmypdf/OCRmyPDF/wiki) for OCRmyPDF's author only.
Links to files hosted elsewhere are perfectly acceptable. You could also look in ``tests/resources`` and see if any of those files reproduce your issue. Links to files hosted elsewhere are perfectly acceptable. You could also look in ``tests/resources`` and see if any of those files reproduce your issue.
+11
View File
@@ -0,0 +1,11 @@
# To get started with Dependabot version updates, you'll need to specify which
# package ecosystems to update and where the package manifests are located.
# Please see the documentation for all configuration options:
# https://docs.github.com/github/administering-a-repository/configuration-options-for-dependency-updates
version: 2
updates:
- package-ecosystem: "github-actions" # See documentation for possible values
directory: "/" # Location of package manifests
schedule:
interval: "weekly"
+19 -19
View File
@@ -31,7 +31,7 @@ jobs:
- os: ubuntu-latest - os: ubuntu-latest
python: "3.9" python: "3.9"
- os: ubuntu-latest - os: ubuntu-latest
python: "pypy-3.7" python: "pypy-3.8"
- os: ubuntu-latest - os: ubuntu-latest
python: "3.9" python: "3.9"
tesseract5: true tesseract5: true
@@ -41,11 +41,11 @@ jobs:
PYTHON: ${{ matrix.python }} PYTHON: ${{ matrix.python }}
steps: steps:
- uses: actions/checkout@v2 - uses: actions/checkout@v3
with: with:
fetch-depth: "0" # 0=all, needed for setuptools-scm to resolve version tags fetch-depth: "0" # 0=all, needed for setuptools-scm to resolve version tags
- uses: actions/setup-python@v2 - uses: actions/setup-python@v4
name: Install Python name: Install Python
with: with:
python-version: ${{ matrix.python }} python-version: ${{ matrix.python }}
@@ -111,7 +111,7 @@ jobs:
python -m pytest --cov-report xml --cov=ocrmypdf --cov=tests/ -n0 tests/ python -m pytest --cov-report xml --cov=ocrmypdf --cov=tests/ -n0 tests/
- name: Upload coverage to Codecov - name: Upload coverage to Codecov
uses: codecov/codecov-action@v1 uses: codecov/codecov-action@v3
with: with:
files: ./coverage.xml files: ./coverage.xml
env_vars: OS,PYTHON env_vars: OS,PYTHON
@@ -129,11 +129,11 @@ jobs:
PYTHON: ${{ matrix.python }} PYTHON: ${{ matrix.python }}
steps: steps:
- uses: actions/checkout@v2 - uses: actions/checkout@v3
with: with:
fetch-depth: "0" # 0=all, needed for setuptools-scm to resolve version tags fetch-depth: "0" # 0=all, needed for setuptools-scm to resolve version tags
- uses: actions/setup-python@v2 - uses: actions/setup-python@v4
name: Install Python name: Install Python
with: with:
python-version: ${{ matrix.python }} python-version: ${{ matrix.python }}
@@ -166,7 +166,7 @@ jobs:
python -m pytest --cov-report xml --cov=ocrmypdf --cov=tests/ -n0 tests/ python -m pytest --cov-report xml --cov=ocrmypdf --cov=tests/ -n0 tests/
- name: Upload coverage to Codecov - name: Upload coverage to Codecov
uses: codecov/codecov-action@v1 uses: codecov/codecov-action@v3
with: with:
files: ./coverage.xml files: ./coverage.xml
env_vars: OS,PYTHON env_vars: OS,PYTHON
@@ -184,11 +184,11 @@ jobs:
PYTHON: ${{ matrix.python }} PYTHON: ${{ matrix.python }}
steps: steps:
- uses: actions/checkout@v2 - uses: actions/checkout@v3
with: with:
fetch-depth: "0" # 0=all, needed for setuptools-scm to resolve version tags fetch-depth: "0" # 0=all, needed for setuptools-scm to resolve version tags
- uses: actions/setup-python@v2 - uses: actions/setup-python@v4
name: Install Python name: Install Python
with: with:
python-version: ${{ matrix.python }} python-version: ${{ matrix.python }}
@@ -196,7 +196,7 @@ jobs:
- name: Install system packages - name: Install system packages
run: | run: |
choco install --yes --no-progress --pre tesseract choco install --yes --no-progress --pre tesseract
choco install --yes --no-progress --ignore-checksums ghostscript pngquant choco install --yes --no-progress --ignore-checksums ghostscript
- name: Install Python packages - name: Install Python packages
run: | run: |
@@ -208,7 +208,7 @@ jobs:
python -m pytest --cov-report xml --cov=ocrmypdf --cov=tests/ -n0 tests/ python -m pytest --cov-report xml --cov=ocrmypdf --cov=tests/ -n0 tests/
- name: Upload coverage to Codecov - name: Upload coverage to Codecov
uses: codecov/codecov-action@v1 uses: codecov/codecov-action@v3
with: with:
files: ./coverage.xml files: ./coverage.xml
env_vars: OS,PYTHON env_vars: OS,PYTHON
@@ -217,11 +217,11 @@ jobs:
name: Build sdist and wheels name: Build sdist and wheels
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: actions/checkout@v2 - uses: actions/checkout@v3
with: with:
fetch-depth: "0" # 0=all, needed for setuptools-scm to resolve version tags fetch-depth: "0" # 0=all, needed for setuptools-scm to resolve version tags
- uses: actions/setup-python@v2 - uses: actions/setup-python@v4
name: Install Python name: Install Python
with: with:
python-version: "3.7" python-version: "3.7"
@@ -232,7 +232,7 @@ jobs:
python setup.py sdist python setup.py sdist
python setup.py bdist_wheel python setup.py bdist_wheel
- uses: actions/upload-artifact@v2 - uses: actions/upload-artifact@v3
with: with:
path: | path: |
./dist/*.whl ./dist/*.whl
@@ -244,7 +244,7 @@ jobs:
runs-on: ubuntu-latest runs-on: ubuntu-latest
if: github.event_name == 'push' && startsWith(github.event.ref, 'refs/tags/v') if: github.event_name == 'push' && startsWith(github.event.ref, 'refs/tags/v')
steps: steps:
- uses: actions/download-artifact@v2 - uses: actions/download-artifact@v3
with: with:
name: artifact name: artifact
path: dist path: dist
@@ -274,22 +274,22 @@ jobs:
- name: Set image name - name: Set image name
run: echo "DOCKER_IMAGE_NAME=ocrmypdf" >> $GITHUB_ENV run: echo "DOCKER_IMAGE_NAME=ocrmypdf" >> $GITHUB_ENV
- uses: actions/checkout@v2 - uses: actions/checkout@v3
with: with:
fetch-depth: "0" # 0=all, needed for setuptools-scm to resolve version tags fetch-depth: "0" # 0=all, needed for setuptools-scm to resolve version tags
- name: Login to Docker Hub - name: Login to Docker Hub
uses: docker/login-action@v1 uses: docker/login-action@v2
with: with:
username: jbarlow83 username: jbarlow83
password: ${{ secrets.DOCKERHUB_TOKEN }} password: ${{ secrets.DOCKERHUB_TOKEN }}
- name: Set up QEMU - name: Set up QEMU
uses: docker/setup-qemu-action@v1 uses: docker/setup-qemu-action@v2
- name: Set up Docker Buildx - name: Set up Docker Buildx
id: buildx id: buildx
uses: docker/setup-buildx-action@v1 uses: docker/setup-buildx-action@v2
- name: Print image tag - name: Print image tag
run: echo "Building image ${DOCKER_REPOSITORY}/${DOCKER_IMAGE_NAME}:${DOCKER_IMAGE_TAG}" run: echo "Building image ${DOCKER_REPOSITORY}/${DOCKER_IMAGE_NAME}:${DOCKER_IMAGE_TAG}"
+6 -6
View File
@@ -1,6 +1,6 @@
repos: repos:
- repo: https://github.com/pre-commit/pre-commit-hooks - repo: https://github.com/pre-commit/pre-commit-hooks
rev: v4.1.0 rev: v4.3.0
hooks: hooks:
- id: check-case-conflict - id: check-case-conflict
- id: check-merge-conflict - id: check-merge-conflict
@@ -11,23 +11,23 @@ repos:
rev: 5.10.1 rev: 5.10.1
hooks: hooks:
- id: isort - id: isort
args: ["--profile", "black"] args: ["--profile", "black", "-a", "from __future__ import annotations"]
- repo: https://github.com/psf/black - repo: https://github.com/psf/black
rev: 22.3.0 rev: 22.6.0
hooks: hooks:
- id: black - id: black
language_version: python language_version: python
- repo: https://github.com/asottile/setup-cfg-fmt - repo: https://github.com/asottile/setup-cfg-fmt
rev: v1.20.1 rev: v1.20.2
hooks: hooks:
- id: setup-cfg-fmt - id: setup-cfg-fmt
- repo: https://github.com/asottile/pyupgrade - repo: https://github.com/asottile/pyupgrade
rev: v2.31.1 rev: v2.37.2
hooks: hooks:
- id: pyupgrade - id: pyupgrade
args: ["--py37-plus"] args: ["--py37-plus"]
- repo: https://github.com/pre-commit/mirrors-mypy - repo: https://github.com/pre-commit/mirrors-mypy
rev: v0.942 rev: v0.971
hooks: hooks:
- id: mypy - id: mypy
additional_dependencies: additional_dependencies:
+4 -5
View File
@@ -1,9 +1,7 @@
<img src="docs/images/logo.svg" width="240" alt="OCRmyPDF"> <img src="docs/images/logo.svg" width="240" alt="OCRmyPDF">
[![Build Status](https://github.com/jbarlow83/OCRmyPDF/actions/workflows/build.yml/badge.svg)](https://github.com/jbarlow83/OCRmyPDF/actions/workflows/build.yml) [![PyPI version][pypi]](https://pypi.org/project/ocrmypdf/) ![Homebrew version][homebrew] ![ReadTheDocs][docs] ![Python versions][pyversions] [![Build Status](https://github.com/ocrmypdf/OCRmyPDF/actions/workflows/build.yml/badge.svg)](https://github.com/ocrmypdf/OCRmyPDF/actions/workflows/build.yml) [![PyPI version][pypi]](https://pypi.org/project/ocrmypdf/) ![Homebrew version][homebrew] ![ReadTheDocs][docs] ![Python versions][pyversions]
[azure]: https://dev.azure.com/jim0585/ocrmypdf/_apis/build/status/jbarlow83.OCRmyPDF?branchName=master
[travis]: https://travis-ci.org/jbarlow83/OCRmyPDF.svg?branch=master "Travis build status"
[pypi]: https://img.shields.io/pypi/v/ocrmypdf.svg "PyPI version" [pypi]: https://img.shields.io/pypi/v/ocrmypdf.svg "PyPI version"
[homebrew]: https://img.shields.io/homebrew/v/ocrmypdf.svg "Homebrew version" [homebrew]: https://img.shields.io/homebrew/v/ocrmypdf.svg "Homebrew version"
[docs]: https://readthedocs.org/projects/ocrmypdf/badge/?version=latest "RTD" [docs]: https://readthedocs.org/projects/ocrmypdf/badge/?version=latest "RTD"
@@ -64,7 +62,8 @@ Linux, Windows, macOS and FreeBSD are supported. Docker images are also availabl
| Debian, Ubuntu | ``apt install ocrmypdf`` | | Debian, Ubuntu | ``apt install ocrmypdf`` |
| Windows Subsystem for Linux | ``apt install ocrmypdf`` | | Windows Subsystem for Linux | ``apt install ocrmypdf`` |
| Fedora | ``dnf install ocrmypdf`` | | Fedora | ``dnf install ocrmypdf`` |
| macOS | ``brew install ocrmypdf`` | | macOS (Homebrew) | ``brew install ocrmypdf`` |
| macOS (nix) | ``nix-env -i ocrmypdf`` |
| LinuxBrew | ``brew install ocrmypdf`` | | LinuxBrew | ``brew install ocrmypdf`` |
| FreeBSD | ``pkg install py37-ocrmypdf`` | | FreeBSD | ``pkg install py37-ocrmypdf`` |
| Conda | ``conda install ocrmypdf`` | | Conda | ``conda install ocrmypdf`` |
@@ -106,7 +105,7 @@ ocrmypdf --help
Our [documentation is served on Read the Docs](https://ocrmypdf.readthedocs.io/en/latest/index.html). Our [documentation is served on Read the Docs](https://ocrmypdf.readthedocs.io/en/latest/index.html).
Please report issues on our [GitHub issues](https://github.com/jbarlow83/OCRmyPDF/issues) page, and follow the issue template for quick response. Please report issues on our [GitHub issues](https://github.com/ocrmypdf/OCRmyPDF/issues) page, and follow the issue template for quick response.
## Requirements ## Requirements
+1 -1
View File
@@ -1,7 +1,7 @@
Format: https://www.debian.org/doc/packaging-manuals/copyright-format/1.0/ Format: https://www.debian.org/doc/packaging-manuals/copyright-format/1.0/
Upstream-Name: OCRmyPDF Upstream-Name: OCRmyPDF
Upstream-Contact: James R. Barlow <barlow.jim@gmail.com> Upstream-Contact: James R. Barlow <barlow.jim@gmail.com>
Source: https://github.com/jbarlow83/OCRmyPDF Source: https://github.com/ocrmypdf/OCRmyPDF
Files: * Files: *
Copyright: Copyright:
+8 -6
View File
@@ -67,11 +67,11 @@ without modifying the PDF. This is to ensure that PDFs that were
previously OCRed or were "born digital" rather than scanned are not previously OCRed or were "born digital" rather than scanned are not
processed. processed.
If ``--skip-text`` is issued, then no OCR will be performed on pages If ``--skip-text`` is issued, then no image processing or OCR will be
that already have text. The page will be copied to the output. This may performed on pages that already have text. The page will be copied to
be useful for documents that contain both "born digital" and scanned the output. This may be useful for documents that contain both "born
content, or to use OCRmyPDF to normalize and convert to PDF/A regardless digital" and scanned content, or to use OCRmyPDF to normalize and
of their contents. convert to PDF/A regardless of their contents.
If ``--redo-ocr`` is issued, then a detailed text analysis is performed. If ``--redo-ocr`` is issued, then a detailed text analysis is performed.
Text is categorized as either visible or invisible. Invisible text (OCR) Text is categorized as either visible or invisible. Invisible text (OCR)
@@ -223,7 +223,9 @@ The ``hocr`` renderer
The ``hocr`` renderer works with older versions of Tesseract. The image The ``hocr`` renderer works with older versions of Tesseract. The image
layer is copied from the original PDF page if possible, avoiding layer is copied from the original PDF page if possible, avoiding
potentially lossy transcoding or loss of other PDF information. If potentially lossy transcoding or loss of other PDF information. If
preprocessing is specified, then the image layer is a new PDF. preprocessing is specified, then the image layer is a new PDF. (You may
need to disable PDF/A conversion nad optimization to eliminate all
lossy transformations.)
Unlike ``sandwich`` this renderer is implemented within OCRmyPDF; anyone Unlike ``sandwich`` this renderer is implemented within OCRmyPDF; anyone
looking to customize how OCR is presented should look here. A major looking to customize how OCR is presented should look here. A major
+20 -12
View File
@@ -36,18 +36,21 @@ Directory trees
=============== ===============
This will walk through a directory tree and run OCR on all files in This will walk through a directory tree and run OCR on all files in
place, printing the output in a way that makes place, and printing each filename in between runs:
.. code-block:: bash .. code-block:: bash
find . -printf '%p' -name '*.pdf' -exec ocrmypdf '{}' '{}' \; find . -printf '%p\n' -name '*.pdf' -exec ocrmypdf '{}' '{}' \;
Alternatively, with a docker container (mounts a volume to the container Alternatively, with a Docker container and streaming the file through
where the PDFs are stored): standard input and output:
.. code-block:: bash .. code-block:: bash
find . -printf '%p' -name '*.pdf' -exec docker run --rm -v <host dir>:<container dir> jbarlow83/ocrmypdf '<container dir>/{}' '<container dir>/{}' \; find . -name '*.pdf' -print0 | xargs -0 | while read pdf; do
pdfout=$(mktemp)
docker run --rm -i jbarlow83/ocrmypdf - - <$pdf >$pdfout && cp $pdfout $pdf
done
This only runs one ``ocrmypdf`` process at a time. This variation uses This only runs one ``ocrmypdf`` process at a time. This variation uses
``find`` to create a directory list and ``parallel`` to parallelize runs ``find`` to create a directory list and ``parallel`` to parallelize runs
@@ -124,7 +127,9 @@ Users may need to customize the script to meet their requirements.
"OCR_INPUT_DIRECTORY", "Set input directory to monitor (recursive)" "OCR_INPUT_DIRECTORY", "Set input directory to monitor (recursive)"
"OCR_OUTPUT_DIRECTORY", "Set output directory (should not be under input)" "OCR_OUTPUT_DIRECTORY", "Set output directory (should not be under input)"
"OCR_ARCHIVE_DIRECTORY", "Set archive directory for processed originals (should not be under input, requires ``OCR_ON_SUCCESS_ARCHIVE`` to be set)"
"OCR_ON_SUCCESS_DELETE", "This will delete the input file if the exit code is 0 (OK)" "OCR_ON_SUCCESS_DELETE", "This will delete the input file if the exit code is 0 (OK)"
"OCR_ON_SUCCESS_ARCHIVE", "This will move the processed orignal file to ``OCR_ARCHIVE_DIRECTORY`` if the exit code is 0 (OK). Note that ``OCR_ON_SUCCESS_DELETE`` takes precedence over this option, i.e. if both options are set, the input file will be deleted."
"OCR_OUTPUT_DIRECTORY_YEAR_MONTH", "This will place files in the output in ``{output}/{year}/{month}/{filename}``" "OCR_OUTPUT_DIRECTORY_YEAR_MONTH", "This will place files in the output in ``{output}/{year}/{month}/{filename}``"
"OCR_DESKEW", "Apply deskew to crooked input PDFs" "OCR_DESKEW", "Apply deskew to crooked input PDFs"
"OCR_JSON_SETTINGS", "A JSON string specifying any other arguments for ``ocrmypdf.ocr``, e.g. ``'OCR_JSON_SETTINGS={""rotate_pages"": true}'``." "OCR_JSON_SETTINGS", "A JSON string specifying any other arguments for ``ocrmypdf.ocr``, e.g. ``'OCR_JSON_SETTINGS={""rotate_pages"": true}'``."
@@ -144,16 +149,18 @@ The watcher service is included in the OCRmyPDF Docker image. To run it:
docker run \ docker run \
-v <path to files to convert>:/input \ -v <path to files to convert>:/input \
-v <path to store results>:/output \ -v <path to store results>:/output \
-v <path to store processed originals>:/archive \
-e OCR_OUTPUT_DIRECTORY_YEAR_MONTH=1 \ -e OCR_OUTPUT_DIRECTORY_YEAR_MONTH=1 \
-e OCR_ON_SUCCESS_DELETE=1 \ -e OCR_ON_SUCCESS_ARCHIVE=1 \
-e OCR_DESKEW=1 \ -e OCR_DESKEW=1 \
-e PYTHONUNBUFFERED=1 \ -e PYTHONUNBUFFERED=1 \
-it --entrypoint python3 \ -it --entrypoint python3 \
jbarlow83/ocrmypdf \ jbarlow83/ocrmypdf \
watcher.py watcher.py
This service will watch for a file that matches ``/input/\*.pdf`` and will This service will watch for a file that matches ``/input/\*.pdf``,
convert it to a OCRed PDF in ``/output/``. The parameters to this image are: convert it to a OCRed PDF in ``/output/``, and move the processed
original to ``/archive``. The parameters to this image are:
.. csv-table:: watcher.py parameters for Docker .. csv-table:: watcher.py parameters for Docker
:header: "Parameter", "Description" :header: "Parameter", "Description"
@@ -161,10 +168,11 @@ convert it to a OCRed PDF in ``/output/``. The parameters to this image are:
"``-v <path to files to convert>:/input``", "Files placed in this location will be OCRed" "``-v <path to files to convert>:/input``", "Files placed in this location will be OCRed"
"``-v <path to store results>:/output``", "This is where OCRed files will be stored" "``-v <path to store results>:/output``", "This is where OCRed files will be stored"
"``-e OCR_OUTPUT_DIRECTORY_YEAR_MONTH=1``", "Define environment variable OCR_OUTPUT_DIRECTORY_YEAR_MONTH=1" "``-v <path to store processed originals>:/archive``", "Archive processed originals here"
"``-e OCR_ON_SUCCESS_DELETE=1``", "Define environment variable" "``-e OCR_OUTPUT_DIRECTORY_YEAR_MONTH=1``", "Define environment variable ``OCR_OUTPUT_DIRECTORY_YEAR_MONTH=1`` to place files in the output in ``{output}/{year}/{month}/{filename}``"
"``-e OCR_DESKEW=1``", "Define environment variable" "``-e OCR_ON_SUCCESS_ARCHIVE=1``", "Define environment variable ``OCR_ON_SUCCESS_ARCHIVE`` to move processed originals"
"``-e PYTHONBUFFERED=1``", "This will force STDOUT to be unbuffered and allow you to see messages in docker logs" "``-e OCR_DESKEW=1``", "Define environment variable ``OCR_DESKEW`` to apply deskew to crooked input PDFs"
"``-e PYTHONBUFFERED=1``", "This will force ``STDOUT`` to be unbuffered and allow you to see messages in docker logs"
This service relies on polling to check for changes to the filesystem. It This service relies on polling to check for changes to the filesystem. It
may not be suitable for some environments, such as filesystems shared on a may not be suitable for some environments, such as filesystems shared on a
+2 -2
View File
@@ -42,7 +42,7 @@ extensions = [
# Extension settings # Extension settings
intersphinx_mapping = {'https://docs.python.org/': None} intersphinx_mapping = {'https://docs.python.org/': None}
napoleon_use_rtype = False napoleon_use_rtype = False
issues_github_path = "jbarlow83/OCRmyPDF" issues_github_path = "ocrmypdf/OCRmyPDF"
# Add any paths that contain templates here, relative to this directory. # Add any paths that contain templates here, relative to this directory.
templates_path = ['_templates'] templates_path = ['_templates']
@@ -63,7 +63,7 @@ master_doc = 'index'
# General information about the project. # General information about the project.
project = 'ocrmypdf' project = 'ocrmypdf'
copyright = ( copyright = (
'2021, James R. Barlow. Licensed under Creative Commons Attribution-ShareAlike 4.0.' '2022, James R. Barlow. Licensed under Creative Commons Attribution-ShareAlike 4.0.'
) )
author = 'James R. Barlow' author = 'James R. Barlow'
+17 -8
View File
@@ -200,7 +200,7 @@ might remove desirable content, especially from poor quality scans.
.. warning:: .. warning::
``--clean-final`` and ``-remove-background`` may leave undesirable ``--clean-final`` and ``--remove-background`` may leave undesirable
visual artifacts in some images where their algorithms have visual artifacts in some images where their algorithms have
shortcomings. Files should be visually reviewed after using these shortcomings. Files should be visually reviewed after using these
options. options.
@@ -243,10 +243,11 @@ You can also optimize all images without performing any OCR:
ocrmypdf --tesseract-timeout=0 --optimize 3 --skip-text input.pdf output.pdf ocrmypdf --tesseract-timeout=0 --optimize 3 --skip-text input.pdf output.pdf
Perform OCR only certain pages Process only certain pages
------------------------------ --------------------------
You can ask OCRmyPDF to only apply OCR to certain pages. You can ask OCRmyPDF to only apply `image processing <#image-processing>`__
and OCR to certain pages.
.. code-block:: bash .. code-block:: bash
@@ -260,10 +261,10 @@ overlap pages. OCRmyPDF does not currently account for document page numbers,
such as an introduction section of a book that uses Roman numerals. It simply such as an introduction section of a book that uses Roman numerals. It simply
counts the number of virtual pieces of paper since the start. counts the number of virtual pieces of paper since the start.
Regardless of the argument to ``--pages``, OCRmyPDF will optimize all pages in Regardless of the argument to ``--pages``, OCRmyPDF will optimize all pages/images
the file and convert it to PDF/A, unless you disable those options. In this in the file and convert it to PDF/A, unless you disable those options. Both of these
example, we want to OCR only the title and otherwise change the PDF as little steps are "whole file" operations. In this example, we want to OCR only the title
as possible: and otherwise change the PDF as little as possible:
.. code-block:: bash .. code-block:: bash
@@ -339,6 +340,9 @@ levels in the GCC compiler.
- Enables lossless optimizations, such as transcoding images to more - Enables lossless optimizations, such as transcoding images to more
efficient formats. Also compress other uncompressed objects in the efficient formats. Also compress other uncompressed objects in the
PDF and enables the more efficient "object streams" within the PDF. PDF and enables the more efficient "object streams" within the PDF.
(If ``--jbig2-lossy`` is issued, then lossy JBIG2 optimization is used.
The decision to use lossy JBIG2 is separate from standard optimization
settings.)
* - ``--optimize 2`` * - ``--optimize 2``
- All of the above, and enables lossy optimizations and color quantization. - All of the above, and enables lossy optimizations and color quantization.
* - ``--optimize 3`` * - ``--optimize 3``
@@ -359,3 +363,8 @@ fo a PDF.
ocrmypdf --optimize 3 in.pdf out.pdf # Make it small ocrmypdf --optimize 3 in.pdf out.pdf # Make it small
Some users may consider enabling lossy JBIG2. See: :ref:`jbig2-lossy`. Some users may consider enabling lossy JBIG2. See: :ref:`jbig2-lossy`.
.. note::
Image processing and PDF/A conversion can also introduce lossy transformations
to your PDF images, even when ``--optimize 1`` is in use.
+239
View File
@@ -0,0 +1,239 @@
<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<svg
width="256"
height="256"
viewBox="0 0 256 256.00001"
version="1.1"
xml:space="preserve"
style="clip-rule:evenodd;fill-rule:evenodd;stroke-linecap:round;stroke-linejoin:round;stroke-miterlimit:1.5"
id="svg270"
sodipodi:docname="logo-square-256.svg"
inkscape:export-filename="/home/jb/src/ocrmypdf/docs/images/logo-square.png"
inkscape:export-xdpi="96"
inkscape:export-ydpi="96"
inkscape:version="1.1.2 (0a00cf5339, 2022-02-04)"
xmlns:inkscape="http://www.inkscape.org/namespaces/inkscape"
xmlns:sodipodi="http://sodipodi.sourceforge.net/DTD/sodipodi-0.dtd"
xmlns="http://www.w3.org/2000/svg"
xmlns:svg="http://www.w3.org/2000/svg"
xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"
xmlns:cc="http://creativecommons.org/ns#"
xmlns:dc="http://purl.org/dc/elements/1.1/"
xmlns:serif="http://www.serif.com/"><metadata
id="metadata276"><rdf:RDF><cc:Work
rdf:about=""><dc:format>image/svg+xml</dc:format><dc:type
rdf:resource="http://purl.org/dc/dcmitype/StillImage" /></cc:Work></rdf:RDF></metadata><defs
id="defs274" /><sodipodi:namedview
pagecolor="#ffffff"
bordercolor="#666666"
borderopacity="1"
objecttolerance="10"
gridtolerance="10"
guidetolerance="10"
inkscape:pageopacity="0"
inkscape:pageshadow="2"
inkscape:window-width="2396"
inkscape:window-height="1691"
id="namedview272"
showgrid="false"
lock-margins="false"
inkscape:zoom="2.0079523"
inkscape:cx="189.74554"
inkscape:cy="54.533168"
inkscape:window-x="26"
inkscape:window-y="23"
inkscape:window-maximized="0"
inkscape:current-layer="svg270"
inkscape:pagecheckerboard="0"
width="256px"
fit-margin-top="0"
fit-margin-left="0"
fit-margin-right="0"
fit-margin-bottom="0" />
<g
id="svg"
transform="matrix(0.48534351,0,0,0.4057699,1.8106874,71.192214)">
<rect
x="0"
y="0"
width="520"
height="280"
style="fill:#ffffff"
id="rect188" />
<g
transform="matrix(1.03522,0,0,1.23823,-69.7528,-83.422)"
id="g267">
<g
transform="translate(243.977,20.0703)"
id="g218">
<g
id="Page">
<g
transform="matrix(0.961773,0,0,1.05962,6.19811,-3.01071)"
id="g192">
<path
d="m 328.5,97.682 c 0,-1.217 -0.517,-2.386 -1.444,-3.264 -7.03,-6.66 -37.614,-35.638 -44.828,-42.474 -0.977,-0.925 -2.327,-1.448 -3.738,-1.448 -13.997,0 -90.407,0 -111.151,0 -2.871,0 -5.198,2.113 -5.198,4.718 0,27.837 0,170.351 0,198.186 0,2.605 2.327,4.717 5.197,4.717 24.904,0 131.821,0 156.2,0 2.74,0 4.962,-2.016 4.962,-4.504 0,-24.345 0,-139.717 0,-155.931 z"
style="fill:#fdfdfd;stroke:#333333;stroke-width:3.95px"
id="path190" />
</g>
<g
id="Dog-ear"
serif:id="Dog ear"
transform="translate(-4,2)">
<path
d="m 277.072,48.496 v 45.352 c 0,1.324 0.526,2.593 1.462,3.529 0.936,0.936 2.205,1.462 3.529,1.462 12.485,0 44.078,0 44.078,0"
style="fill:#f5f5f5;stroke:#333333;stroke-width:4px"
id="path194" />
</g>
</g>
<g
transform="translate(-29.6816,-0.395178)"
id="g216">
<g
transform="matrix(1.00243,0,0,1.11818,-144.72,-8.80181)"
id="g200">
<path
d="m 465.73,119.654 c 0,-2.049 -1.856,-3.713 -4.142,-3.713 H 310.259 c -2.286,0 -4.142,1.664 -4.142,3.713 v 63.454 c 0,2.049 1.856,3.713 4.142,3.713 h 151.329 c 2.286,0 4.142,-1.664 4.142,-3.713 z"
style="fill:#f80000;stroke:#ffffff;stroke-width:3.77px"
id="path198" />
</g>
<g
transform="matrix(1.24571,0,0,1.35864,116.812,84.3924)"
id="g214">
<g
transform="matrix(64,0,0,64,42.1437,77.6203)"
id="g204">
<path
d="m 0.084,0 v -0.68 h 0.213 c 0.074,0 0.137,0.017 0.19,0.05 0.053,0.034 0.079,0.09 0.079,0.168 0,0.077 -0.028,0.134 -0.085,0.17 -0.057,0.037 -0.121,0.055 -0.193,0.055 H 0.213 V 0 Z m 0.209,-0.572 h -0.08 v 0.228 h 0.082 c 0.039,0 0.07,-0.009 0.094,-0.027 0.024,-0.017 0.037,-0.045 0.04,-0.083 0,-0.044 -0.012,-0.075 -0.036,-0.092 -0.024,-0.017 -0.057,-0.026 -0.1,-0.026 z"
style="fill:#ffffff;fill-rule:nonzero"
id="path202" />
</g>
<g
transform="matrix(64,0,0,64,79.7117,77.6203)"
id="g208">
<path
d="M 0.332,0 H 0.084 v -0.68 h 0.252 c 0.105,0 0.182,0.032 0.233,0.095 0.051,0.063 0.076,0.144 0.076,0.241 0,0.105 -0.027,0.189 -0.082,0.251 C 0.508,-0.031 0.431,0 0.332,0 Z M 0.337,-0.57 H 0.213 v 0.461 H 0.33 c 0.055,0 0.099,-0.018 0.132,-0.054 C 0.495,-0.199 0.511,-0.259 0.511,-0.344 0.511,-0.415 0.497,-0.47 0.469,-0.51 0.441,-0.55 0.397,-0.57 0.337,-0.57 Z"
style="fill:#ffffff;fill-rule:nonzero"
id="path206" />
</g>
<g
transform="matrix(64,0,0,64,123.424,77.6203)"
id="g212">
<path
d="M 0.405,-0.288 H 0.213 V 0 H 0.084 v -0.68 h 0.385 l 0.02,0.102 H 0.213 v 0.189 h 0.173 z"
style="fill:#ffffff;fill-rule:nonzero"
id="path210" />
</g>
</g>
</g>
</g>
<g
transform="matrix(1,0,0,1.52217,67.3796,10.7507)"
id="g222">
<rect
x="23.500999"
y="81.300003"
width="162.30499"
height="61.77"
style="fill:#b4d5ff"
id="rect220" />
</g>
<g
transform="matrix(0.967536,0,0,0.961535,5.90498,47.9703)"
id="g236">
<g
transform="matrix(90.4804,0,0,90.4804,82.6698,167.705)"
id="g226">
<path
d="m 0.057,-0.337 c 0,-0.105 0.027,-0.19 0.082,-0.257 0.055,-0.066 0.132,-0.1 0.231,-0.102 0.107,0 0.186,0.034 0.237,0.103 0.051,0.069 0.077,0.152 0.077,0.249 0,0.105 -0.027,0.191 -0.082,0.258 -0.055,0.067 -0.133,0.1 -0.232,0.1 C 0.264,0.014 0.185,-0.02 0.134,-0.089 0.083,-0.157 0.057,-0.24 0.057,-0.337 Z m 0.135,-0.001 c 0,0.071 0.014,0.13 0.043,0.175 0.029,0.045 0.073,0.068 0.134,0.068 0.055,0 0.098,-0.02 0.131,-0.061 0.033,-0.041 0.049,-0.103 0.049,-0.188 0,-0.071 -0.014,-0.129 -0.043,-0.174 -0.029,-0.045 -0.073,-0.068 -0.134,-0.068 -0.053,0 -0.097,0.022 -0.13,0.067 -0.033,0.045 -0.05,0.105 -0.05,0.181 z"
style="fill:#333333;fill-rule:nonzero"
id="path224" />
</g>
<g
transform="matrix(90.4804,0,0,90.4804,147.906,167.705)"
id="g230">
<path
d="M 0.505,-0.557 C 0.473,-0.567 0.448,-0.574 0.429,-0.579 0.41,-0.583 0.388,-0.585 0.361,-0.585 c -0.054,0 -0.096,0.022 -0.125,0.066 -0.029,0.044 -0.044,0.104 -0.044,0.181 0,0.066 0.012,0.123 0.037,0.171 0.025,0.048 0.066,0.072 0.124,0.072 0.029,0 0.056,-0.003 0.081,-0.009 0.025,-0.006 0.047,-0.013 0.068,-0.022 L 0.551,-0.03 C 0.525,-0.017 0.494,-0.006 0.457,0.002 0.42,0.01 0.388,0.014 0.36,0.014 0.254,0.014 0.177,-0.02 0.129,-0.088 0.081,-0.156 0.057,-0.239 0.057,-0.337 c 0,-0.105 0.027,-0.19 0.08,-0.257 0.053,-0.067 0.129,-0.1 0.228,-0.1 0.02,0 0.048,0.003 0.083,0.01 0.035,0.007 0.068,0.018 0.097,0.034 z"
style="fill:#333333;fill-rule:nonzero"
id="path228" />
</g>
<g
transform="matrix(90.4804,0,0,90.4804,199.751,167.705)"
id="g234">
<path
d="m 0.293,-0.572 h -0.08 v 0.208 h 0.082 c 0.039,0 0.071,-0.008 0.096,-0.024 0.025,-0.015 0.038,-0.041 0.038,-0.077 0,-0.038 -0.012,-0.065 -0.036,-0.082 -0.024,-0.017 -0.057,-0.025 -0.1,-0.025 z M 0.479,0 0.335,-0.26 C 0.328,-0.259 0.32,-0.259 0.312,-0.259 0.304,-0.258 0.296,-0.258 0.288,-0.258 H 0.213 V 0 H 0.084 v -0.68 h 0.213 c 0.074,0 0.137,0.017 0.19,0.051 0.053,0.034 0.079,0.087 0.079,0.158 0,0.042 -0.011,0.078 -0.032,0.108 -0.022,0.031 -0.05,0.054 -0.084,0.071 L 0.617,0 Z"
style="fill:#333333;fill-rule:nonzero"
id="path232" />
</g>
</g>
<g
transform="matrix(0.916882,0,0,1,121.475,-32.6535)"
id="g246">
<g
transform="matrix(86.953,0,0,86.953,152.996,241.878)"
id="g240">
<path
d="M 0.479,-0.428 C 0.5,-0.451 0.527,-0.47 0.562,-0.484 c 0.034,-0.013 0.065,-0.02 0.092,-0.02 0.066,0 0.113,0.019 0.141,0.058 0.027,0.039 0.041,0.086 0.041,0.142 V 0 H 0.705 v -0.298 c 0,-0.031 -0.007,-0.054 -0.022,-0.071 -0.015,-0.016 -0.036,-0.024 -0.064,-0.024 -0.019,0 -0.038,0.005 -0.059,0.015 -0.021,0.01 -0.039,0.021 -0.056,0.034 0.001,0.007 0.001,0.013 0.002,0.02 0.001,0.007 0.001,0.013 0.001,0.02 V 0 H 0.376 v -0.298 c 0,-0.031 -0.007,-0.054 -0.022,-0.071 -0.015,-0.016 -0.036,-0.024 -0.063,-0.024 -0.017,0 -0.033,0.003 -0.05,0.01 -0.017,0.007 -0.034,0.016 -0.049,0.027 V 0 H 0.062 V -0.485 H 0.13 l 0.032,0.044 c 0.022,-0.02 0.049,-0.035 0.08,-0.047 0.031,-0.011 0.058,-0.016 0.083,-0.016 0.038,0 0.07,0.007 0.095,0.02 0.025,0.014 0.045,0.033 0.059,0.056 z"
style="fill:#333333;fill-rule:nonzero"
id="path238" />
</g>
<g
transform="matrix(86.953,0,0,86.953,228.906,241.878)"
id="g244">
<path
d="M 0.156,0.023 0.179,-0.034 0.006,-0.467 0.14,-0.485 0.252,-0.191 0.358,-0.485 H 0.495 L 0.278,0.064 C 0.263,0.103 0.236,0.137 0.197,0.165 0.158,0.193 0.118,0.212 0.075,0.222 L 0.029,0.115 C 0.052,0.105 0.077,0.093 0.104,0.079 0.13,0.064 0.147,0.046 0.156,0.023 Z"
style="fill:#333333;fill-rule:nonzero"
id="path242" />
</g>
</g>
<g
id="Selectors"
transform="matrix(0.965977,0,0,0.807602,67.3796,67.3718)">
<g
id="Right-selector"
serif:id="Right selector">
<g
transform="matrix(1.03522,0,0,1.23823,2.07044,0)"
id="g250">
<path
d="M 185.806,161.156 V 67.132"
style="fill:none;stroke:#4c9fff;stroke-width:4px;stroke-linecap:butt"
id="path248" />
</g>
<g
transform="matrix(1.03522,0,0,1.23823,161.788,169.469)"
id="g254">
<circle
cx="31.523001"
cy="34.313999"
r="10.021"
style="fill:#4c9fff;stroke:#4c9fff;stroke-width:4px;stroke-linecap:butt"
id="circle252" />
</g>
</g>
<g
id="Left-selector"
serif:id="Left selector">
<g
transform="matrix(1.03522,0,0,1.23823,-170.092,0)"
id="g259">
<path
d="M 185.806,161.156 V 67.132"
style="fill:none;stroke:#4c9fff;stroke-width:4px;stroke-linecap:butt"
id="path257" />
</g>
<g
transform="matrix(1.03522,0,0,1.23823,-10.3742,28.2274)"
id="g263">
<circle
cx="31.523001"
cy="34.313999"
r="10.021"
style="fill:#4c9fff;stroke:#4c9fff;stroke-width:4px;stroke-linecap:butt"
id="circle261" />
</g>
</g>
</g>
</g>
</g>
</svg>

After

Width:  |  Height:  |  Size: 11 KiB

+73 -110
View File
@@ -12,21 +12,23 @@ system/platform. This version may be out of date, however.
These platforms have one-liner installs: These platforms have one-liner installs:
+-------------------------------+-------------------------------+ +-------------------------------+-----------------------------------------+
| Debian, Ubuntu | ``apt install ocrmypdf`` | | Debian, Ubuntu | ``apt install ocrmypdf`` |
+-------------------------------+-------------------------------+ +-------------------------------+-----------------------------------------+
| Windows Subsystem for Linux | ``apt install ocrmypdf`` | | Windows Subsystem for Linux | ``apt install ocrmypdf`` |
+-------------------------------+-------------------------------+ +-------------------------------+-----------------------------------------+
| Fedora | ``dnf install ocrmypdf`` | | Fedora | ``dnf install ocrmypdf`` |
+-------------------------------+-------------------------------+ +-------------------------------+-----------------------------------------+
| macOS | ``brew install ocrmypdf`` | | macOS | ``brew install ocrmypdf`` |
+-------------------------------+-------------------------------+ +-------------------------------+-----------------------------------------+
| LinuxBrew | ``brew install ocrmypdf`` | | LinuxBrew | ``brew install ocrmypdf`` |
+-------------------------------+-------------------------------+ +-------------------------------+-----------------------------------------+
| FreeBSD | ``pkg install py38-ocrmypdf`` | | FreeBSD | ``pkg install textproc/py-ocrmypdf`` |
+-------------------------------+-------------------------------+ +-------------------------------+-----------------------------------------+
| Conda (WSL, macOS, Linux) | ``conda install ocrmypdf`` | | Conda (WSL, macOS, Linux) | ``conda install ocrmypdf`` |
+-------------------------------+-------------------------------+ +-------------------------------+-----------------------------------------+
| Snap (snapcraft packaging) | ``snap install ocrmypdf`` |
+-------------------------------+-----------------------------------------+
More detailed procedures are outlined below. If you want to do a manual More detailed procedures are outlined below. If you want to do a manual
install, or install a more recent version than your platform provides, read on. install, or install a more recent version than your platform provides, read on.
@@ -41,11 +43,11 @@ Installing on Linux
Debian and Ubuntu 18.04 or newer Debian and Ubuntu 18.04 or newer
-------------------------------- --------------------------------
.. |deb-stable| image:: https://repology.org/badge/version-for-repo/debian_stable/ocrmypdf.svg .. |deb-11| image:: https://repology.org/badge/version-for-repo/debian_11/ocrmypdf.svg
:alt: Debian 9 stable ("stretch") :alt: Debian 11
.. |deb-testing| image:: https://repology.org/badge/version-for-repo/debian_testing/ocrmypdf.svg .. |deb-12| image:: https://repology.org/badge/version-for-repo/debian_12/ocrmypdf.svg
:alt: Debian 10 testing ("buster") :alt: Debian 12
.. |deb-unstable| image:: https://repology.org/badge/version-for-repo/debian_unstable/ocrmypdf.svg .. |deb-unstable| image:: https://repology.org/badge/version-for-repo/debian_unstable/ocrmypdf.svg
:alt: Debian unstable :alt: Debian unstable
@@ -56,17 +58,17 @@ Debian and Ubuntu 18.04 or newer
.. |ubu-2004| image:: https://repology.org/badge/version-for-repo/ubuntu_20_04/ocrmypdf.svg .. |ubu-2004| image:: https://repology.org/badge/version-for-repo/ubuntu_20_04/ocrmypdf.svg
:alt: Ubuntu 20.04 LTS :alt: Ubuntu 20.04 LTS
.. |ubu-2110| image:: https://repology.org/badge/version-for-repo/ubuntu_21_10/ocrmypdf.svg .. |ubu-2204| image:: https://repology.org/badge/version-for-repo/ubuntu_22_04/ocrmypdf.svg
:alt: Ubuntu 21.10 :alt: Ubuntu 22.04 LTS
+-----------------------------------------------+ +-----------------------------------------------+
| **OCRmyPDF versions in Debian & Ubuntu** | | **OCRmyPDF versions in Debian & Ubuntu** |
+-----------------------------------------------+ +-----------------------------------------------+
| |latest| | | |latest| |
+-----------------------------------------------+ +-----------------------------------------------+
| |deb-stable| |deb-testing| |deb-unstable| | | |deb-11| |deb-12| |deb-unstable| |
+-----------------------------------------------+ +-----------------------------------------------+
| |ubu-1804| |ubu-2004| |ubu-2110| | | |ubu-1804| |ubu-2004| |ubu-2204| |
+-----------------------------------------------+ +-----------------------------------------------+
Users of Debian 9 ("stretch") or later, or Ubuntu 18.04 or later, including users Users of Debian 9 ("stretch") or later, or Ubuntu 18.04 or later, including users
@@ -80,8 +82,7 @@ As indicated in the table above, Debian and Ubuntu releases may lag
behind the latest version. If the version available for your platform is behind the latest version. If the version available for your platform is
out of date, you could opt to install the latest version from source. out of date, you could opt to install the latest version from source.
See `Installing HEAD revision from See `Installing HEAD revision from
sources <#installing-head-revision-from-sources>`__. Ubuntu 16.10 to 17.10 sources <#installing-head-revision-from-sources>`__.
inclusive also had ocrmypdf, but these versions are end of life.
For full details on version availability for your platform, check the For full details on version availability for your platform, check the
`Debian Package Tracker <https://tracker.debian.org/pkg/ocrmypdf>`__ or `Debian Package Tracker <https://tracker.debian.org/pkg/ocrmypdf>`__ or
@@ -91,19 +92,19 @@ For full details on version availability for your platform, check the
OCRmyPDF for Debian and Ubuntu currently omit the JBIG2 encoder. OCRmyPDF for Debian and Ubuntu currently omit the JBIG2 encoder.
OCRmyPDF works fine without it but will produce larger output files. OCRmyPDF works fine without it but will produce larger output files.
If you build jbig2enc from source, ocrmypdf 7.0.0 and later will If you build jbig2enc from source, ocrmypdf will
automatically detect it (specifically the ``jbig2`` binary) on the automatically detect it (specifically the ``jbig2`` binary) on the
``PATH``. To add JBIG2 encoding, see :ref:`jbig2`. ``PATH``. To add JBIG2 encoding, see :ref:`jbig2`.
Fedora Fedora
------ ------
.. |fedora-34| image:: https://repology.org/badge/version-for-repo/fedora_34/ocrmypdf.svg
:alt: Fedora 34
.. |fedora-35| image:: https://repology.org/badge/version-for-repo/fedora_35/ocrmypdf.svg .. |fedora-35| image:: https://repology.org/badge/version-for-repo/fedora_35/ocrmypdf.svg
:alt: Fedora 35 :alt: Fedora 35
.. |fedora-36| image:: https://repology.org/badge/version-for-repo/fedora_36/ocrmypdf.svg
:alt: Fedora 36
.. |fedora-rawhide| image:: https://repology.org/badge/version-for-repo/fedora_rawhide/ocrmypdf.svg .. |fedora-rawhide| image:: https://repology.org/badge/version-for-repo/fedora_rawhide/ocrmypdf.svg
:alt: Fedore Rawhide :alt: Fedore Rawhide
@@ -112,7 +113,7 @@ Fedora
+-----------------------------------------------+ +-----------------------------------------------+
| |latest| | | |latest| |
+-----------------------------------------------+ +-----------------------------------------------+
| |fedora-34| |fedora-35| |fedora-rawhide| | | |fedora-35| |fedora-36| |fedora-rawhide| |
+-----------------------------------------------+ +-----------------------------------------------+
Users of Fedora 29 or later may simply Users of Fedora 29 or later may simply
@@ -138,9 +139,29 @@ from sources <#installing-head-revision-from-sources>`__.
.. _ubuntu-lts-latest: .. _ubuntu-lts-latest:
Installing the latest version on Ubuntu 20.04 LTS Installing the latest version on Ubuntu 22.04 LTS
------------------------------------------------- -------------------------------------------------
Ubuntu 22.04 includes ocrmypdf 13.4.0 - you can install that with
``apt install ocrmypdf``. To install a more recent version for the current
user, follow these steps:
.. code-block:: bash
sudo apt-get update
sudo apt-get -y install ocrmypdf python3-pip
pip install --user --upgrade ocrmypdf
If you get the message ``WARNING: The script ocrmypdf is installed in
'/home/$USER/.local/bin' which is not on PATH.``, you may need to re-login
or open a new shell, or manually add this to your user's PATH.
To add JBIG2 encoding, see :ref:`jbig2`.
Ubuntu 20.04 LTS
----------------
Ubuntu 20.04 includes ocrmypdf 9.6.0 - you can install that with ``apt``. To Ubuntu 20.04 includes ocrmypdf 9.6.0 - you can install that with ``apt``. To
install a more recent version, uninstall the system-provided version of install a more recent version, uninstall the system-provided version of
ocrmypdf, and install the following dependencies: ocrmypdf, and install the following dependencies:
@@ -171,6 +192,8 @@ To install for the current user only:
export PATH=$HOME/.local/bin:$PATH export PATH=$HOME/.local/bin:$PATH
pip3 install --user ocrmypdf pip3 install --user ocrmypdf
To add JBIG2 encoding, see :ref:`jbig2`.
Ubuntu 18.04 LTS Ubuntu 18.04 LTS
---------------- ----------------
@@ -291,46 +314,6 @@ To install OCRmyPDF for Alpine Linux:
apk add ocrmypdf apk add ocrmypdf
Mageia 7
--------
There is no OS-level packaging available for Mageia, so you must install the
dependencies:
.. code-block:: bash
# As root user
urpmi.update -a
urpmi \
ghostscript \
icc-profiles-openicc \
jbig2dec \
pngquant \
python3-pip \
python3-distutils-extra \
python3-pkg-resources \
python3-reportlab \
qpdf \
tesseract \
tesseract-osd \
tesseract-eng \
tesseract-fra
To install ocrmypdf for the system:
.. code-block:: bash
# As root user
pip3 install ocrmypdf
ldconfig
Or, to install for the current user only:
.. code-block:: bash
export PATH=$HOME/.local/bin:$PATH
pip3 install --user ocrmypdf
Other Linux packages Other Linux packages
-------------------- --------------------
@@ -365,19 +348,6 @@ languages you can optionally install them all:
brew install tesseract-lang # Optional: Install all language packs brew install tesseract-lang # Optional: Install all language packs
.. note::
Users who previously installed OCRmyPDF on macOS using
``pip install ocrmypdf`` should remove the pip version
(``pip3 uninstall ocrmypdf``) before switching to the Homebrew
version.
.. note::
Users who previously installed OCRmyPDF from the private tap should
switch to the mainline version (``brew untap jbarlow83/ocrmypdf``)
and install from there.
Manual installation on macOS Manual installation on macOS
---------------------------- ----------------------------
@@ -411,19 +381,19 @@ Update the homebrew pip:
.. code-block:: bash .. code-block:: bash
pip3 install --upgrade pip pip install --upgrade pip
You can then install OCRmyPDF from PyPI, for the current user: You can then install OCRmyPDF from PyPI, for the current user:
.. code-block:: bash .. code-block:: bash
pip3 install --user ocrmypdf pip install --user ocrmypdf
or system-wide: or system-wide:
.. code-block:: bash .. code-block:: bash
pip3 install ocrmypdf pip install ocrmypdf
The command line program should now be available: The command line program should now be available:
@@ -485,8 +455,8 @@ to change the PATH.
Windows Subsystem for Linux Windows Subsystem for Linux
--------------------------- ---------------------------
#. Install Ubuntu 20.04 for Windows Subsystem for Linux, if not already installed. #. Install Ubuntu 22.04 for Windows Subsystem for Linux, if not already installed.
#. Follow the procedure to install :ref:`OCRmyPDF on Ubuntu 20.04 <ubuntu-lts-latest>`. #. Follow the procedure to install :ref:`OCRmyPDF on Ubuntu 22.04 <ubuntu-lts-latest>`.
#. Open the Windows command prompt and create a symlink: #. Open the Windows command prompt and create a symlink:
.. code-block:: powershell .. code-block:: powershell
@@ -558,16 +528,13 @@ your command prompt can run the docker "hello world" container.
Installing on FreeBSD Installing on FreeBSD
===================== =====================
.. image:: https://repology.org/badge/version-for-repo/freebsd/python:ocrmypdf.svg .. image:: https://repology.org/badge/version-for-repo/freebsd/ocrmypdf.svg
:alt: FreeBSD :alt: FreeBSD
:target: https://repology.org/project/python:ocrmypdf/versions :target: https://repology.org/project/ocrmypdf/versions
FreeBSD 11.3, 12.0, 12.1-RELEASE and 13.0-CURRENT are supported. Other
versions likely work but have not been tested.
.. code-block:: bash .. code-block:: bash
pkg install py38-ocrmypdf pkg install textproc/py-ocrmypdf
To install a more recent version, you could attempt to first install the system To install a more recent version, you could attempt to first install the system
version with ``pkg``, then use ``pip install --user ocrmypdf``. version with ``pkg``, then use ``pip install --user ocrmypdf``.
@@ -616,18 +583,18 @@ try:
.. code-block:: bash .. code-block:: bash
pip3 install --user ocrmypdf pip install --user ocrmypdf
You should then be able to run ``ocrmypdf --version`` and see that the You should then be able to run ``ocrmypdf --version`` and see that the
latest version was located. latest version was located.
Since ``pip3 install --user`` does not work correctly on some platforms, Since ``pip install --user`` does not work correctly on some platforms,
notably Ubuntu 16.04 and older, and the Homebrew version of Python, notably Ubuntu 16.04 and older, and the Homebrew version of Python,
instead use this for a system wide installation: instead use this for a system wide installation:
.. code-block:: bash .. code-block:: bash
pip3 install ocrmypdf pip install ocrmypdf
.. note:: .. note::
@@ -643,13 +610,9 @@ OCRmyPDF currently requires these external programs and libraries to be
installed, and must be satisfied using the operating system package installed, and must be satisfied using the operating system package
manager. ``pip`` cannot provide them. manager. ``pip`` cannot provide them.
The following versions are required:
- Python 3.7 or newer - Python 3.7 or newer
- Ghostscript 9.15 or newer
- Tesseract 4.0.0-beta or newer
As of ocrmypdf 7.2.1, the following versions are recommended:
- Python 3.9 or newer
- Ghostscript 9.23 or newer - Ghostscript 9.23 or newer
- Tesseract 4.0.0 or newer - Tesseract 4.0.0 or newer
- jbig2enc 0.29 or newer - jbig2enc 0.29 or newer
@@ -696,7 +659,7 @@ environment:
.. code-block:: bash .. code-block:: bash
pip3 install git+https://github.com/jbarlow83/OCRmyPDF.git pip install git+https://github.com/ocrmypdf/OCRmyPDF.git
Or, to install in `development Or, to install in `development
mode <https://pythonhosted.org/setuptools/setuptools.html#development-mode>`__, mode <https://pythonhosted.org/setuptools/setuptools.html#development-mode>`__,
@@ -704,18 +667,18 @@ allowing customization of OCRmyPDF, use the ``-e`` flag:
.. code-block:: bash .. code-block:: bash
pip3 install -e git+https://github.com/jbarlow83/OCRmyPDF.git pip install -e git+https://github.com/ocrmypdf/OCRmyPDF.git
You may find it easiest to install in a virtual environment, rather than You may find it easiest to install in a virtual environment, rather than
system-wide: system-wide:
.. code-block:: bash .. code-block:: bash
git clone -b master https://github.com/jbarlow83/OCRmyPDF.git git clone -b master https://github.com/ocrmypdf/OCRmyPDF.git
python3 -m venv python3 -m venv
source venv/bin/activate source venv/bin/activate
cd OCRmyPDF cd OCRmyPDF
pip3 install . pip install .
However, ``ocrmypdf`` will only be accessible on the system PATH when However, ``ocrmypdf`` will only be accessible on the system PATH when
you activate the virtual environment. you activate the virtual environment.
@@ -738,8 +701,8 @@ To install all of the development and test requirements:
.. code-block:: bash .. code-block:: bash
git clone -b master https://github.com/jbarlow83/OCRmyPDF.git git clone -b master https://github.com/ocrmypdf/OCRmyPDF.git
python3 -m venv python -m venv
source venv/bin/activate source venv/bin/activate
cd OCRmyPDF cd OCRmyPDF
pip install -e .[test] pip install -e .[test]
+3
View File
@@ -32,6 +32,9 @@ For all other Linux, you must build a JBIG2 encoder from source:
.. _jbig2-lossy: .. _jbig2-lossy:
Dependencies include libtoolize and libleptonica, which on Ubuntu systems
are packaged as libtool and libleptonica-dev.
Lossy mode JBIG2 Lossy mode JBIG2
================ ================
+27
View File
@@ -54,6 +54,33 @@ to what languages it should search for. Multiple languages can be
requested using either ``-l eng+fra`` (English and French) or requested using either ``-l eng+fra`` (English and French) or
``-l eng -l fra``. ``-l eng -l fra``.
Gentoo users
============
On Gentoo the package ``app-text/tessdata_fast``, which ``app-text/tesseract`` depends on, handles Tesseract languages.
It accepts USE flags to select what languages should be installed, these can be set in ``/etc/portage/package.use``.
Alternatively one can globally set the `L10N use extension <https://wiki.gentoo.org/wiki/Localization/Guide#L10N>`__ in ``/etc/portage/make.conf``.
This enables these languages for all packages (e.g. including aspell).
.. code-block:: bash
# Display a list of all Tesseract language packs
equery uses app-text/tessdata_fast
# Add English and German language support for Tesseract only
echo 'app-text/tessdata_fast l10n_de l10n_en' >> /etc/portage/package.use
# Add global English and German language support (the `l10n_` from equery has to be omited)
echo L10N="de en" >> /etc/portage/make.conf
# update system to reflect changed USE flags
emerge --update --deep --newuse @world
You can then pass the ``-l LANG`` argument to OCRmyPDF to give a hint as
to what languages it should search for. Multiple languages can be
requested using either ``-l eng+fra`` (English and French) or
``-l eng -l fra``.
macOS users macOS users
=========== ===========
+5 -5
View File
@@ -2,7 +2,7 @@
Maintainer notes Maintainer notes
================ ================
This is for those who package OCRmyPDF for downstream use. (Thank you This is for those who package OCRmyPDF for downstream use. (Thank you
for your hard work.) for your hard work.)
Known ports/packagers Known ports/packagers
@@ -25,7 +25,7 @@ Non-Python dependencies
Note that we have non-Python dependencies. In particular, OCRmyPDF requires Note that we have non-Python dependencies. In particular, OCRmyPDF requires
Ghostscript and Tesseract OCR to be installed and needs to be able to locate their Ghostscript and Tesseract OCR to be installed and needs to be able to locate their
binaries on the system PATH. On Windows, OCRmyPDF will also check the registry binaries on the system PATH. On Windows, OCRmyPDF will also check the registry
for their locations. for their locations.
Tesseract OCR relies on SIMD for performance and only has proper support for this Tesseract OCR relies on SIMD for performance and only has proper support for this
@@ -38,13 +38,13 @@ OCRmyPDF uses setuptools-scm for versioning, which derives the version from
Git as a single source of truth. This may be unsuitable for some distributions, e.g. Git as a single source of truth. This may be unsuitable for some distributions, e.g.
to indicate that your distribution modifies OCRmyPDF in some way. to indicate that your distribution modifies OCRmyPDF in some way.
You can patch the ``__version__`` variable in ``src/ocrmypdf/_version.py`` if You can patch the ``__version__`` variable in ``src/ocrmypdf/_version.py`` if
necessary. necessary.
OCRmyPDF uses setuptools-scm-git-archive to ensure that tarballs downloaded from OCRmyPDF uses setuptools-scm-git-archive to ensure that tarballs downloaded from
GitHub contain version information. Unfortunately, these tarballs are not always GitHub contain version information. Unfortunately, these tarballs are not always
deterministic. See this deterministic. See this
`issue <https://github.com/jbarlow83/OCRmyPDF/issues/841#issuecomment-936562696>`_. `issue <https://github.com/ocrmypdf/OCRmyPDF/issues/841#issuecomment-936562696>`_.
jbig2enc jbig2enc
-------- --------
+13 -1
View File
@@ -162,6 +162,11 @@ Examples
management system. management system.
Suppressing or overriding other plugins
---------------------------------------
.. autofunction:: ocrmypdf.pluginspec.initialize
Custom command line arguments Custom command line arguments
----------------------------- -----------------------------
@@ -172,7 +177,7 @@ Custom command line arguments
Execution and progress reporting Execution and progress reporting
-------------------------------- --------------------------------
.. autoclass: ocrmypdf.pluginspec.Executor .. autoclass:: ocrmypdf.pluginspec.Executor
:members: :members:
.. autofunction:: ocrmypdf.pluginspec.get_logging_console .. autofunction:: ocrmypdf.pluginspec.get_logging_console
@@ -216,3 +221,10 @@ PDF/A production
---------------- ----------------
.. autofunction:: ocrmypdf.pluginspec.generate_pdfa .. autofunction:: ocrmypdf.pluginspec.generate_pdfa
PDF optimization
----------------
.. autofunction:: ocrmypdf.pluginspec.optimize_pdf
.. autofunction:: ocrmypdf.pluginspec.is_optimization_enabled
+73
View File
@@ -16,8 +16,81 @@ The most recent release of OCRmyPDF is |OCRmyPDF PyPI|. Any newer versions
referred to in these notes may exist the main branch but have not been referred to in these notes may exist the main branch but have not been
tagged yet. tagged yet.
.. note::
Attention maintainers: that these release notes may be updated with information
about a forthcoming release that has not been tagged yet. A release is only
official when it's tagged and posted to PyPI.
.. |OCRmyPDF PyPI| image:: https://img.shields.io/pypi/v/ocrmypdf.svg .. |OCRmyPDF PyPI| image:: https://img.shields.io/pypi/v/ocrmypdf.svg
v13.6.2
=======
- Added a shim to prevent an "error during error handling" for Python 3.7 and 3.8.
- Modernized some type annotations.
- Improved annotations on our _windows module to help IDEs and mypy figure out what
we're doing.
v13.6.1
=======
- Require setuptools-scm 7.0.5 to avoid possible issues with source distributions in
earlier versions of setuptools-scm.
- Suppress a spurious warning, improve tests, improve typing and other miscellany.
v13.6.0
=======
- Added a new ``initialize`` plugin hook, making it possible to suppress built-in
plugins more easily, among other possibilities.
- Fixed an issue where unpaper would exit with a "wrong stream" error, probably
related to images with an odd integer width. :issue:`887, 665`
v13.5.0
=======
- Added a new ``optimize_pdf`` plugin hook, making it possible to create plugins that
replace or enhance OCRmyPDF's PDF optimizer.
- Removed all max version restrictions. Our new policy is to blacklist known-bad releases
and only block known-bad versions of dependencies.
- The naming schema for object that holds all OCR text that OCRmyPDF inserts has
changed. This has always been an implementation detail (and remains so), but possibly,
someone was relying on it and would appreciate the heads-up.
- Cleanup.
v13.4.7
=======
- Fixed PermissionError when cleaning up temporary files in rare cases. :issue:`974`
- Fixed PermissionError when calling ``os.nice`` on platforms that lack it. :issue:`973`
- Suppressed some warnings from libxmp during tests.
v13.4.6
=======
- Convert error on corrupt ICC profiles into a warning. Thanks to @oscherler.
v13.4.5
=======
- Remove upper bound on pdfminer.six version.
- Documentation.
v13.4.4
=======
- Updated pdfminer.six version.
- Docker image changed to Ubuntu 22.04 now that it is released and provides the
dependencies we need. This seems more consistent than our recent change to
Debian.
v13.4.3
=======
- Fix error on pytest.skip() with older versions of pytest.
- Documentation updates.
v13.4.2 v13.4.2
======= =======
+2 -1
View File
@@ -19,8 +19,9 @@
# OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE # OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
# SOFTWARE. # SOFTWARE.
# This script must be edited to meet your needs. from __future__ import annotations
# This script must be edited to meet your needs.
import logging import logging
import os import os
import sys import sys
+2
View File
@@ -37,6 +37,8 @@ To use this as an API:
) )
""" """
from __future__ import annotations
import logging import logging
from PIL import Image from PIL import Image
+2 -1
View File
@@ -19,8 +19,9 @@
# OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE # OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
# SOFTWARE. # SOFTWARE.
# This script must be edited to meet your needs. from __future__ import annotations
# This script must be edited to meet your needs.
import logging import logging
import os import os
import shutil import shutil
+16 -4
View File
@@ -20,9 +20,12 @@
# OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE # OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
# SOFTWARE. # SOFTWARE.
from __future__ import annotations
import json import json
import logging import logging
import os import os
import shutil
import sys import sys
import time import time
from datetime import datetime from datetime import datetime
@@ -44,8 +47,10 @@ def getenv_bool(name: str, default: str = 'False'):
INPUT_DIRECTORY = os.getenv('OCR_INPUT_DIRECTORY', '/input') INPUT_DIRECTORY = os.getenv('OCR_INPUT_DIRECTORY', '/input')
OUTPUT_DIRECTORY = os.getenv('OCR_OUTPUT_DIRECTORY', '/output') OUTPUT_DIRECTORY = os.getenv('OCR_OUTPUT_DIRECTORY', '/output')
ARCHIVE_DIRECTORY = os.getenv('OCR_ARCHIVE_DIRECTORY', '/processed')
OUTPUT_DIRECTORY_YEAR_MONTH = getenv_bool('OCR_OUTPUT_DIRECTORY_YEAR_MONTH') OUTPUT_DIRECTORY_YEAR_MONTH = getenv_bool('OCR_OUTPUT_DIRECTORY_YEAR_MONTH')
ON_SUCCESS_DELETE = getenv_bool('OCR_ON_SUCCESS_DELETE') ON_SUCCESS_DELETE = getenv_bool('OCR_ON_SUCCESS_DELETE')
ON_SUCCESS_ARCHIVE = getenv_bool('OCR_ON_SUCCESS_ARCHIVE')
DESKEW = getenv_bool('OCR_DESKEW') DESKEW = getenv_bool('OCR_DESKEW')
OCR_JSON_SETTINGS = json.loads(os.getenv('OCR_JSON_SETTINGS', '{}')) OCR_JSON_SETTINGS = json.loads(os.getenv('OCR_JSON_SETTINGS', '{}'))
POLL_NEW_FILE_SECONDS = int(os.getenv('OCR_POLL_NEW_FILE_SECONDS', '1')) POLL_NEW_FILE_SECONDS = int(os.getenv('OCR_POLL_NEW_FILE_SECONDS', '1'))
@@ -108,9 +113,13 @@ def execute_ocrmypdf(file_path):
deskew=DESKEW, deskew=DESKEW,
**OCR_JSON_SETTINGS, **OCR_JSON_SETTINGS,
) )
if exit_code == 0 and ON_SUCCESS_DELETE: if exit_code == 0:
log.info(f'OCR is done. Deleting: {file_path}') if ON_SUCCESS_DELETE:
file_path.unlink() log.info(f'OCR is done. Deleting: {file_path}')
file_path.unlink()
elif ON_SUCCESS_ARCHIVE:
log.info(f'OCR is done. Archiving {file_path.name} to {ARCHIVE_DIRECTORY}')
shutil.move(file_path, f'{ARCHIVE_DIRECTORY}/{file_path.name}')
else: else:
log.info('OCR is done') log.info('OCR is done')
@@ -135,13 +144,16 @@ def main():
f"Starting OCRmyPDF watcher with config:\n" f"Starting OCRmyPDF watcher with config:\n"
f"Input Directory: {INPUT_DIRECTORY}\n" f"Input Directory: {INPUT_DIRECTORY}\n"
f"Output Directory: {OUTPUT_DIRECTORY}\n" f"Output Directory: {OUTPUT_DIRECTORY}\n"
f"Output Directory Year & Month: {OUTPUT_DIRECTORY_YEAR_MONTH}" f"Output Directory Year & Month: {OUTPUT_DIRECTORY_YEAR_MONTH}\n"
f"Archive Directory: {ARCHIVE_DIRECTORY}"
) )
log.debug( log.debug(
f"INPUT_DIRECTORY: {INPUT_DIRECTORY}\n" f"INPUT_DIRECTORY: {INPUT_DIRECTORY}\n"
f"OUTPUT_DIRECTORY: {OUTPUT_DIRECTORY}\n" f"OUTPUT_DIRECTORY: {OUTPUT_DIRECTORY}\n"
f"OUTPUT_DIRECTORY_YEAR_MONTH: {OUTPUT_DIRECTORY_YEAR_MONTH}\n" f"OUTPUT_DIRECTORY_YEAR_MONTH: {OUTPUT_DIRECTORY_YEAR_MONTH}\n"
f"ARCHIVE_DIRECTORY: {ARCHIVE_DIRECTORY}\n"
f"ON_SUCCESS_DELETE: {ON_SUCCESS_DELETE}\n" f"ON_SUCCESS_DELETE: {ON_SUCCESS_DELETE}\n"
f"ON_SUCCESS_ARCHIVE: {ON_SUCCESS_ARCHIVE}\n"
f"DESKEW: {DESKEW}\n" f"DESKEW: {DESKEW}\n"
f"ARGS: {OCR_JSON_SETTINGS}\n" f"ARGS: {OCR_JSON_SETTINGS}\n"
f"POLL_NEW_FILE_SECONDS: {POLL_NEW_FILE_SECONDS}\n" f"POLL_NEW_FILE_SECONDS: {POLL_NEW_FILE_SECONDS}\n"
+2
View File
@@ -24,6 +24,8 @@ to emphasize that SaaS deployments should make sure they comply with
Ghostscript's license as well as OCRmyPDF's. Ghostscript's license as well as OCRmyPDF's.
""" """
from __future__ import annotations
import os import os
import shlex import shlex
from subprocess import PIPE, run from subprocess import PIPE, run
+10 -6
View File
@@ -1,14 +1,12 @@
[build-system] [build-system]
requires = [ requires = [
"setuptools >= 30.3.0", "setuptools >= 52",
"wheel", "setuptools_scm[toml] >= 7.0.5",
"setuptools_scm[toml] >= 3.4", "wheel"
"setuptools_scm_git_archive"
] ]
build-backend = "setuptools.build_meta" build-backend = "setuptools.build_meta"
[tool.setuptools_scm] [tool.setuptools_scm]
version_scheme = "post-release"
[tool.black] [tool.black]
line-length = 88 line-length = 88
@@ -96,6 +94,12 @@ module = [
'pdfminer.*', 'pdfminer.*',
'reportlab.*', 'reportlab.*',
'fitz', 'fitz',
'libxmp.utils' 'libxmp.utils',
'importlib_metadata'
] ]
ignore_missing_imports = true ignore_missing_imports = true
[tool.pylint.basic]
good-names = ["i", "j", "k", "ex", "Run", "_", "e", "p", "im", "w", "h", "m", "x", "y", "a", "b", "fp", "n", "f", "s", "v", "q", "dx", "dy"]
logging-format-style = "old"
disable = ["raw-checker-failed", "bad-inline-option", "locally-disabled", "file-ignored", "suppressed-message", "useless-suppression", "deprecated-pragma", "use-symbolic-message-instead", "logging-fstring-interpolation", "missing-function-docstring", "too-few-public-methods"]
+11 -8
View File
@@ -3,7 +3,7 @@ name = ocrmypdf
description = OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched description = OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
long_description = file: README.md long_description = file: README.md
long_description_content_type = text/markdown long_description_content_type = text/markdown
url = https://github.com/jbarlow83/OCRmyPDF url = https://github.com/ocrmypdf/OCRmyPDF
author = James R. Barlow author = James R. Barlow
author_email = james@purplerock.ca author_email = james@purplerock.ca
license = MPL-2.0 license = MPL-2.0
@@ -39,23 +39,24 @@ keywords =
scanning scanning
project_urls = project_urls =
Documentation = https://ocrmypdf.readthedocs.io/ Documentation = https://ocrmypdf.readthedocs.io/
Source = https://github.com/jbarlow83/ocrmypdf Source = https://github.com/ocrmypdf/OCRmyPDF
Tracker = https://github.com/jbarlow83/ocrmypdf/issues Tracker = https://github.com/ocrmypdf/OCRmyPDF/issues
[options] [options]
packages = find: packages = find:
install_requires = install_requires =
Pillow>=8.2.0 Pillow>=8.2.0
coloredlogs>=14.0 # strictly optional coloredlogs>=14.0 # strictly optional
img2pdf>=0.3.0,<0.5 # pure Python img2pdf>=0.3.0 # pure Python
packaging>=20 packaging>=20
pdfminer.six!=20200720,>=20191110,<=20220319 pdfminer.six!=20200720,>=20191110
pikepdf!=5.0.0,>=4.0.0 pikepdf!=5.0.0,>=4.0.0
pluggy>=0.13.0,<2 pluggy>=0.13.0
reportlab>=3.5.66 reportlab>=3.5.66
tqdm>=4 tqdm>=4
importlib-metadata>=4;python_version<'3.8' # until Python 3.8 importlib-metadata>=4;python_version<'3.8' # until Python 3.8
importlib-resources>=5;python_version<'3.9' # until Python 3.9 importlib-resources>=5;python_version<'3.9' # until Python 3.9
typing-extensions>=4;python_version<'3.10'
python_requires = >=3.7 python_requires = >=3.7
include_package_data = True include_package_data = True
package_dir = package_dir =
@@ -86,10 +87,12 @@ test =
pytest-cov>=2.11.1 pytest-cov>=2.11.1
pytest-xdist>=2.2.0 pytest-xdist>=2.2.0
python-xmp-toolkit==2.0.1 # also requires apt-get install libexempi3 python-xmp-toolkit==2.0.1 # also requires apt-get install libexempi3
types-Pillow
types-humanfriendly
watcher = watcher =
watchdog>=1.0.2,<3 watchdog>=1.0.2
webservice = webservice =
Flask>=1,<3 Flask>=1
[options.package_data] [options.package_data]
ocrmypdf = ocrmypdf =
+4 -8
View File
@@ -4,14 +4,10 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""setup.py to support older setuptools and pip."""
from __future__ import annotations
from setuptools import setup from setuptools import setup
# Minimal setup to support older setuptools/setuptools_scm setup()
setup(
setup_requires=[ # can be removed whenever we can drop pip 9 support
'setuptools_scm', # so that version will work
'setuptools_scm_git_archive', # enable version from github tarballs
],
use_scm_version={'version_scheme': 'post-release'},
)
+72
View File
@@ -0,0 +1,72 @@
name: ocrmypdf
title: OCRmyPDF
base: core20
version: git
summary: OCRmyPDF adds optical character recognition (OCR) to PDFs
description: OCRmyPDF packaged for snap
grade: stable
confinement: strict
icon: docs/images/logo-square-256.svg
license: MPL-2.0
architectures: [amd64]
environment:
TESSDATA_PREFIX: $SNAP/usr/share/tesseract-ocr/4.00/tessdata
GS_LIB: $SNAP/usr/share/ghostscript/9.50/Resource/Init
GS_FONTPATH: $SNAP/usr/share/ghostscript/9.50/Resource/Font
LD_LIBRARY_PATH: $SNAP/usr/lib/x86_64-linux-gnu
apps:
ocrmypdf:
command: usr/bin/snapcraft-preload python3 -m ocrmypdf
plugs:
- desktop
- desktop-legacy
- wayland
- x11
- home
- removable-media
parts:
snapcraft-preload:
source: https://github.com/sergiusens/snapcraft-preload.git
plugin: cmake
cmake-parameters:
- -DCMAKE_INSTALL_PREFIX=/usr -DLIBPATH=/usr/lib
build-packages:
- on amd64:
- gcc-multilib
- g++-multilib
stage-packages:
- lib32stdc++6
ocrmypdf:
plugin: python
source: https://github.com/ocrmypdf/OCRmyPDF.git
stage-packages:
- ghostscript
- icc-profiles-free
- liblept5
- libxml2
- pngquant
- tesseract-ocr-all
- unpaper
- qpdf
- zlib1g
python-packages:
- cffi
- pdfminer.six
- pikepdf
- Pillow
- pluggy
- reportlab
- setuptools
- tqdm
- pipe
override-build: |
snapcraftctl build
ln -sf ../usr/lib/libsnapcraft-preload.so $SNAPCRAFT_PART_INSTALL/lib/libsnapcraft-preload.so
+3
View File
@@ -4,6 +4,9 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Adds OCR layer to PDFs."""
from __future__ import annotations
from pluggy import HookimplMarker as _HookimplMarker from pluggy import HookimplMarker as _HookimplMarker
+6 -2
View File
@@ -5,11 +5,15 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""ocrmypdf command line entrypoint."""
from __future__ import annotations
import logging import logging
import os import os
import signal import signal
import sys import sys
from contextlib import suppress
from multiprocessing import set_start_method from multiprocessing import set_start_method
from ocrmypdf import __version__ from ocrmypdf import __version__
@@ -34,7 +38,7 @@ def sigbus(*args):
def run(args=None): def run(args=None):
_parser, options, plugin_manager = get_parser_options_plugins(args=args) _parser, options, plugin_manager = get_parser_options_plugins(args=args)
if hasattr(os, 'nice'): with suppress(AttributeError, PermissionError):
os.nice(5) os.nice(5)
verbosity = options.verbose verbosity = options.verbose
@@ -62,7 +66,7 @@ def run(args=None):
log.error(e) log.error(e)
return ExitCode.missing_dependency return ExitCode.missing_dependency
if hasattr(signal, 'SIGBUS'): with suppress(AttributeError, OSError):
signal.signal(signal.SIGBUS, sigbus) signal.signal(signal.SIGBUS, sigbus)
result = run_pipeline(options=options, plugin_manager=plugin_manager) result = run_pipeline(options=options, plugin_manager=plugin_manager)
+13 -5
View File
@@ -4,9 +4,13 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""OCRmyPDF concurrency abstractions."""
from __future__ import annotations
import threading import threading
from abc import ABC, abstractmethod from abc import ABC, abstractmethod
from typing import Callable, Iterable, Optional from typing import Callable, Iterable
def _task_noop(*_args, **_kwargs): def _task_noop(*_args, **_kwargs):
@@ -14,6 +18,8 @@ def _task_noop(*_args, **_kwargs):
class NullProgressBar: class NullProgressBar:
"""Progress bar API that takes no actions."""
def __init__(self, **kwargs): def __init__(self, **kwargs):
pass pass
@@ -28,6 +34,8 @@ class NullProgressBar:
class Executor(ABC): class Executor(ABC):
"""Abstract concurrent executor."""
pool_lock = threading.Lock() pool_lock = threading.Lock()
pbar_class = NullProgressBar pbar_class = NullProgressBar
@@ -41,10 +49,10 @@ class Executor(ABC):
use_threads: bool, use_threads: bool,
max_workers: int, max_workers: int,
tqdm_kwargs: dict, tqdm_kwargs: dict,
worker_initializer: Optional[Callable] = None, worker_initializer: Callable | None = None,
task: Optional[Callable] = None, task: Callable | None = None,
task_arguments: Optional[Iterable] = None, task_arguments: Iterable | None = None,
task_finished: Optional[Callable] = None, task_finished: Callable | None = None,
) -> None: ) -> None:
""" """
Set up parallel execution and progress reporting. Set up parallel execution and progress reporting.
+2
View File
@@ -6,3 +6,5 @@
"""Manage third party executables""" """Manage third party executables"""
from __future__ import annotations
+17 -26
View File
@@ -7,6 +7,8 @@
"""Interface to Ghostscript executable""" """Interface to Ghostscript executable"""
from __future__ import annotations
import logging import logging
import os import os
import re import re
@@ -14,13 +16,11 @@ import sys
from io import BytesIO from io import BytesIO
from os import fspath from os import fspath
from pathlib import Path from pathlib import Path
from shutil import which
from subprocess import PIPE, CalledProcessError from subprocess import PIPE, CalledProcessError
from typing import Optional
from PIL import Image, UnidentifiedImageError from PIL import Image, UnidentifiedImageError
from ocrmypdf.exceptions import MissingDependencyError, SubprocessOutputError from ocrmypdf.exceptions import SubprocessOutputError
from ocrmypdf.helpers import Resolution from ocrmypdf.helpers import Resolution
from ocrmypdf.subprocess import get_version, run, run_polling_stderr from ocrmypdf.subprocess import get_version, run, run_polling_stderr
@@ -33,29 +33,18 @@ except AttributeError:
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
missing_gs_error = """
---------------------------------------------------------------------
This error normally occurs when ocrmypdf find can't Ghostscript.
Please ensure Ghostscript is installed and its location is added to
the system PATH environment variable.
For details see:
https://ocrmypdf.readthedocs.io/en/latest/installation.html
---------------------------------------------------------------------
"""
# Most reliable what to get the bitness of Python interpreter, according to Python docs # Most reliable what to get the bitness of Python interpreter, according to Python docs
_is_64bit = sys.maxsize > 2 ** 32 _IS_64BIT = sys.maxsize > 2**32
_gswin = None _GSWIN = None
if os.name == 'nt': if os.name == 'nt':
if _is_64bit: if _IS_64BIT:
_gswin = 'gswin64c' _GSWIN = 'gswin64c'
else: else:
_gswin = 'gswin32c' _GSWIN = 'gswin32c'
GS = _gswin if _gswin else 'gs' GS = _GSWIN if _GSWIN else 'gs'
del _gswin del _GSWIN
def version(): def version():
@@ -89,8 +78,8 @@ def rasterize_pdf(
raster_device: str, raster_device: str,
raster_dpi: Resolution, raster_dpi: Resolution,
pageno: int = 1, pageno: int = 1,
page_dpi: Optional[Resolution] = None, page_dpi: Resolution | None = None,
rotation: Optional[int] = None, rotation: int | None = None,
filter_vector: bool = False, filter_vector: bool = False,
): ):
"""Rasterize one page of a PDF at resolution raster_dpi in canvas units.""" """Rasterize one page of a PDF at resolution raster_dpi in canvas units."""
@@ -115,7 +104,7 @@ def rasterize_pdf(
+ [ + [
'-o', '-o',
'-', '-',
'-sstdout=%stderr', '-sstdout=%stderr', # Literal %s, not string interpolation
'-dAutoRotatePages=/None', # Probably has no effect on raster '-dAutoRotatePages=/None', # Probably has no effect on raster
'-f', '-f',
fspath(input_file), fspath(input_file),
@@ -126,7 +115,7 @@ def rasterize_pdf(
p = run(args_gs, stdout=PIPE, stderr=PIPE, check=True) p = run(args_gs, stdout=PIPE, stderr=PIPE, check=True)
except CalledProcessError as e: except CalledProcessError as e:
log.error(e.stderr.decode(errors='replace')) log.error(e.stderr.decode(errors='replace'))
raise SubprocessOutputError('Ghostscript rasterizing failed') raise SubprocessOutputError('Ghostscript rasterizing failed') from e
else: else:
stderr = p.stderr.decode(errors='replace') stderr = p.stderr.decode(errors='replace')
if _gs_error_reported(stderr): if _gs_error_reported(stderr):
@@ -156,6 +145,8 @@ def rasterize_pdf(
class GhostscriptFollower: class GhostscriptFollower:
"""Parses the output of Ghostscript and uses it to update the progress bar."""
re_process = re.compile(r"Processing pages \d+ through (\d+).") re_process = re.compile(r"Processing pages \d+ through (\d+).")
re_page = re.compile(r"Page (\d+)") re_page = re.compile(r"Page (\d+)")
@@ -251,7 +242,7 @@ def generate_pdfa(
"-dPDFACompatibilityPolicy=1", "-dPDFACompatibilityPolicy=1",
"-o", "-o",
"-", "-",
"-sstdout=%stderr", "-sstdout=%stderr", # Literal %s, not string interpolation
] ]
) )
args_gs.extend(fspath(s) for s in pdf_pages) # Stringify Path objs args_gs.extend(fspath(s) for s in pdf_pages) # Stringify Path objs
+2
View File
@@ -7,6 +7,8 @@
"""Interface to jbig2 executable""" """Interface to jbig2 executable"""
from __future__ import annotations
from subprocess import PIPE from subprocess import PIPE
from ocrmypdf.exceptions import MissingDependencyError from ocrmypdf.exceptions import MissingDependencyError
+2
View File
@@ -7,6 +7,8 @@
"""Interface to pngquant executable""" """Interface to pngquant executable"""
from __future__ import annotations
from contextlib import contextmanager from contextlib import contextmanager
from io import BytesIO from io import BytesIO
from pathlib import Path from pathlib import Path
+19 -15
View File
@@ -7,13 +7,14 @@
"""Interface to Tesseract executable""" """Interface to Tesseract executable"""
from __future__ import annotations
import logging import logging
import re import re
from math import pi from math import pi
from os import fspath from os import fspath
from pathlib import Path from pathlib import Path
from subprocess import PIPE, STDOUT, CalledProcessError, TimeoutExpired from subprocess import PIPE, STDOUT, CalledProcessError, TimeoutExpired
from typing import Dict, Iterator, List, Optional
from packaging.version import Version from packaging.version import Version
from PIL import Image from PIL import Image
@@ -46,7 +47,7 @@ HOCR_TEMPLATE = """<?xml version="1.0" encoding="UTF-8"?>
</html> </html>
""" """
TESSERACT_THRESHOLDING_METHODS: Dict[str, int] = { TESSERACT_THRESHOLDING_METHODS: dict[str, int] = {
'auto': 0, 'auto': 0,
'otsu': 0, 'otsu': 0,
'adaptive-otsu': 1, 'adaptive-otsu': 1,
@@ -55,9 +56,11 @@ TESSERACT_THRESHOLDING_METHODS: Dict[str, int] = {
class TesseractLoggerAdapter(logging.LoggerAdapter): class TesseractLoggerAdapter(logging.LoggerAdapter):
"Prepend [tesseract] to messages emitted from tesseract"
def process(self, msg, kwargs): def process(self, msg, kwargs):
kwargs['extra'] = self.extra kwargs['extra'] = self.extra
return '[tesseract] %s' % (msg), kwargs return f'[tesseract] {msg}', kwargs
TESSERACT_VERSION_PATTERN = r""" TESSERACT_VERSION_PATTERN = r"""
@@ -105,6 +108,7 @@ TESSERACT_VERSION_PATTERN = r"""
class TesseractVersion(Version): class TesseractVersion(Version):
"Modify standard packaging.Version regex to support Tesseract idiosyncracies."
_regex = re.compile( _regex = re.compile(
r"^\s*" + TESSERACT_VERSION_PATTERN + r"\s*$", re.VERBOSE | re.IGNORECASE r"^\s*" + TESSERACT_VERSION_PATTERN + r"\s*$", re.VERBOSE | re.IGNORECASE
) )
@@ -159,7 +163,7 @@ def get_languages():
return {lang.strip() for lang in rest} return {lang.strip() for lang in rest}
def tess_base_args(langs: List[str], engine_mode: Optional[int]) -> List[str]: def tess_base_args(langs: list[str], engine_mode: int | None) -> list[str]:
args = ['tesseract'] args = ['tesseract']
if langs: if langs:
args.extend(['-l', '+'.join(langs)]) args.extend(['-l', '+'.join(langs)])
@@ -168,19 +172,19 @@ def tess_base_args(langs: List[str], engine_mode: Optional[int]) -> List[str]:
return args return args
def _parse_tesseract_output(binary_output: bytes) -> Dict[str, str]: def _parse_tesseract_output(binary_output: bytes) -> dict[str, str]:
def g(): def gen():
for line in binary_output.decode().splitlines(): for line in binary_output.decode().splitlines():
line = line.strip() line = line.strip()
parts = line.split(':', maxsplit=2) parts = line.split(':', maxsplit=2)
if len(parts) == 2: if len(parts) == 2:
yield parts[0].strip(), parts[1].strip() yield parts[0].strip(), parts[1].strip()
return {k: v for k, v in g()} return dict(gen())
def get_orientation( def get_orientation(
input_file: Path, engine_mode: Optional[int], timeout: float input_file: Path, engine_mode: int | None, timeout: float
) -> OrientationConfidence: ) -> OrientationConfidence:
args_tesseract = tess_base_args(['osd'], engine_mode) + [ args_tesseract = tess_base_args(['osd'], engine_mode) + [
'--psm', '--psm',
@@ -205,14 +209,14 @@ def get_orientation(
osd = _parse_tesseract_output(p.stdout) osd = _parse_tesseract_output(p.stdout)
angle = int(osd.get('Orientation in degrees', 0)) angle = int(osd.get('Orientation in degrees', 0))
oc = OrientationConfidence( orient_conf = OrientationConfidence(
angle=angle, confidence=float(osd.get('Orientation confidence', 0)) angle=angle, confidence=float(osd.get('Orientation confidence', 0))
) )
return oc return orient_conf
def get_deskew( def get_deskew(
input_file: Path, languages: List[str], engine_mode: Optional[int], timeout: float input_file: Path, languages: list[str], engine_mode: int | None, timeout: float
) -> float: ) -> float:
"""Gets angle to deskew this page, in degrees.""" """Gets angle to deskew this page, in degrees."""
args_tesseract = tess_base_args(languages, engine_mode) + [ args_tesseract = tess_base_args(languages, engine_mode) + [
@@ -303,9 +307,9 @@ def generate_hocr(
input_file: Path, input_file: Path,
output_hocr: Path, output_hocr: Path,
output_text: Path, output_text: Path,
languages: List[str], languages: list[str],
engine_mode: int, engine_mode: int,
tessconfig: List[str], tessconfig: list[str],
timeout: float, timeout: float,
pagesegmode: int, pagesegmode: int,
thresholding: int, thresholding: int,
@@ -369,9 +373,9 @@ def generate_pdf(
input_file: Path, input_file: Path,
output_pdf: Path, output_pdf: Path,
output_text: Path, output_text: Path,
languages: List[str], languages: list[str],
engine_mode: int, engine_mode: int,
tessconfig: List[str], tessconfig: list[str],
timeout: float, timeout: float,
pagesegmode: int, pagesegmode: int,
thresholding: int, thresholding: int,
+53 -27
View File
@@ -5,26 +5,48 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
# unpaper documentation: # unpaper documentation:
# https://github.com/Flameeyes/unpaper/blob/master/doc/basic-concepts.md # https://github.com/Flameeyes/unpaper/blob/master/doc/basic-concepts.md
"""Interface to unpaper executable""" """Interface to unpaper executable"""
import logging import logging
import os import os
import shlex import shlex
import sys
from contextlib import contextmanager from contextlib import contextmanager
from decimal import Decimal from decimal import Decimal
from pathlib import Path from pathlib import Path
from subprocess import PIPE, STDOUT from subprocess import PIPE, STDOUT
from tempfile import TemporaryDirectory from typing import Iterator, Union
from typing import Iterator, List, Optional, Tuple, Union
from PIL import Image from PIL import Image
from ocrmypdf.exceptions import MissingDependencyError, SubprocessOutputError from ocrmypdf.exceptions import MissingDependencyError, SubprocessOutputError
from ocrmypdf.subprocess import get_version, run from ocrmypdf.subprocess import get_version, run
if sys.version_info >= (3, 10):
from tempfile import TemporaryDirectory
else:
from tempfile import TemporaryDirectory as _TemporaryDirectory
class TemporaryDirectory(_TemporaryDirectory):
"""Shim to consume ignore_cleanup_errors kwarg on Python 3.9 and older.
The argument is consumed without action. If users are getting errors related
to temporary file cleanup, they should upgrade to Python 3.10 which properly
cleans up temporary directories on Windows.
See: https://github.com/python/cpython/pull/24793
"""
def __init__(self, ignore_cleanup_errors=False, **kwargs):
super().__init__(**kwargs)
del _TemporaryDirectory
UNPAPER_IMAGE_PIXEL_LIMIT = 256 * 1024 * 1024 UNPAPER_IMAGE_PIXEL_LIMIT = 256 * 1024 * 1024
DecFloat = Union[Decimal, float] DecFloat = Union[Decimal, float]
@@ -33,6 +55,8 @@ log = logging.getLogger(__name__)
class UnpaperImageTooLargeError(Exception): class UnpaperImageTooLargeError(Exception):
"""To capture details when an image is too large for unpaper."""
def __init__( def __init__(
self, self,
w, w,
@@ -49,11 +73,13 @@ def version() -> str:
return get_version('unpaper') return get_version('unpaper')
def _convert_image(im: Image.Image) -> Tuple[Image.Image, bool, str]: SUPPORTED_MODES = {'1', 'L', 'RGB'}
SUFFIXES = {'1': '.pbm', 'L': '.pgm', 'RGB': '.ppm'}
def _convert_image(im: Image.Image) -> tuple[Image.Image, bool]:
im_modified = False im_modified = False
if im.mode not in SUFFIXES: if im.mode not in SUPPORTED_MODES:
log.info("Converting image to other colorspace") log.info("Converting image to other colorspace")
try: try:
if im.mode == 'P' and len(im.getcolors()) == 2: if im.mode == 'P' and len(im.getcolors()) == 2:
@@ -66,41 +92,41 @@ def _convert_image(im: Image.Image) -> Tuple[Image.Image, bool, str]:
) from e ) from e
else: else:
im_modified = True im_modified = True
try: if im.mode not in SUPPORTED_MODES:
suffix = SUFFIXES[im.mode] raise MissingDependencyError(
except KeyError: "Failed to convert image to a supported format."
raise MissingDependencyError( ) from None
"Failed to convert image to a supported format." return im, im_modified
) from None
return im, im_modified, suffix
@contextmanager @contextmanager
def _setup_unpaper_io(input_file: Path) -> Iterator[Tuple[Path, Path, Path]]: def _setup_unpaper_io(input_file: Path) -> Iterator[tuple[Path, Path, Path]]:
with Image.open(input_file) as im: with Image.open(input_file) as im:
if im.width * im.height >= UNPAPER_IMAGE_PIXEL_LIMIT: if im.width * im.height >= UNPAPER_IMAGE_PIXEL_LIMIT:
raise UnpaperImageTooLargeError(w=im.width, h=im.height) raise UnpaperImageTooLargeError(w=im.width, h=im.height)
im, im_modified, suffix = _convert_image(im) im, im_modified = _convert_image(im)
with TemporaryDirectory() as tmpdir: with TemporaryDirectory(ignore_cleanup_errors=True) as tmpdir:
tmppath = Path(tmpdir) tmppath = Path(tmpdir)
if im_modified or input_file.suffix != '.pnm': if im_modified or input_file.suffix != '.png':
input_pnm = tmppath / 'input.pnm' input_png = tmppath / 'input.png'
im.save(input_pnm, format='PPM') im.save(input_png, format='PNG')
else: else:
# No changes, PNG input, just use the file we already have # No changes, PNG input, just use the file we already have
input_pnm = input_file input_png = input_file
output_pnm = tmppath / f'output{suffix}' # unpaper can write .png too, but it seems to write them slowly
yield input_pnm, output_pnm, tmppath # adds a few seconds to test suite - so just use pnm
output_pnm = tmppath / 'output.pnm'
yield input_png, output_pnm, tmppath
def run_unpaper( def run_unpaper(
input_file: Path, output_file: Path, *, dpi: DecFloat, mode_args: List[str] input_file: Path, output_file: Path, *, dpi: DecFloat, mode_args: list[str]
) -> None: ) -> None:
args_unpaper = ['unpaper', '-v', '--dpi', str(round(dpi, 6))] + mode_args args_unpaper = ['unpaper', '-v', '--dpi', str(round(dpi, 6))] + mode_args
with _setup_unpaper_io(input_file) as (input_pnm, output_pnm, tmpdir): with _setup_unpaper_io(input_file) as (input_png, output_pnm, tmpdir):
# To prevent any shenanigans from accepting arbitrary parameters in # To prevent any shenanigans from accepting arbitrary parameters in
# --unpaper-args, we: # --unpaper-args, we:
# 1) run with cwd set to a tmpdir with only unpaper's files # 1) run with cwd set to a tmpdir with only unpaper's files
@@ -108,7 +134,7 @@ def run_unpaper(
# 3) append absolute paths for the input and output file # 3) append absolute paths for the input and output file
# This should ensure that a user cannot clobber some other file with # This should ensure that a user cannot clobber some other file with
# their unpaper arguments (whether intentionally or otherwise) # their unpaper arguments (whether intentionally or otherwise)
args_unpaper.extend([os.fspath(input_pnm), os.fspath(output_pnm)]) args_unpaper.extend([os.fspath(input_png), os.fspath(output_pnm)])
run( run(
args_unpaper, args_unpaper,
close_fds=True, close_fds=True,
@@ -129,7 +155,7 @@ def run_unpaper(
) from e ) from e
def validate_custom_args(args: str) -> List[str]: def validate_custom_args(args: str) -> list[str]:
unpaper_args = shlex.split(args) unpaper_args = shlex.split(args)
if any(('/' in arg or arg == '.' or arg == '..') for arg in unpaper_args): if any(('/' in arg or arg == '.' or arg == '..') for arg in unpaper_args):
raise ValueError('No filenames allowed in --unpaper-args') raise ValueError('No filenames allowed in --unpaper-args')
@@ -141,7 +167,7 @@ def clean(
output_file: Path, output_file: Path,
*, *,
dpi: DecFloat, dpi: DecFloat,
unpaper_args: Optional[List[str]] = None, unpaper_args: list[str] | None = None,
) -> Path: ) -> Path:
default_args = [ default_args = [
'--layout', '--layout',
+10 -6
View File
@@ -4,19 +4,19 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""For grafting text-only PDF pages onto freeform PDF pages."""
from __future__ import annotations
import logging import logging
import uuid
from contextlib import suppress from contextlib import suppress
from pathlib import Path from pathlib import Path
from typing import Optional
from pikepdf import ( from pikepdf import (
Dictionary, Dictionary,
Name, Name,
Object, Object,
Operator, Operator,
Page,
Pdf, Pdf,
PdfError, PdfError,
PdfMatrix, PdfMatrix,
@@ -81,6 +81,8 @@ def strip_invisible_text(pdf, page):
class OcrGrafter: class OcrGrafter:
"""Manages grafting text-only PDFs onto regular PDFs."""
def __init__(self, context): def __init__(self, context):
self.context = context self.context = context
self.path_base = context.origin self.path_base = context.origin
@@ -102,8 +104,8 @@ class OcrGrafter:
self, self,
*, *,
pageno: int, pageno: int,
image: Optional[Path], image: Path | None,
textpdf: Optional[Path], textpdf: Path | None,
autorotate_correction: int, autorotate_correction: int,
): ):
if textpdf and not self.font: if textpdf and not self.font:
@@ -236,6 +238,8 @@ class OcrGrafter:
): ):
"""Insert the text layer from text page 0 on to pdf_base at page_num""" """Insert the text layer from text page 0 on to pdf_base at page_num"""
# pylint: disable=invalid-name
log.debug("Grafting") log.debug("Grafting")
if Path(textpdf).stat().st_size == 0: if Path(textpdf).stat().st_size == 0:
return return
@@ -282,7 +286,7 @@ class OcrGrafter:
base_resources = _ensure_dictionary(base_page, Name.Resources) base_resources = _ensure_dictionary(base_page, Name.Resources)
base_xobjs = _ensure_dictionary(base_resources, Name.XObject) base_xobjs = _ensure_dictionary(base_resources, Name.XObject)
text_xobj_name = Name('/' + str(uuid.uuid4())) text_xobj_name = Name.random(prefix="OCR-")
xobj = self.pdf_base.make_stream(pdf_text_contents) xobj = self.pdf_base.make_stream(pdf_text_contents)
base_xobjs[text_xobj_name] = xobj base_xobjs[text_xobj_name] = xobj
xobj.Type = Name.XObject xobj.Type = Name.XObject
+5 -2
View File
@@ -4,6 +4,9 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Defines context objects that are passed to child processes/threads."""
from __future__ import annotations
import os import os
import shutil import shutil
@@ -49,7 +52,7 @@ class PdfContext:
""" """
return self.work_folder / name return self.work_folder / name
def get_page_contexts(self) -> Iterator['PageContext']: def get_page_contexts(self) -> Iterator[PageContext]:
"""Get all ``PageContext`` for this PDF.""" """Get all ``PageContext`` for this PDF."""
npages = len(self.pdfinfo) npages = len(self.pdfinfo)
for n in range(npages): for n in range(npages):
@@ -83,7 +86,7 @@ class PageContext:
The path will be based in a common temporary folder and have a prefix based The path will be based in a common temporary folder and have a prefix based
on the page number. on the page number.
""" """
return self.work_folder / ("%06d_%s" % (self.pageno + 1, name)) return self.work_folder / f"{(self.pageno + 1):06d}_{name}"
def __getstate__(self): def __getstate__(self):
state = self.__dict__.copy() state = self.__dict__.copy()
+5 -1
View File
@@ -4,15 +4,19 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Logging support classes."""
from __future__ import annotations
import logging import logging
import sys
from contextlib import suppress from contextlib import suppress
from tqdm import tqdm from tqdm import tqdm
class PageNumberFilter(logging.Filter): class PageNumberFilter(logging.Filter):
"""Insert PDF page number that emitted log message to log record."""
def filter(self, record): def filter(self, record):
pageno = getattr(record, 'pageno', None) pageno = getattr(record, 'pageno', None)
if isinstance(pageno, int): if isinstance(pageno, int):
+46 -33
View File
@@ -4,6 +4,9 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""OCRmyPDF page processing pipeline functions."""
from __future__ import annotations
import logging import logging
import os import os
@@ -13,7 +16,7 @@ from contextlib import suppress
from datetime import datetime, timezone from datetime import datetime, timezone
from pathlib import Path from pathlib import Path
from shutil import copyfileobj from shutil import copyfileobj
from typing import Dict, Iterable, Optional from typing import Iterable
import img2pdf import img2pdf
import pikepdf import pikepdf
@@ -34,14 +37,13 @@ from ocrmypdf.exceptions import (
) )
from ocrmypdf.helpers import IMG2PDF_KWARGS, Resolution, safe_symlink from ocrmypdf.helpers import IMG2PDF_KWARGS, Resolution, safe_symlink
from ocrmypdf.hocrtransform import HocrTransform from ocrmypdf.hocrtransform import HocrTransform
from ocrmypdf.optimize import optimize
from ocrmypdf.pdfa import generate_pdfa_ps from ocrmypdf.pdfa import generate_pdfa_ps
from ocrmypdf.pdfinfo import Colorspace, Encoding, PdfInfo from ocrmypdf.pdfinfo import Colorspace, Encoding, PdfInfo
# Remove this workaround when we require Pillow >= 10 # Remove this workaround when we require Pillow >= 10
try: try:
BICUBIC = Image.Resampling.BICUBIC # type: ignore BICUBIC = Image.Resampling.BICUBIC # type: ignore
except AttributeError: except AttributeError: # pragma: no cover
# Pillow 9 shim # Pillow 9 shim
BICUBIC = Image.BICUBIC # type: ignore BICUBIC = Image.BICUBIC # type: ignore
@@ -332,7 +334,8 @@ def is_ocr_required(page_context: PageContext):
ocr_required = False ocr_required = False
log.warning( log.warning(
"page too big, skipping OCR " "page too big, skipping OCR "
f"({(pixel_count / 1_000_000):.1f} MPixels > {options.skip_big:.1f} MPixels --skip-big)" f"({(pixel_count / 1_000_000):.1f} MPixels > "
f"{options.skip_big:.1f} MPixels --skip-big)"
) )
return ocr_required return ocr_required
@@ -430,8 +433,8 @@ def rasterize(
output_file = page_context.get_path(f'rasterize{output_tag}.png') output_file = page_context.get_path(f'rasterize{output_tag}.png')
pageinfo = page_context.pageinfo pageinfo = page_context.pageinfo
def at_least(cs): def at_least(colorspace):
return max(device_idx, colorspaces.index(cs)) return max(device_idx, colorspaces.index(colorspace))
for image in pageinfo.images: for image in pageinfo.images:
if image.type_ != 'image': if image.type_ != 'image':
@@ -471,10 +474,10 @@ def rasterize(
def preprocess_remove_background(input_file: Path, page_context: PageContext): def preprocess_remove_background(input_file: Path, page_context: PageContext):
if any(image.bpc > 1 for image in page_context.pageinfo.images): if any(image.bpc > 1 for image in page_context.pageinfo.images):
output_file = page_context.get_path('pp_rm_bg.png')
# leptonica.remove_background(input_file, output_file)
raise NotImplementedError("--remove-background is temporarily not implemented") raise NotImplementedError("--remove-background is temporarily not implemented")
return output_file # output_file = page_context.get_path('pp_rm_bg.png')
# leptonica.remove_background(input_file, output_file)
# return output_file
else: else:
log.info("background removal skipped on mono page") log.info("background removal skipped on mono page")
return input_file return input_file
@@ -493,7 +496,7 @@ def preprocess_deskew(input_file: Path, page_context: PageContext):
deskewed = im.rotate( deskewed = im.rotate(
deskew_angle_degrees, deskew_angle_degrees,
resample=BICUBIC, resample=BICUBIC,
fillcolor=ImageColor.getcolor('white', mode=im.mode), fillcolor=ImageColor.getcolor('white', mode=im.mode), # type: ignore
) )
deskewed.save(output_file, dpi=dpi) deskewed.save(output_file, dpi=dpi)
@@ -664,7 +667,7 @@ def ocr_engine_textonly_pdf(input_image: Path, page_context: PageContext):
return (output_pdf, output_text) return (output_pdf, output_text)
def get_docinfo(base_pdf: pikepdf.Pdf, context: PdfContext) -> Dict[str, str]: def get_docinfo(base_pdf: pikepdf.Pdf, context: PdfContext) -> dict[str, str]:
options = context.options options = context.options
def from_document_info(key): def from_document_info(key):
@@ -678,15 +681,14 @@ def get_docinfo(base_pdf: pikepdf.Pdf, context: PdfContext) -> Dict[str, str]:
k: from_document_info(k) k: from_document_info(k)
for k in ('/Title', '/Author', '/Keywords', '/Subject', '/CreationDate') for k in ('/Title', '/Author', '/Keywords', '/Subject', '/CreationDate')
} }
if options is not None: if options.title:
if options.title: pdfmark['/Title'] = options.title
pdfmark['/Title'] = options.title if options.author:
if options.author: pdfmark['/Author'] = options.author
pdfmark['/Author'] = options.author if options.keywords:
if options.keywords: pdfmark['/Keywords'] = options.keywords
pdfmark['/Keywords'] = options.keywords if options.subject:
if options.subject: pdfmark['/Subject'] = options.subject
pdfmark['/Subject'] = options.subject
creator_tag = context.plugin_manager.hook.get_ocr_engine().creator_tag(options) creator_tag = context.plugin_manager.hook.get_ocr_engine().creator_tag(options)
@@ -816,13 +818,14 @@ def metadata_fixup(working_file: Path, context: PdfContext):
missing = set(meta_original.keys()) - set(meta.keys()) missing = set(meta_original.keys()) - set(meta.keys())
report_on_metadata(missing) report_on_metadata(missing)
optimizing = context.plugin_manager.hook.is_optimization_enabled(
context=context
)
pdf.save( pdf.save(
output_file, output_file,
**get_pdf_save_settings(options.output_type), **get_pdf_save_settings(options.output_type),
linearize=( # Don't linearize if optimize() will be linearizing too linearize=( # Don't linearize if optimize() will be linearizing too
should_linearize(working_file, context) not optimizing and should_linearize(working_file, context)
if options.optimize == 0
else False
), ),
) )
@@ -831,12 +834,22 @@ def metadata_fixup(working_file: Path, context: PdfContext):
def optimize_pdf(input_file: Path, context: PdfContext, executor: Executor): def optimize_pdf(input_file: Path, context: PdfContext, executor: Executor):
output_file = context.get_path('optimize.pdf') output_file = context.get_path('optimize.pdf')
save_settings = dict( output_pdf, messages = context.plugin_manager.hook.optimize_pdf(
input_pdf=input_file,
output_pdf=output_file,
context=context,
executor=executor,
linearize=should_linearize(input_file, context), linearize=should_linearize(input_file, context),
**get_pdf_save_settings(context.options.output_type),
) )
optimize(input_file, output_file, context, save_settings, executor)
return output_file input_size = input_file.stat().st_size
output_size = output_file.stat().st_size
if output_size > 0:
ratio = input_size / output_size
savings = 1 - output_size / input_size
log.info(f"Optimize ratio: {ratio:.2f} savings: {(savings):.1%}")
return output_pdf, messages
def enumerate_compress_ranges(iterable): def enumerate_compress_ranges(iterable):
@@ -855,11 +868,11 @@ def enumerate_compress_ranges(iterable):
yield (skipped_from, index), None yield (skipped_from, index), None
def merge_sidecars(txt_files: Iterable[Optional[Path]], context: PdfContext): def merge_sidecars(txt_files: Iterable[Path | None], context: PdfContext):
output_file = context.get_path('sidecar.txt') output_file = context.get_path('sidecar.txt')
with open(output_file, 'w', encoding="utf-8") as stream: with open(output_file, 'w', encoding="utf-8") as stream:
for (frm, to), txt_file in enumerate_compress_ranges(txt_files): for (from_, to_), txt_file in enumerate_compress_ranges(txt_files):
if frm != 1: if from_ != 1:
stream.write('\f') # Form feed between pages stream.write('\f') # Form feed between pages
if txt_file: if txt_file:
with open(txt_file, encoding="utf-8") as in_: with open(txt_file, encoding="utf-8") as in_:
@@ -872,10 +885,10 @@ def merge_sidecars(txt_files: Iterable[Optional[Path]], context: PdfContext):
else: else:
stream.write(txt) stream.write(txt)
else: else:
if frm != to: if from_ != to_:
pages = f'{frm}-{to}' pages = f'{from_}-{to_}'
else: else:
pages = f'{frm}' pages = f'{from_}'
stream.write(f'[OCR skipped on page(s) {pages}]') stream.write(f'[OCR skipped on page(s) {pages}]')
return output_file return output_file
+14 -6
View File
@@ -4,6 +4,9 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Plugin manager using pluggy."""
from __future__ import annotations
import argparse import argparse
import importlib import importlib
@@ -11,7 +14,7 @@ import importlib.util
import pkgutil import pkgutil
import sys import sys
from pathlib import Path from pathlib import Path
from typing import List, Sequence, Tuple, Union from typing import Sequence
import pluggy import pluggy
@@ -33,7 +36,7 @@ class OcrmypdfPluginManager(pluggy.PluginManager):
def __init__( def __init__(
self, self,
*args, *args,
plugins: List[Union[str, Path]], plugins: list[str | Path],
builtins: bool = True, builtins: bool = True,
**kwargs, **kwargs,
): ):
@@ -100,23 +103,28 @@ class OcrmypdfPluginManager(pluggy.PluginManager):
self.register(module) self.register(module)
def get_plugin_manager(plugins: List[Union[str, Path]], builtins=True): def get_plugin_manager(plugins: list[str | Path], builtins=True):
pm = OcrmypdfPluginManager( return OcrmypdfPluginManager(
project_name='ocrmypdf', project_name='ocrmypdf',
plugins=plugins, plugins=plugins,
builtins=builtins, builtins=builtins,
) )
return pm
def get_parser_options_plugins( def get_parser_options_plugins(
args: Sequence[str], args: Sequence[str],
) -> Tuple[argparse.ArgumentParser, argparse.Namespace, pluggy.PluginManager]: ) -> tuple[argparse.ArgumentParser, argparse.Namespace, pluggy.PluginManager]:
pre_options, _unused = plugins_only_parser.parse_known_args(args=args) pre_options, _unused = plugins_only_parser.parse_known_args(args=args)
plugin_manager = get_plugin_manager(pre_options.plugins) plugin_manager = get_plugin_manager(pre_options.plugins)
parser = get_parser() parser = get_parser()
plugin_manager.hook.initialize( # pylint: disable=no-member
plugin_manager=plugin_manager
)
plugin_manager.hook.add_options(parser=parser) # pylint: disable=no-member plugin_manager.hook.add_options(parser=parser) # pylint: disable=no-member
options = parser.parse_args(args=args) options = parser.parse_args(args=args)
return parser, options, plugin_manager return parser, options, plugin_manager
__all__ = ['OcrmypdfPluginManager', 'get_plugin_manager', 'get_parser_options_plugins']
+27 -15
View File
@@ -4,6 +4,10 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Implements the concurrent and page synchronous parts of the pipeline."""
from __future__ import annotations
import argparse import argparse
import logging import logging
@@ -16,7 +20,7 @@ from concurrent.futures.thread import BrokenThreadPool
from functools import partial from functools import partial
from pathlib import Path from pathlib import Path
from tempfile import mkdtemp from tempfile import mkdtemp
from typing import List, NamedTuple, Optional, Tuple, cast from typing import NamedTuple, Sequence, cast
import PIL import PIL
@@ -68,11 +72,13 @@ from ocrmypdf.pdfa import file_claims_pdfa
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
class PageResult(NamedTuple): # pylint: disable=inherit-non-class class PageResult(NamedTuple):
"""Result when a page is finished processing."""
pageno: int pageno: int
pdf_page_from_image: Optional[Path] pdf_page_from_image: Path | None
ocr: Optional[Path] ocr: Path | None
text: Optional[Path] text: Path | None
orientation_correction: int orientation_correction: int
@@ -111,7 +117,7 @@ def preprocess(
def make_intermediate_images( def make_intermediate_images(
page_context: PageContext, orientation_correction: int page_context: PageContext, orientation_correction: int
) -> Tuple[Path, Optional[Path]]: ) -> tuple[Path, Path | None]:
options = page_context.options options = page_context.options
ocr_image = preprocess_out = None ocr_image = preprocess_out = None
@@ -226,7 +232,9 @@ def exec_page_sync(page_context: PageContext) -> PageResult:
) )
def post_process(pdf_file: Path, context: PdfContext, executor: Executor) -> Path: def post_process(
pdf_file: Path, context: PdfContext, executor: Executor
) -> tuple[Path, Sequence[str]]:
pdf_out = pdf_file pdf_out = pdf_file
if context.options.output_type.startswith('pdfa'): if context.options.output_type.startswith('pdfa'):
ps_stub_out = generate_postscript_stub(context) ps_stub_out = generate_postscript_stub(context)
@@ -244,7 +252,7 @@ def worker_init(max_pixels: int) -> None:
pikepdf_enable_mmap() pikepdf_enable_mmap()
def exec_concurrent(context: PdfContext, executor: Executor) -> None: def exec_concurrent(context: PdfContext, executor: Executor) -> Sequence[str]:
"""Execute the pipeline concurrently""" """Execute the pipeline concurrently"""
# Run exec_page_sync on every page context # Run exec_page_sync on every page context
@@ -253,7 +261,7 @@ def exec_concurrent(context: PdfContext, executor: Executor) -> None:
if max_workers > 1: if max_workers > 1:
log.info("Start processing %d pages concurrently", max_workers) log.info("Start processing %d pages concurrently", max_workers)
sidecars: List[Optional[Path]] = [None] * len(context.pdfinfo) sidecars: list[Path | None] = [None] * len(context.pdfinfo)
ocrgraft = OcrGrafter(context) ocrgraft = OcrGrafter(context)
def update_page(result: PageResult, pbar): def update_page(result: PageResult, pbar):
@@ -296,13 +304,15 @@ def exec_concurrent(context: PdfContext, executor: Executor) -> None:
# Merge layers to one single pdf # Merge layers to one single pdf
pdf = ocrgraft.finalize() pdf = ocrgraft.finalize()
messages: Sequence[str] = []
if options.output_type != 'none': if options.output_type != 'none':
# PDF/A and metadata # PDF/A and metadata
log.info("Postprocessing...") log.info("Postprocessing...")
pdf = post_process(pdf, context, executor) pdf, messages = post_process(pdf, context, executor)
# Copy PDF file to destination # Copy PDF file to destination
copy_final(pdf, options.output_file, context) copy_final(pdf, options.output_file, context)
return messages
def configure_debug_logging( def configure_debug_logging(
@@ -329,7 +339,7 @@ def configure_debug_logging(
def run_pipeline( def run_pipeline(
options: argparse.Namespace, options: argparse.Namespace,
*, *,
plugin_manager: Optional[OcrmypdfPluginManager], plugin_manager: OcrmypdfPluginManager | None,
api: bool = False, api: bool = False,
) -> ExitCode: ) -> ExitCode:
# Any changes to options will not take effect for options that are already # Any changes to options will not take effect for options that are already
@@ -382,7 +392,7 @@ def run_pipeline(
validate_pdfinfo_options(context) validate_pdfinfo_options(context)
# Execute the pipeline # Execute the pipeline
exec_concurrent(context, executor) optimize_messages = exec_concurrent(context, executor)
if options.output_file == '-': if options.output_file == '-':
log.info("Output sent to stdout") log.info("Output sent to stdout")
@@ -408,7 +418,9 @@ def run_pipeline(
if not check_pdf(options.output_file): if not check_pdf(options.output_file):
log.warning('Output file: The generated PDF is INVALID') log.warning('Output file: The generated PDF is INVALID')
return ExitCode.invalid_output_pdf return ExitCode.invalid_output_pdf
report_output_file_size(options, start_input_file, options.output_file) report_output_file_size(
options, start_input_file, options.output_file, optimize_messages
)
except (KeyboardInterrupt if not api else NeverRaise): except (KeyboardInterrupt if not api else NeverRaise):
if options.verbose >= 1: if options.verbose >= 1:
@@ -425,7 +437,7 @@ def run_pipeline(
else: else:
log.error(type(e).__name__) log.error(type(e).__name__)
return e.exit_code return e.exit_code
except (PIL.Image.DecompressionBombError if not api else NeverRaise) as e: except (PIL.Image.DecompressionBombError if not api else NeverRaise):
log.exception( log.exception(
"A decompression bomb error was encountered while executing the " "A decompression bomb error was encountered while executing the "
"pipeline. Use the argument --max-image-mpixels to raise the maximum " "pipeline. Use the argument --max-image-mpixels to raise the maximum "
@@ -435,7 +447,7 @@ def run_pipeline(
except ( except (
BrokenProcessPool if not api else NeverRaise, BrokenProcessPool if not api else NeverRaise,
BrokenThreadPool if not api else NeverRaise, BrokenThreadPool if not api else NeverRaise,
) as e: ):
log.exception( log.exception(
"A worker process was terminated unexpectedly. This is known to occur if " "A worker process was terminated unexpectedly. This is known to occur if "
"processing your file takes all available swap space and RAM. It may " "processing your file takes all available swap space and RAM. It may "
+34 -67
View File
@@ -5,6 +5,9 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Validate a work order from API or command line."""
from __future__ import annotations
import locale import locale
import logging import logging
@@ -13,19 +16,19 @@ import sys
import unicodedata import unicodedata
from pathlib import Path from pathlib import Path
from shutil import copyfileobj from shutil import copyfileobj
from typing import List, Set, Tuple from typing import Sequence
import pikepdf import pikepdf
import PIL import PIL
from ocrmypdf._exec import jbig2enc, pngquant, unpaper from ocrmypdf._exec import unpaper
from ocrmypdf.exceptions import ( from ocrmypdf.exceptions import (
BadArgsError, BadArgsError,
InputFileError, InputFileError,
MissingDependencyError, MissingDependencyError,
OutputFileAccessError, OutputFileAccessError,
) )
from ocrmypdf.helpers import is_file_writable, monotonic, safe_symlink, samefile from ocrmypdf.helpers import is_file_writable, monotonic, safe_symlink
from ocrmypdf.hocrtransform import HOCR_OK_LANGS from ocrmypdf.hocrtransform import HOCR_OK_LANGS
from ocrmypdf.subprocess import check_external_program from ocrmypdf.subprocess import check_external_program
@@ -143,16 +146,16 @@ def check_options_preprocessing(options):
raise BadArgsError("--unpaper-args: " + str(e)) from e raise BadArgsError("--unpaper-args: " + str(e)) from e
def _pages_from_ranges(ranges: str) -> Set[int]: def _pages_from_ranges(ranges: str) -> set[int]:
pages: List[int] = [] pages: list[int] = []
page_groups = ranges.replace(' ', '').split(',') page_groups = ranges.replace(' ', '').split(',')
for g in page_groups: for group in page_groups:
if not g: if not group:
continue continue
try: try:
start, end = g.split('-') start, end = group.split('-')
except ValueError: except ValueError:
pages.append(int(g) - 1) pages.append(int(group) - 1)
else: else:
try: try:
new_pages = list(range(int(start) - 1, int(end))) new_pages = list(range(int(start) - 1, int(end)))
@@ -162,7 +165,7 @@ def _pages_from_ranges(ranges: str) -> Set[int]:
) from None ) from None
pages.extend(new_pages) pages.extend(new_pages)
except ValueError: except ValueError:
raise BadArgsError(f"invalid page subrange '{g}'") from None raise BadArgsError(f"invalid page subrange '{group}'") from None
if not pages: if not pages:
raise BadArgsError( raise BadArgsError(
@@ -193,37 +196,6 @@ def check_options_ocr_behavior(options):
options.pages = _pages_from_ranges(options.pages) options.pages = _pages_from_ranges(options.pages)
def check_options_optimizing(options):
if options.optimize >= 2:
check_external_program(
program='pngquant',
package='pngquant',
version_checker=pngquant.version,
need_version='2.0.1',
required_for='--optimize {2,3}',
)
if options.optimize >= 2:
# Although we use JBIG2 for optimize=1, don't nag about it unless the
# user is asking for more optimization
check_external_program(
program='jbig2',
package='jbig2enc',
version_checker=jbig2enc.version,
need_version='0.28',
required_for='--optimize {2,3} | --jbig2-lossy',
recommended=True if not options.jbig2_lossy else False,
)
if options.optimize == 0 and any(
[options.jbig2_lossy, options.png_quality, options.jpeg_quality]
):
log.warning(
"The arguments --jbig2-lossy, --png-quality, and --jpeg-quality "
"will be ignored because --optimize=0."
)
def check_options_advanced(options): def check_options_advanced(options):
if options.pdfa_image_compression != 'auto' and not options.output_type.startswith( if options.pdfa_image_compression != 'auto' and not options.output_type.startswith(
'pdfa' 'pdfa'
@@ -237,13 +209,13 @@ def check_options_advanced(options):
def check_options_metadata(options): def check_options_metadata(options):
docinfo = [options.title, options.author, options.keywords, options.subject] docinfo = [options.title, options.author, options.keywords, options.subject]
for s in (m for m in docinfo if m): for s in (m for m in docinfo if m):
for c in s: for char in s:
if unicodedata.category(c) == 'Co' or ord(c) >= 0x10000: if unicodedata.category(char) == 'Co' or ord(char) >= 0x10000:
hexchar = hex(ord(char))[2:].upper()
raise ValueError( raise ValueError(
"One of the metadata strings contains " "One of the metadata strings contains "
"an unsupported Unicode character: '{}' (U+{})".format( "an unsupported Unicode character: "
c, hex(ord(c))[2:].upper() f"{char} (U+{hexchar})"
)
) )
@@ -261,7 +233,6 @@ def _check_options(options, plugin_manager, ocr_engine_languages):
check_options_sidecar(options) check_options_sidecar(options)
check_options_preprocessing(options) check_options_preprocessing(options)
check_options_ocr_behavior(options) check_options_ocr_behavior(options)
check_options_optimizing(options)
check_options_advanced(options) check_options_advanced(options)
check_options_pillow(options) check_options_pillow(options)
plugin_manager.hook.check_options(options=options) plugin_manager.hook.check_options(options=options)
@@ -272,7 +243,7 @@ def check_options(options, plugin_manager):
_check_options(options, plugin_manager, ocr_engine_languages) _check_options(options, plugin_manager, ocr_engine_languages)
def create_input_file(options, work_folder: Path) -> Tuple[Path, str]: def create_input_file(options, work_folder: Path) -> tuple[Path, str]:
if options.input_file == '-': if options.input_file == '-':
# stdin # stdin
log.info('reading file from standard input') log.info('reading file from standard input')
@@ -293,7 +264,7 @@ def create_input_file(options, work_folder: Path) -> Tuple[Path, str]:
target = work_folder / 'origin' target = work_folder / 'origin'
safe_symlink(options.input_file, target) safe_symlink(options.input_file, target)
return target, os.fspath(options.input_file) return target, os.fspath(options.input_file)
except FileNotFoundError: except FileNotFoundError as e:
msg = f"File not found - {options.input_file}" msg = f"File not found - {options.input_file}"
if Path('/.dockerenv').exists(): # pragma: no cover if Path('/.dockerenv').exists(): # pragma: no cover
msg += ( msg += (
@@ -304,7 +275,7 @@ def create_input_file(options, work_folder: Path) -> Tuple[Path, str]:
"\n" "\n"
"\tdocker run -i --rm jbarlow83/ocrmypdf - - <input.pdf >output.pdf\n" "\tdocker run -i --rm jbarlow83/ocrmypdf - - <input.pdf >output.pdf\n"
) )
raise InputFileError(msg) raise InputFileError(msg) from e
def check_requested_output_file(options): def check_requested_output_file(options):
@@ -324,7 +295,16 @@ def check_requested_output_file(options):
) )
def report_output_file_size(options, input_file, output_file): def report_output_file_size(
options,
input_file: Path,
output_file: Path,
optimize_messages: Sequence[str] | None = None,
file_overhead: int = 4000,
page_overhead: int = 3000,
):
if optimize_messages is None:
optimize_messages = []
try: try:
output_size = Path(output_file).stat().st_size output_size = Path(output_file).stat().st_size
input_size = Path(input_file).stat().st_size input_size = Path(input_file).stat().st_size
@@ -333,9 +313,7 @@ def report_output_file_size(options, input_file, output_file):
with pikepdf.open(output_file) as p: with pikepdf.open(output_file) as p:
# Overhead constants obtained by estimating amount of data added by OCR # Overhead constants obtained by estimating amount of data added by OCR
# PDF/A conversion, and possible XMP metadata addition, with compression # PDF/A conversion, and possible XMP metadata addition, with compression
FILE_OVERHEAD = 4000 reasonable_overhead = file_overhead + page_overhead * len(p.pages)
OCR_PER_PAGE_OVERHEAD = 3000
reasonable_overhead = FILE_OVERHEAD + OCR_PER_PAGE_OVERHEAD * len(p.pages)
ratio = output_size / input_size ratio = output_size / input_size
reasonable_ratio = output_size / (input_size + reasonable_overhead) reasonable_ratio = output_size / (input_size + reasonable_overhead)
if reasonable_ratio < 1.35 or input_size < 25000: if reasonable_ratio < 1.35 or input_size < 25000:
@@ -355,19 +333,8 @@ def report_output_file_size(options, input_file, output_file):
f"The argument --{arg.replace('_', '-')} was issued, causing transcoding." f"The argument --{arg.replace('_', '-')} was issued, causing transcoding."
) )
if options.optimize == 0: reasons.extend(optimize_messages)
reasons.append("Optimization was disabled.")
else:
image_optimizers = {
'jbig2': jbig2enc.available(),
'pngquant': pngquant.available(),
}
for name, available in image_optimizers.items():
if not available:
reasons.append(
f"The optional dependency '{name}' was not found, so some image "
f"optimizations could not be attempted."
)
if options.output_type.startswith('pdfa'): if options.output_type.startswith('pdfa'):
reasons.append("PDF/A conversion was enabled. (Try `--output-type pdf`.)") reasons.append("PDF/A conversion was enabled. (Try `--output-type pdf`.)")
if options.plugins: if options.plugins:
+8 -2
View File
@@ -4,11 +4,17 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Get version by introspecting package information.
OCRmyPDF uses setuptools_scm to derive version from git tags.
"""
from __future__ import annotations
try: try:
from importlib_metadata import version as _package_version
except ImportError:
from importlib.metadata import version as _package_version from importlib.metadata import version as _package_version
except ImportError:
from importlib_metadata import version as _package_version # type: ignore
PROGRAM_NAME = 'ocrmypdf' PROGRAM_NAME = 'ocrmypdf'
+19 -13
View File
@@ -4,6 +4,9 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Functions for using ocrmypdf as an API."""
from __future__ import annotations
import logging import logging
import os import os
@@ -12,7 +15,7 @@ import threading
from enum import IntEnum from enum import IntEnum
from io import IOBase from io import IOBase
from pathlib import Path from pathlib import Path
from typing import AnyStr, BinaryIO, Iterable, Optional, Union from typing import AnyStr, BinaryIO, Iterable, Union
from warnings import warn from warnings import warn
from ocrmypdf._logging import PageNumberFilter, TqdmConsole from ocrmypdf._logging import PageNumberFilter, TqdmConsole
@@ -25,7 +28,10 @@ from ocrmypdf.helpers import is_iterable_notstr
try: try:
import coloredlogs import coloredlogs
except ModuleNotFoundError: except ModuleNotFoundError:
coloredlogs = None coloredlogs = None # pylint: disable=invalid-name
if coloredlogs:
from humanfriendly.terminal import enable_ansi_support
StrPath = Union[Path, AnyStr] StrPath = Union[Path, AnyStr]
@@ -37,6 +43,7 @@ _api_lock = threading.Lock()
class Verbosity(IntEnum): class Verbosity(IntEnum):
"""Verbosity level for configure_logging.""" """Verbosity level for configure_logging."""
# pylint: disable=invalid-name
quiet = -1 #: Suppress most messages quiet = -1 #: Suppress most messages
default = 0 #: Default level of logging default = 0 #: Default level of logging
debug = 1 #: Output ocrmypdf debug messages debug = 1 #: Output ocrmypdf debug messages
@@ -116,16 +123,15 @@ def configure_logging(
fmt = '%(pageno)s%(message)s' fmt = '%(pageno)s%(message)s'
use_colors = progress_bar_friendly use_colors = progress_bar_friendly
if not coloredlogs: formatter = None
use_colors = False if coloredlogs and use_colors:
if use_colors: use_colors = enable_ansi_support()
if os.name == 'nt':
use_colors = coloredlogs.enable_ansi_support()
if use_colors: if use_colors:
use_colors = coloredlogs.terminal_supports_colors() use_colors = coloredlogs.terminal_supports_colors()
if use_colors: if use_colors:
formatter = coloredlogs.ColoredFormatter(fmt=fmt) formatter = coloredlogs.ColoredFormatter(fmt=fmt)
else:
if not formatter:
formatter = logging.Formatter(fmt=fmt) formatter = logging.Formatter(fmt=fmt)
console.setFormatter(formatter) console.setFormatter(formatter)
@@ -193,7 +199,7 @@ def create_options(
else: else:
cmdline.append(os.fspath(output_file)) cmdline.append(os.fspath(output_file))
parser._api_mode = True parser.enable_api_mode()
options = parser.parse_args(cmdline) options = parser.parse_args(cmdline)
for keyword, val in deferred: for keyword, val in deferred:
setattr(options, keyword, val) setattr(options, keyword, val)
@@ -213,7 +219,7 @@ def ocr( # pylint: disable=unused-argument
language: Iterable[str] = None, language: Iterable[str] = None,
image_dpi: int = None, image_dpi: int = None,
output_type=None, output_type=None,
sidecar: Optional[StrPath] = None, sidecar: StrPath | None = None,
jobs: int = None, jobs: int = None,
use_threads: bool = None, use_threads: bool = None,
title: str = None, title: str = None,
@@ -295,7 +301,7 @@ def ocr( # pylint: disable=unused-argument
text already, and settings did not tell us to proceed. text already, and settings did not tell us to proceed.
ocrmypdf.InputFileError: Any other problem with the input file. ocrmypdf.InputFileError: Any other problem with the input file.
ocrmypdf.SubprocessOutputError: Any error related to executing a subprocess. ocrmypdf.SubprocessOutputError: Any error related to executing a subprocess.
ocrmypdf.EncryptedPdfERror: If the input PDF is encrypted (password protected). ocrmypdf.EncryptedPdfError: If the input PDF is encrypted (password protected).
OCRmyPDF does not remove passwords. OCRmyPDF does not remove passwords.
ocrmypdf.TesseractConfigError: If Tesseract reported its configuration was not ocrmypdf.TesseractConfigError: If Tesseract reported its configuration was not
valid. valid.
+2
View File
@@ -4,6 +4,8 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
# This file exists only mark builtin_plugins as a package. # This file exists only mark builtin_plugins as a package.
# The plugin manager will not load it, so anything defined here may not be # The plugin manager will not load it, so anything defined here may not be
# processed as a module. # processed as a module.
+20 -4
View File
@@ -4,12 +4,15 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
# © 2020 James R. Barlow: github.com/jbarlow83 # © 2020 James R. Barlow: github.com/jbarlow83
# #
# This Source Code Form is subject to the terms of the Mozilla Public # This Source Code Form is subject to the terms of the Mozilla Public
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""OCRmyPDF's multiprocessing/multithreading abstraction layer."""
import logging import logging
import logging.handlers import logging.handlers
@@ -21,7 +24,6 @@ import sys
import threading import threading
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor, as_completed from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor, as_completed
from contextlib import suppress from contextlib import suppress
from multiprocessing.pool import Pool, ThreadPool
from typing import Callable, Iterable, Type, Union from typing import Callable, Iterable, Type, Union
from tqdm import tqdm from tqdm import tqdm
@@ -44,7 +46,8 @@ def log_listener(q: Queue):
should actually write to sys.stderr or whatever we're using, so if this is should actually write to sys.stderr or whatever we're using, so if this is
made into a process the main application needs to be directed to it. made into a process the main application needs to be directed to it.
See https://docs.python.org/3/howto/logging-cookbook.html#logging-to-a-single-file-from-multiple-processes See:
https://docs.python.org/3/howto/logging-cookbook.html#logging-to-a-single-file-from-multiple-processes
""" """
while True: while True:
@@ -89,6 +92,8 @@ def process_init(q: Queue, user_init: UserInit, loglevel) -> None:
def thread_init(q: Queue, user_init: UserInit, loglevel) -> None: def thread_init(q: Queue, user_init: UserInit, loglevel) -> None:
del q # unused but required argument
del loglevel # unused but required argument
# As a thread, block SIGBUS so the main thread deals with it... # As a thread, block SIGBUS so the main thread deals with it...
with suppress(AttributeError): with suppress(AttributeError):
signal.pthread_sigmask(signal.SIG_BLOCK, {signal.SIGBUS}) signal.pthread_sigmask(signal.SIG_BLOCK, {signal.SIGBUS})
@@ -98,6 +103,17 @@ def thread_init(q: Queue, user_init: UserInit, loglevel) -> None:
class StandardExecutor(Executor): class StandardExecutor(Executor):
"""Standard OCRmyPDF concurrent task executor."""
def _cancel_futures_kwargs(self):
"""Shim older Pythons that do not have Executor.shutdown(...cancel_futures=).
Remove this code when support for Python 3.8 is dropped.
"""
if sys.version_info[:2] < (3, 9):
return {}
return dict(cancel_futures=True)
def _execute( def _execute(
self, self,
*, *,
@@ -136,7 +152,7 @@ class StandardExecutor(Executor):
task_finished(result, pbar) task_finished(result, pbar)
except KeyboardInterrupt: except KeyboardInterrupt:
# Terminate pool so we exit instantly # Terminate pool so we exit instantly
executor.shutdown(wait=False, cancel_futures=True) executor.shutdown(wait=False, **self._cancel_futures_kwargs())
raise raise
except Exception: except Exception:
if not os.environ.get("PYTEST_CURRENT_TEST", ""): if not os.environ.get("PYTEST_CURRENT_TEST", ""):
@@ -145,7 +161,7 @@ class StandardExecutor(Executor):
# results will be discard. But if the condition above is True, # results will be discard. But if the condition above is True,
# then we are running in pytest, and we want everything to exit # then we are running in pytest, and we want everything to exit
# as cleanly as possible so that we get good error messages. # as cleanly as possible so that we get good error messages.
executor.shutdown(wait=False, cancel_futures=True) executor.shutdown(wait=False, **self._cancel_futures_kwargs())
raise raise
finally: finally:
# Terminate log listener # Terminate log listener
@@ -4,6 +4,10 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""OCRmyPDF automatically installs these filters as plugins."""
from __future__ import annotations
from ocrmypdf import hookimpl from ocrmypdf import hookimpl
@@ -5,6 +5,10 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Built-in plugin to implement PDF page rasterization and PDF/A production."""
from __future__ import annotations
import logging import logging
from ocrmypdf import hookimpl from ocrmypdf import hookimpl
+160
View File
@@ -0,0 +1,160 @@
# © 2022 James R. Barlow: github.com/jbarlow83
#
# This Source Code Form is subject to the terms of the Mozilla Public
# License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Built-in plugin to implement PDF page optimization."""
from __future__ import annotations
import argparse
import logging
from pathlib import Path
from typing import Sequence
from ocrmypdf import Executor, PdfContext, hookimpl
from ocrmypdf._exec import jbig2enc, pngquant
from ocrmypdf._pipeline import get_pdf_save_settings
from ocrmypdf.cli import numeric
from ocrmypdf.optimize import optimize
from ocrmypdf.subprocess import check_external_program
log = logging.getLogger(__name__)
@hookimpl
def add_options(parser):
optimizing = parser.add_argument_group(
"Optimization options", "Control how the PDF is optimized after OCR"
)
optimizing.add_argument(
'-O',
'--optimize',
type=int,
choices=range(0, 4),
default=1,
help=(
"Control how PDF is optimized after processing:"
"0 - do not optimize; "
"1 - do safe, lossless optimizations (default); "
"2 - do lossy JPEG and JPEG2000 optimizations; "
"3 - do more aggressive lossy JPEG and JPEG2000 optimizations. "
"To enable lossy JBIG2, see --jbig2-lossy."
),
)
optimizing.add_argument(
'--jpeg-quality',
type=numeric(int, 0, 100),
default=0,
metavar='Q',
help=(
"Adjust JPEG quality level for JPEG optimization. "
"100 is best quality and largest output size; "
"1 is lowest quality and smallest output; "
"0 uses the default."
),
)
optimizing.add_argument(
'--jpg-quality',
type=numeric(int, 0, 100),
default=0,
metavar='Q',
dest='jpeg_quality',
help=argparse.SUPPRESS, # Alias for --jpeg-quality
)
optimizing.add_argument(
'--png-quality',
type=numeric(int, 0, 100),
default=0,
metavar='Q',
help=(
"Adjust PNG quality level to use when quantizing PNGs. "
"Values have same meaning as with --jpeg-quality"
),
)
optimizing.add_argument(
'--jbig2-lossy',
action='store_true',
help=(
"Enable JBIG2 lossy mode (better compression, not suitable for some "
"use cases - see documentation). Only takes effect if --optimize 1 or "
"higher is also enabled."
),
)
optimizing.add_argument(
'--jbig2-page-group-size',
type=numeric(int, 1, 10000),
default=0,
metavar='N',
# Adjust number of pages to consider at once for JBIG2 compression
help=argparse.SUPPRESS,
)
@hookimpl
def check_options(options):
if options.optimize >= 2:
check_external_program(
program='pngquant',
package='pngquant',
version_checker=pngquant.version,
need_version='2.0.1',
required_for='--optimize {2,3}',
)
if options.optimize >= 2:
# Although we use JBIG2 for optimize=1, don't nag about it unless the
# user is asking for more optimization
check_external_program(
program='jbig2',
package='jbig2enc',
version_checker=jbig2enc.version,
need_version='0.28',
required_for='--optimize {2,3} | --jbig2-lossy',
recommended=True if not options.jbig2_lossy else False,
)
if options.optimize == 0 and any(
[options.jbig2_lossy, options.png_quality, options.jpeg_quality]
):
log.warning(
"The arguments --jbig2-lossy, --png-quality, and --jpeg-quality "
"will be ignored because --optimize=0."
)
@hookimpl
def optimize_pdf(
input_pdf: Path,
output_pdf: Path,
context: PdfContext,
executor: Executor,
linearize: bool,
) -> tuple[Path, Sequence[str]]:
save_settings = dict(
linearize=linearize,
**get_pdf_save_settings(context.options.output_type),
)
result_path = optimize(input_pdf, output_pdf, context, save_settings, executor)
messages = []
if context.options.optimize == 0:
messages.append("Optimization was disabled.")
else:
image_optimizers = {
'jbig2': jbig2enc.available(),
'pngquant': pngquant.available(),
}
for name, available in image_optimizers.items():
if not available:
messages.append(
f"The optional dependency '{name}' was not found, so some image "
f"optimizations could not be attempted."
)
return result_path, messages
@hookimpl
def is_optimization_enabled(context: PdfContext) -> bool:
return context.options.optimize != 0
@@ -4,6 +4,10 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Built-in plugin to implement OCR using Tesseract."""
from __future__ import annotations
import logging import logging
import os import os
@@ -138,6 +142,8 @@ def validate(pdfinfo, options):
class TesseractOcrEngine(OcrEngine): class TesseractOcrEngine(OcrEngine):
"""Implements OCR with Tesseract."""
@staticmethod @staticmethod
def version(): def version():
return tesseract.version() return tesseract.version()
+16 -71
View File
@@ -4,9 +4,12 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Command line interface customization and validation."""
from __future__ import annotations
import argparse import argparse
from typing import Any, Callable, Mapping, Optional, TypeVar from typing import Any, Callable, Mapping, TypeVar
from ocrmypdf._version import PROGRAM_NAME as _PROGRAM_NAME from ocrmypdf._version import PROGRAM_NAME as _PROGRAM_NAME
from ocrmypdf._version import __version__ as _VERSION from ocrmypdf._version import __version__ as _VERSION
@@ -14,9 +17,7 @@ from ocrmypdf._version import __version__ as _VERSION
T = TypeVar('T', int, float) T = TypeVar('T', int, float)
def numeric( def numeric(basetype: Callable[[Any], T], min_: T | None = None, max_: T | None = None):
basetype: Callable[[Any], T], min_: Optional[T] = None, max_: Optional[T] = None
):
"""Validator for numeric params""" """Validator for numeric params"""
min_ = basetype(min_) if min_ is not None else None min_ = basetype(min_) if min_ is not None else None
max_ = basetype(max_) if max_ is not None else None max_ = basetype(max_) if max_ is not None else None
@@ -42,7 +43,7 @@ def str_to_int(mapping: Mapping[str, int]):
except KeyError: except KeyError:
raise argparse.ArgumentTypeError( raise argparse.ArgumentTypeError(
f"{s!r} must be one of: {', '.join(mapping.keys())}" f"{s!r} must be one of: {', '.join(mapping.keys())}"
) ) from None
return _str_to_int return _str_to_int
@@ -51,12 +52,20 @@ class ArgumentParser(argparse.ArgumentParser):
"""Override parser's default behavior of calling sys.exit() """Override parser's default behavior of calling sys.exit()
https://stackoverflow.com/questions/5943249/python-argparse-and-controlling-overriding-the-exit-status-code https://stackoverflow.com/questions/5943249/python-argparse-and-controlling-overriding-the-exit-status-code
OCRmyPDF began as a CLI but eventually acquired an API. The API works inside out,
by synthesizing a command line argument. So we subclass the standard parser with
one that doesn't call sys.exit(). Obviously this is not the ideal way to do things
but it works for us.
""" """
def __init__(self, *args, **kwargs): def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs) super().__init__(*args, **kwargs)
self._api_mode = False self._api_mode = False
def enable_api_mode(self):
self._api_mode = True
def error(self, message): def error(self, message):
if not self._api_mode: if not self._api_mode:
super().error(message) super().error(message)
@@ -65,6 +74,8 @@ class ArgumentParser(argparse.ArgumentParser):
class LanguageSetAction(argparse.Action): class LanguageSetAction(argparse.Action):
"""Manages a list of languages."""
def __init__(self, option_strings, dest, default=None, **kwargs): def __init__(self, option_strings, dest, default=None, **kwargs):
if default is None: if default is None:
default = set() default = set()
@@ -341,72 +352,6 @@ Online documentation is located at:
"but include skipped pages in final output", "but include skipped pages in final output",
) )
optimizing = parser.add_argument_group(
"Optimization options", "Control how the PDF is optimized after OCR"
)
optimizing.add_argument(
'-O',
'--optimize',
type=int,
choices=range(0, 4),
default=1,
help=(
"Control how PDF is optimized after processing:"
"0 - do not optimize; "
"1 - do safe, lossless optimizations (default); "
"2 - do lossy JPEG and JPEG2000 optimizations; "
"3 - do more aggressive lossy JPEG and JPEG2000 optimizations. "
"To enable lossy JBIG2, see --jbig2-lossy."
),
)
optimizing.add_argument(
'--jpeg-quality',
type=numeric(int, 0, 100),
default=0,
metavar='Q',
help=(
"Adjust JPEG quality level for JPEG optimization. "
"100 is best quality and largest output size; "
"1 is lowest quality and smallest output; "
"0 uses the default."
),
)
optimizing.add_argument(
'--jpg-quality',
type=numeric(int, 0, 100),
default=0,
metavar='Q',
dest='jpeg_quality',
help=argparse.SUPPRESS, # Alias for --jpeg-quality
)
optimizing.add_argument(
'--png-quality',
type=numeric(int, 0, 100),
default=0,
metavar='Q',
help=(
"Adjust PNG quality level to use when quantizing PNGs. "
"Values have same meaning as with --jpeg-quality"
),
)
optimizing.add_argument(
'--jbig2-lossy',
action='store_true',
help=(
"Enable JBIG2 lossy mode (better compression, not suitable for some "
"use cases - see documentation). Only takes effect if --optimize 1 or "
"higher is also enabled."
),
)
optimizing.add_argument(
'--jbig2-page-group-size',
type=numeric(int, 1, 10000),
default=0,
metavar='N',
# Adjust number of pages to consider at once for JBIG2 compression
help=argparse.SUPPRESS,
)
advanced = parser.add_argument_group( advanced = parser.add_argument_group(
"Advanced", "Advanced options to control OCRmyPDF" "Advanced", "Advanced options to control OCRmyPDF"
) )
+2
View File
@@ -6,3 +6,5 @@
"""Data files used to generate certain PDFs.""" """Data files used to generate certain PDFs."""
from __future__ import annotations
+35 -2
View File
@@ -4,12 +4,18 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""OCRmyPDF's exceptions."""
from __future__ import annotations
from enum import IntEnum from enum import IntEnum
from textwrap import dedent from textwrap import dedent
class ExitCode(IntEnum): class ExitCode(IntEnum):
"""OCRmyPDF's exit codes."""
# pylint: disable=invalid-name
ok = 0 ok = 0
bad_args = 1 bad_args = 1
input_file = 2 input_file = 2
@@ -26,6 +32,8 @@ class ExitCode(IntEnum):
class ExitCodeException(Exception): class ExitCodeException(Exception):
"""An exception which should return an exit code with sys.exit()."""
exit_code = ExitCode.other_error exit_code = ExitCode.other_error
message = "" message = ""
@@ -37,17 +45,24 @@ class ExitCodeException(Exception):
class BadArgsError(ExitCodeException): class BadArgsError(ExitCodeException):
"""Invalid arguments on the command line or API."""
exit_code = ExitCode.bad_args exit_code = ExitCode.bad_args
class PdfMergeFailedError(ExitCodeException): class PdfMergeFailedError(ExitCodeException): # deprecated
"""An intermediate PDF can't be merged.
No longer in use.
"""
exit_code = ExitCode.input_file exit_code = ExitCode.input_file
message = dedent( message = dedent(
'''\ '''\
Failed to merge PDF image layer with OCR layer Failed to merge PDF image layer with OCR layer
Usually this happens because the input PDF file is malformed and Usually this happens because the input PDF file is malformed and
ocrmypdf cannot automatically correct the problem on its own. ocrmypdf cannot correct the problem on its own.
Try using Try using
ocrmypdf --pdf-renderer sandwich [..other args..] ocrmypdf --pdf-renderer sandwich [..other args..]
@@ -56,34 +71,50 @@ class PdfMergeFailedError(ExitCodeException):
class MissingDependencyError(ExitCodeException): class MissingDependencyError(ExitCodeException):
"""A third-party dependency is missing."""
exit_code = ExitCode.missing_dependency exit_code = ExitCode.missing_dependency
class UnsupportedImageFormatError(ExitCodeException): class UnsupportedImageFormatError(ExitCodeException):
"""The image format is not supported."""
exit_code = ExitCode.input_file exit_code = ExitCode.input_file
class DpiError(ExitCodeException): class DpiError(ExitCodeException):
"""Missing information about input image DPI."""
exit_code = ExitCode.input_file exit_code = ExitCode.input_file
class OutputFileAccessError(ExitCodeException): class OutputFileAccessError(ExitCodeException):
"""Cannot access the intended output file path."""
exit_code = ExitCode.file_access_error exit_code = ExitCode.file_access_error
class PriorOcrFoundError(ExitCodeException): class PriorOcrFoundError(ExitCodeException):
"""This file already has OCR."""
exit_code = ExitCode.already_done_ocr exit_code = ExitCode.already_done_ocr
class InputFileError(ExitCodeException): class InputFileError(ExitCodeException):
"""Something is wrong with the input file."""
exit_code = ExitCode.input_file exit_code = ExitCode.input_file
class SubprocessOutputError(ExitCodeException): class SubprocessOutputError(ExitCodeException):
"""A subprocess returned an unexpected error."""
exit_code = ExitCode.child_process_error exit_code = ExitCode.child_process_error
class EncryptedPdfError(ExitCodeException): class EncryptedPdfError(ExitCodeException):
"""Input PDF is encrypted."""
exit_code = ExitCode.encrypted_pdf exit_code = ExitCode.encrypted_pdf
message = dedent( message = dedent(
'''\ '''\
@@ -100,5 +131,7 @@ class EncryptedPdfError(ExitCodeException):
class TesseractConfigError(ExitCodeException): class TesseractConfigError(ExitCodeException):
"""Tesseract config can't be parsed."""
exit_code = ExitCode.invalid_config exit_code = ExitCode.invalid_config
message = "Error occurred while parsing a Tesseract configuration file" message = "Error occurred while parsing a Tesseract configuration file"
+20 -12
View File
@@ -20,6 +20,8 @@ be guaranteed, some workers may end up with too much work while others are idle.
It is less efficient than the standard implementation, so not th edefault. It is less efficient than the standard implementation, so not th edefault.
""" """
from __future__ import annotations
import logging import logging
import logging.handlers import logging.handlers
import signal import signal
@@ -28,7 +30,7 @@ from enum import Enum, auto
from itertools import islice, repeat, takewhile, zip_longest from itertools import islice, repeat, takewhile, zip_longest
from multiprocessing import Pipe, Process from multiprocessing import Pipe, Process
from multiprocessing.connection import Connection, wait from multiprocessing.connection import Connection, wait
from typing import Callable, Iterable, Iterator, List from typing import Callable, Iterable, Iterator
from ocrmypdf import Executor, hookimpl from ocrmypdf import Executor, hookimpl
from ocrmypdf._concurrent import NullProgressBar from ocrmypdf._concurrent import NullProgressBar
@@ -37,9 +39,11 @@ from ocrmypdf.helpers import remove_all_log_handlers
class MessageType(Enum): class MessageType(Enum):
exception = auto() """Implement basic IPC messaging."""
result = auto()
complete = auto() exception = auto() # pylint: disable=invalid-name
result = auto() # pylint: disable=invalid-name
complete = auto() # pylint: disable=invalid-name
def split_every(n: int, iterable: Iterable) -> Iterator: def split_every(n: int, iterable: Iterable) -> Iterator:
@@ -59,6 +63,8 @@ def process_sigbus(*args):
class ConnectionLogHandler(logging.handlers.QueueHandler): class ConnectionLogHandler(logging.handlers.QueueHandler):
"""Handler used by child processes to forward log messages to parent."""
def __init__(self, conn: Connection) -> None: def __init__(self, conn: Connection) -> None:
# sets the parent's queue to None - parent only touches queue # sets the parent's queue to None - parent only touches queue
# in enqueue() which we override # in enqueue() which we override
@@ -91,7 +97,7 @@ def process_loop(
for args in task_args: for args in task_args:
try: try:
result = task(args) result = task(args)
except Exception as e: except Exception as e: # pylint: disable=broad-except
conn.send((MessageType.exception, e)) conn.send((MessageType.exception, e))
break break
else: else:
@@ -103,6 +109,8 @@ def process_loop(
class LambdaExecutor(Executor): class LambdaExecutor(Executor):
"""Executor for AWS Lambda or similar environments that lack semaphores."""
def _execute( def _execute(
self, self,
*, *,
@@ -128,8 +136,8 @@ class LambdaExecutor(Executor):
if not grouped_args: if not grouped_args:
return return
processes: List[Process] = [] processes: list[Process] = []
connections: List[Connection] = [] connections: list[Connection] = []
for chunk in grouped_args: for chunk in grouped_args:
parent_conn, child_conn = Pipe() parent_conn, child_conn = Pipe()
@@ -153,13 +161,13 @@ class LambdaExecutor(Executor):
with self.pbar_class(**tqdm_kwargs) as pbar: with self.pbar_class(**tqdm_kwargs) as pbar:
while connections: while connections:
for r in wait(connections): for result in wait(connections):
if not isinstance(r, Connection): if not isinstance(result, Connection):
raise NotImplementedError("We only support Connection()") raise NotImplementedError("We only support Connection()")
try: try:
msg_type, msg = r.recv() msg_type, msg = result.recv()
except EOFError: except EOFError:
connections.remove(r) connections.remove(result)
continue continue
if msg_type == MessageType.result: if msg_type == MessageType.result:
@@ -170,7 +178,7 @@ class LambdaExecutor(Executor):
logger = logging.getLogger(record.name) logger = logging.getLogger(record.name)
logger.handle(record) logger.handle(record)
elif msg_type == MessageType.complete: elif msg_type == MessageType.complete:
connections.remove(r) connections.remove(result)
elif msg_type == MessageType.exception: elif msg_type == MessageType.exception:
for process in processes: for process in processes:
process.terminate() process.terminate()
+18 -15
View File
@@ -4,6 +4,9 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Support functions."""
from __future__ import annotations
import logging import logging
import multiprocessing import multiprocessing
@@ -137,11 +140,11 @@ def safe_symlink(input_file: os.PathLike, soft_link_name: os.PathLike):
os.symlink(os.path.abspath(input_file), soft_link_name) os.symlink(os.path.abspath(input_file), soft_link_name)
def samefile(f1: os.PathLike, f2: os.PathLike): def samefile(file1: os.PathLike, file2: os.PathLike):
if os.name == 'nt': if os.name == 'nt':
return f1 == f2 return file1 == file2
else: else:
return os.path.samefile(f1, f2) return os.path.samefile(file1, file2)
def is_iterable_notstr(thing: Any) -> bool: def is_iterable_notstr(thing: Any) -> bool:
@@ -149,9 +152,9 @@ def is_iterable_notstr(thing: Any) -> bool:
return isinstance(thing, Iterable) and not isinstance(thing, str) return isinstance(thing, Iterable) and not isinstance(thing, str)
def monotonic(L: Sequence) -> bool: def monotonic(seq: Sequence) -> bool:
"""Does this sequence increase monotonically?""" """Does this sequence increase monotonically?"""
return all(b > a for a, b in zip(L, L[1:])) return all(b > a for a, b in zip(seq, seq[1:]))
def page_number(input_file: os.PathLike) -> int: def page_number(input_file: os.PathLike) -> int:
@@ -166,7 +169,7 @@ def available_cpu_count() -> int:
except NotImplementedError: except NotImplementedError:
pass pass
warnings.warn( warnings.warn(
"Could not get CPU count. Assuming one (1) CPU." "Use -j N to set manually." "Could not get CPU count. Assuming one (1) CPU. Use -j N to set manually."
) )
return 1 return 1
@@ -190,16 +193,16 @@ def is_file_writable(test_file: os.PathLike) -> bool:
os.W_OK, os.W_OK,
effective_ids=(os.access in os.supports_effective_ids), effective_ids=(os.access in os.supports_effective_ids),
) )
try:
fp = p.open('wb')
except OSError:
return False
else: else:
try: fp.close()
fp = p.open('wb') with suppress(OSError):
except OSError: p.unlink()
return False return True
else:
fp.close()
with suppress(OSError):
p.unlink()
return True
except (OSError, RuntimeError) as e: except (OSError, RuntimeError) as e:
log.debug(e) log.debug(e)
log.error(str(e)) log.error(str(e))
+20 -11
View File
@@ -28,17 +28,26 @@
# TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE # TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE
# SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. # SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
"""Transform .hocr and page image to text PDF."""
from __future__ import annotations
import argparse import argparse
import os import os
import re import re
import warnings
from math import atan, cos, sin from math import atan, cos, sin
from pathlib import Path from pathlib import Path
from typing import Any, NamedTuple, Optional, Tuple, Union from typing import Any, NamedTuple, Optional, Tuple, Union
from xml.etree import ElementTree from xml.etree import ElementTree
from reportlab.lib.colors import black, cyan, magenta, red with warnings.catch_warnings():
from reportlab.lib.units import inch warnings.filterwarnings(
from reportlab.pdfgen.canvas import Canvas 'ignore', category=DeprecationWarning, message=r".*load_module.*"
)
from reportlab.lib.colors import black, cyan, magenta, red
from reportlab.lib.units import inch
from reportlab.pdfgen.canvas import Canvas
# According to Wikipedia these languages are supported in the ISO-8859-1 character # According to Wikipedia these languages are supported in the ISO-8859-1 character
# set, meaning reportlab can generate them and they are compatible with hocr, # set, meaning reportlab can generate them and they are compatible with hocr,
@@ -99,7 +108,7 @@ HOCR_OK_LANGS = frozenset(
Element = ElementTree.Element Element = ElementTree.Element
class Rect(NamedTuple): # pylint: disable=inherit-non-class class Rect(NamedTuple):
"""A rectangle for managing PDF coordinates.""" """A rectangle for managing PDF coordinates."""
x1: Any x1: Any
@@ -109,7 +118,7 @@ class Rect(NamedTuple): # pylint: disable=inherit-non-class
class HocrTransformError(Exception): class HocrTransformError(Exception):
pass """Error while applying hOCR transform."""
class HocrTransform: class HocrTransform:
@@ -132,7 +141,7 @@ class HocrTransform:
{'': 'ff', '': 'ffi', '': 'ffl', '': 'fi', '': 'fl'} {'': 'ff', '': 'ffi', '': 'ffl', '': 'fi', '': 'fl'}
) )
def __init__(self, *, hocr_filename: Union[str, Path], dpi: float): def __init__(self, *, hocr_filename: str | Path, dpi: float):
self.dpi = dpi self.dpi = dpi
self.hocr = ElementTree.parse(os.fspath(hocr_filename)) self.hocr = ElementTree.parse(os.fspath(hocr_filename))
@@ -196,7 +205,7 @@ class HocrTransform:
return out return out
@classmethod @classmethod
def baseline(cls, element: Element) -> Tuple[float, float]: def baseline(cls, element: Element) -> tuple[float, float]:
""" """
Returns a tuple containing the baseline slope and intercept. Returns a tuple containing the baseline slope and intercept.
""" """
@@ -212,7 +221,7 @@ class HocrTransform:
""" """
return Rect._make((c / self.dpi * inch) for c in pxl) return Rect._make((c / self.dpi * inch) for c in pxl)
def _child_xpath(self, html_tag: str, html_class: Optional[str] = None) -> str: def _child_xpath(self, html_tag: str, html_class: str | None = None) -> str:
xpath = f".//{self.xmlns}{html_tag}" xpath = f".//{self.xmlns}{html_tag}"
if html_class: if html_class:
xpath += f"[@class='{html_class}']" xpath += f"[@class='{html_class}']"
@@ -239,7 +248,7 @@ class HocrTransform:
self, self,
*, *,
out_filename: Path, out_filename: Path,
image_filename: Optional[Path] = None, image_filename: Path | None = None,
show_bounding_boxes: bool = False, show_bounding_boxes: bool = False,
fontname: str = "Helvetica", fontname: str = "Helvetica",
invisible_text: bool = False, invisible_text: bool = False,
@@ -287,7 +296,7 @@ class HocrTransform:
continue continue
pxl_coords = self.element_coordinates(elem) pxl_coords = self.element_coordinates(elem)
pt = self.pt_from_pixel(pxl_coords) pt = self.pt_from_pixel(pxl_coords) # pylint: disable=invalid-name
# draw the bbox border # draw the bbox border
if show_bounding_boxes: # pragma: no cover if show_bounding_boxes: # pragma: no cover
@@ -342,7 +351,7 @@ class HocrTransform:
def _do_line( def _do_line(
self, self,
pdf: Canvas, pdf: Canvas,
line: Optional[Element], line: Element | None,
elemclass: str, elemclass: str,
fontname: str, fontname: str,
invisible_text: bool, invisible_text: bool,
+34 -39
View File
@@ -4,6 +4,10 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Post-processing image optimization of OCR PDFs."""
from __future__ import annotations
import logging import logging
import sys import sys
@@ -12,18 +16,7 @@ import threading
from collections import defaultdict from collections import defaultdict
from os import fspath from os import fspath
from pathlib import Path from pathlib import Path
from typing import ( from typing import Callable, Iterator, MutableSet, NamedTuple, NewType, Sequence
Callable,
Dict,
Iterator,
List,
MutableSet,
NamedTuple,
NewType,
Optional,
Sequence,
Tuple,
)
from zlib import compress from zlib import compress
import img2pdf import img2pdf
@@ -55,7 +48,9 @@ DEFAULT_PNG_QUALITY = 70
Xref = NewType('Xref', int) Xref = NewType('Xref', int)
class XrefExt(NamedTuple): # pylint: disable=inherit-non-class class XrefExt(NamedTuple):
"""A PDF xref and image extension pair."""
xref: Xref xref: Xref
ext: str ext: str
@@ -74,7 +69,7 @@ def jpg_name(root: Path, xref: Xref) -> Path:
def extract_image_filter( def extract_image_filter(
pike: Pdf, root: Path, image: Stream, xref: Xref pike: Pdf, root: Path, image: Stream, xref: Xref
) -> Optional[Tuple[PdfImage, Tuple[Name, Object]]]: ) -> tuple[PdfImage, tuple[Name, Object]] | None:
del pike # unused args del pike # unused args
del root del root
@@ -131,7 +126,7 @@ def extract_image_filter(
def extract_image_jbig2( def extract_image_jbig2(
*, pike: Pdf, root: Path, image: Stream, xref: Xref, options *, pike: Pdf, root: Path, image: Stream, xref: Xref, options
) -> Optional[XrefExt]: ) -> XrefExt | None:
del options # unused arg del options # unused arg
result = extract_image_filter(pike, root, image, xref) result = extract_image_filter(pike, root, image, xref)
@@ -172,7 +167,7 @@ def extract_image_jbig2(
def extract_image_generic( def extract_image_generic(
*, pike: Pdf, root: Path, image: Stream, xref: Xref, options *, pike: Pdf, root: Path, image: Stream, xref: Xref, options
) -> Optional[XrefExt]: ) -> XrefExt | None:
result = extract_image_filter(pike, root, image, xref) result = extract_image_filter(pike, root, image, xref)
if result is None: if result is None:
return None return None
@@ -236,8 +231,8 @@ def extract_images(
pike: Pdf, pike: Pdf,
root: Path, root: Path,
options, options,
extract_fn: Callable[..., Optional[XrefExt]], extract_fn: Callable[..., XrefExt | None],
) -> Iterator[Tuple[int, XrefExt]]: ) -> Iterator[tuple[int, XrefExt]]:
"""Extract image using extract_fn """Extract image using extract_fn
Enumerate images on each page, lookup their xref/ID number in the PDF. Enumerate images on each page, lookup their xref/ID number in the PDF.
@@ -266,7 +261,7 @@ def extract_images(
if image.objgen[1] != 0: if image.objgen[1] != 0:
continue # Ignore images in an incremental PDF continue # Ignore images in an incremental PDF
xref = Xref(image.objgen[0]) xref = Xref(image.objgen[0])
if hasattr(image, 'SMask'): if Name.SMask in image:
# Ignore soft masks # Ignore soft masks
smask_xref = Xref(image.SMask.objgen[0]) smask_xref = Xref(image.SMask.objgen[0])
exclude_xrefs.add(smask_xref) exclude_xrefs.add(smask_xref)
@@ -296,7 +291,7 @@ def extract_images(
def extract_images_generic( def extract_images_generic(
pike: Pdf, root: Path, options pike: Pdf, root: Path, options
) -> Tuple[List[Xref], List[Xref]]: ) -> tuple[list[Xref], list[Xref]]:
"""Extract any >=2bpp image we think we can improve""" """Extract any >=2bpp image we think we can improve"""
jpegs = [] jpegs = []
@@ -311,7 +306,7 @@ def extract_images_generic(
return jpegs, pngs return jpegs, pngs
def extract_images_jbig2(pike: Pdf, root: Path, options) -> Dict[int, List[XrefExt]]: def extract_images_jbig2(pike: Pdf, root: Path, options) -> dict[int, list[XrefExt]]:
"""Extract any bitonal image that we think we can improve as JBIG2""" """Extract any bitonal image that we think we can improve as JBIG2"""
jbig2_groups = defaultdict(list) jbig2_groups = defaultdict(list)
@@ -324,11 +319,11 @@ def extract_images_jbig2(pike: Pdf, root: Path, options) -> Dict[int, List[XrefE
def _produce_jbig2_images( def _produce_jbig2_images(
jbig2_groups: Dict[int, List[XrefExt]], root: Path, options, executor: Executor jbig2_groups: dict[int, list[XrefExt]], root: Path, options, executor: Executor
) -> None: ) -> None:
"""Produce JBIG2 images from their groups""" """Produce JBIG2 images from their groups"""
def jbig2_group_args(root: Path, groups: Dict[int, List[XrefExt]]): def jbig2_group_args(root: Path, groups: dict[int, list[XrefExt]]):
for group, xref_exts in groups.items(): for group, xref_exts in groups.items():
prefix = f'group{group:08d}' prefix = f'group{group:08d}'
yield ( yield (
@@ -337,7 +332,7 @@ def _produce_jbig2_images(
prefix, # =out_prefix prefix, # =out_prefix
) )
def jbig2_single_args(root, groups: Dict[int, List[XrefExt]]): def jbig2_single_args(root, groups: dict[int, list[XrefExt]]):
for group, xref_exts in groups.items(): for group, xref_exts in groups.items():
prefix = f'group{group:08d}' prefix = f'group{group:08d}'
# Second loop is to ensure multiple images per page are unpacked # Second loop is to ensure multiple images per page are unpacked
@@ -372,7 +367,7 @@ def _produce_jbig2_images(
def convert_to_jbig2( def convert_to_jbig2(
pike: Pdf, pike: Pdf,
jbig2_groups: Dict[int, List[XrefExt]], jbig2_groups: dict[int, list[XrefExt]],
root: Path, root: Path,
options, options,
executor: Executor, executor: Executor,
@@ -389,7 +384,7 @@ def convert_to_jbig2(
When the JBIG2 symbolic coder is not used, each JBIG2 stands on its own When the JBIG2 symbolic coder is not used, each JBIG2 stands on its own
and needs no dictionary. Currently this must be lossless JBIG2. and needs no dictionary. Currently this must be lossless JBIG2.
""" """
jbig2_globals_dict: Optional[Dictionary] jbig2_globals_dict: Dictionary | None
_produce_jbig2_images(jbig2_groups, root, options, executor) _produce_jbig2_images(jbig2_groups, root, options, executor)
@@ -415,7 +410,7 @@ def convert_to_jbig2(
) )
def _optimize_jpeg(args: Tuple[Xref, Path, Path, int]) -> Tuple[Xref, Optional[Path]]: def _optimize_jpeg(args: tuple[Xref, Path, Path, int]) -> tuple[Xref, Path | None]:
xref, in_jpg, opt_jpg, jpeg_quality = args xref, in_jpg, opt_jpg, jpeg_quality = args
with Image.open(in_jpg) as im: with Image.open(in_jpg) as im:
@@ -431,13 +426,13 @@ def _optimize_jpeg(args: Tuple[Xref, Path, Path, int]) -> Tuple[Xref, Optional[P
def transcode_jpegs( def transcode_jpegs(
pike: Pdf, jpegs: Sequence[Xref], root: Path, options, executor: Executor pike: Pdf, jpegs: Sequence[Xref], root: Path, options, executor: Executor
) -> None: ) -> None:
def jpeg_args() -> Iterator[Tuple[Xref, Path, Path, int]]: def jpeg_args() -> Iterator[tuple[Xref, Path, Path, int]]:
for xref in jpegs: for xref in jpegs:
in_jpg = jpg_name(root, xref) in_jpg = jpg_name(root, xref)
opt_jpg = in_jpg.with_suffix('.opt.jpg') opt_jpg = in_jpg.with_suffix('.opt.jpg')
yield xref, in_jpg, opt_jpg, options.jpeg_quality yield xref, in_jpg, opt_jpg, options.jpeg_quality
def finish_jpeg(result: Tuple[Xref, Optional[Path]], pbar): def finish_jpeg(result: tuple[Xref, Path | None], pbar):
xref, opt_jpg = result xref, opt_jpg = result
if opt_jpg: if opt_jpg:
compdata = opt_jpg.read_bytes() # JPEG can inserted into PDF as is compdata = opt_jpg.read_bytes() # JPEG can inserted into PDF as is
@@ -462,11 +457,11 @@ def transcode_jpegs(
def _find_deflatable_jpeg( def _find_deflatable_jpeg(
*, pike: Pdf, root: Path, image: Stream, xref: Xref, options *, pike: Pdf, root: Path, image: Stream, xref: Xref, options
) -> Optional[XrefExt]: ) -> XrefExt | None:
result = extract_image_filter(pike, root, image, xref) result = extract_image_filter(pike, root, image, xref)
if result is None: if result is None:
return None return None
pim, filtdp = result _pim, filtdp = result
if filtdp[0] == Name.DCTDecode and not filtdp[1] and options.optimize >= 1: if filtdp[0] == Name.DCTDecode and not filtdp[1] and options.optimize >= 1:
return XrefExt(xref, '.memory') return XrefExt(xref, '.memory')
@@ -474,7 +469,7 @@ def _find_deflatable_jpeg(
return None return None
def _deflate_jpeg(args: Tuple[Pdf, threading.Lock, Xref, int]) -> Tuple[Xref, bytes]: def _deflate_jpeg(args: tuple[Pdf, threading.Lock, Xref, int]) -> tuple[Xref, bytes]:
pike, lock, xref, complevel = args pike, lock, xref, complevel = args
with lock: with lock:
xobj = pike.get_object(xref, 0) xobj = pike.get_object(xref, 0)
@@ -621,11 +616,11 @@ def optimize(
context, context,
save_settings, save_settings,
executor: Executor = DEFAULT_EXECUTOR, executor: Executor = DEFAULT_EXECUTOR,
) -> None: ) -> Path:
options = context.options options = context.options
if options.optimize == 0: if options.optimize == 0:
safe_symlink(input_file, output_file) safe_symlink(input_file, output_file)
return return output_file
if options.jpeg_quality == 0: if options.jpeg_quality == 0:
options.jpeg_quality = DEFAULT_JPEG_QUALITY if options.optimize < 3 else 40 options.jpeg_quality = DEFAULT_JPEG_QUALITY if options.optimize < 3 else 40
@@ -660,9 +655,7 @@ def optimize(
f"Output file not created after optimizing. We probably ran " f"Output file not created after optimizing. We probably ran "
f"out of disk space in the temporary folder: {tempfile.gettempdir()}." f"out of disk space in the temporary folder: {tempfile.gettempdir()}."
) )
ratio = input_size / output_size
savings = 1 - output_size / input_size savings = 1 - output_size / input_size
log.info(f"Optimize ratio: {ratio:.2f} savings: {(savings):.1%}")
if savings < 0: if savings < 0:
log.info( log.info(
@@ -676,6 +669,8 @@ def optimize(
else: else:
safe_symlink(target_file, output_file) safe_symlink(target_file, output_file)
return output_file
def main(infile, outfile, level, jobs=1): def main(infile, outfile, level, jobs=1):
from shutil import copy # pylint: disable=import-outside-toplevel from shutil import copy # pylint: disable=import-outside-toplevel
@@ -707,9 +702,9 @@ def main(infile, outfile, level, jobs=1):
jb2lossy=False, jb2lossy=False,
) )
with TemporaryDirectory() as td: with TemporaryDirectory() as tmpdir:
context = PdfContext(options, td, infile, None, None) context = PdfContext(options, tmpdir, infile, None, None)
tmpout = Path(td) / 'out.pdf' tmpout = Path(tmpdir) / 'out.pdf'
optimize( optimize(
infile, infile,
tmpout, tmpout,
+11 -8
View File
@@ -9,14 +9,16 @@
Utilities for PDF/A production and confirmation with Ghostspcript. Utilities for PDF/A production and confirmation with Ghostspcript.
""" """
from __future__ import annotations
import base64 import base64
from pathlib import Path from pathlib import Path
from typing import Dict, Iterator, Union from typing import Iterator
try: try:
from importlib_resources import files as package_files
except ImportError:
from importlib.resources import files as package_files from importlib.resources import files as package_files
except ImportError:
from importlib_resources import files as package_files # type: ignore
import pikepdf import pikepdf
@@ -25,7 +27,7 @@ SRGB_ICC_PROFILE_NAME = 'sRGB.icc'
def _postscript_objdef( def _postscript_objdef(
alias: str, alias: str,
dictionary: Dict[str, str], dictionary: dict[str, str],
*, *,
stream_name: str = None, stream_name: str = None,
stream_data: bytes = None, stream_data: bytes = None,
@@ -97,7 +99,8 @@ def generate_pdfa_ps(target_filename: Path, icc: str = 'sRGB'):
target_filename: filename to save target_filename: filename to save
icc: ICC identifier such as 'sRGB' icc: ICC identifier such as 'sRGB'
References: References:
Adobe PDFMARK Reference: https://www.adobe.com/content/dam/acom/en/devnet/acrobat/pdfs/pdfmark_reference.pdf Adobe PDFMARK Reference:
https://www.adobe.com/content/dam/acom/en/devnet/acrobat/pdfs/pdfmark_reference.pdf
""" """
if icc != 'sRGB': if icc != 'sRGB':
raise NotImplementedError("Only supporting sRGB") raise NotImplementedError("Only supporting sRGB")
@@ -105,11 +108,11 @@ def generate_pdfa_ps(target_filename: Path, icc: str = 'sRGB'):
bytes_icc_profile = ( bytes_icc_profile = (
package_files('ocrmypdf.data') / SRGB_ICC_PROFILE_NAME package_files('ocrmypdf.data') / SRGB_ICC_PROFILE_NAME
).read_bytes() ).read_bytes()
ps = '\n'.join(_make_postscript(icc, bytes_icc_profile, 3)) postscript = '\n'.join(_make_postscript(icc, bytes_icc_profile, 3))
# We should have encoded everything to pure ASCII by this point, and # We should have encoded everything to pure ASCII by this point, and
# to be safe, only allow ASCII in PostScript # to be safe, only allow ASCII in PostScript
Path(target_filename).write_text(ps, encoding='ascii') Path(target_filename).write_text(postscript, encoding='ascii')
return target_filename return target_filename
@@ -130,7 +133,7 @@ def file_claims_pdfa(filename: Path):
} }
valid_part_conforms = {'1A', '1B', '2A', '2B', '2U', '3A', '3B', '3U'} valid_part_conforms = {'1A', '1B', '2A', '2B', '2U', '3A', '3B', '3U'}
conformance = f'PDF/A-{pdfmeta.pdfa_status}' conformance = f'PDF/A-{pdfmeta.pdfa_status}'
pdfa_dict: Dict[str, Union[str, bool]] = {} pdfa_dict: dict[str, str | bool] = {}
if pdfmeta.pdfa_status in valid_part_conforms: if pdfmeta.pdfa_status in valid_part_conforms:
pdfa_dict['pass'] = True pdfa_dict['pass'] = True
pdfa_dict['output'] = 'pdfa' pdfa_dict['output'] = 'pdfa'
+4
View File
@@ -6,4 +6,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""For extracting information about PDFs prior to OCR."""
from __future__ import annotations
from ocrmypdf.pdfinfo.info import Colorspace, Encoding, PdfInfo from ocrmypdf.pdfinfo.info import Colorspace, Encoding, PdfInfo
+110 -67
View File
@@ -6,28 +6,30 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Extract information about the content of a PDF."""
from __future__ import annotations
import atexit import atexit
import logging import logging
import re import re
from collections import defaultdict from collections import defaultdict
from contextlib import ExitStack from contextlib import ExitStack
from decimal import Decimal from decimal import Decimal
from enum import Enum from enum import Enum, auto
from functools import partial from functools import partial
from math import hypot, inf, isclose from math import hypot, inf, isclose
from os import PathLike from os import PathLike
from pathlib import Path from pathlib import Path
from typing import ( from typing import (
Container, Container,
Dict, Iterable,
Iterator, Iterator,
List,
Mapping, Mapping,
NamedTuple, NamedTuple,
Optional, Optional,
Sequence, Sequence,
Tuple, Tuple,
Union,
) )
from warnings import warn from warnings import warn
@@ -37,6 +39,7 @@ from pikepdf import (
PdfImage, PdfImage,
PdfInlineImage, PdfInlineImage,
PdfMatrix, PdfMatrix,
UnsupportedImageTypeError,
parse_content_stream, parse_content_stream,
) )
@@ -47,13 +50,41 @@ from ocrmypdf.pdfinfo.layout import get_page_analysis, get_text_boxes
logger = logging.getLogger() logger = logging.getLogger()
Colorspace = Enum('Colorspace', 'gray rgb cmyk lab icc index sep devn pattern jpeg2000')
Encoding = Enum( class Colorspace(Enum):
'Encoding', 'ccitt jpeg jpeg2000 jbig2 asciihex ascii85 lzw flate runlength' """Description of common image colorspaces in a PDF."""
)
FRIENDLY_COLORSPACE: Dict[str, Colorspace] = { # pylint: disable=invalid-name
gray = auto()
rgb = auto()
cmyk = auto()
lab = auto()
icc = auto()
index = auto()
sep = auto()
devn = auto()
pattern = auto()
jpeg2000 = auto()
class Encoding(Enum):
"""Description of common image encodings in a PDF."""
# pylint: disable=invalid-name
ccitt = auto()
jpeg = auto()
jpeg2000 = auto()
jbig2 = auto()
asciihex = auto()
ascii85 = auto()
lzw = auto()
flate = auto()
runlength = auto()
FloatRect = Tuple[float, float, float, float]
FRIENDLY_COLORSPACE: dict[str, Colorspace] = {
'/DeviceGray': Colorspace.gray, '/DeviceGray': Colorspace.gray,
'/CalGray': Colorspace.gray, '/CalGray': Colorspace.gray,
'/DeviceRGB': Colorspace.rgb, '/DeviceRGB': Colorspace.rgb,
@@ -71,7 +102,7 @@ FRIENDLY_COLORSPACE: Dict[str, Colorspace] = {
'/I': Colorspace.index, '/I': Colorspace.index,
} }
FRIENDLY_ENCODING: Dict[str, Encoding] = { FRIENDLY_ENCODING: dict[str, Encoding] = {
'/CCITTFaxDecode': Encoding.ccitt, '/CCITTFaxDecode': Encoding.ccitt,
'/DCTDecode': Encoding.jpeg, '/DCTDecode': Encoding.jpeg,
'/JPXDecode': Encoding.jpeg2000, '/JPXDecode': Encoding.jpeg2000,
@@ -85,7 +116,7 @@ FRIENDLY_ENCODING: Dict[str, Encoding] = {
'/RL': Encoding.runlength, '/RL': Encoding.runlength,
} }
FRIENDLY_COMP: Dict[Colorspace, int] = { FRIENDLY_COMP: dict[Colorspace, int] = {
Colorspace.gray: 1, Colorspace.gray: 1,
Colorspace.rgb: 3, Colorspace.rgb: 3,
Colorspace.cmyk: 4, Colorspace.cmyk: 4,
@@ -104,37 +135,45 @@ def _is_unit_square(shorthand):
class XobjectSettings(NamedTuple): class XobjectSettings(NamedTuple):
"""Info about an XObject found in a PDF."""
name: str name: str
shorthand: Tuple[float, float, float, float, float, float] shorthand: tuple[float, float, float, float, float, float]
stack_depth: int stack_depth: int
class InlineSettings(NamedTuple): class InlineSettings(NamedTuple):
"""Info about an inline image found in a PDF."""
iimage: PdfInlineImage iimage: PdfInlineImage
shorthand: Tuple[float, float, float, float, float, float] shorthand: tuple[float, float, float, float, float, float]
stack_depth: int stack_depth: int
class ContentsInfo(NamedTuple): class ContentsInfo(NamedTuple):
xobject_settings: List[XobjectSettings] """Info about various objects found in a PDF."""
inline_images: List[InlineSettings]
xobject_settings: list[XobjectSettings]
inline_images: list[InlineSettings]
found_vector: bool found_vector: bool
found_text: bool found_text: bool
name_index: Mapping[str, List[XobjectSettings]] name_index: Mapping[str, list[XobjectSettings]]
class TextboxInfo(NamedTuple): class TextboxInfo(NamedTuple):
bbox: Tuple[float, float, float, float] """Info about a text box found in a PDF."""
bbox: tuple[float, float, float, float]
is_visible: bool is_visible: bool
is_corrupt: bool is_corrupt: bool
class VectorMarker: class VectorMarker:
pass """Sentinel indicating vector drawing operations were found on a page."""
class TextMarker: class TextMarker:
pass """Sentinel indicating text drawing operations were found on a page."""
def _normalize_stack(graphobjs): def _normalize_stack(graphobjs):
@@ -177,8 +216,8 @@ def _interpret_contents(contentstream: Object, initial_shorthand=UNIT_SQUARE):
stack = [] stack = []
ctm = PdfMatrix(initial_shorthand) ctm = PdfMatrix(initial_shorthand)
xobject_settings: List[XobjectSettings] = [] xobject_settings: list[XobjectSettings] = []
inline_images: List[InlineSettings] = [] inline_images: list[InlineSettings] = []
name_index = defaultdict(lambda: []) name_index = defaultdict(lambda: [])
found_vector = False found_vector = False
found_text = False found_text = False
@@ -196,7 +235,7 @@ def _interpret_contents(contentstream: Object, initial_shorthand=UNIT_SQUARE):
if len(stack) > 32: # See docstring if len(stack) > 32: # See docstring
if len(stack) > 128: if len(stack) > 128:
raise RuntimeError( raise RuntimeError(
"PDF graphics stack overflowed hard limit, operator %i" % n f"PDF graphics stack overflowed hard limit at operator {n}"
) )
warn("PDF graphics stack overflowed spec limit") warn("PDF graphics stack overflowed spec limit")
elif operator == 'Q': elif operator == 'Q':
@@ -282,7 +321,7 @@ def _get_dpi(ctm_shorthand, image_size) -> Resolution:
""" """
a, b, c, d, _, _ = ctm_shorthand a, b, c, d, _, _ = ctm_shorthand # pylint: disable=invalid-name
# Calculate the width and height of the image in PDF units # Calculate the width and height of the image in PDF units
image_drawn = hypot(a, b), hypot(c, d) image_drawn = hypot(a, b), hypot(c, d)
@@ -298,23 +337,25 @@ def _get_dpi(ctm_shorthand, image_size) -> Resolution:
class ImageInfo: class ImageInfo:
"""Information about an image found in a PDF."""
DPI_PREC = Decimal('1.000') DPI_PREC = Decimal('1.000')
_comp: Optional[int] _comp: int | None
_name: str _name: str
def __init__( def __init__(
self, self,
*, *,
name='', name='',
pdfimage: Optional[Object] = None, pdfimage: Object | None = None,
inline: Optional[PdfInlineImage] = None, inline: PdfInlineImage | None = None,
shorthand=None, shorthand=None,
): ):
self._name = str(name) self._name = str(name)
self._shorthand = shorthand self._shorthand = shorthand
pim: Union[PdfInlineImage, PdfImage] pim: PdfInlineImage | PdfImage
if inline is not None: if inline is not None:
self._origin = 'inline' self._origin = 'inline'
@@ -350,13 +391,20 @@ class ImageInfo:
if self._color == Colorspace.icc: if self._color == Colorspace.icc:
# Check the ICC profile to determine actual colorspace # Check the ICC profile to determine actual colorspace
pim_icc = pim.icc try:
if pim_icc.profile.xcolor_space == 'GRAY': pim_icc = pim.icc
self._comp = 1 if pim_icc.profile.xcolor_space == 'GRAY':
elif pim_icc.profile.xcolor_space == 'CMYK': self._comp = 1
self._comp = 4 elif pim_icc.profile.xcolor_space == 'CMYK':
else: self._comp = 4
self._comp = 3 else:
self._comp = 3
except (AttributeError, UnsupportedImageTypeError) as ex:
self._comp = None
logger.warning(
f"An image with a corrupt or unreadable ICC profile was found. "
f"The output PDF may not match the input PDF visually: {ex}. {self}"
)
else: else:
if isinstance(self._color, Colorspace): if isinstance(self._color, Colorspace):
self._comp = FRIENDLY_COMP.get(self._color) self._comp = FRIENDLY_COMP.get(self._color)
@@ -409,15 +457,10 @@ class ImageInfo:
return _get_dpi(self._shorthand, (self._width, self._height)) return _get_dpi(self._shorthand, (self._width, self._height))
def __repr__(self): def __repr__(self):
class_locals = {
attr: getattr(self, attr, None)
for attr in dir(self)
if not attr.startswith('_')
}
return ( return (
"<ImageInfo '{name}' {type_} {width}x{height} {color} " f"<ImageInfo '{self.name}' {self.type_} {self.width}x{self.height} "
"{comp} {bpc} {enc} {dpi}>" f"{self.color} {self.comp} {self.bpc} {self.enc} {self.dpi}>"
).format(**class_locals) )
def _find_inline_images(contentsinfo: ContentsInfo) -> Iterator[ImageInfo]: def _find_inline_images(contentsinfo: ContentsInfo) -> Iterator[ImageInfo]:
@@ -425,11 +468,11 @@ def _find_inline_images(contentsinfo: ContentsInfo) -> Iterator[ImageInfo]:
for n, inline in enumerate(contentsinfo.inline_images): for n, inline in enumerate(contentsinfo.inline_images):
yield ImageInfo( yield ImageInfo(
name='inline-%02d' % n, shorthand=inline.shorthand, inline=inline.iimage name=f'inline-{n:02d}', shorthand=inline.shorthand, inline=inline.iimage
) )
def _image_xobjects(container) -> Iterator[Tuple[Object, str]]: def _image_xobjects(container) -> Iterator[tuple[Object, str]]:
"""Search for all XObject-based images in the container """Search for all XObject-based images in the container
Usually the container is a page, but it could also be a Form XObject Usually the container is a page, but it could also be a Form XObject
@@ -517,7 +560,7 @@ def _find_form_xobject_images(pdf: Pdf, container: Object, contentsinfo: Content
def _process_content_streams( def _process_content_streams(
*, pdf: Pdf, container: Object, shorthand=None *, pdf: Pdf, container: Object, shorthand=None
) -> Iterator[Union[VectorMarker, TextMarker, ImageInfo]]: ) -> Iterator[VectorMarker | TextMarker | ImageInfo]:
"""Find all individual instances of images drawn in the container """Find all individual instances of images drawn in the container
Usually the container is a page, but it may also be a Form XObject. Usually the container is a page, but it may also be a Form XObject.
@@ -566,10 +609,10 @@ def _process_content_streams(
yield from _find_form_xobject_images(pdf, container, contentsinfo) yield from _find_form_xobject_images(pdf, container, contentsinfo)
def _page_has_text(text_blocks, page_width, page_height) -> bool: def _page_has_text(text_blocks: Iterable[FloatRect], page_width, page_height) -> bool:
"""Smarter text detection that ignores text in margins""" """Smarter text detection that ignores text in margins"""
pw, ph = float(page_width), float(page_height) pw, ph = float(page_width), float(page_height) # pylint: disable=invalid-name
margin_ratio = 0.125 margin_ratio = 0.125
interior_bbox = ( interior_bbox = (
@@ -579,7 +622,7 @@ def _page_has_text(text_blocks, page_width, page_height) -> bool:
margin_ratio * ph, # bottom (first quadrant: bottom < top) margin_ratio * ph, # bottom (first quadrant: bottom < top)
) )
def rects_intersect(a, b) -> bool: def rects_intersect(a: FloatRect, b: FloatRect) -> bool:
""" """
Where (a,b) are 4-tuple rects (left-0, top-1, right-2, bottom-3) Where (a,b) are 4-tuple rects (left-0, top-1, right-2, bottom-3)
https://stackoverflow.com/questions/306316/determine-if-two-rectangles-overlap-each-other https://stackoverflow.com/questions/306316/determine-if-two-rectangles-overlap-each-other
@@ -601,19 +644,19 @@ def simplify_textboxes(miner, textbox_getter) -> Iterator[TextboxInfo]:
We do this to save memory and ensure that our objects are pickleable. We do this to save memory and ensure that our objects are pickleable.
""" """
for box in textbox_getter(miner): for box in textbox_getter(miner):
first_line = box._objs[0] first_line = box._objs[0] # pylint: disable=protected-access
first_char = first_line._objs[0] first_char = first_line._objs[0] # pylint: disable=protected-access
visible = first_char.rendermode != 3 visible = first_char.rendermode != 3
corrupt = first_char.get_text() == '\ufffd' corrupt = first_char.get_text() == '\ufffd'
yield TextboxInfo(box.bbox, visible, corrupt) yield TextboxInfo(box.bbox, visible, corrupt)
worker_pdf = None worker_pdf = None # pylint: disable=invalid-name
def _pdf_pageinfo_sync_init(pdf: Pdf, infile: Path, pdfminer_loglevel): def _pdf_pageinfo_sync_init(pdf: Pdf, infile: Path, pdfminer_loglevel):
global worker_pdf # pylint: disable=global-statement global worker_pdf # pylint: disable=global-statement,invalid-name
pikepdf_enable_mmap() pikepdf_enable_mmap()
logging.getLogger('pdfminer').setLevel(pdfminer_loglevel) logging.getLogger('pdfminer').setLevel(pdfminer_loglevel)
@@ -647,8 +690,8 @@ def _pdf_pageinfo_concurrent(
max_workers, max_workers,
check_pages, check_pages,
detailed_analysis=False, detailed_analysis=False,
) -> Sequence[Optional['PageInfo']]: ) -> Sequence[PageInfo | None]:
pages: Sequence[Optional['PageInfo']] = [None] * len(pdf.pages) pages: Sequence[PageInfo | None] = [None] * len(pdf.pages)
def update_pageinfo(result, pbar): def update_pageinfo(result, pbar):
page = result page = result
@@ -698,9 +741,11 @@ def _pdf_pageinfo_concurrent(
class PageInfo: class PageInfo:
_has_text: Optional[bool] """Information about type of contents on each page in a PDF."""
_has_vector: Optional[bool]
_images: List[ImageInfo] _has_text: bool | None
_has_vector: bool | None
_images: list[ImageInfo]
def __init__( def __init__(
self, self,
@@ -759,15 +804,15 @@ class PageInfo:
self._has_vector = False self._has_vector = False
self._has_text = False self._has_text = False
self._images = [] self._images = []
for ci in _process_content_streams( for info in _process_content_streams(
pdf=pdf, container=page, shorthand=userunit_shorthand pdf=pdf, container=page, shorthand=userunit_shorthand
): ):
if isinstance(ci, VectorMarker): if isinstance(info, VectorMarker):
self._has_vector = True self._has_vector = True
elif isinstance(ci, TextMarker): elif isinstance(info, TextMarker):
self._has_text = True self._has_text = True
elif isinstance(ci, ImageInfo): elif isinstance(info, ImageInfo):
self._images.append(ci) self._images.append(info)
else: else:
raise NotImplementedError() raise NotImplementedError()
else: else:
@@ -833,9 +878,7 @@ class PageInfo:
def images(self): def images(self):
return self._images return self._images
def get_textareas( def get_textareas(self, visible: bool | None = None, corrupt: bool | None = None):
self, visible: Optional[bool] = None, corrupt: Optional[bool] = None
):
def predicate(obj, want_visible, want_corrupt): def predicate(obj, want_visible, want_corrupt):
result = True result = True
if want_visible is not None: if want_visible is not None:
@@ -919,7 +962,7 @@ class PdfInfo:
self._has_acroform = True self._has_acroform = True
@property @property
def pages(self) -> Sequence[Optional[PageInfo]]: def pages(self) -> Sequence[PageInfo | None]:
return self._pages return self._pages
@property @property
@@ -936,7 +979,7 @@ class PdfInfo:
return self._has_acroform return self._has_acroform
@property @property
def filename(self) -> Union[str, Path]: def filename(self) -> str | Path:
if not isinstance(self._infile, (str, Path)): if not isinstance(self._infile, (str, Path)):
raise NotImplementedError("can't get filename from stream") raise NotImplementedError("can't get filename from stream")
return self._infile return self._infile
+2
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import re import re
from math import copysign from math import copysign
from pathlib import Path from pathlib import Path
+99 -11
View File
@@ -4,16 +4,19 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""OCRmyPDF pluggy plugin specification."""
from __future__ import annotations
from abc import ABC, abstractmethod from abc import ABC, abstractmethod
from argparse import ArgumentParser, Namespace from argparse import ArgumentParser, Namespace
from logging import Handler from logging import Handler
from pathlib import Path from pathlib import Path
from typing import TYPE_CHECKING, AbstractSet, List, NamedTuple, Optional from typing import TYPE_CHECKING, AbstractSet, NamedTuple, Sequence
import pluggy import pluggy
from ocrmypdf._concurrent import Executor from ocrmypdf import Executor, PdfContext
from ocrmypdf.helpers import Resolution from ocrmypdf.helpers import Resolution
if TYPE_CHECKING: if TYPE_CHECKING:
@@ -42,6 +45,33 @@ def get_logging_console() -> Handler:
""" """
@hookspec
def initialize(plugin_manager: pluggy.PluginManager):
"""Called when this plugin is first loaded into OCRmyPDF.
The primary intended use of this is for plugins to check compatibility with other
plugins and possibly block other blocks, a plugin that wishes to block ocrmypdf's
built-in optimize plugin could do:
.. code-block::
plugin_manager.set_blocked('ocrmypdf.builtin_plugins.optimize')
It would also be reasonable for an plugin implementation to check if it is unable
to proceed, for example, because a required dependency is missing. (If the plugin's
ability to proceed depends on options and arguments, use ``validate`` instead.)
Raises:
ocrmypdf.exceptions.ExitCodeException: If options are not acceptable
and the application should terminate gracefully with an informative
message and error code.
Note:
This hook will be called from the main process, and may modify global state
before child worker processes are forked.
"""
@hookspec @hookspec
def add_options(parser: ArgumentParser) -> None: def add_options(parser: ArgumentParser) -> None:
"""Allows the plugin to add its own command line and API arguments. """Allows the plugin to add its own command line and API arguments.
@@ -141,7 +171,7 @@ def get_progressbar_class():
@hookspec @hookspec
def validate(pdfinfo: 'PdfInfo', options: Namespace) -> None: def validate(pdfinfo: PdfInfo, options: Namespace) -> None:
"""Called to give a plugin an opportunity to review *options* and *pdfinfo*. """Called to give a plugin an opportunity to review *options* and *pdfinfo*.
*options* contains the "work order" to process a particular file. *pdfinfo* *options* contains the "work order" to process a particular file. *pdfinfo*
@@ -167,8 +197,8 @@ def rasterize_pdf_page(
raster_device: str, raster_device: str,
raster_dpi: Resolution, raster_dpi: Resolution,
pageno: int, pageno: int,
page_dpi: Optional[Resolution], page_dpi: Resolution | None,
rotation: Optional[int], rotation: int | None,
filter_vector: bool, filter_vector: bool,
) -> Path: ) -> Path:
"""Rasterize one page of a PDF at resolution raster_dpi in canvas units. """Rasterize one page of a PDF at resolution raster_dpi in canvas units.
@@ -197,7 +227,7 @@ def rasterize_pdf_page(
@hookspec(firstresult=True) @hookspec(firstresult=True)
def filter_ocr_image(page: 'PageContext', image: 'Image.Image') -> 'Image.Image': def filter_ocr_image(page: PageContext, image: Image.Image) -> Image.Image:
"""Called to filter the image before it is sent to OCR. """Called to filter the image before it is sent to OCR.
This is the image that OCR sees, not what the user sees when they view the This is the image that OCR sees, not what the user sees when they view the
@@ -224,7 +254,7 @@ def filter_ocr_image(page: 'PageContext', image: 'Image.Image') -> 'Image.Image'
@hookspec(firstresult=True) @hookspec(firstresult=True)
def filter_page_image(page: 'PageContext', image_filename: Path) -> Path: def filter_page_image(page: PageContext, image_filename: Path) -> Path:
"""Called to filter the whole page before it is inserted into the PDF. """Called to filter the whole page before it is inserted into the PDF.
A whole page image is only produced when preprocessing command line arguments A whole page image is only produced when preprocessing command line arguments
@@ -264,9 +294,7 @@ def filter_page_image(page: 'PageContext', image_filename: Path) -> Path:
@hookspec(firstresult=True) @hookspec(firstresult=True)
def filter_pdf_page( def filter_pdf_page(page: PageContext, image_filename: Path, output_pdf: Path) -> Path:
page: 'PageContext', image_filename: Path, output_pdf: Path
) -> Path:
"""Called to convert a filtered whole page image into a PDF. """Called to convert a filtered whole page image into a PDF.
A whole page image is only produced when preprocessing command line arguments A whole page image is only produced when preprocessing command line arguments
@@ -409,7 +437,7 @@ def get_ocr_engine() -> OcrEngine:
@hookspec(firstresult=True) @hookspec(firstresult=True)
def generate_pdfa( def generate_pdfa(
pdf_pages: List[Path], pdf_pages: list[Path],
pdfmark: Path, pdfmark: Path,
output_file: Path, output_file: Path,
compression: str, compression: str,
@@ -456,3 +484,63 @@ def generate_pdfa(
See also: See also:
https://github.com/tqdm/tqdm https://github.com/tqdm/tqdm
""" """
@hookspec(firstresult=True)
def optimize_pdf(
input_pdf: Path,
output_pdf: Path,
context: PdfContext,
executor: Executor,
linearize: bool,
) -> tuple[Path, Sequence[str]]:
"""Optimize a PDF after image, OCR and metadata processing.
If the input_pdf is a PDF/A, the plugin should modify input_pdf in a way
that preserves the PDF/A status, or report to the user when this is not possible.
If the implementation fails to produce a smaller file than the input file, it
should return input_pdf instead.
A plugin that implements a new optimizer may need to suppress the built-in
optimizer by implementing an ``initialize`` hook.
Arguments:
input_pdf: The input PDF, which has OCR added.
output_pdf: The requested filename of the output PDF which should be created
by this optimization hook.
context: The current context.
executor: An initialized executor which may be used during optimization,
to distribute optimization tasks.
linearize: If True, OCRmyPDF requires ``optimize_pdf`` to return a linearized,
also known as fast web view PDF.
Returns:
Path: If optimization is successful, the hook should return ``output_file``.
If optimization does not produce a smaller file, the hook should return
``input_file``.
Sequence[str]: Any comments that the plugin wishes to report to the user,
especially reasons it was not able to further optimize the file. For
example, the plugin could report that a required third party was not
installed, so a specific optimization was not attempted.
Note:
This is a :ref:`firstresult hook<firstresult>`.
"""
@hookspec(firstresult=True)
def is_optimization_enabled(context: PdfContext) -> bool:
"""For a given PdfContext, OCRmyPDF asks the plugin if optimization is enabled.
An optimization plugin might be installed and active but could be disabled by
user settings.
If this returns False, OCRmyPDF will take certain actions to finalize the PDF.
Returns:
True if the plugin's optimization is enabled.
Note:
This is a :ref:`firstresult hook<firstresult>`.
"""
+2
View File
@@ -8,6 +8,8 @@
"""Utilities to measure OCR quality""" """Utilities to measure OCR quality"""
from __future__ import annotations
import re import re
from typing import Iterable from typing import Iterable
+85 -43
View File
@@ -7,16 +7,18 @@
"""Wrappers to manage subprocess calls""" """Wrappers to manage subprocess calls"""
from __future__ import annotations
import logging import logging
import os import os
import re import re
import sys import sys
from collections.abc import Mapping
from contextlib import suppress from contextlib import suppress
from functools import lru_cache from functools import lru_cache
from pathlib import Path
from subprocess import PIPE, STDOUT, CalledProcessError, CompletedProcess, Popen from subprocess import PIPE, STDOUT, CalledProcessError, CompletedProcess, Popen
from subprocess import run as subprocess_run from subprocess import run as subprocess_run
from typing import Callable, Optional, Type, Union from typing import Callable, Mapping, Sequence, Union
from packaging.version import Version from packaging.version import Version
@@ -26,9 +28,17 @@ from ocrmypdf.exceptions import MissingDependencyError
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
Args = Sequence[Union[Path, str]]
OsEnviron = os._Environ # pylint: disable=protected-access
def run( def run(
args, *, env=None, logs_errors_to_stdout: bool = False, **kwargs args: Args,
*,
env: OsEnviron | None = None,
logs_errors_to_stdout: bool = False,
check: bool = False,
**kwargs,
) -> CompletedProcess: ) -> CompletedProcess:
"""Wrapper around :py:func:`subprocess.run` """Wrapper around :py:func:`subprocess.run`
@@ -50,7 +60,7 @@ def run(
stderr = None stderr = None
stderr_name = 'stderr' if not logs_errors_to_stdout else 'stdout' stderr_name = 'stderr' if not logs_errors_to_stdout else 'stdout'
try: try:
proc = subprocess_run(args, env=env, **kwargs) proc = subprocess_run(args, env=env, check=check, **kwargs)
except CalledProcessError as e: except CalledProcessError as e:
stderr = getattr(e, stderr_name, None) stderr = getattr(e, stderr_name, None)
raise raise
@@ -68,7 +78,12 @@ def run(
def run_polling_stderr( def run_polling_stderr(
args, *, callback: Callable[[str], None], check: bool = False, env=None, **kwargs args: Args,
*,
callback: Callable[[str], None],
check: bool = False,
env: OsEnviron | None = None,
**kwargs,
) -> CompletedProcess: ) -> CompletedProcess:
"""Run a process like ``ocrmypdf.subprocess.run``, and poll stderr. """Run a process like ``ocrmypdf.subprocess.run``, and poll stderr.
@@ -101,7 +116,9 @@ def run_polling_stderr(
return CompletedProcess(args, proc.returncode, None, stderr=stderr) return CompletedProcess(args, proc.returncode, None, stderr=stderr)
def _fix_process_args(args, env, kwargs): def _fix_process_args(
args: Args, env: OsEnviron | None, kwargs
) -> tuple[Args, OsEnviron, logging.Logger, bool]:
assert 'universal_newlines' not in kwargs, "Use text= instead of universal_newlines" assert 'universal_newlines' not in kwargs, "Use text= instead of universal_newlines"
if not env: if not env:
@@ -110,21 +127,26 @@ def _fix_process_args(args, env, kwargs):
# Search in spoof path if necessary # Search in spoof path if necessary
program = str(args[0]) program = str(args[0])
if os.name == 'nt': if sys.platform == 'win32':
# pylint: disable=import-outside-toplevel
from ocrmypdf.subprocess._windows import fix_windows_args from ocrmypdf.subprocess._windows import fix_windows_args
args = fix_windows_args(program, args, env) args = fix_windows_args(program, args, env)
log.debug("Running: %s", args) log.debug("Running: %s", args)
process_log = log.getChild(os.path.basename(program)) process_log = log.getChild(os.path.basename(program))
text = kwargs.get('text', False) text = bool(kwargs.get('text', False))
return args, env, process_log, text return args, env, process_log, text
@lru_cache(maxsize=None) @lru_cache(maxsize=None)
def get_version( def get_version(
program: str, *, version_arg: str = '--version', regex=r'(\d+(\.\d+)*)', env=None program: str,
*,
version_arg: str = '--version',
regex=r'(\d+(\.\d+)*)',
env: OsEnviron | None = None,
) -> str: ) -> str:
"""Get the version of the specified program """Get the version of the specified program
@@ -171,42 +193,42 @@ def get_version(
return version return version
missing_program = ''' MISSING_PROGRAM = '''
The program '{program}' could not be executed or was not found on your The program '{program}' could not be executed or was not found on your
system PATH. system PATH.
''' '''
missing_optional_program = ''' MISSING_OPTIONAL_PROGRAM = '''
The program '{program}' could not be executed or was not found on your The program '{program}' could not be executed or was not found on your
system PATH. This program is required when you use the system PATH. This program is required when you use the
{required_for} arguments. You could try omitting these arguments, or install {required_for} arguments. You could try omitting these arguments, or install
the package. the package.
''' '''
missing_recommend_program = ''' MISSING_RECOMMEND_PROGRAM = '''
The program '{program}' could not be executed or was not found on your The program '{program}' could not be executed or was not found on your
system PATH. This program is recommended when using the {required_for} arguments, system PATH. This program is recommended when using the {required_for} arguments,
but not required, so we will proceed. For best results, install the program. but not required, so we will proceed. For best results, install the program.
''' '''
old_version = ''' OLD_VERSION = '''
OCRmyPDF requires '{program}' {need_version} or higher. Your system appears OCRmyPDF requires '{program}' {need_version} or higher. Your system appears
to have {found_version}. Please update this program. to have {found_version}. Please update this program.
''' '''
old_version_required_for = ''' OLD_VERSION_REQUIRED_FOR = '''
OCRmyPDF requires '{program}' {need_version} or higher when run with the OCRmyPDF requires '{program}' {need_version} or higher when run with the
{required_for} arguments. If you omit these arguments, OCRmyPDF may be able to {required_for} arguments. If you omit these arguments, OCRmyPDF may be able to
proceed. For best results, install the program. proceed. For best results, install the program.
''' '''
osx_install_advice = ''' OSX_INSTALL_ADVICE = '''
If you have homebrew installed, try these command to install the missing If you have homebrew installed, try these command to install the missing
package: package:
brew install {package} brew install {package}
''' '''
linux_install_advice = ''' LINUX_INSTALL_ADVICE = '''
On systems with the aptitude package manager (Debian, Ubuntu), try these On systems with the aptitude package manager (Debian, Ubuntu), try these
commands: commands:
sudo apt-get update sudo apt-get update
@@ -216,14 +238,14 @@ On RPM-based systems (Red Hat, Fedora), search for instructions on
installing the RPM for {program}. installing the RPM for {program}.
''' '''
windows_install_advice = ''' WINDOWS_INSTALL_ADVICE = '''
If not already installed, install the Chocolatey package manager. Then use If not already installed, install the Chocolatey package manager. Then use
a command prompt to install the missing package: a command prompt to install the missing package:
choco install {package} choco install {package}
''' '''
def _get_platform(): def _get_platform() -> str:
if sys.platform.startswith('freebsd'): if sys.platform.startswith('freebsd'):
return 'freebsd' return 'freebsd'
elif sys.platform.startswith('linux'): elif sys.platform.startswith('linux'):
@@ -233,46 +255,66 @@ def _get_platform():
return sys.platform return sys.platform
def _error_trailer(program, package, **kwargs): def _error_trailer(program: str, package: str | Mapping[str, str], **kwargs) -> None:
del kwargs
if isinstance(package, Mapping): if isinstance(package, Mapping):
package = package.get(_get_platform(), program) package = package.get(_get_platform(), program)
if _get_platform() == 'darwin': if _get_platform() == 'darwin':
log.info(osx_install_advice.format(**locals())) log.info(OSX_INSTALL_ADVICE.format(**locals()))
elif _get_platform() == 'linux': elif _get_platform() == 'linux':
log.info(linux_install_advice.format(**locals())) log.info(LINUX_INSTALL_ADVICE.format(**locals()))
elif _get_platform() == 'windows': elif _get_platform() == 'windows':
log.info(windows_install_advice.format(**locals())) log.info(WINDOWS_INSTALL_ADVICE.format(**locals()))
def _error_missing_program(program, package, required_for, recommended): def _error_missing_program(
program: str, package: str, required_for: str | None, recommended: bool
) -> None:
# pylint: disable=unused-argument
if recommended: if recommended:
log.warning(missing_recommend_program.format(**locals())) log.warning(MISSING_RECOMMEND_PROGRAM.format(**locals()))
elif required_for: elif required_for:
log.error(missing_optional_program.format(**locals())) log.error(MISSING_OPTIONAL_PROGRAM.format(**locals()))
else: else:
log.error(missing_program.format(**locals())) log.error(MISSING_PROGRAM.format(**locals()))
_error_trailer(**locals()) _error_trailer(**locals())
def _error_old_version(program, package, need_version, found_version, required_for): def _error_old_version(
program: str,
package: str,
need_version: str,
found_version: str,
required_for: str | None,
) -> None:
# pylint: disable=unused-argument
if required_for: if required_for:
log.error(old_version_required_for.format(**locals())) log.error(OLD_VERSION_REQUIRED_FOR.format(**locals()))
else: else:
log.error(old_version.format(**locals())) log.error(OLD_VERSION.format(**locals()))
_error_trailer(**locals()) _error_trailer(**locals())
def _remove_leading_v(s: str) -> str:
if sys.version_info >= (3, 9):
return s.removeprefix('v')
if s.startswith('v'):
return s[1:]
return s
def check_external_program( def check_external_program(
*, *,
program: str, program: str,
package: str, package: str,
version_checker: Callable, version_checker: Callable[[], str],
need_version: str, need_version: str,
required_for: Optional[str] = None, required_for: str | None = None,
recommended=False, recommended: bool = False,
version_parser: Type[Version] = Version, version_parser: type[Version] = Version,
): ) -> None:
"""Check for required version of external program and raise exception if not. """Check for required version of external program and raise exception if not.
Args: Args:
@@ -294,19 +336,19 @@ def check_external_program(
found_version = version_checker() found_version = version_checker()
else: # deprecated else: # deprecated
found_version = version_checker found_version = version_checker
except (CalledProcessError, FileNotFoundError, MissingDependencyError): except (CalledProcessError, FileNotFoundError) as e:
_error_missing_program(program, package, required_for, recommended) _error_missing_program(program, package, required_for, recommended)
if not recommended: if not recommended:
raise MissingDependencyError(program) raise MissingDependencyError(program) from e
return
except MissingDependencyError:
_error_missing_program(program, package, required_for, recommended)
if not recommended:
raise
return return
def remove_leading_v(s): found_version = _remove_leading_v(found_version)
if s.startswith('v'): need_version = _remove_leading_v(need_version)
return s[1:]
return s
found_version = remove_leading_v(found_version)
need_version = remove_leading_v(need_version)
if found_version and version_parser(found_version) < version_parser(need_version): if found_version and version_parser(found_version) < version_parser(need_version):
_error_old_version(program, package, need_version, found_version, required_for) _error_old_version(program, package, need_version, found_version, required_for)
+33 -19
View File
@@ -3,9 +3,9 @@
# This Source Code Form is subject to the terms of the Mozilla Public # This Source Code Form is subject to the terms of the Mozilla Public
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Find Tesseract and Ghostscript binaries on Windows using the registry."""
# type: ignore from __future__ import annotations
# Non-Windows mypy now breaks when trying to typecheck winreg
import logging import logging
import os import os
@@ -13,12 +13,26 @@ import shutil
import sys import sys
from itertools import chain from itertools import chain
from pathlib import Path from pathlib import Path
from typing import Any, Callable, Iterable, Iterator, Set, Tuple, TypeVar from typing import Any, Callable, Iterable, Iterator, TypeVar
try: if sys.version_info >= (3, 10):
import winreg from typing import TypeAlias
except ModuleNotFoundError as e: else:
raise ModuleNotFoundError("This module is for Windows only") from e from typing_extensions import TypeAlias # pragma: no cover
if sys.platform == 'win32':
# mypy understands 'if sys.platform' better than try/except ModuleNotFoundError
import winreg # pylint: disable=import-error
HKEYType: TypeAlias = winreg.HKEYType
else:
from unittest.mock import Mock
winreg = Mock(
spec=['HKEYType', 'EnumKey', 'EnumValue', 'HKEY_LOCAL_MACHINE', 'OpenKey']
)
# mypy does not understand winreg.HKeyType where winreg is a Mock (fair enough!)
HKEYType: TypeAlias = Any
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@@ -26,7 +40,7 @@ log = logging.getLogger(__name__)
T = TypeVar('T') T = TypeVar('T')
def ghostscript_version_key(s: str) -> Tuple[int, int, int]: def ghostscript_version_key(s: str) -> tuple[int, int, int]:
"""Compare Ghostscript version numbers.""" """Compare Ghostscript version numbers."""
try: try:
release = [int(elem) for elem in s.split('.', maxsplit=3)] release = [int(elem) for elem in s.split('.', maxsplit=3)]
@@ -37,30 +51,29 @@ def ghostscript_version_key(s: str) -> Tuple[int, int, int]:
return (0, 0, 0) return (0, 0, 0)
def registry_enum( def registry_enum(key: HKEYType, enum_fn: Callable[[HKEYType, int], T]) -> Iterator[T]:
key: winreg.HKEYType, enum_fn: Callable[[winreg.HKEYType, int], T] limit = 999
) -> Iterator[T]:
LIMIT = 999
n = 0 n = 0
while n < LIMIT: while n < limit:
try: try:
yield enum_fn(key, n) yield enum_fn(key, n)
n += 1 n += 1
except OSError: except OSError:
break break
if n == LIMIT: if n == limit:
raise ValueError(f"Too many registry keys under {key}") raise ValueError(f"Too many registry keys under {key}")
def registry_subkeys(key: winreg.HKEYType) -> Iterator[str]: def registry_subkeys(key: HKEYType) -> Iterator[str]:
return registry_enum(key, winreg.EnumKey) return registry_enum(key, winreg.EnumKey)
def registry_values(key: winreg.HKEYType) -> Iterator[Tuple[str, Any, int]]: def registry_values(key: HKEYType) -> Iterator[tuple[str, Any, int]]:
return registry_enum(key, winreg.EnumValue) return registry_enum(key, winreg.EnumValue)
def registry_path_ghostscript(env=None) -> Iterator[Path]: def registry_path_ghostscript(env=None) -> Iterator[Path]:
del env # unused (but needed for protocol)
try: try:
with winreg.OpenKey( with winreg.OpenKey(
winreg.HKEY_LOCAL_MACHINE, r"SOFTWARE\Artifex\GPL Ghostscript" winreg.HKEY_LOCAL_MACHINE, r"SOFTWARE\Artifex\GPL Ghostscript"
@@ -71,13 +84,14 @@ def registry_path_ghostscript(env=None) -> Iterator[Path]:
with winreg.OpenKey( with winreg.OpenKey(
winreg.HKEY_LOCAL_MACHINE, fr"SOFTWARE\Artifex\GPL Ghostscript\{latest_gs}" winreg.HKEY_LOCAL_MACHINE, fr"SOFTWARE\Artifex\GPL Ghostscript\{latest_gs}"
) as k: ) as k:
_, gs_path, _ = next(registry_values(k)) for _, gs_path, _ in registry_values(k):
yield Path(gs_path) / 'bin' yield Path(gs_path) / 'bin'
except OSError as e: except OSError as e:
log.warning(e) log.warning(e)
def registry_path_tesseract(env=None) -> Iterator[Path]: def registry_path_tesseract(env=None) -> Iterator[Path]:
del env # unused (but needed for protocol)
try: try:
with winreg.OpenKey(winreg.HKEY_LOCAL_MACHINE, r"SOFTWARE\Tesseract-OCR") as k: with winreg.OpenKey(winreg.HKEY_LOCAL_MACHINE, r"SOFTWARE\Tesseract-OCR") as k:
for subkey, val, _valtype in registry_values(k): for subkey, val, _valtype in registry_values(k):
@@ -157,7 +171,7 @@ def unique_everseen(iterable: Iterable[T], key: Callable[[T], T]) -> Iterator[T]
"List unique elements, preserving order." "List unique elements, preserving order."
# unique_everseen('AAAABBBCCDAABBB') --> A B C D # unique_everseen('AAAABBBCCDAABBB') --> A B C D
# unique_everseen('ABBCcAD', str.lower) --> A B C D # unique_everseen('ABBCcAD', str.lower) --> A B C D
seen: Set[T] = set() seen: set[T] = set()
seen_add = seen.add seen_add = seen.add
for element in iterable: for element in iterable:
k = key(element) k = key(element)
+2
View File
@@ -4,4 +4,6 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
# Empty __init__.py file # Empty __init__.py file
@@ -0,0 +1 @@
Tesseract Open Source OCR Engine v4.1.1 with Leptonica
@@ -0,0 +1,13 @@
Portez ce vieux whisky au juge
blond qui fume sur son Ile
interieure, a cöte de l'alcöve
ovoide, oU les büches se
consument dans l'ätre, ce qui
lui permet de penser & la
caenogenese de |'etre dont il
est question dans la cause
ambigu& entendue a MoY, dans
un capharnaüm qui, pense-t-il,
diminue ca et la la qualite de son
ceuvre.
+1
View File
@@ -82,3 +82,4 @@
{"tesseract_version": "4.1.1", "system": "Linux", "python": "3.9.5", "argv_slug": "__-l__eng__--oem__1__000001_ocr.png__000001_ocr_tess__pdf__txt", "sourcefile": "resources/trivial.pdf", "args": ["-l", "eng", "--oem", "1", "-c", "textonly_pdf=1", "$TMPDIR/000001_ocr.png", "$TMPDIR/000001_ocr_tess", "pdf", "txt"]} {"tesseract_version": "4.1.1", "system": "Linux", "python": "3.9.5", "argv_slug": "__-l__eng__--oem__1__000001_ocr.png__000001_ocr_tess__pdf__txt", "sourcefile": "resources/trivial.pdf", "args": ["-l", "eng", "--oem", "1", "-c", "textonly_pdf=1", "$TMPDIR/000001_ocr.png", "$TMPDIR/000001_ocr_tess", "pdf", "txt"]}
{"tesseract_version": "5.0.0", "system": "Linux", "python": "3.9.5", "argv_slug": "__-l__eng__thresholding_method=1__000001_ocr.png__000001_ocr_tess__pdf__txt", "sourcefile": "resources/trivial.pdf", "args": ["-l", "eng", "-c", "textonly_pdf=1", "-c", "thresholding_method=1", "$TMPDIR/000001_ocr.png", "$TMPDIR/000001_ocr_tess", "pdf", "txt"]} {"tesseract_version": "5.0.0", "system": "Linux", "python": "3.9.5", "argv_slug": "__-l__eng__thresholding_method=1__000001_ocr.png__000001_ocr_tess__pdf__txt", "sourcefile": "resources/trivial.pdf", "args": ["-l", "eng", "-c", "textonly_pdf=1", "-c", "thresholding_method=1", "$TMPDIR/000001_ocr.png", "$TMPDIR/000001_ocr_tess", "pdf", "txt"]}
{"tesseract_version": "5.0.0", "system": "Linux", "python": "3.9.5", "argv_slug": "__-l__eng__thresholding_method=2__000001_ocr.png__000001_ocr_tess__pdf__txt", "sourcefile": "resources/trivial.pdf", "args": ["-l", "eng", "-c", "textonly_pdf=1", "-c", "thresholding_method=2", "$TMPDIR/000001_ocr.png", "$TMPDIR/000001_ocr_tess", "pdf", "txt"]} {"tesseract_version": "5.0.0", "system": "Linux", "python": "3.9.5", "argv_slug": "__-l__eng__thresholding_method=2__000001_ocr.png__000001_ocr_tess__pdf__txt", "sourcefile": "resources/trivial.pdf", "args": ["-l", "eng", "-c", "textonly_pdf=1", "-c", "thresholding_method=2", "$TMPDIR/000001_ocr.png", "$TMPDIR/000001_ocr_tess", "pdf", "txt"]}
{"tesseract_version": "4.1.1", "system": "Linux", "python": "3.10.4", "argv_slug": "__-l__deu__000001_ocr.png__000001_ocr_tess__pdf__txt", "sourcefile": "resources/francais.pdf", "args": ["-l", "deu", "-c", "textonly_pdf=1", "$TMPDIR/000001_ocr.png", "$TMPDIR/000001_ocr_tess", "pdf", "txt"]}
+3 -1
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import os import os
import platform import platform
import sys import sys
@@ -52,7 +54,7 @@ def resources() -> Path:
@pytest.fixture @pytest.fixture
def ocrmypdf_exec() -> List[str]: def ocrmypdf_exec() -> list[str]:
return [sys.executable, '-m', 'ocrmypdf'] return [sys.executable, '-m', 'ocrmypdf']
+2
View File
@@ -19,6 +19,8 @@
# TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE # TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE
# SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. # SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
from __future__ import annotations
from unittest.mock import patch from unittest.mock import patch
from ocrmypdf import hookimpl from ocrmypdf import hookimpl
+2
View File
@@ -19,6 +19,8 @@
# TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE # TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE
# SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. # SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
from __future__ import annotations
from unittest.mock import patch from unittest.mock import patch
from ocrmypdf import hookimpl from ocrmypdf import hookimpl
+2
View File
@@ -19,6 +19,8 @@
# TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE # TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE
# SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. # SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
from __future__ import annotations
from pathlib import Path from pathlib import Path
from subprocess import CalledProcessError from subprocess import CalledProcessError
from unittest.mock import patch from unittest.mock import patch
+2
View File
@@ -19,6 +19,8 @@
# TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE # TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE
# SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. # SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
from __future__ import annotations
from subprocess import CalledProcessError from subprocess import CalledProcessError
from unittest.mock import patch from unittest.mock import patch
+2
View File
@@ -26,6 +26,8 @@ that is not UTF-8 compatible, so we are forced to check that we can convert it
and present it to the user. and present it to the user.
""" """
from __future__ import annotations
from contextlib import contextmanager from contextlib import contextmanager
from subprocess import CalledProcessError from subprocess import CalledProcessError
from unittest.mock import patch from unittest.mock import patch
@@ -19,6 +19,8 @@
# TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE # TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE
# SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. # SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
from __future__ import annotations
from contextlib import contextmanager from contextlib import contextmanager
from subprocess import CalledProcessError from subprocess import CalledProcessError
from unittest.mock import patch from unittest.mock import patch
+20 -6
View File
@@ -44,17 +44,19 @@ Assumes Tesseract 4.0.0-alpha or higher.
""" """
from __future__ import annotations
import argparse import argparse
import json import json
import logging import logging
import platform import platform
import re import re
import shutil import shutil
import threading
from functools import partial from functools import partial
from pathlib import Path from pathlib import Path
from subprocess import PIPE, CalledProcessError, CompletedProcess from subprocess import PIPE, CalledProcessError, CompletedProcess
from unittest.mock import patch from unittest.mock import patch
import threading
from ocrmypdf import hookimpl from ocrmypdf import hookimpl
from ocrmypdf.builtin_plugins.tesseract_ocr import TesseractOcrEngine from ocrmypdf.builtin_plugins.tesseract_ocr import TesseractOcrEngine
@@ -177,28 +179,40 @@ def cached_run(options, run_args, **run_kwargs):
class CacheOcrEngine(TesseractOcrEngine): class CacheOcrEngine(TesseractOcrEngine):
# Concurrent threads (with --use-threads) might try to use different parts
# of the OcrEngine, so we need a lock to protect the state of patched
# module whenever it's patched. Should refactor ocrmypdf._exec.tesseract so that
# it does not to be patched at all for testing.
lock = threading.Lock() lock = threading.Lock()
@staticmethod @staticmethod
def get_orientation(input_file, options): def get_orientation(input_file, options):
with CacheOcrEngine.lock, patch('ocrmypdf._exec.tesseract.run', new=partial(cached_run, options)): with CacheOcrEngine.lock, patch(
return TesseractOcrEngine.get_orientation(input_file, options) 'ocrmypdf._exec.tesseract.run', new=partial(cached_run, options)
):
return TesseractOcrEngine.get_orientation(input_file, options)
@staticmethod @staticmethod
def get_deskew(input_file, options) -> float: def get_deskew(input_file, options) -> float:
with CacheOcrEngine.lock, patch('ocrmypdf._exec.tesseract.run', new=partial(cached_run, options)): with CacheOcrEngine.lock, patch(
'ocrmypdf._exec.tesseract.run', new=partial(cached_run, options)
):
return TesseractOcrEngine.get_deskew(input_file, options) return TesseractOcrEngine.get_deskew(input_file, options)
@staticmethod @staticmethod
def generate_hocr(input_file, output_hocr, output_text, options): def generate_hocr(input_file, output_hocr, output_text, options):
with CacheOcrEngine.lock, patch('ocrmypdf._exec.tesseract.run', new=partial(cached_run, options)): with CacheOcrEngine.lock, patch(
'ocrmypdf._exec.tesseract.run', new=partial(cached_run, options)
):
TesseractOcrEngine.generate_hocr( TesseractOcrEngine.generate_hocr(
input_file, output_hocr, output_text, options input_file, output_hocr, output_text, options
) )
@staticmethod @staticmethod
def generate_pdf(input_file, output_pdf, output_text, options): def generate_pdf(input_file, output_pdf, output_text, options):
with CacheOcrEngine.lock, patch('ocrmypdf._exec.tesseract.run', new=partial(cached_run, options)): with CacheOcrEngine.lock, patch(
'ocrmypdf._exec.tesseract.run', new=partial(cached_run, options)
):
TesseractOcrEngine.generate_pdf( TesseractOcrEngine.generate_pdf(
input_file, output_pdf, output_text, options input_file, output_pdf, output_text, options
) )
+2
View File
@@ -19,6 +19,8 @@
# TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE # TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE
# SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. # SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
from __future__ import annotations
import signal import signal
from contextlib import contextmanager from contextlib import contextmanager
from subprocess import CalledProcessError from subprocess import CalledProcessError
+2
View File
@@ -31,6 +31,8 @@ In 'pdf' mode, convert the image to PDF using another program.
In orientation check mode, report 0, 90, 180, 270... based on page number. In orientation check mode, report 0, 90, 180, 270... based on page number.
""" """
from __future__ import annotations
import pikepdf import pikepdf
from PIL import Image from PIL import Image
+2
View File
@@ -30,6 +30,8 @@ In 'pdf' mode, convert the image to PDF using another program.
In orientation check mode, report the orientation is upright. In orientation check mode, report the orientation is upright.
""" """
from __future__ import annotations
import pikepdf import pikepdf
from PIL import Image from PIL import Image
@@ -19,8 +19,6 @@
# TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE # TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE
# SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. # SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
# type: ignore
"""Tesseract no-op plugin that simulates the OOM killer on page 4. """Tesseract no-op plugin that simulates the OOM killer on page 4.
OCRmyPDF can use a lot of memory, even that it might trigger the OCRmyPDF can use a lot of memory, even that it might trigger the
@@ -30,14 +28,18 @@ ensure we fail with an error rather than deadlock in such cases.
Page 4 was chosen because of this number's association with bad luck Page 4 was chosen because of this number's association with bad luck
in many East Asian cultures. in many East Asian cultures.
""" """
# type: ignore
from __future__ import annotations
import os import os
import signal import signal
import sys
from pathlib import Path from pathlib import Path
from ocrmypdf import hookimpl from ocrmypdf import hookimpl
# type: ignore
# Ugly hack that let us use the NoopOcrEngine without setting up packaging for our # Ugly hack that let us use the NoopOcrEngine without setting up packaging for our
# tests. # tests.
# This hack also requires us to set type: ignore # This hack also requires us to set type: ignore
@@ -47,7 +49,7 @@ exec(parent)
NoopOcrEngine = locals()['NoopOcrEngine'] NoopOcrEngine = locals()['NoopOcrEngine']
class Page4Engine(NoopOcrEngine): class Page4Engine(NoopOcrEngine): # type: ignore
def __str__(self): def __str__(self):
return f"NO-OP Page 4 {NoopOcrEngine.version()}" return f"NO-OP Page 4 {NoopOcrEngine.version()}"
+2
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import logging import logging
import pytest import pytest
+2
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import logging import logging
from io import BytesIO, StringIO from io import BytesIO, StringIO
+2
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import pytest import pytest
from ocrmypdf.helpers import check_pdf from ocrmypdf.helpers import check_pdf
+2
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import os import os
from subprocess import PIPE, run from subprocess import PIPE, run
+2
View File
@@ -4,6 +4,8 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import os import os
import pytest import pytest
+2
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import logging import logging
import subprocess import subprocess
from decimal import Decimal from decimal import Decimal
+2
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
from unittest.mock import patch from unittest.mock import patch
import pikepdf import pikepdf
+2
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import logging import logging
import multiprocessing import multiprocessing
import os import os
+2
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import re import re
from io import StringIO from io import StringIO
+2
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
from unittest.mock import patch from unittest.mock import patch
import img2pdf import img2pdf
+2
View File
@@ -4,6 +4,8 @@
# License, v. 2.0. If a copy of the MPL was not distributed with this # License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import logging import logging
import pytest import pytest
+52 -11
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import os import os
import shutil import shutil
from math import isclose from math import isclose
@@ -465,12 +467,18 @@ def test_overlay(resources, outpdf):
) )
def test_destination_not_writable(resources, outdir): @pytest.fixture
if os.name != 'nt' and (os.getuid() == 0 or os.geteuid() == 0): def protected_file(outdir):
pytest.xfail(reason="root can write to anything")
protected_file = outdir / 'protected.pdf' protected_file = outdir / 'protected.pdf'
protected_file.touch() protected_file.touch()
protected_file.chmod(0o400) # Read-only protected_file.chmod(0o400) # Read-only
yield protected_file
@pytest.mark.skipif(
os.name == 'nt' or os.geteuid() == 0, reason="root can write to anything"
)
def test_destination_not_writable(resources, protected_file):
p = run_ocrmypdf( p = run_ocrmypdf(
resources / 'jbig2.pdf', resources / 'jbig2.pdf',
protected_file, protected_file,
@@ -480,7 +488,8 @@ def test_destination_not_writable(resources, outdir):
assert p.returncode == ExitCode.file_access_error, "Expected error" assert p.returncode == ExitCode.file_access_error, "Expected error"
def test_tesseract_config_valid(resources, outdir): @pytest.fixture
def valid_tess_config(outdir):
cfg_file = outdir / 'test.cfg' cfg_file = outdir / 'test.cfg'
with cfg_file.open('w') as f: with cfg_file.open('w') as f:
f.write( f.write(
@@ -490,20 +499,22 @@ language_model_penalty_non_dict_word 0
language_model_penalty_non_freq_dict_word 0 language_model_penalty_non_freq_dict_word 0
''' '''
) )
yield cfg_file
def test_tesseract_config_valid(resources, valid_tess_config, outpdf):
check_ocrmypdf( check_ocrmypdf(
resources / '3small.pdf', resources / '3small.pdf',
outdir / 'out.pdf', outpdf,
'--tesseract-config', '--tesseract-config',
cfg_file, valid_tess_config,
'--pages', '--pages',
'1', '1',
) )
@pytest.mark.slow # This test sometimes times out in CI @pytest.fixture
@pytest.mark.parametrize('renderer', RENDERERS) def invalid_tess_config(outdir):
def test_tesseract_config_invalid(renderer, resources, outdir):
cfg_file = outdir / 'test.cfg' cfg_file = outdir / 'test.cfg'
with cfg_file.open('w') as f: with cfg_file.open('w') as f:
f.write( f.write(
@@ -511,14 +522,19 @@ def test_tesseract_config_invalid(renderer, resources, outdir):
THIS FILE IS INVALID THIS FILE IS INVALID
''' '''
) )
yield cfg_file
@pytest.mark.slow # This test sometimes times out in CI
@pytest.mark.parametrize('renderer', RENDERERS)
def test_tesseract_config_invalid(renderer, resources, invalid_tess_config, outpdf):
p = run_ocrmypdf( p = run_ocrmypdf(
resources / 'ccitt.pdf', resources / 'ccitt.pdf',
outdir / 'out.pdf', outpdf,
'--pdf-renderer', '--pdf-renderer',
renderer, renderer,
'--tesseract-config', '--tesseract-config',
cfg_file, invalid_tess_config,
) )
assert ( assert (
"parameter not found" in p.stderr.lower() "parameter not found" in p.stderr.lower()
@@ -801,6 +817,9 @@ def test_text_curves(resources, outpdf):
info = PdfInfo(outpdf) info = PdfInfo(outpdf)
assert len(info.pages[0].images) == 0, "added images to the vector PDF" assert len(info.pages[0].images) == 0, "added images to the vector PDF"
def test_text_curves_force(resources, outpdf):
with patch('ocrmypdf._pipeline.VECTOR_PAGE_DPI', 100):
check_ocrmypdf( check_ocrmypdf(
resources / 'vector.pdf', resources / 'vector.pdf',
outpdf, outpdf,
@@ -922,3 +941,25 @@ def test_outputtype_none(resources, outtxt):
'tests/plugins/tesseract_noop.py', 'tests/plugins/tesseract_noop.py',
) )
assert p.returncode == ExitCode.ok assert p.returncode == ExitCode.ok
@pytest.fixture
def graph_bad_icc(resources, outdir):
synth_input_file = outdir / 'graph-bad-icc.pdf'
with pikepdf.open(resources / 'graph.pdf') as pdf:
icc = pdf.make_stream(
b'invalid icc profile', N=3, Alternate=pikepdf.Name.DeviceRGB
)
pdf.pages[0].Resources.XObject['/Im0'].ColorSpace = pikepdf.Array(
[pikepdf.Name.ICCBased, icc]
)
pdf.save(synth_input_file)
yield synth_input_file
def test_corrupt_icc(graph_bad_icc, outpdf, caplog):
result = run_ocrmypdf_api(graph_bad_icc, outpdf)
assert result == ExitCode.ok
assert any(
'corrupt or unreadable ICC profile' in rec.message for rec in caplog.records
)
+28 -14
View File
@@ -5,11 +5,12 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
import datetime import datetime
import warnings
from datetime import timezone from datetime import timezone
from os import fspath
from shutil import copyfile from shutil import copyfile
from unittest.mock import patch
import pikepdf import pikepdf
import pytest import pytest
@@ -17,7 +18,7 @@ from pikepdf.models.metadata import decode_pdf_date
from ocrmypdf._jobcontext import PdfContext from ocrmypdf._jobcontext import PdfContext
from ocrmypdf._pipeline import convert_to_pdfa, metadata_fixup from ocrmypdf._pipeline import convert_to_pdfa, metadata_fixup
from ocrmypdf._plugin_manager import get_plugin_manager from ocrmypdf._plugin_manager import get_parser_options_plugins, get_plugin_manager
from ocrmypdf.cli import get_parser from ocrmypdf.cli import get_parser
from ocrmypdf.exceptions import ExitCode from ocrmypdf.exceptions import ExitCode
from ocrmypdf.pdfa import file_claims_pdfa, generate_pdfa_ps from ocrmypdf.pdfa import file_claims_pdfa, generate_pdfa_ps
@@ -173,6 +174,19 @@ def test_creation_date_preserved(output_type, resources, infile, outpdf):
assert seconds_between_dates(date_after, datetime.datetime.now(timezone.utc)) < 1000 assert seconds_between_dates(date_after, datetime.datetime.now(timezone.utc)) < 1000
@pytest.fixture
def libxmp_file_to_dict():
try:
with warnings.catch_warnings():
warnings.simplefilter("ignore", DeprecationWarning)
from libxmp.utils import (
file_to_dict, # pylint: disable=import-outside-toplevel
)
except Exception: # pylint: disable=broad-except
pytest.skip("libxmp not available or libexempi3 not installed")
return file_to_dict
@pytest.mark.parametrize( @pytest.mark.parametrize(
'test_file,output_type', 'test_file,output_type',
[ [
@@ -182,15 +196,12 @@ def test_creation_date_preserved(output_type, resources, infile, outpdf):
('3small.pdf', 'pdfa'), ('3small.pdf', 'pdfa'),
], ],
) )
def test_xml_metadata_preserved(test_file, output_type, resources, outpdf): def test_xml_metadata_preserved(
libxmp_file_to_dict, test_file, output_type, resources, outpdf
):
input_file = resources / test_file input_file = resources / test_file
try: before = libxmp_file_to_dict(str(input_file))
from libxmp.utils import file_to_dict # pylint: disable=import-outside-toplevel
except Exception: # pylint: disable=broad-except
pytest.skip(reason="libxmp not available or libexempi3 not installed")
before = file_to_dict(str(input_file))
check_ocrmypdf( check_ocrmypdf(
input_file, input_file,
@@ -202,7 +213,7 @@ def test_xml_metadata_preserved(test_file, output_type, resources, outpdf):
'tests/plugins/tesseract_noop.py', 'tests/plugins/tesseract_noop.py',
) )
after = file_to_dict(str(outpdf)) after = libxmp_file_to_dict(str(outpdf))
equal_properties = [ equal_properties = [
'dc:contributor', 'dc:contributor',
@@ -290,8 +301,8 @@ def test_kodak_toc(resources, outpdf):
def test_metadata_fixup_warning(resources, outdir, caplog): def test_metadata_fixup_warning(resources, outdir, caplog):
options = get_parser().parse_args( _parser, options, _pm = get_parser_options_plugins(
args=['--output-type', 'pdfa-2', 'graph.pdf', 'out.pdf'] ['--output-type', 'pdfa-2', 'graph.pdf', 'out.pdf']
) )
copyfile(resources / 'graph.pdf', outdir / 'graph.pdf') copyfile(resources / 'graph.pdf', outdir / 'graph.pdf')
@@ -316,6 +327,9 @@ def test_metadata_fixup_warning(resources, outdir, caplog):
assert any(record.levelname == 'WARNING' for record in caplog.records) assert any(record.levelname == 'WARNING' for record in caplog.records)
XMP_MAGIC = b'W5M0MpCehiHzreSzNTczkc9d'
def test_prevent_gs_invalid_xml(resources, outdir): def test_prevent_gs_invalid_xml(resources, outdir):
generate_pdfa_ps(outdir / 'pdfa.ps') generate_pdfa_ps(outdir / 'pdfa.ps')
copyfile(resources / 'trivial.pdf', outdir / 'layers.rendered.pdf') copyfile(resources / 'trivial.pdf', outdir / 'layers.rendered.pdf')
@@ -342,7 +356,7 @@ def test_prevent_gs_invalid_xml(resources, outdir):
contents = (outdir / 'pdfa.pdf').read_bytes() contents = (outdir / 'pdfa.pdf').read_bytes()
# Since the XML may be invalid, we scan instead of actually feeding it # Since the XML may be invalid, we scan instead of actually feeding it
# to a parser. # to a parser.
XMP_MAGIC = b'W5M0MpCehiHzreSzNTczkc9d'
xmp_start = contents.find(XMP_MAGIC) xmp_start = contents.find(XMP_MAGIC)
xmp_end = contents.rfind(b'<?xpacket end', xmp_start) xmp_end = contents.rfind(b'<?xpacket end', xmp_start)
assert 0 < xmp_start < xmp_end assert 0 < xmp_start < xmp_end
+2
View File
@@ -5,6 +5,8 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from __future__ import annotations
from os import fspath from os import fspath
from pathlib import Path from pathlib import Path
from unittest.mock import patch from unittest.mock import patch

Some files were not shown because too many files have changed in this diff Show More