Compare commits

...
27 Commits
Author SHA1 Message Date
James R. Barlow 0e013df161 v16.2.0 release notes 2024-04-16 00:37:03 -07:00
James R. Barlow 9ba4e3ab46 Log unusual exceptions when trying to obtain a version
Fixes #1262
2024-04-07 14:39:08 -07:00
James R. Barlow 5fdcb7602b Make downsampling large images that Tesseract would otherwise error on into default behavior
Fixes #1281
2024-04-07 13:43:20 -07:00
James R. Barlow b4db1b741f optimize: fix handling of [/FlateDecode none] - type images
Closes #1271
2024-04-07 01:44:08 -07:00
James R. Barlow 7a8cc21e31 Add support for sidecar output to io.BytesIO
Closes #1252
2024-04-07 01:38:55 -07:00
James R. Barlow 0674829d8f Remove tool.black config 2024-04-07 00:36:52 -07:00
James R. Barlow 315aa0474b Merge branch 'main' of github.com:ocrmypdf/OCRmyPDF 2024-04-07 00:34:51 -07:00
Ben BeasleyandGitHub df3451e779 Update the typer[all] dependency to typer-slim[standard] (#1287)
In 0.12.1, Typer was significantly reorganized.

- `typer-slim` is the library (for `import typer`)
- `typer-slim[standard]` adds optional dependencies (currently `rich`
  and `shellingham`, basically equivalent to the old `typer[all]`)
- `typer-cli` is the `typer` command-line tool
- `typer` is now basically a metapackage that brings in *all of the
  above*, and it no longer has an `all` extra

Pip will warn about this and proceed,

```
WARNING: typer 0.12.1 does not provide the extra 'all'
```

but there are other tools that will fail hard when asked to resolve a
(now) nonexistent extra.

Since this project doesn’t need the `typer` command-line tool, it looks
like changing the dependency to `typer-slim[standard]` is the best way
forward.

See https://typer.tiangolo.com/release-notes/#0121 and
tiangolo/typer#785 for further discussion
and details.
2024-04-07 00:34:34 -07:00
akierigandGitHub 3ba42802d1 added Macports install information (#1286) 2024-04-07 00:33:57 -07:00
James R. Barlow d6342cb8c2 Add heif/heic input image support 2024-04-07 00:33:13 -07:00
James R. Barlow 065bddbc6c Reformat with ruff format 2024-04-07 00:25:32 -07:00
James R. Barlow 067f429dde Merge branch 'main' of github.com:ocrmypdf/OCRmyPDF 2024-03-26 15:34:00 -07:00
Daniel LovegroveandGitHub 6895c2d70f Fix Broken Documentation Links (#1275)
* Update URL for PDFMARK documentation

For reference, here is a link to the old PDF:
https://web.archive.org/web/20190806035303/https://www.adobe.com/content/dam/acom/en/devnet/acrobat/pdfs/pdfmark_reference.pdf

It appears Adobe converted the PDF into a webpage-based document, the
wording seems to almost identical b/w the PDF and the website.

* Fix cross-references to JBIG2 page

* Fix links for Fedora + Arch + HEAD revision install

Fedora 39 has been released, and the package tracker no longer includes
a release overview for Fedora 37 hence why it was removed here.
2024-03-22 14:38:52 -07:00
James R. Barlow 686481982a Fix naming of hOCR rendered files 2024-03-22 13:27:20 -07:00
James R. Barlow a9e1d19b78 v16.1.2 release notes 2024-03-20 12:56:13 -07:00
James R. Barlow f95aa63718 Merge branch 'main' of github.com:ocrmypdf/OCRmyPDF 2024-03-20 12:26:02 -07:00
James Barlow 855de287b2 Fix test suite failure with Ghostscript >= 10.3
Ghostscript is more picky about a specific case with SMask that cannot be converted to PDF/A

Details here
https://github.com/ArtifexSoftware/ghostpdl/commit/4dcfae36bb4dcbc4ef3b5e5afc98bcde0d6b9ddc
2024-03-19 17:20:33 -07:00
NilsRoandGitHub feeb9f213f batch example: added archive, small corrections and optimizations (#1277)
* Added archive, small corrections

Added a function to archive originals and avoid calling ocrmypdf if they are still is PDF/A.

* Added Copyright
2024-03-18 13:22:24 -07:00
Emiel MolenaarandGitHub e7eb8fa805 Update Dockerfile.alpine (#1268)
Use Alpine 3.19 as base image to ensure we get GhostScript 10.2.1 to eliminate serious regressions that corrupt PDFs with existing text.
2024-03-13 14:49:42 -07:00
James R. Barlow 8a747f005a pixels -> megapixels
Fixes #1265
2024-02-29 15:31:07 -08:00
James R. Barlow 16ab4a8b4e Fix error message about missing Python exec
Message is
unable to start container process: exec: "python": executable file not found in $PATH: unknown.

Closes #1260
2024-02-21 23:54:41 -08:00
James R. Barlow 8d30cff4ef Undo future annotations from watcher.py till Typer fixes its issue
Fixes #1258
2024-02-20 19:14:39 -08:00
James R. Barlow 59d5b0d1bd v16.1.1 release notes 2024-02-15 16:56:25 -08:00
James R. Barlow 9ec0745ab8 Try pypy3.10 2024-02-14 14:25:13 -08:00
James R. Barlow 3a3635f7f9 Python 3.10 cleanup, manual fixes 2024-02-14 12:48:17 -08:00
James R. Barlow 6a746a1cbb ruff linting/Python 3.10 cleanup 2024-02-14 12:41:51 -08:00
James R. Barlow 906c130f96 Update rust toml settings 2024-02-14 12:32:26 -08:00
49 changed files with 297 additions and 183 deletions
+1
View File
@@ -64,6 +64,7 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
img2pdf \ img2pdf \
libsm6 libxext6 libxrender-dev \ libsm6 libxext6 libxrender-dev \
pngquant \ pngquant \
python-is-python3 \
tesseract-ocr \ tesseract-ocr \
tesseract-ocr-chi-sim \ tesseract-ocr-chi-sim \
tesseract-ocr-deu \ tesseract-ocr-deu \
+1 -1
View File
@@ -1,7 +1,7 @@
# SPDX-FileCopyrightText: 2023 James R. Barlow # SPDX-FileCopyrightText: 2023 James R. Barlow
# SPDX-License-Identifier: MPL-2.0 # SPDX-License-Identifier: MPL-2.0
FROM alpine:3.18 as base FROM alpine:3.19 as base
ENV LANG=C.UTF-8 ENV LANG=C.UTF-8
ENV TZ=UTC ENV TZ=UTC
+2 -2
View File
@@ -32,8 +32,8 @@ jobs:
- os: ubuntu-latest - os: ubuntu-latest
python: "3.12" python: "3.12"
tesseract5: true tesseract5: true
# - os: ubuntu-latest - os: ubuntu-latest
# python: "pypy3.10" python: "pypy3.10"
env: env:
OS: ${{ matrix.os }} OS: ${{ matrix.os }}
+1
View File
@@ -70,6 +70,7 @@ Linux, Windows, macOS and FreeBSD are supported. Docker images are also availabl
| Windows Subsystem for Linux | ``apt install ocrmypdf`` | | Windows Subsystem for Linux | ``apt install ocrmypdf`` |
| Fedora | ``dnf install ocrmypdf`` | | Fedora | ``dnf install ocrmypdf`` |
| macOS (Homebrew) | ``brew install ocrmypdf`` | | macOS (Homebrew) | ``brew install ocrmypdf`` |
| macOS (MacPorts) | ``port install ocrmypdf`` |
| macOS (nix) | ``nix-env -i ocrmypdf`` | | macOS (nix) | ``nix-env -i ocrmypdf`` |
| LinuxBrew | ``brew install ocrmypdf`` | | LinuxBrew | ``brew install ocrmypdf`` |
| FreeBSD | ``pkg install py-ocrmypdf`` | | FreeBSD | ``pkg install py-ocrmypdf`` |
+7 -6
View File
@@ -118,15 +118,16 @@ OCR for huge images
------------------- -------------------
Tesseract has internal limits on the size Tesseract has internal limits on the size
of images it will process. If you issue of images it will process. By default,
``--tesseract-downsample-large-images``, OCRmyPDF will downsample images ``--tesseract-downsample-large-images`` is enabled, and OCRmyPDF will
to fit Tesseract limits. (The limits are usually entered only for scanned downsample images to fit Tesseract limits. (The limits are usually encountered
images of oversized media, such as large maps or blueprints exceeding only for scanned images of oversized media, such as large maps or blueprints exceeding
110 cm or 43 inches in either dimension, and at high DPI.) 110 cm or 43 inches in either dimension, and at high DPI.) This feature can disabled
using ``--no-tesseract-downsample-large-images``.
``--tesseract-downsample-above Npixels`` adjusts the threshold at which images ``--tesseract-downsample-above Npixels`` adjusts the threshold at which images
will be downsampled. By default, only images that exceed any of Tesseract's will be downsampled. By default, only images that exceed any of Tesseract's
internal limits are downsampled. internal limits are downsampled (32767 pixels on either dimension).
You will also need to set ``--tesseract-timeout`` high enough to allow You will also need to set ``--tesseract-timeout`` high enough to allow
for processing. for processing.
+29 -11
View File
@@ -23,7 +23,9 @@ These platforms have one-liner installs:
+-------------------------------+-----------------------------------------+ +-------------------------------+-----------------------------------------+
| Fedora | ``dnf install ocrmypdf tesseract-osd`` | | Fedora | ``dnf install ocrmypdf tesseract-osd`` |
+-------------------------------+-----------------------------------------+ +-------------------------------+-----------------------------------------+
| macOS | ``brew install ocrmypdf`` | | macOS (Homebrew) | ``brew install ocrmypdf`` |
+-------------------------------+-----------------------------------------+
| macOS (MacPorts) | ``port install ocrmypdf`` |
+-------------------------------+-----------------------------------------+ +-------------------------------+-----------------------------------------+
| LinuxBrew | ``brew install ocrmypdf`` | | LinuxBrew | ``brew install ocrmypdf`` |
+-------------------------------+-----------------------------------------+ +-------------------------------+-----------------------------------------+
@@ -99,12 +101,12 @@ For full details on version availability for your platform, check the
Fedora Fedora
------ ------
.. |fedora-37| image:: https://repology.org/badge/version-for-repo/fedora_37/ocrmypdf.svg
:alt: Fedora 37
.. |fedora-38| image:: https://repology.org/badge/version-for-repo/fedora_38/ocrmypdf.svg .. |fedora-38| image:: https://repology.org/badge/version-for-repo/fedora_38/ocrmypdf.svg
:alt: Fedora 38 :alt: Fedora 38
.. |fedora-39| image:: https://repology.org/badge/version-for-repo/fedora_39/ocrmypdf.svg
:alt: Fedora 39
.. |fedora-rawhide| image:: https://repology.org/badge/version-for-repo/fedora_rawhide/ocrmypdf.svg .. |fedora-rawhide| image:: https://repology.org/badge/version-for-repo/fedora_rawhide/ocrmypdf.svg
:alt: Fedore Rawhide :alt: Fedore Rawhide
@@ -113,7 +115,7 @@ Fedora
+-----------------------------------------------+ +-----------------------------------------------+
| |latest| | | |latest| |
+-----------------------------------------------+ +-----------------------------------------------+
| |fedora-37| |fedora-38| |fedora-rawhide| | | |fedora-38| |fedora-39| |fedora-rawhide| |
+-----------------------------------------------+ +-----------------------------------------------+
Users of Fedora may simply Users of Fedora may simply
@@ -123,7 +125,7 @@ Users of Fedora may simply
dnf install ocrmypdf tesseract-osd dnf install ocrmypdf tesseract-osd
For full details on version availability, check the `Fedora Package For full details on version availability, check the `Fedora Package
Tracker <https://apps.fedoraproject.org/packages/ocrmypdf>`__. Tracker <https://packages.fedoraproject.org/pkgs/ocrmypdf/ocrmypdf/>`__.
If the version available for your platform is out of date, you could opt If the version available for your platform is out of date, you could opt
to install the latest version from source. See `Installing HEAD revision to install the latest version from source. See `Installing HEAD revision
@@ -135,7 +137,7 @@ from sources <#installing-head-revision-from-sources>`__.
issues. OCRmyPDF works fine without it but will produce larger output issues. OCRmyPDF works fine without it but will produce larger output
files. If you build jbig2enc from source, ocrmypdf 7.0.0 and later files. If you build jbig2enc from source, ocrmypdf 7.0.0 and later
will automatically detect it on the ``PATH``. To add JBIG2 encoding, will automatically detect it on the ``PATH``. To add JBIG2 encoding,
see `Installing the JBIG2 encoder <jbig2>`__. see :ref:`Installing the JBIG2 encoder <jbig2>`.
.. _ubuntu-lts-latest: .. _ubuntu-lts-latest:
@@ -160,7 +162,7 @@ and build ocrmypdf in virtual environment:
python3.11 -m venv .venv python3.11 -m venv .venv
To add JBIG2 encoding, see `Installing the JBIG2 encoder <jbig2>`__. To add JBIG2 encoding, see :ref:`Installing the JBIG2 encoder <jbig2>`.
Note Fedora packages for language data haven't been branched for RHEL/EPEL, but you can get traineddata files directly from `tesseract Note Fedora packages for language data haven't been branched for RHEL/EPEL, but you can get traineddata files directly from `tesseract
<https://github.com/tesseract-ocr/tessdata/>`__ and place them in ``/usr/share/tesseract/tessdata``. <https://github.com/tesseract-ocr/tessdata/>`__ and place them in ``/usr/share/tesseract/tessdata``.
@@ -217,7 +219,7 @@ you are using a VM image, such as `the official Vagrant image
be completed for you. be completed for you.
Next you should install the `base-devel package group Next you should install the `base-devel package group
<https://www.archlinux.org/groups/x86_64/base-devel/>`__. This includes the <https://archlinux.org/packages/core/any/base-devel/>`__. This includes the
standard tooling needed to build packages, such as a compiler and binary tools. standard tooling needed to build packages, such as a compiler and binary tools.
.. code-block:: bash .. code-block:: bash
@@ -260,7 +262,7 @@ page.
<https://aur.archlinux.org/packages/jbig2enc-git/>`__ and may be installed <https://aur.archlinux.org/packages/jbig2enc-git/>`__ and may be installed
using the same series of steps as for the installation OCRmyPDF AUR using the same series of steps as for the installation OCRmyPDF AUR
package. Alternatively, it may be built manually from source following the package. Alternatively, it may be built manually from source following the
instructions in `Installing the JBIG2 encoder <jbig2>`__. If JBIG2 is instructions in :ref:`Installing the JBIG2 encoder <jbig2>`. If JBIG2 is
installed, OCRmyPDF 7.0.0 and later will automatically detect it. installed, OCRmyPDF 7.0.0 and later will automatically detect it.
Alpine Linux Alpine Linux
@@ -325,6 +327,22 @@ languages you can optionally install them all:
brew install tesseract-lang # Optional: Install all language packs brew install tesseract-lang # Optional: Install all language packs
MacPorts
--------
.. image:: https://img.shields.io/badge/dynamic/json?url=https%3A%2F%2Fports.macports.org%2Fapi%2Fv1%2Fports%2Focrmypdf%2F%3Fformat%3Djson&query=version&label=MacPorts
:alt: Macports Version Information
:target: https://ports.macports.org/port/ocrmypdf
OCRmyPDF is includes in MacPorts:
.. code-block:: bash
sudo port install ocrmypdf
Note that while this will install tesseract you will need to install
the appropriate tesseract `language ports <https://ports.macports.org/search/?selected_facets=categories_exact%3Atextproc&installed_file=&q=tesseract&name=on>`__.
Manual installation on macOS Manual installation on macOS
---------------------------- ----------------------------
@@ -623,7 +641,7 @@ environment:
pip install git+https://github.com/ocrmypdf/OCRmyPDF.git pip install git+https://github.com/ocrmypdf/OCRmyPDF.git
Or, to install in `development Or, to install in `development
mode <https://pythonhosted.org/setuptools/setuptools.html#development-mode>`__, mode <https://packaging.python.org/en/latest/guides/distributing-packages-using-setuptools/#working-in-development-mode>`__,
allowing customization of OCRmyPDF, use the ``-e`` flag: allowing customization of OCRmyPDF, use the ``-e`` flag:
.. code-block:: bash .. code-block:: bash
+9
View File
@@ -65,3 +65,12 @@ installation documentation.
If you maintain a Linux distribution that supports 32-bit x86 or ARM, OCRmyPDF If you maintain a Linux distribution that supports 32-bit x86 or ARM, OCRmyPDF
should continue to work as long as all of its dependencies continue to be should continue to work as long as all of its dependencies continue to be
available in 32-bit form. Please note we do not test on 32-bit platforms. available in 32-bit form. Please note we do not test on 32-bit platforms.
HEIF/HEIC
---------
OCRmyPDF defaults to installing the pi-heif PyPI package, which supports converting
HEIF (High Efficiency Image File Format) images to PDF from the command line.
If your distribution does not have this library available, you can exclude it and
OCRmyPDF will gracefully degrade automatically, losing only support for this
feature.
+30
View File
@@ -30,6 +30,36 @@ OCRmyPDF typically supports the three most recent Python versions.
.. |OCRmyPDF PyPI| image:: https://img.shields.io/pypi/v/ocrmypdf.svg .. |OCRmyPDF PyPI| image:: https://img.shields.io/pypi/v/ocrmypdf.svg
v16.2.0
=======
- Fixed issue 'NoneType' object has no attribute 'get' when optimizing certain PDFs.
:issue:`1293,1271`
- Switched formatting from black to ruff.
- Added support for sending sidecar output to io.BytesIO.
- Added support for converting HEIF/HEIC images (the native image of iPhones and
some other devices) to PDFs, when the appropriate pi-hief library is installed.
This library is marked as a dependency, but maintainers may opt out if needed.
- We now default to downsampling large images that would exceed Tesseract's internal
limits, but only if it cause processing to fail. Previously, this behavior only
occurred if specifically requested on command line. It can still be configured
and disabled. See the --tesseract command line options.
- Added Macports install instructions. Thanks @akierig.
- Improved logging output when an unexpected error occurs while trying to obtain
the version of a third party program.
v16.1.2
=======
- Fixed test suite failure when using Ghostscript 10.3.
- Other minor corrections.
v16.1.1
=======
- Fixed PyPy 3.10 support.
v16.1.0 v16.1.0
======= =======
+48 -10
View File
@@ -1,5 +1,6 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
# SPDX-FileCopyrightText: 2016 findingorder <https://github.com/findingorder> # SPDX-FileCopyrightText: 2016 findingorder <https://github.com/findingorder>
# SPDX-FileCopyrightText: 2024 nilsro <https://github.com/nilsro>
# SPDX-License-Identifier: MIT # SPDX-License-Identifier: MIT
"""Example of using ocrmypdf as a library in a script. """Example of using ocrmypdf as a library in a script.
@@ -15,6 +16,10 @@ from __future__ import annotations
import logging import logging
import sys import sys
import os
import posixpath
import shutil
import filecmp
from pathlib import Path from pathlib import Path
import ocrmypdf import ocrmypdf
@@ -22,32 +27,65 @@ import ocrmypdf
# pylint: disable=logging-format-interpolation # pylint: disable=logging-format-interpolation
# pylint: disable=logging-not-lazy # pylint: disable=logging-not-lazy
def filecompare(a, b):
try:
return filecmp.cmp(a, b, shallow=True)
except FileNotFoundError:
return False
script_dir = Path(__file__).parent script_dir = Path(__file__).parent
# set archive_dir to a path for backup original documents. Leave empty if not required.
archive_dir = "/pdfbak"
if len(sys.argv) > 1: if len(sys.argv) > 1:
start_dir = Path(sys.argv[1]) start_dir = Path(sys.argv[1])
else: else:
start_dir = Path('.') start_dir = Path(".")
if len(sys.argv) > 2: if len(sys.argv) > 2:
log_file = Path(sys.argv[2]) log_file = Path(sys.argv[2])
else: else:
log_file = script_dir.with_name('ocr-tree.log') log_file = script_dir.with_name("ocr-tree.log")
logging.basicConfig( logging.basicConfig(
level=logging.INFO, level=logging.INFO,
format='%(asctime)s %(message)s', format="%(asctime)s %(message)s",
filename=log_file, filename=log_file,
filemode='a', filemode="a",
) )
logging.info(f"Start directory {start_dir}")
ocrmypdf.configure_logging(ocrmypdf.Verbosity.default) ocrmypdf.configure_logging(ocrmypdf.Verbosity.default)
for filename in start_dir.glob("**/*.py"): for filename in start_dir.glob("**/*.pdf"):
logging.info(f"Processing {filename}") logging.info(f"Processing {filename}")
result = ocrmypdf.ocr(filename, filename, deskew=True) if ocrmypdf.pdfa.file_claims_pdfa(filename)["pass"]:
if result == ocrmypdf.ExitCode.already_done_ocr: logging.info("Skipped document because it already contained text")
logging.error("Skipped document because it already contained text") else:
elif result == ocrmypdf.ExitCode.ok: archive_filename = archive_dir + str(filename)
if len(archive_dir) > 0 and not filecompare(filename, archive_filename):
logging.info(f"Archiving document to {archive_filename}")
try:
shutil.copy2(filename, posixpath.dirname(archive_filename))
except IOError as io_err:
os.makedirs(posixpath.dirname(archive_filename))
shutil.copy2(filename, posixpath.dirname(archive_filename))
try:
result = ocrmypdf.ocr(filename, filename, deskew=True)
logging.info(result)
except ocrmypdf.exceptions.EncryptedPdfError:
logging.info("Skipped document because it is encrypted")
except ocrmypdf.exceptions.PriorOcrFoundError:
logging.info("Skipped document because it already contained text")
except ocrmypdf.exceptions.DigitalSignatureError:
logging.info("Skipped document because it has a digital signature")
except ocrmypdf.exceptions.TaggedPDFError:
logging.info(
"Skipped document because it does not need ocr as it is tagged"
)
except:
logging.error("Unhandled error occured")
logging.info("OCR complete") logging.info("OCR complete")
logging.info(result)
+4 -3
View File
@@ -53,9 +53,10 @@ for dir_name, _subdirs, file_list in os.walk(start_dir):
] ]
logging.info(cmd) logging.info(cmd)
full_path_ocr = os.path.join(dir_name, filename_ocr) full_path_ocr = os.path.join(dir_name, filename_ocr)
with open(filename, 'rb') as input_file, open( with (
full_path_ocr, 'wb' open(filename, 'rb') as input_file,
) as output_file: open(full_path_ocr, 'wb') as output_file,
):
proc = subprocess.run( proc = subprocess.run(
cmd, cmd,
stdin=input_file, stdin=input_file,
+2 -3
View File
@@ -7,7 +7,6 @@
# Do not enable annotations! # Do not enable annotations!
# https://github.com/tiangolo/typer/discussions/598 # https://github.com/tiangolo/typer/discussions/598
# from __future__ import annotations
import json import json
import logging import logging
@@ -131,7 +130,7 @@ def execute_ocrmypdf(
class HandleObserverEvent(PatternMatchingEventHandler): class HandleObserverEvent(PatternMatchingEventHandler):
def __init__( def __init__( # noqa: D107
self, self,
patterns=None, patterns=None,
ignore_patterns=None, ignore_patterns=None,
@@ -191,7 +190,7 @@ def main(
bool, bool,
typer.Option( typer.Option(
envvar='OCR_OUTPUT_DIRECTORY_YEAR_MONTH', envvar='OCR_OUTPUT_DIRECTORY_YEAR_MONTH',
help='Create a subdirectory in the output directory for each year and month', help='Create a subdirectory in the output directory for each year/month',
), ),
] = False, ] = False,
on_success_delete: Annotated[ on_success_delete: Annotated[
+13 -31
View File
@@ -12,12 +12,13 @@ readme = "README.md"
license = { text = "MPL-2.0" } license = { text = "MPL-2.0" }
requires-python = ">=3.10" requires-python = ">=3.10"
dependencies = [ dependencies = [
"Pillow>=10.0.1",
"deprecation>=2.1.0", "deprecation>=2.1.0",
"img2pdf>=0.5", "img2pdf>=0.5",
"packaging>=20", "packaging>=20",
"pdfminer.six>=20220319", "pdfminer.six>=20220319",
"pi-heif", # Heif image format - maintainers: if this is removed, it will NOT break
"pikepdf>=8.10.1", "pikepdf>=8.10.1",
"Pillow>=10.0.1",
"pluggy>=1", "pluggy>=1",
"rich>=13", "rich>=13",
] ]
@@ -60,7 +61,7 @@ test = [
"types-Pillow", "types-Pillow",
"types-humanfriendly", "types-humanfriendly",
] ]
watcher = ["watchdog>=1.0.2", "typer[all]", "python-dotenv"] watcher = ["watchdog>=1.0.2", "typer-slim[standard]", "python-dotenv"]
webservice = ["Flask>=2.0.1"] webservice = ["Flask>=2.0.1"]
[project.scripts] [project.scripts]
@@ -78,29 +79,6 @@ namespaces = false
[tool.distutils.bdist_wheel] [tool.distutils.bdist_wheel]
python-tag = "py310" python-tag = "py310"
[tool.black]
line-length = 88
target-version = ["py310", "py311", "py312"]
skip-string-normalization = true
include = '\.pyi?$'
exclude = '''
/(
\.eggs
| \.git
| \.hg
| \.mypy_cache
| \.tox
| \.venv
| _build
| buck-out
| build
| dist
| docs
| misc
| \.egg-info
)/
'''
[tool.coverage.run] [tool.coverage.run]
branch = true branch = true
parallel = true parallel = true
@@ -153,7 +131,10 @@ module = [
ignore_missing_imports = true ignore_missing_imports = true
[tool.ruff] [tool.ruff]
select = [ target-version = "py310"
[tool.ruff.lint]
"select" = [
"D", # pydocstyle "D", # pydocstyle
"E", # pycodestyle "E", # pycodestyle
"W", # pycodestyle "W", # pycodestyle
@@ -161,17 +142,18 @@ select = [
"I001", # isort "I001", # isort
"UP", # pyupgrade "UP", # pyupgrade
] ]
target-version = "py310"
[tool.ruff.isort] [tool.ruff.lint.isort]
known-first-party = ["ocrmypdf"] known-first-party = ["ocrmypdf"]
required-imports = ["from __future__ import annotations"]
[tool.ruff.pydocstyle] [tool.ruff.lint.pydocstyle]
convention = "google" convention = "google"
[tool.ruff.per-file-ignores] [tool.ruff.lint.per-file-ignores]
"docs/conf.py" = ["D100", "D101", "D105"] "docs/conf.py" = ["D100", "D101", "D105"]
"tests/*.py" = ["D100", "D101", "D102", "D103", "D105"] "tests/*.py" = ["D100", "D101", "D102", "D103", "D105"]
"misc/*.py" = ["D103", "D101", "D102"] "misc/*.py" = ["D103", "D101", "D102"]
"src/ocrmypdf/builtin_plugins/*.py" = ["D103", "D102", "D105"] "src/ocrmypdf/builtin_plugins/*.py" = ["D103", "D102", "D105"]
[tool.ruff.format]
quote-style = "preserve"
+2 -2
View File
@@ -7,8 +7,8 @@ from __future__ import annotations
import threading import threading
from abc import ABC, abstractmethod from abc import ABC, abstractmethod
from collections.abc import Iterable from collections.abc import Callable, Iterable
from typing import Any, Callable, TypeVar from typing import Any, TypeVar
from ocrmypdf._progressbar import NullProgressBar, ProgressBar from ocrmypdf._progressbar import NullProgressBar, ProgressBar
+2 -1
View File
@@ -220,7 +220,8 @@ def get_deskew(
def tesseract_log_output(stream: bytes) -> None: def tesseract_log_output(stream: bytes) -> None:
tlog = TesseractLoggerAdapter( tlog = TesseractLoggerAdapter(
log, extra=log.extra if hasattr(log, 'extra') else None # type: ignore log,
extra=log.extra if hasattr(log, 'extra') else None, # type: ignore
) )
if not stream: if not stream:
+2 -4
View File
@@ -8,19 +8,17 @@ from __future__ import annotations
import logging import logging
import os import os
import shlex import shlex
import sys
from collections.abc import Iterator from collections.abc import Iterator
from contextlib import contextmanager from contextlib import contextmanager
from decimal import Decimal from decimal import Decimal
from pathlib import Path from pathlib import Path
from subprocess import PIPE, STDOUT from subprocess import PIPE, STDOUT
from tempfile import TemporaryDirectory from tempfile import TemporaryDirectory
from typing import Union
from packaging.version import Version from packaging.version import Version
from PIL import Image from PIL import Image
from ocrmypdf.exceptions import MissingDependencyError, SubprocessOutputError from ocrmypdf.exceptions import SubprocessOutputError
from ocrmypdf.subprocess import get_version, run from ocrmypdf.subprocess import get_version, run
# unpaper documentation: # unpaper documentation:
@@ -29,7 +27,7 @@ from ocrmypdf.subprocess import get_version, run
UNPAPER_IMAGE_PIXEL_LIMIT = 256 * 1024 * 1024 UNPAPER_IMAGE_PIXEL_LIMIT = 256 * 1024 * 1024
DecFloat = Union[Decimal, float] DecFloat = Decimal | float
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
+2 -2
View File
@@ -93,8 +93,8 @@ class PageContext:
state = self.__dict__.copy() state = self.__dict__.copy()
state['options'] = copy(self.options) state['options'] = copy(self.options)
if not isinstance(state['options'].input_file, (str, bytes, os.PathLike)): if not isinstance(state['options'].input_file, str | bytes | os.PathLike):
state['options'].input_file = 'stream' state['options'].input_file = 'stream'
if not isinstance(state['options'].output_file, (str, bytes, os.PathLike)): if not isinstance(state['options'].output_file, str | bytes | os.PathLike):
state['options'].output_file = 'stream' state['options'].output_file = 'stream'
return state return state
+6 -3
View File
@@ -166,9 +166,12 @@ def metadata_fixup(
with Pdf.open(context.origin) as original, Pdf.open(working_file) as pdf: with Pdf.open(context.origin) as original, Pdf.open(working_file) as pdf:
docinfo = get_docinfo(original, context) docinfo = get_docinfo(original, context)
with original.open_metadata( with (
set_pikepdf_as_editor=False, update_docinfo=False, strict=False original.open_metadata(
) as meta_original, pdf.open_metadata() as meta_pdf: set_pikepdf_as_editor=False, update_docinfo=False, strict=False
) as meta_original,
pdf.open_metadata() as meta_pdf,
):
meta_pdf.load_from_docinfo( meta_pdf.load_from_docinfo(
docinfo, delete_missing=False, raise_failure=False docinfo, delete_missing=False, raise_failure=False
) )
+16 -3
View File
@@ -14,7 +14,7 @@ from collections.abc import Iterable, Iterator, Sequence
from contextlib import suppress from contextlib import suppress
from io import BytesIO from io import BytesIO
from pathlib import Path from pathlib import Path
from shutil import copyfileobj, copystat from shutil import copyfileobj
from typing import Any, BinaryIO, TypeVar, cast from typing import Any, BinaryIO, TypeVar, cast
import img2pdf import img2pdf
@@ -41,12 +41,23 @@ from ocrmypdf.pdfa import generate_pdfa_ps
from ocrmypdf.pdfinfo import Colorspace, Encoding, PageInfo, PdfInfo from ocrmypdf.pdfinfo import Colorspace, Encoding, PageInfo, PdfInfo
from ocrmypdf.pluginspec import OrientationConfidence from ocrmypdf.pluginspec import OrientationConfidence
try:
from pi_heif import register_heif_opener
except ImportError:
def register_heif_opener():
pass
T = TypeVar("T") T = TypeVar("T")
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
VECTOR_PAGE_DPI = 400 VECTOR_PAGE_DPI = 400
register_heif_opener()
def triage_image_file(input_file: Path, output_file: Path, options) -> None: def triage_image_file(input_file: Path, output_file: Path, options) -> None:
"""Triage the input image file. """Triage the input image file.
@@ -131,7 +142,7 @@ def _pdf_guess_version(input_file: Path, search_window=1024) -> str:
""" """
with open(input_file, 'rb') as f: with open(input_file, 'rb') as f:
signature = f.read(search_window) signature = f.read(search_window)
m = re.search(br'%PDF-(\d\.\d)', signature) m = re.search(rb'%PDF-(\d\.\d)', signature)
if m: if m:
return m.group(1).decode('ascii') return m.group(1).decode('ascii')
return '' return ''
@@ -767,7 +778,9 @@ def render_hocr_page(hocr: Path, page_context: PageContext) -> Path:
font=Courier(), font=Courier(),
) )
HocrTransform( HocrTransform(
hocr_filename=hocr, dpi=dpi.to_scalar(), **debug_kwargs # square hocr_filename=hocr,
dpi=dpi.to_scalar(),
**debug_kwargs, # square
).to_pdf( ).to_pdf(
out_filename=output_file, out_filename=output_file,
image_filename=None, image_filename=None,
+2 -2
View File
@@ -11,13 +11,13 @@ import os
import shutil import shutil
import sys import sys
import threading import threading
from collections.abc import Sequence from collections.abc import Callable, Sequence
from concurrent.futures.process import BrokenProcessPool from concurrent.futures.process import BrokenProcessPool
from concurrent.futures.thread import BrokenThreadPool from concurrent.futures.thread import BrokenThreadPool
from contextlib import contextmanager from contextlib import contextmanager
from dataclasses import dataclass from dataclasses import dataclass
from pathlib import Path from pathlib import Path
from typing import Callable, NamedTuple, cast from typing import NamedTuple, cast
import PIL import PIL
@@ -4,7 +4,6 @@
"""Implements the concurrent and page synchronous parts of the pipeline.""" """Implements the concurrent and page synchronous parts of the pipeline."""
from __future__ import annotations from __future__ import annotations
import argparse import argparse
+7 -7
View File
@@ -4,7 +4,6 @@
"""Implements the concurrent and page synchronous parts of the pipeline.""" """Implements the concurrent and page synchronous parts of the pipeline."""
from __future__ import annotations from __future__ import annotations
import argparse import argparse
@@ -155,12 +154,13 @@ def _run_pipeline(
options: argparse.Namespace, options: argparse.Namespace,
plugin_manager: OcrmypdfPluginManager, plugin_manager: OcrmypdfPluginManager,
) -> ExitCode: ) -> ExitCode:
with manage_work_folder( with (
work_folder=Path(mkdtemp(prefix="ocrmypdf.io.")), manage_work_folder(
retain=options.keep_temporary_files, work_folder=Path(mkdtemp(prefix="ocrmypdf.io.")),
print_location=options.keep_temporary_files, retain=options.keep_temporary_files,
) as work_folder, manage_debug_log_handler( print_location=options.keep_temporary_files,
options=options, work_folder=work_folder ) as work_folder,
manage_debug_log_handler(options=options, work_folder=work_folder),
): ):
executor = setup_pipeline(options, plugin_manager) executor = setup_pipeline(options, plugin_manager)
check_requested_output_file(options) check_requested_output_file(options)
-1
View File
@@ -4,7 +4,6 @@
"""Implements the concurrent and page synchronous parts of the pipeline.""" """Implements the concurrent and page synchronous parts of the pipeline."""
from __future__ import annotations from __future__ import annotations
import argparse import argparse
+13 -8
View File
@@ -14,7 +14,7 @@ from collections.abc import Iterable, Sequence
from enum import IntEnum from enum import IntEnum
from io import IOBase from io import IOBase
from pathlib import Path from pathlib import Path
from typing import AnyStr, BinaryIO, Union from typing import AnyStr, BinaryIO
from warnings import warn from warnings import warn
import pluggy import pluggy
@@ -28,8 +28,8 @@ from ocrmypdf._validation import check_options
from ocrmypdf.cli import ArgumentParser, get_parser from ocrmypdf.cli import ArgumentParser, get_parser
from ocrmypdf.helpers import is_iterable_notstr from ocrmypdf.helpers import is_iterable_notstr
StrPath = Union[Path, AnyStr] StrPath = Path | AnyStr
PathOrIO = Union[BinaryIO, StrPath] PathOrIO = BinaryIO | StrPath
# Installing plugins affects the global state of the Python interpreter, # Installing plugins affects the global state of the Python interpreter,
# so we need to use a lock to prevent multiple threads from installing # so we need to use a lock to prevent multiple threads from installing
@@ -169,7 +169,7 @@ def _kwargs_to_cmdline(
# We have a parameter # We have a parameter
cmdline.append(f"--{cmd_style_arg}") cmdline.append(f"--{cmd_style_arg}")
if isinstance(val, (int, float)): if isinstance(val, int | float):
cmdline.append(str(val)) cmdline.append(str(val))
elif isinstance(val, str): elif isinstance(val, str):
cmdline.append(val) cmdline.append(val)
@@ -201,14 +201,17 @@ def create_options(
defer_kwargs={'progress_bar', 'plugins', 'parser', 'input_file', 'output_file'}, defer_kwargs={'progress_bar', 'plugins', 'parser', 'input_file', 'output_file'},
**kwargs, **kwargs,
) )
if isinstance(input_file, (BinaryIO, IOBase)): if isinstance(input_file, BinaryIO | IOBase):
cmdline.append('stream://input_file') cmdline.append('stream://input_file')
else: else:
cmdline.append(os.fspath(input_file)) cmdline.append(os.fspath(input_file))
if isinstance(output_file, (BinaryIO, IOBase)): if isinstance(output_file, BinaryIO | IOBase):
cmdline.append('stream://output_file') cmdline.append('stream://output_file')
else: else:
cmdline.append(os.fspath(output_file)) cmdline.append(os.fspath(output_file))
if 'sidecar' in kwargs and isinstance(kwargs['sidecar'], BinaryIO | IOBase):
cmdline.append('--sidecar')
cmdline.append('stream://sidecar')
parser.enable_api_mode() parser.enable_api_mode()
options = parser.parse_args(cmdline) options = parser.parse_args(cmdline)
@@ -219,6 +222,8 @@ def create_options(
options.input_file = input_file options.input_file = input_file
if options.output_file == 'stream://output_file': if options.output_file == 'stream://output_file':
options.output_file = output_file options.output_file = output_file
if options.sidecar == 'stream://sidecar':
options.sidecar = kwargs['sidecar']
return options return options
@@ -230,7 +235,7 @@ def ocr( # noqa: D417
language: Iterable[str] | None = None, language: Iterable[str] | None = None,
image_dpi: int | None = None, image_dpi: int | None = None,
output_type: str | None = None, output_type: str | None = None,
sidecar: StrPath | None = None, sidecar: PathOrIO | None = None,
jobs: int | None = None, jobs: int | None = None,
use_threads: bool | None = None, use_threads: bool | None = None,
title: str | None = None, title: str | None = None,
@@ -343,7 +348,7 @@ def ocr( # noqa: D417
if not plugins: if not plugins:
plugins = [] plugins = []
elif isinstance(plugins, (str, Path)): elif isinstance(plugins, str | Path):
plugins = [plugins] plugins = [plugins]
else: else:
plugins = list(plugins) plugins = list(plugins)
+14 -9
View File
@@ -12,10 +12,10 @@ import queue
import signal import signal
import sys import sys
import threading import threading
from collections.abc import Iterable from collections.abc import Callable, Iterable
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor, as_completed from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor, as_completed
from contextlib import suppress from contextlib import suppress
from typing import Callable, Union from typing import Union
from rich.console import Console as RichConsole from rich.console import Console as RichConsole
@@ -25,8 +25,10 @@ from ocrmypdf._progressbar import RichProgressBar
from ocrmypdf.exceptions import InputFileError from ocrmypdf.exceptions import InputFileError
from ocrmypdf.helpers import remove_all_log_handlers from ocrmypdf.helpers import remove_all_log_handlers
FuturesExecutorClass = Union[type[ThreadPoolExecutor], type[ProcessPoolExecutor]] FuturesExecutorClass = Union[ # noqa: UP007
Queue = Union[multiprocessing.Queue, queue.Queue] type[ThreadPoolExecutor], type[ProcessPoolExecutor]
]
Queue = Union[multiprocessing.Queue, queue.Queue] # noqa: UP007
UserInit = Callable[[], None] UserInit = Callable[[], None]
WorkerInit = Callable[[Queue, UserInit, int], None] WorkerInit = Callable[[Queue, UserInit, int], None]
@@ -128,11 +130,14 @@ class StandardExecutor(Executor):
listener = threading.Thread(target=log_listener, args=(log_queue,)) listener = threading.Thread(target=log_listener, args=(log_queue,))
listener.start() listener.start()
with self.pbar_class(**progress_kwargs) as pbar, executor_class( with (
max_workers=max_workers, self.pbar_class(**progress_kwargs) as pbar,
initializer=initializer, executor_class(
initargs=(log_queue, worker_initializer, logging.getLogger("").level), max_workers=max_workers,
) as executor: initializer=initializer,
initargs=(log_queue, worker_initializer, logging.getLogger("").level),
) as executor,
):
futures = [executor.submit(task, *args) for args in task_arguments] futures = [executor.submit(task, *args) for args in task_arguments]
try: try:
for future in as_completed(futures): for future in as_completed(futures):
@@ -8,7 +8,5 @@ from ocrmypdf import hookimpl
@hookimpl @hookimpl
def filter_pdf_page( def filter_pdf_page(page, image_filename, output_pdf): # pylint: disable=unused-argument
page, image_filename, output_pdf
): # pylint: disable=unused-argument
return output_pdf return output_pdf
@@ -2,9 +2,9 @@
# SPDX-License-Identifier: MPL-2.0 # SPDX-License-Identifier: MPL-2.0
"""Built-in plugin to implement OCR using Tesseract.""" """Built-in plugin to implement OCR using Tesseract."""
from __future__ import annotations from __future__ import annotations
import argparse
import logging import logging
import os import os
@@ -94,7 +94,8 @@ def add_options(parser):
) )
tess.add_argument( tess.add_argument(
'--tesseract-downsample-large-images', '--tesseract-downsample-large-images',
action='store_true', action=argparse.BooleanOptionalAction,
default=True,
help=( help=(
"Downsample large images before OCR. Tesseract has an upper limit on the " "Downsample large images before OCR. Tesseract has an upper limit on the "
"size images it will support. If this argument is given, OCRmyPDF will " "size images it will support. If this argument is given, OCRmyPDF will "
@@ -220,7 +221,7 @@ class TesseractOcrEngine(OcrEngine):
@staticmethod @staticmethod
def creator_tag(options): def creator_tag(options):
tag = '-PDF' if options.pdf_renderer == 'sandwich' else 'hOCR' tag = '-PDF' if options.pdf_renderer == 'sandwich' else '-hOCR'
return f"Tesseract OCR{tag} {TesseractOcrEngine.version()}" return f"Tesseract OCR{tag} {TesseractOcrEngine.version()}"
def __str__(self): def __str__(self):
+3 -3
View File
@@ -6,8 +6,8 @@
from __future__ import annotations from __future__ import annotations
import argparse import argparse
from collections.abc import Mapping from collections.abc import Callable, Mapping
from typing import Any, Callable, TypeVar from typing import Any, TypeVar
from ocrmypdf._version import PROGRAM_NAME as _PROGRAM_NAME from ocrmypdf._version import PROGRAM_NAME as _PROGRAM_NAME
from ocrmypdf._version import __version__ as _VERSION from ocrmypdf._version import __version__ as _VERSION
@@ -390,7 +390,7 @@ Online documentation is located at:
action='store', action='store',
type=numeric(float, 0), type=numeric(float, 0),
metavar='MPixels', metavar='MPixels',
help="Set maximum number of pixels to unpack before treating an image as a " help="Set maximum number of megapixels to unpack before treating an image as a "
"decompression bomb", "decompression bomb",
default=250.0, default=250.0,
) )
+1 -2
View File
@@ -20,13 +20,12 @@ from __future__ import annotations
import logging import logging
import logging.handlers import logging.handlers
import signal import signal
from collections.abc import Iterable, Iterator from collections.abc import Callable, Iterable, Iterator
from contextlib import suppress from contextlib import suppress
from enum import Enum, auto from enum import Enum, auto
from itertools import islice, repeat, takewhile, zip_longest from itertools import islice, repeat, takewhile, zip_longest
from multiprocessing import Pipe, Process from multiprocessing import Pipe, Process
from multiprocessing.connection import Connection, wait from multiprocessing.connection import Connection, wait
from typing import Callable
from ocrmypdf import Executor, hookimpl from ocrmypdf import Executor, hookimpl
from ocrmypdf._concurrent import NullProgressBar from ocrmypdf._concurrent import NullProgressBar
+1 -2
View File
@@ -10,7 +10,7 @@ import multiprocessing
import os import os
import shutil import shutil
import warnings import warnings
from collections.abc import Iterable, Sequence from collections.abc import Callable, Iterable, Sequence
from contextlib import suppress from contextlib import suppress
from decimal import Decimal from decimal import Decimal
from io import StringIO from io import StringIO
@@ -19,7 +19,6 @@ from pathlib import Path
from statistics import harmonic_mean from statistics import harmonic_mean
from typing import ( from typing import (
Any, Any,
Callable,
Generic, Generic,
TypeVar, TypeVar,
) )
-1
View File
@@ -84,7 +84,6 @@ class HocrTransform:
debug_render_options: DebugRenderOptions | None = None, debug_render_options: DebugRenderOptions | None = None,
): ):
"""Initialize the HocrTransform object.""" """Initialize the HocrTransform object."""
if debug: if debug:
log.warning("Use debug_render_options instead", DeprecationWarning) log.warning("Use debug_render_options instead", DeprecationWarning)
self.render_options = DebugRenderOptions( self.render_options = DebugRenderOptions(
+5 -3
View File
@@ -7,12 +7,12 @@ Derived from
https://www.loc.gov/standards/iso639-2/ascii_8bits.html https://www.loc.gov/standards/iso639-2/ascii_8bits.html
""" """
from typing import NamedTuple from typing import NamedTuple
class ISOCodeData(NamedTuple): class ISOCodeData(NamedTuple):
"""Data for a single ISO 639 code.""" """Data for a single ISO 639 code."""
alt: str alt: str
alpha_2: str alpha_2: str
english: str english: str
@@ -168,8 +168,10 @@ ISO_639_3 = {
'chu': ISOCodeData( 'chu': ISOCodeData(
'', '',
'cu', 'cu',
('Church Slavic; Old Slavonic; Church Slavonic;' (
' Old Bulgarian; Old Church Slavonic'), 'Church Slavic; Old Slavonic; Church Slavonic;'
' Old Bulgarian; Old Church Slavonic'
),
"slavon d'église; vieux slave; slavon liturgique; vieux bulgare", "slavon d'église; vieux slave; slavon liturgique; vieux bulgare",
), ),
'chv': ISOCodeData('', 'cv', 'Chuvash', 'tchouvache'), 'chv': ISOCodeData('', 'cv', 'Chuvash', 'tchouvache'),
+3 -4
View File
@@ -3,7 +3,6 @@
"""Post-processing image optimization of OCR PDFs.""" """Post-processing image optimization of OCR PDFs."""
from __future__ import annotations from __future__ import annotations
import logging import logging
@@ -11,11 +10,10 @@ import sys
import tempfile import tempfile
import threading import threading
from collections import defaultdict from collections import defaultdict
from collections.abc import Iterator, MutableSet, Sequence from collections.abc import Callable, Iterator, MutableSet, Sequence
from os import fspath from os import fspath
from pathlib import Path from pathlib import Path
from typing import Any, Callable, NamedTuple, NewType from typing import Any, NamedTuple, NewType
from warnings import warn
from zlib import compress from zlib import compress
import img2pdf import img2pdf
@@ -91,6 +89,7 @@ def extract_image_filter(
if ( if (
len(pim.filter_decodeparms) == 2 len(pim.filter_decodeparms) == 2
and first_filtdp[0] == Name.FlateDecode and first_filtdp[0] == Name.FlateDecode
and first_filtdp[1] is not None
and first_filtdp[1].get(Name.Predictor, 1) == 1 and first_filtdp[1].get(Name.Predictor, 1) == 1
and second_filtdp[0] == Name.DCTDecode and second_filtdp[0] == Name.DCTDecode
and not second_filtdp[1] and not second_filtdp[1]
+1 -1
View File
@@ -90,7 +90,7 @@ def generate_pdfa_ps(target_filename: Path, icc: str = 'sRGB'):
icc: ICC identifier such as 'sRGB' icc: ICC identifier such as 'sRGB'
References: References:
Adobe PDFMARK Reference: Adobe PDFMARK Reference:
https://www.adobe.com/content/dam/acom/en/devnet/acrobat/pdfs/pdfmark_reference.pdf https://opensource.adobe.com/dc-acrobat-sdk-docs/library/pdfmark/
""" """
if icc != 'sRGB': if icc != 'sRGB':
raise NotImplementedError("Only supporting sRGB") raise NotImplementedError("Only supporting sRGB")
+4 -10
View File
@@ -10,9 +10,8 @@ import atexit
import logging import logging
import re import re
import statistics import statistics
import sys
from collections import defaultdict from collections import defaultdict
from collections.abc import Container, Iterable, Iterator, Mapping, Sequence from collections.abc import Callable, Container, Iterable, Iterator, Mapping, Sequence
from contextlib import contextmanager from contextlib import contextmanager
from decimal import Decimal from decimal import Decimal
from enum import Enum, auto from enum import Enum, auto
@@ -20,7 +19,7 @@ from functools import partial
from math import hypot, inf, isclose from math import hypot, inf, isclose
from os import PathLike from os import PathLike
from pathlib import Path from pathlib import Path
from typing import Callable, NamedTuple from typing import NamedTuple
from warnings import warn from warnings import warn
from pdfminer.layout import LTPage, LTTextBox from pdfminer.layout import LTPage, LTTextBox
@@ -1060,12 +1059,7 @@ class PageInfo:
weights = [area / total_drawn_area for area in image_areas] weights = [area / total_drawn_area for area in image_areas]
# Calculate harmonic mean of DPIs weighted by area # Calculate harmonic mean of DPIs weighted by area
if sys.version_info >= (3, 10): weighted_dpi = statistics.harmonic_mean(image_dpis, weights)
weighted_dpi = statistics.harmonic_mean(image_dpis, weights)
else:
weighted_dpi = sum(weights) / sum(
weight / dpi for weight, dpi in zip(weights, image_dpis)
)
max_dpi = max(image_dpis) max_dpi = max(image_dpis)
dpi_average_max_ratio = weighted_dpi / max_dpi dpi_average_max_ratio = weighted_dpi / max_dpi
@@ -1176,7 +1170,7 @@ class PdfInfo:
@property @property
def filename(self) -> str | Path: def filename(self) -> str | Path:
"""Return filename of PDF.""" """Return filename of PDF."""
if not isinstance(self._infile, (str, Path)): if not isinstance(self._infile, str | Path):
raise NotImplementedError("can't get filename from stream") raise NotImplementedError("can't get filename from stream")
return self._infile return self._infile
+2 -2
View File
@@ -5,12 +5,12 @@
from __future__ import annotations from __future__ import annotations
import re import re
from collections.abc import Mapping from collections.abc import Iterator, Mapping
from contextlib import contextmanager from contextlib import contextmanager
from math import copysign from math import copysign
from os import PathLike from os import PathLike
from pathlib import Path from pathlib import Path
from typing import Any, Iterator from typing import Any
from unittest.mock import patch from unittest.mock import patch
import pdfminer import pdfminer
-1
View File
@@ -3,7 +3,6 @@
"""Utilities to measure OCR quality.""" """Utilities to measure OCR quality."""
from __future__ import annotations from __future__ import annotations
import re import re
+3 -3
View File
@@ -8,12 +8,11 @@ import logging
import os import os
import re import re
import sys import sys
from collections.abc import Mapping, Sequence from collections.abc import Callable, Mapping, Sequence
from contextlib import suppress from contextlib import suppress
from pathlib import Path from pathlib import Path
from subprocess import PIPE, STDOUT, CalledProcessError, CompletedProcess, Popen from subprocess import PIPE, STDOUT, CalledProcessError, CompletedProcess, Popen
from subprocess import run as subprocess_run from subprocess import run as subprocess_run
from typing import Callable, Union
from packaging.version import Version from packaging.version import Version
@@ -23,7 +22,7 @@ from ocrmypdf.exceptions import MissingDependencyError
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
Args = Sequence[Union[Path, str]] Args = Sequence[Path | str]
OsEnviron = os._Environ # pylint: disable=protected-access OsEnviron = os._Environ # pylint: disable=protected-access
@@ -172,6 +171,7 @@ def get_version(
) from e ) from e
except CalledProcessError as e: except CalledProcessError as e:
if e.returncode != 0: if e.returncode != 0:
log.exception(e)
raise MissingDependencyError( raise MissingDependencyError(
f"Ran program '{program}' but it exited with an error:\n{e.output}" f"Ran program '{program}' but it exited with an error:\n{e.output}"
) from e ) from e
+3 -8
View File
@@ -9,18 +9,13 @@ import os
import re import re
import shutil import shutil
import sys import sys
from collections.abc import Iterable, Iterator from collections.abc import Callable, Iterable, Iterator
from itertools import chain from itertools import chain
from pathlib import Path from pathlib import Path
from typing import Any, Callable, TypeVar from typing import Any, TypeAlias, TypeVar
from packaging.version import InvalidVersion, Version from packaging.version import InvalidVersion, Version
if sys.version_info >= (3, 10):
from typing import TypeAlias
else:
from typing_extensions import TypeAlias # pragma: no cover
if sys.platform == 'win32': if sys.platform == 'win32':
# mypy understands 'if sys.platform' better than try/except ModuleNotFoundError # mypy understands 'if sys.platform' better than try/except ModuleNotFoundError
import winreg # pylint: disable=import-error import winreg # pylint: disable=import-error
@@ -84,7 +79,7 @@ def registry_path_ghostscript(env=None) -> Iterator[Path]:
registry_subkeys(k), key=ghostscript_version_key, default=(0, 0, 0) registry_subkeys(k), key=ghostscript_version_key, default=(0, 0, 0)
) )
with winreg.OpenKey( with winreg.OpenKey(
winreg.HKEY_LOCAL_MACHINE, fr"SOFTWARE\Artifex\GPL Ghostscript\{latest_gs}" winreg.HKEY_LOCAL_MACHINE, rf"SOFTWARE\Artifex\GPL Ghostscript\{latest_gs}"
) as k: ) as k:
for _, gs_path, _ in registry_values(k): for _, gs_path, _ in registry_values(k):
yield Path(gs_path) / 'bin' yield Path(gs_path) / 'bin'
-1
View File
@@ -3,7 +3,6 @@
from __future__ import annotations from __future__ import annotations
from pathlib import Path
from subprocess import CalledProcessError from subprocess import CalledProcessError
from unittest.mock import patch from unittest.mock import patch
+12 -8
View File
@@ -169,22 +169,25 @@ class CacheOcrEngine(TesseractOcrEngine):
@staticmethod @staticmethod
def get_orientation(input_file, options): def get_orientation(input_file, options):
with CacheOcrEngine.lock, patch( with (
'ocrmypdf._exec.tesseract.run', new=partial(cached_run, options) CacheOcrEngine.lock,
patch('ocrmypdf._exec.tesseract.run', new=partial(cached_run, options)),
): ):
return TesseractOcrEngine.get_orientation(input_file, options) return TesseractOcrEngine.get_orientation(input_file, options)
@staticmethod @staticmethod
def get_deskew(input_file, options) -> float: def get_deskew(input_file, options) -> float:
with CacheOcrEngine.lock, patch( with (
'ocrmypdf._exec.tesseract.run', new=partial(cached_run, options) CacheOcrEngine.lock,
patch('ocrmypdf._exec.tesseract.run', new=partial(cached_run, options)),
): ):
return TesseractOcrEngine.get_deskew(input_file, options) return TesseractOcrEngine.get_deskew(input_file, options)
@staticmethod @staticmethod
def generate_hocr(input_file, output_hocr, output_text, options): def generate_hocr(input_file, output_hocr, output_text, options):
with CacheOcrEngine.lock, patch( with (
'ocrmypdf._exec.tesseract.run', new=partial(cached_run, options) CacheOcrEngine.lock,
patch('ocrmypdf._exec.tesseract.run', new=partial(cached_run, options)),
): ):
TesseractOcrEngine.generate_hocr( TesseractOcrEngine.generate_hocr(
input_file, output_hocr, output_text, options input_file, output_hocr, output_text, options
@@ -192,8 +195,9 @@ class CacheOcrEngine(TesseractOcrEngine):
@staticmethod @staticmethod
def generate_pdf(input_file, output_pdf, output_text, options): def generate_pdf(input_file, output_pdf, output_text, options):
with CacheOcrEngine.lock, patch( with (
'ocrmypdf._exec.tesseract.run', new=partial(cached_run, options) CacheOcrEngine.lock,
patch('ocrmypdf._exec.tesseract.run', new=partial(cached_run, options)),
): ):
TesseractOcrEngine.generate_pdf( TesseractOcrEngine.generate_pdf(
input_file, output_pdf, output_text, options input_file, output_pdf, output_text, options
+4 -3
View File
@@ -72,9 +72,10 @@ class FixedRotateNoopOcrEngine(OcrEngine):
@staticmethod @staticmethod
def generate_hocr(input_file, output_hocr, output_text, options): def generate_hocr(input_file, output_hocr, output_text, options):
with Image.open(input_file) as im, open( with (
output_hocr, 'w', encoding='utf-8' Image.open(input_file) as im,
) as f: open(output_hocr, 'w', encoding='utf-8') as f,
):
w, h = im.size w, h = im.size
f.write(HOCR_TEMPLATE.format(str(w), str(h))) f.write(HOCR_TEMPLATE.format(str(w), str(h)))
with open(output_text, 'w') as f: with open(output_text, 'w') as f:
+4 -3
View File
@@ -70,9 +70,10 @@ class NoopOcrEngine(OcrEngine):
@staticmethod @staticmethod
def generate_hocr(input_file, output_hocr, output_text, options): def generate_hocr(input_file, output_hocr, output_text, options):
with Image.open(input_file) as im, open( with (
output_hocr, 'w', encoding='utf-8' Image.open(input_file) as im,
) as f: open(output_hocr, 'w', encoding='utf-8') as f,
):
w, h = im.size w, h = im.size
f.write(HOCR_TEMPLATE.format(str(w), str(h))) f.write(HOCR_TEMPLATE.format(str(w), str(h)))
with open(output_text, 'w') as f: with open(output_text, 'w') as f:
@@ -9,6 +9,7 @@ ensure we fail with an error rather than deadlock in such cases.
Page 4 was chosen because of this number's association with bad luck Page 4 was chosen because of this number's association with bad luck
in many East Asian cultures. in many East Asian cultures.
""" """
# type: ignore # type: ignore
from __future__ import annotations from __future__ import annotations
+13
View File
@@ -29,6 +29,18 @@ def test_stream_api(resources: Path):
assert b'%PDF' in out.read(1024) assert b'%PDF' in out.read(1024)
def test_sidecar_stringio(resources: Path, outdir: Path, outpdf: Path):
s = BytesIO()
ocrmypdf.ocr(
resources / 'ccitt.pdf',
outpdf,
plugins=['tests/plugins/tesseract_cache.py'],
sidecar=s
)
s.seek(0)
assert b'the' in s.getvalue()
def test_hocr_api_multipage(resources: Path, outdir: Path, outpdf: Path): def test_hocr_api_multipage(resources: Path, outdir: Path, outpdf: Path):
ocrmypdf.api._pdf_to_hocr( ocrmypdf.api._pdf_to_hocr(
resources / 'multipage.pdf', resources / 'multipage.pdf',
@@ -62,3 +74,4 @@ def test_hocr_to_pdf_api(resources: Path, outdir: Path, outpdf: Path):
text = extract_text(outpdf) text = extract_text(outpdf)
assert 'hocr' in text and 'the' not in text assert 'hocr' in text and 'the' not in text
+5 -3
View File
@@ -11,9 +11,11 @@ import ocrmypdf
def test_no_glyphless_graft(resources, outdir): def test_no_glyphless_graft(resources, outdir):
with pikepdf.open(resources / 'francais.pdf') as pdf, pikepdf.open( with (
resources / 'aspect.pdf' pikepdf.open(resources / 'francais.pdf') as pdf,
) as pdf_aspect, pikepdf.open(resources / 'cmyk.pdf') as pdf_cmyk: pikepdf.open(resources / 'aspect.pdf') as pdf_aspect,
pikepdf.open(resources / 'cmyk.pdf') as pdf_cmyk,
):
pdf.pages.extend(pdf_aspect.pages) pdf.pages.extend(pdf_aspect.pages)
pdf.pages.extend(pdf_cmyk.pages) pdf.pages.extend(pdf_cmyk.pages)
pdf.save(outdir / 'test.pdf') pdf.save(outdir / 'test.pdf')
+1 -1
View File
@@ -17,7 +17,7 @@ from PIL import Image
import ocrmypdf import ocrmypdf
from ocrmypdf._exec import tesseract from ocrmypdf._exec import tesseract
from ocrmypdf.exceptions import ExitCode, MissingDependencyError, OutputFileAccessError from ocrmypdf.exceptions import ExitCode, MissingDependencyError
from ocrmypdf.pdfa import file_claims_pdfa from ocrmypdf.pdfa import file_claims_pdfa
from ocrmypdf.pdfinfo import Colorspace, Encoding, PdfInfo from ocrmypdf.pdfinfo import Colorspace, Encoding, PdfInfo
from ocrmypdf.subprocess import get_version from ocrmypdf.subprocess import get_version
+4 -3
View File
@@ -35,9 +35,10 @@ def test_preserve_docinfo(output_type, resources, outpdf):
'--plugin', '--plugin',
'tests/plugins/tesseract_noop.py', 'tests/plugins/tesseract_noop.py',
) )
with pikepdf.open(resources / 'graph.pdf') as pdf_before, pikepdf.open( with (
output pikepdf.open(resources / 'graph.pdf') as pdf_before,
) as pdf_after: pikepdf.open(output) as pdf_after,
):
for key in ('/Title', '/Author'): for key in ('/Title', '/Author'):
assert pdf_before.docinfo[key] == pdf_after.docinfo[key] assert pdf_before.docinfo[key] == pdf_after.docinfo[key]
pdfa_info = file_claims_pdfa(str(output)) pdfa_info = file_claims_pdfa(str(output))
+7 -2
View File
@@ -9,10 +9,11 @@ import pytest
from PIL import Image from PIL import Image
from ocrmypdf._exec import ghostscript, tesseract from ocrmypdf._exec import ghostscript, tesseract
from ocrmypdf.exceptions import ExitCode
from ocrmypdf.helpers import Resolution from ocrmypdf.helpers import Resolution
from ocrmypdf.pdfinfo import PdfInfo from ocrmypdf.pdfinfo import PdfInfo
from .conftest import check_ocrmypdf, have_unpaper from .conftest import check_ocrmypdf, have_unpaper, run_ocrmypdf
RENDERERS = ['hocr', 'sandwich'] RENDERERS = ['hocr', 'sandwich']
@@ -107,7 +108,7 @@ def test_non_square_resolution(renderer, resources, outpdf):
in_pageinfo = PdfInfo(resources / 'aspect.pdf') in_pageinfo = PdfInfo(resources / 'aspect.pdf')
assert in_pageinfo[0].dpi.x != in_pageinfo[0].dpi.y assert in_pageinfo[0].dpi.x != in_pageinfo[0].dpi.y
check_ocrmypdf( proc = run_ocrmypdf(
resources / 'aspect.pdf', resources / 'aspect.pdf',
outpdf, outpdf,
'--pdf-renderer', '--pdf-renderer',
@@ -115,6 +116,10 @@ def test_non_square_resolution(renderer, resources, outpdf):
'--plugin', '--plugin',
'tests/plugins/tesseract_cache.py', 'tests/plugins/tesseract_cache.py',
) )
# PDF/A conversion can fail for this file if Ghostscript >= 10.3, so don't test
# exit code in that case
if proc.returncode != ExitCode.pdfa_conversion_failed:
proc.check_returncode()
out_pageinfo = PdfInfo(outpdf) out_pageinfo = PdfInfo(outpdf)
+1 -2
View File
@@ -7,10 +7,9 @@ from math import isclose
import pytest import pytest
from ocrmypdf.exceptions import ExitCode
from ocrmypdf.pdfinfo import PdfInfo from ocrmypdf.pdfinfo import PdfInfo
from .conftest import check_ocrmypdf, run_ocrmypdf_api from .conftest import check_ocrmypdf
# pylint: disable=redefined-outer-name # pylint: disable=redefined-outer-name