Compare commits

..
66 Commits
Author SHA1 Message Date
James R. Barlow 42713b77d7 v12.7.0 release notes 2021-10-12 13:39:49 -07:00
James R. Barlow 690f88119d Fix test failures on pikepdf 3.2.0 + pybind11 2.8.0
When compiled without pybind11 2.8.0, pikepdf supplies a shim to implement
pikepdf._ObjectMapping.values() which has subtly different semantics
from a true dict-like objects; in particular it supports
next(objectmap.values())
where a standard dict requires
next(iter(objectmap.values()).

pybind11 2.8.0 now implements .values() properly, meaning some misuses of
protocol  in ocrmypdf fail.

If pybind11 < 2.8.0, pikepdf will
continue to offer its shim. If pybind11 >= 2.8.0, pikepdf does not add its shim.

Consequently no changes were needed in pikepdf.

Closes #843
2021-10-12 13:38:52 -07:00
James R. Barlow 78f391536b Offer hint to user to use --max-image-mpixels after decompression bob error
Closes #801
2021-10-06 00:19:11 -07:00
mara004andGitHub 7bdd1828a9 [ci skip] docs/conf.py: add intersphinx mapping to make external links work (#838) 2021-10-04 00:34:59 -07:00
mara004andGitHub a8f513eeeb [ci skip] Update api.rst (#839) 2021-10-04 00:34:39 -07:00
fedeliallalineaandGitHub af18bc0684 fixs importlib.{metadata,resource} for new python version (#840)
Signed-off-by: Marco Genasci <fedeliallalinea@gmail.com>
2021-10-03 23:30:11 -07:00
James R. Barlow b621df6947 v12.6.0 release notes 2021-10-02 01:02:17 -07:00
James R. Barlow 313c9e7dc1 docs: add missing sphinx extensions 2021-09-26 23:41:25 -07:00
James R. Barlow 9d04795f7f docs: fix package version error 2021-09-26 23:41:17 -07:00
James R. Barlow 9a08e71e7f docs: show logo 2021-09-26 23:40:54 -07:00
James R. Barlow 790d3022f6 Implement --output-type=none to skip producing the PDF and use only the sidecar
Closes #787
2021-09-26 01:07:34 -07:00
James R. Barlow ec311af796 typing: subprocess 2021-09-22 17:18:59 -07:00
James R. Barlow c725bf79da flake8 delinting 2021-09-21 16:37:03 -07:00
James R. Barlow 9559f76fae optimize: fix typo in debug msg 2021-09-19 16:31:00 -07:00
James R. Barlow 45736b7c2b cli: clarify text to more accurately describe behavior of --jbig2-lossy 2021-09-16 16:03:39 -07:00
James R. Barlow 5629e960b9 v12.5.0 release notes update 2021-09-15 00:27:32 -07:00
James R. Barlow 79fd8d01a5 info: fix incorrect handling of inline images and other typing fixes 2021-09-15 00:26:15 -07:00
James R. Barlow 79fe7a0a85 graft: remove separate implementation of unparse_content_stream 2021-09-15 00:25:06 -07:00
James R. Barlow b4b32a35b5 optimize: fix typing consistency 2021-09-15 00:09:08 -07:00
James R. Barlow 4634b3db55 v12.5.0 release notes 2021-09-14 17:26:07 -07:00
James R. Barlow f5053158d4 hocrtransform: fix regression causing hocr text to be not rendered
Fixes #828
2021-09-14 17:23:09 -07:00
James R. Barlow dfa4ce1612 pyproject: configure mypy 2021-09-14 17:22:41 -07:00
James R. Barlow 585595a98e pre-commit: autoupdate 2021-09-14 17:22:29 -07:00
James R. Barlow f6396fbaac pre-commit: drop isort mirror 2021-09-14 17:20:43 -07:00
James R. Barlow 3859bae85e Merge branch 'feature/drop-pkg-resources' 2021-09-14 00:31:36 -07:00
James R. Barlow ee1a7baae7 requirements: drop setuptools
With importlib_* we no longer setuptools's pkg_resources.
2021-09-14 00:28:42 -07:00
James R. Barlow a4da05b66b docs: various fixes
As suggested by @Chealer

Closes #829, #830, #831, #832
2021-09-14 00:24:18 -07:00
James R. Barlow 4d67812d51 importlib helpers don't provide importlib.thing, but importlib_thing
Fix everywhere.
2021-09-14 00:15:07 -07:00
James R. Barlow 3534742ef9 ci: Enable build of feature branches 2021-09-13 01:35:14 -07:00
James R. Barlow 8bfd46c80d docs: update copyright 2021-09-13 01:10:49 -07:00
James R. Barlow cc6e9cecc0 Replace pkg_resources version lookup with importlib.metadata 2021-09-13 01:10:49 -07:00
James R. Barlow 208657f840 pdfa: replace pkg_resources with importlib.resources 2021-09-13 01:10:49 -07:00
James R. Barlow f3de980447 Introduce importlib-resources,metadata backports for Python < 3.9 2021-09-13 01:10:43 -07:00
James R. Barlow eb8992e58b Update release notes 2021-09-09 15:53:08 -07:00
James R. Barlow 72ad618ae6 Merge branch 'master' of github.com:jbarlow83/OCRmyPDF 2021-09-09 15:25:54 -07:00
James R. Barlow f07d0c39bb Require pikepdf<3 on PyPy 3.6
Because cibuildwheel does not build wheels for PyPy 3.6 anymore,
so pikepdf does not offer one.
2021-09-09 15:25:38 -07:00
James R. Barlow 9c5c7d9be0 release notes: mention Py3.6 EOL 2021-09-08 00:04:21 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
0b19b084e2 build(deps): bump pillow from 8.2.0 to 8.3.2 in /requirements/main.txt (#825)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 8.2.0 to 8.3.2.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/master/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/8.2.0...8.3.2)

---
updated-dependencies:
- dependency-name: pillow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2021-09-08 00:03:35 -07:00
James R. Barlow 9b4516af7a ci : add packages for PyPy build 2021-08-31 02:44:07 -07:00
James R. Barlow 1eb45de5c9 v12.4.0 release notes 2021-08-31 02:35:39 -07:00
James R. Barlow 390b9924f5 ci: Add PyPy to matrix 2021-08-31 02:34:12 -07:00
James R. Barlow c28858a099 leptonica: fix a PyPy-specific error
Error is:
TypeError: from_buffer() got a 'memoryview' object, which supports the buffer interface but cannot be rendered as a plain raw address on PyPy

PyPy is happy to access a bytes() copy of the memoryview.
2021-08-31 02:31:58 -07:00
James R. Barlow f00b3c00cd Merge remote-tracking branch 'origin/master' 2021-08-31 02:19:09 -07:00
James R. Barlow 4e4f0bfa1f graft: use faster unparse_content_stream if available 2021-08-31 02:16:25 -07:00
Elliott Sales de AndradeandGitHub 0a31acf888 Allow pluggy v1. (#822)
There are breaking changes, but I could not find any reference to them
in the code.
2021-08-31 02:15:43 -07:00
James R. Barlow b91096c615 hoctransform: fix deprecation warning 2021-08-28 02:11:33 -07:00
James R. Barlow 95d9e8d91a info: inconsistent types used in ContentsInfo.name_index
This broke PyPy but CPython is fine with it.
2021-08-28 00:18:14 -07:00
James R. Barlow cb6c1939e9 typing: fix runtime issues 2021-08-27 02:18:54 -07:00
James R. Barlow 3764ee872a typing: refactor namedtuples in info 2021-08-27 00:51:11 -07:00
James R. Barlow e402d5cb4b typing: fix path and numeric issues 2021-08-27 00:23:38 -07:00
James R. Barlow 53cd04799a typing: fix Pillow usage, leptonica 2021-08-27 00:06:38 -07:00
James R. Barlow f2545d4496 typing: remove deprecated abstractstaticmethod 2021-08-26 23:51:11 -07:00
James R. Barlow 9b81e76ed4 typing: fix issues; avoid some magic literals 2021-08-26 23:48:27 -07:00
James R. Barlow 0956fc81aa typing: improvements for concurrency files 2021-08-26 23:47:53 -07:00
James R. Barlow 72279e7759 typing: confirmed that _pages_from_ranges(not-str) is never used 2021-08-26 23:47:11 -07:00
James R. Barlow 6f9b948064 typing: fix some trivial issues 2021-08-26 23:46:44 -07:00
James R. Barlow 4eca0a165b pre-commit: pyupgrade modernizing 2021-08-26 18:04:38 -07:00
James R. Barlow 067e61e03a pre-commit: auto update 2021-08-26 18:00:37 -07:00
James R. Barlow 1b46481f7e pre-commit: add setup.cfg fmt 2021-08-26 17:59:40 -07:00
James R. Barlow d8d9c41abb v12.3.3 release notes 2021-08-21 18:06:14 -07:00
James R. Barlow a8cad72f72 Merge branch 'master' of github.com:jbarlow83/OCRmyPDF 2021-08-21 17:38:19 -07:00
James R. Barlow 86c04305f4 readme: confirm Tesseract 5 support 2021-08-21 17:37:10 -07:00
James R. Barlow 0a110fac55 watcher: fix bool not working as expecting
Closes #821
2021-08-21 17:30:14 -07:00
James R. Barlow f8970ad862 Update CI to test Tesseract 5 and more Linux versions 2021-08-17 02:30:24 -07:00
James R. Barlow fcfc78b7ee Move tool config to pyproject 2021-08-17 01:46:22 -07:00
8bb244df24 [ci skip] Update jbig2.rst (#817)
Co-authored-by: jbarlow83 <jbarlow83@users.noreply.github.com>
2021-08-14 20:05:27 -07:00
47 changed files with 653 additions and 370 deletions
+32 -2
View File
@@ -6,6 +6,7 @@ on:
- master - master
- ci - ci
- release/* - release/*
- feature/*
tags: tags:
- v* - v*
paths-ignore: paths-ignore:
@@ -18,8 +19,24 @@ jobs:
runs-on: ${{ matrix.os }} runs-on: ${{ matrix.os }}
strategy: strategy:
matrix: matrix:
os: [ubuntu-18.04] #, ubuntu-20.04] include:
python: ["3.6"] #, "3.7", "3.8", "3.9"] - os: ubuntu-18.04
python: 3.6
- os: ubuntu-18.04
python: 3.7
- os: ubuntu-20.04
python: 3.8
- os: ubuntu-20.04
python: 3.9
- os: ubuntu-latest
python: 3.9
- os: ubuntu-20.04
python: "pypy-3.6"
- os: ubuntu-latest
python: "pypy-3.7"
- os: ubuntu-latest
python: 3.9
tesseract5: true
env: env:
OS: ${{ matrix.os }} OS: ${{ matrix.os }}
@@ -35,6 +52,11 @@ jobs:
with: with:
python-version: ${{ matrix.python }} python-version: ${{ matrix.python }}
- name: Install Tesseract 5
if: matrix.tesseract5
run: |
sudo add-apt-repository ppa:alex-p/tesseract-ocr-devel
- name: Install common packages - name: Install common packages
run: | run: |
sudo apt-get update sudo apt-get update
@@ -65,6 +87,14 @@ jobs:
sudo apt-get install -y --no-install-recommends \ sudo apt-get install -y --no-install-recommends \
libexempi8 libexempi8
- name: Install Ubuntu packages for PyPy
if: startsWith(matrix.python, 'pypy')
run: |
sudo apt-get install -y --no-install-recommends \
libxml2-dev \
libxslt1-dev \
pypy3-dev
- name: Install Python packages - name: Install Python packages
run: | run: |
python -m pip install .[test] python -m pip install .[test]
+23 -8
View File
@@ -1,23 +1,38 @@
repos: repos:
- repo: https://github.com/pre-commit/pre-commit-hooks - repo: https://github.com/pre-commit/pre-commit-hooks
rev: v3.4.0 rev: v4.0.1
hooks: hooks:
- id: check-case-conflict - id: check-case-conflict
- id: check-merge-conflict - id: check-merge-conflict
- id: check-toml - id: check-toml
- id: check-yaml - id: check-yaml
- id: debug-statements - id: debug-statements
- repo: https://github.com/asottile/seed-isort-config - repo: https://github.com/pycqa/isort
rev: v2.2.0 rev: 5.9.3
hooks:
- id: seed-isort-config
- repo: https://github.com/pre-commit/mirrors-isort
rev: v5.7.0 # pick the isort version you'd like to use from https://github.com/pre-commit/mirrors-isort/releases
hooks: hooks:
- id: isort - id: isort
args: ["--profile", "black"]
- repo: https://github.com/psf/black - repo: https://github.com/psf/black
rev: 20.8b1 rev: 21.9b0
hooks: hooks:
- id: black - id: black
language_version: python language_version: python
exclude: ^src/ocrmypdf/lib/_leptonica.py exclude: ^src/ocrmypdf/lib/_leptonica.py
- repo: https://github.com/asottile/setup-cfg-fmt
rev: v1.17.0
hooks:
- id: setup-cfg-fmt
- repo: https://github.com/asottile/pyupgrade
rev: v2.26.0
hooks:
- id: pyupgrade
args: ["--py36-plus"]
- repo: https://github.com/pre-commit/mirrors-mypy
rev: v0.910
hooks:
- id: mypy
additional_dependencies:
- types-toml
- types-setuptools
- types-requests
- types-Pillow
+6 -1
View File
@@ -91,6 +91,11 @@ brew install tesseract-lang
You can then pass the `-l LANG` argument to OCRmyPDF to give a hint as to what languages it should search for. Multiple languages can be requested. You can then pass the `-l LANG` argument to OCRmyPDF to give a hint as to what languages it should search for. Multiple languages can be requested.
OCRmyPDF supports Tesseract 4.0 and the beta versions of Tesseract 5.0. It will
automatically use whichever version it finds first on the `PATH` environment
variable. On Windows, if `PATH` does not provide a Tesseract binary, we use
the highest version number that is installed according to the Windows Registry.
## Documentation and support ## Documentation and support
Once OCRmyPDF is installed, the built-in help which explains the command syntax and options can be accessed via: Once OCRmyPDF is installed, the built-in help which explains the command syntax and options can be accessed via:
@@ -115,7 +120,7 @@ In addition to the required Python version (3.6+), OCRmyPDF requires external pr
- [heise Open Source, 09/2014: Texterkennung mit OCRmyPDF](https://heise.de/-2356670) - [heise Open Source, 09/2014: Texterkennung mit OCRmyPDF](https://heise.de/-2356670)
- [heise Durchsuchbare PDF-Dokumente mit OCRmyPDF erstellen](https://www.heise.de/ratgeber/Durchsuchbare-PDF-Dokumente-mit-OCRmyPDF-erstellen-4607592.html) - [heise Durchsuchbare PDF-Dokumente mit OCRmyPDF erstellen](https://www.heise.de/ratgeber/Durchsuchbare-PDF-Dokumente-mit-OCRmyPDF-erstellen-4607592.html)
- [Excellent Utilities: OCRmyPDF](https://www.linuxlinks.com/excellent-utilities-ocrmypdf-add-ocr-text-layer-scanned-pdfs/) - [Excellent Utilities: OCRmyPDF](https://www.linuxlinks.com/excellent-utilities-ocrmypdf-add-ocr-text-layer-scanned-pdfs/)
- [LinuxUser Texterkennung mit OCRmyPDF und Scanbd automatisieren](https://www.linux-community.de/ausgaben/linuxuser/2021/06/texterkennung-mit-ocrmypdf-und-scanbd-automatisieren/) - [LinuxUser Texterkennung mit OCRmyPDF und Scanbd automatisieren](https://www.linux-community.de/ausgaben/linuxuser/2021/06/texterkennung-mit-ocrmypdf-und-scanbd-automatisieren/)
## Business enquiries ## Business enquiries
+8 -8
View File
@@ -12,7 +12,7 @@ subprocess call anyway, as this provides isolation of its activities.
Example Example
======= =======
OCRmyPDF one high-level function to run its main engine from an OCRmyPDF provides one high-level function to run its main engine from an
application. The parameters are symmetric to the command line arguments application. The parameters are symmetric to the command line arguments
and largely have the same functions. and largely have the same functions.
@@ -23,7 +23,7 @@ and largely have the same functions.
if __name__ == '__main__': # To ensure correct behavior on Windows and macOS if __name__ == '__main__': # To ensure correct behavior on Windows and macOS
ocrmypdf.ocr('input.pdf', 'output.pdf', deskew=True) ocrmypdf.ocr('input.pdf', 'output.pdf', deskew=True)
With a few exceptions, all of the command line arguments are available With some exceptions, all of the command line arguments are available
and may be passed as equivalent keywords. and may be passed as equivalent keywords.
A few differences are that ``verbose`` and ``quiet`` are not available. A few differences are that ``verbose`` and ``quiet`` are not available.
@@ -41,29 +41,29 @@ execution. To do this, it will:
- manage the signal flags of its worker processes - manage the signal flags of its worker processes
- execute other subprocesses (forking and executing other programs) - execute other subprocesses (forking and executing other programs)
The Python process that calls ``ocrmypdf.ocr()`` must be sufficiently The Python process that calls :func:`ocrmypdf.ocr()` must be sufficiently
privileged to perform these actions. privileged to perform these actions.
There currently is no option to manage how jobs are scheduled other There currently is no option to manage how jobs are scheduled other
than the argument ``jobs=`` which will limit the number of worker than the argument ``jobs=`` which will limit the number of worker
processes. processes.
Creating a child process to call ``ocrmypdf.ocr()`` is suggested. That Creating a child process to call :func:`ocrmypdf.ocr()` is suggested. That
way your application will survive and remain interactive even if way your application will survive and remain interactive even if
OCRmyPDF fails for any reason. OCRmyPDF fails for any reason.
Programs that call ``ocrmypdf.ocr()`` should also install a SIGBUS signal Programs that call :func:`ocrmypdf.ocr()` should also install a SIGBUS signal
handler (except on Windows), to raise an exception if access to a memory handler (except on Windows), to raise an exception if access to a memory
mapped file fails. OCRmyPDF may use memory mapping. mapped file fails. OCRmyPDF may use memory mapping.
``ocrmypdf.ocr()`` will take a threading lock to prevent multiple runs of itself :func:`ocrmypdf.ocr()` will take a threading lock to prevent multiple runs of itself
in the same Python interpreter process. This is not thread-safe, because of how in the same Python interpreter process. This is not thread-safe, because of how
OCRmyPDF's plugins and Python's library import system work. If you need to parallelize OCRmyPDF's plugins and Python's library import system work. If you need to parallelize
OCRmyPDF, use processes. OCRmyPDF, use processes.
.. warning:: .. warning::
On Windows and macOS, the script that calls ``ocrmypdf.ocr()`` must be On Windows and macOS, the script that calls :func:`ocrmypdf.ocr()` must be
protected by an "ifmain" guard (``if __name__ == '__main__'``). If you do protected by an "ifmain" guard (``if __name__ == '__main__'``). If you do
not take at least one of these steps, process semantics will prevent not take at least one of these steps, process semantics will prevent
OCRmyPDF from working correctly. OCRmyPDF from working correctly.
@@ -96,7 +96,7 @@ Exceptions
OCRmyPDF may throw standard Python exceptions, ``ocrmypdf.exceptions.*`` OCRmyPDF may throw standard Python exceptions, ``ocrmypdf.exceptions.*``
exceptions, some exceptions related to multiprocessing, and exceptions, some exceptions related to multiprocessing, and
``KeyboardInterrupt``. The parent process should provide an exception :exc:`KeyboardInterrupt`. The parent process should provide an exception
handler. OCRmyPDF will clean up its temporary files and worker processes handler. OCRmyPDF will clean up its temporary files and worker processes
automatically when an exception occurs. automatically when an exception occurs.
+12 -6
View File
@@ -31,9 +31,16 @@
# Add any Sphinx extension module names here, as strings. They can be # Add any Sphinx extension module names here, as strings. They can be
# extensions coming with Sphinx (named 'sphinx.ext.*') or your custom # extensions coming with Sphinx (named 'sphinx.ext.*') or your custom
# ones. # ones.
extensions = ['sphinx.ext.napoleon', 'sphinx_issues'] extensions = [
'sphinx.ext.autodoc',
'sphinx.ext.intersphinx',
'sphinx.ext.autosummary',
'sphinx.ext.napoleon',
'sphinx_issues',
]
# Extension settings # Extension settings
intersphinx_mapping = {'https://docs.python.org/': None}
napoleon_use_rtype = False napoleon_use_rtype = False
issues_github_path = "jbarlow83/OCRmyPDF" issues_github_path = "jbarlow83/OCRmyPDF"
@@ -56,7 +63,7 @@ master_doc = 'index'
# General information about the project. # General information about the project.
project = 'ocrmypdf' project = 'ocrmypdf'
copyright = ( copyright = (
'2020, James R. Barlow. Licensed under Creative Commons Attribution-ShareAlike 4.0.' '2021, James R. Barlow. Licensed under Creative Commons Attribution-ShareAlike 4.0.'
) )
author = 'James R. Barlow' author = 'James R. Barlow'
@@ -88,11 +95,10 @@ if on_rtd:
] ]
sys.modules.update((mod_name, Mock()) for mod_name in MOCK_MODULES) sys.modules.update((mod_name, Mock()) for mod_name in MOCK_MODULES)
from importlib_metadata import version as package_version
from pkg_resources import get_distribution, DistributionNotFound
# The full version, including alpha/beta/rc tags. # The full version, including alpha/beta/rc tags.
release = get_distribution('ocrmypdf').version release = package_version('ocrmypdf')
version = '.'.join(release.split('.')[:2]) version = '.'.join(release.split('.')[:2])
@@ -275,7 +281,7 @@ htmlhelp_basename = 'ocrmypdfdoc'
# -- Options for LaTeX output --------------------------------------------- # -- Options for LaTeX output ---------------------------------------------
latex_elements = { latex_elements = { # type: ignore
# The paper size ('letterpaper' or 'a4paper'). # The paper size ('letterpaper' or 'a4paper').
# #
# 'papersize': 'letterpaper', # 'papersize': 'letterpaper',
+2
View File
@@ -1,6 +1,8 @@
OCRmyPDF documentation OCRmyPDF documentation
====================== ======================
.. figure:: images/logo.svg
OCRmyPDF adds an optical character recognition (OCR) text layer to scanned PDF OCRmyPDF adds an optical character recognition (OCR) text layer to scanned PDF
files, allowing them to be searched. files, allowing them to be searched.
+14 -9
View File
@@ -2,7 +2,12 @@
Introduction Introduction
============ ============
OCRmyPDF is a Python 3 application and library that adds OCR layers to PDFs. OCRmyPDF is an application and library that adds text "layers" to images
in PDFs, making scanned image PDFs searchable. It uses OCR to guess what text
is contained in images. It is written in Python. OCRmyPDF supports plugins
that allow customization of its processing steps, and is very tolerant of
PDFs that contain scanned images and "born digital" content that needs no
text recognition.
About OCR About OCR
========= =========
@@ -26,7 +31,7 @@ exactly. They contain `vector
graphics <http://vector-conversions.com/vectorizing/raster_vs_vector.html>`__ graphics <http://vector-conversions.com/vectorizing/raster_vs_vector.html>`__
that can contain raster objects such as scanned images. Because PDFs can that can contain raster objects such as scanned images. Because PDFs can
contain multiple pages (unlike many image formats) and can contain fonts contain multiple pages (unlike many image formats) and can contain fonts
and text, it is a good formats for exchanging scanned documents. and text, it is a good format for exchanging scanned documents.
|image| |image|
@@ -35,9 +40,9 @@ have one image. Some scanners or scanning software will segment pages
into monochromatic text and color regions for example, to improve the into monochromatic text and color regions for example, to improve the
compression ratio and appearance of the page. compression ratio and appearance of the page.
Rasterizing a PDF is the process of generating an image suitable for Rasterizing a PDF is the process of generating corresponding raster images.
display or analyzing with an OCR engine. OCR engines like Tesseract work OCR engines like Tesseract work with images, not scalable vector graphics
with images, not vector objects. or mixed raster-vector-text graphics such as PDF.
About PDF/A About PDF/A
=========== ===========
@@ -76,7 +81,7 @@ OCRmyPDF analyzes each page of a PDF to determine the colorspace and
resolution (DPI) needed to capture all of the information on that page resolution (DPI) needed to capture all of the information on that page
without losing content. It uses without losing content. It uses
`Ghostscript <http://ghostscript.com/>`__ to rasterize the page, and `Ghostscript <http://ghostscript.com/>`__ to rasterize the page, and
then performs on OCR on the rasterized image to create an OCR "layer". then performs on OCR the rasterized image to create an OCR "layer".
The layer is then grafted back onto the original PDF. The layer is then grafted back onto the original PDF.
While one can use a program like Ghostscript or ImageMagick to get an While one can use a program like Ghostscript or ImageMagick to get an
@@ -84,9 +89,9 @@ image and put the image through Tesseract, that actually creates a new
PDF and many details may be lost. OCRmyPDF can produce a minimally PDF and many details may be lost. OCRmyPDF can produce a minimally
changed PDF as output. changed PDF as output.
OCRmyPDF also some image processing options like deskew which improve OCRmyPDF also provides some image processing options, like deskew, which
the appearance of files and quality of OCR. When these are used, the OCR improves the appearance of files and quality of OCR. When these are used,
layer is grafted onto the processed image instead. the OCR layer is grafted onto the processed image instead.
By default, OCRmyPDF produces archival PDFs PDF/A, which are a By default, OCRmyPDF produces archival PDFs PDF/A, which are a
stricter subset of PDF features designed for long term archives. If stricter subset of PDF features designed for long term archives. If
+3 -3
View File
@@ -9,11 +9,11 @@ encoding was patented for a long time. All known JBIG2 US patents have
expired as of 2017, but it is possible that unknown patents exist. expired as of 2017, but it is possible that unknown patents exist.
JBIG2 encoding is recommended for OCRmyPDF and is used to losslessly JBIG2 encoding is recommended for OCRmyPDF and is used to losslessly
create smaller PDFs. If JBIG2 encoding not available, lower quality create smaller PDFs. If JBIG2 encoding is not available, lower quality
encodings will be used. encodings will be used.
JBIG2 decoding is not patented and is performed automatically by most JBIG2 decoding is not patented and is performed automatically by most
PDF viewers. It is widely supported has been part of the PDF PDF viewers. It is widely supported and has been part of the PDF
specification since 2001. specification since 2001.
On macOS, Homebrew packages jbig2enc and OCRmyPDF includes it by On macOS, Homebrew packages jbig2enc and OCRmyPDF includes it by
@@ -37,7 +37,7 @@ Lossy mode JBIG2
OCRmyPDF provides lossy mode JBIG2 as an advanced feature. Users should OCRmyPDF provides lossy mode JBIG2 as an advanced feature. Users should
`review the technical concerns with JBIG2 in lossy `review the technical concerns with JBIG2 in lossy
mode <https://abbyy.technology/en:kb:tip:jbig2_compression_and_ocr>`__ mode <https://en.wikipedia.org/wiki/JBIG2#Disadvantages>`__
and decide if this feature is acceptable for their use case. and decide if this feature is acceptable for their use case.
JBIG2 lossy mode does achieve higher compression ratios than any other JBIG2 lossy mode does achieve higher compression ratios than any other
+58
View File
@@ -12,6 +12,64 @@ may be unreliable. Use the API to depend on precise behavior.
The public API may be useful in scripts that launch OCRmyPDF processes or that The public API may be useful in scripts that launch OCRmyPDF processes or that
wish to use some of its features for working with PDFs. wish to use some of its features for working with PDFs.
.. note::
Python 3.6 reaches end of life on December 23, 2021. We will end support
for Python 3.6 around that time. The change will be marked with a major
release.
v12.7.0
=======
- Fixed test suite failure when using pikepdf 3.2.0 that was compiled with pybind11
2.8.0. :issue:`843`
- Improve advice to user about using ``--max-image-mpixels`` if OCR fails for this
reason.
- Minor documentation fixes. (Thanks to @mara004.)
- Don't require importlib-metadata and importlib-resources backports on versions of
Python where the standard library implementation is sufficient.
(Thanks to Marco Genasci.)
v12.6.0
=======
- Implemented ``--output-type=none`` to skip producing PDFs for applications that
only want sidecar files (:issue:`787`).
- Fixed ambiguities in descriptions of behavior of ``--jbig2-lossy``.
- Various improvements to documentation.
v12.5.0
=======
- Fixed build failure for the combination of PyPy 3.6 and pikepdf 3.0. This
combination can work in a source build but does not work with wheels.
- Accepted bot that wanted to upgrade our deprecated requirements.txt.
- Documentation updates.
- Replace pkg_resources and install dependency on setuptools with
importlib-metadata and importlib-resources.
- Fixed regression in hocrtransform causing text to be omitted when this
renderer was used.
- Fixed some typing errors.
v12.4.0
=======
- When grafting text layers, use pikepdf's ``unparse_content_stream`` if available.
- Confirmed support for pluggy 1.0. (Thanks @QuLogic.)
- Fixed some typing issues, improved pre-commit settings, and fixed issues
flagged by linters.
- PyPy 7.3.3 (=Python 3.6) is now supported. Note that PyPy does not necessarily
run faster, because the vast majority of OCRmyPDF's execution time is spent
running OCR or generally executing native code. However, PyPy may bring speed
improvements in some areas.
v12.3.3
=======
- watcher.py: fixed interpretation of boolean env vars (:issue:`821`).
- Adjust CI scripts to test Tesseract 5 betas.
- Document our support for the Tesseract 5 betas.
v12.3.2 v12.3.2
======= =======
+1 -1
View File
@@ -52,7 +52,7 @@ logging.basicConfig(
ocrmypdf.configure_logging(ocrmypdf.Verbosity.default) ocrmypdf.configure_logging(ocrmypdf.Verbosity.default)
for dir_name, subdirs, file_list in os.walk(start_dir): for dir_name, _subdirs, file_list in os.walk(start_dir):
logging.info(dir_name + '\n') logging.info(dir_name + '\n')
os.chdir(dir_name) os.chdir(dir_name)
for filename in file_list: for filename in file_list:
+1
View File
@@ -54,6 +54,7 @@ function __fish_ocrmypdf_output_type
echo -e "pdfa-1\t"(_ "output a PDF/A-1b") echo -e "pdfa-1\t"(_ "output a PDF/A-1b")
echo -e "pdfa-2\t"(_ "output a PDF/A-2b") echo -e "pdfa-2\t"(_ "output a PDF/A-2b")
echo -e "pdfa-3\t"(_ "output a PDF/A-3b") echo -e "pdfa-3\t"(_ "output a PDF/A-3b")
echo -e "none\t"(_ "do not produce an output PDF (for example, if you only care about --sidecar)")
end end
complete -c ocrmypdf -x -l output-type -a '(__fish_ocrmypdf_output_type)' -d "select PDF output options" complete -c ocrmypdf -x -l output-type -a '(__fish_ocrmypdf_output_type)' -d "select PDF output options"
+1 -1
View File
@@ -46,7 +46,7 @@ if len(sys.argv) > 1:
else: else:
start_dir = '.' start_dir = '.'
for dir_name, subdirs, file_list in os.walk(start_dir): for dir_name, _subdirs, file_list in os.walk(start_dir):
logging.info(dir_name) logging.info(dir_name)
os.chdir(dir_name) os.chdir(dir_name)
for filename in file_list: for filename in file_list:
+9 -4
View File
@@ -37,14 +37,19 @@ import ocrmypdf
# pylint: disable=logging-format-interpolation # pylint: disable=logging-format-interpolation
def getenv_bool(name: str, default: str = 'False'):
return os.getenv(name, default).lower() in ('true', 'yes', 'y', '1')
INPUT_DIRECTORY = os.getenv('OCR_INPUT_DIRECTORY', '/input') INPUT_DIRECTORY = os.getenv('OCR_INPUT_DIRECTORY', '/input')
OUTPUT_DIRECTORY = os.getenv('OCR_OUTPUT_DIRECTORY', '/output') OUTPUT_DIRECTORY = os.getenv('OCR_OUTPUT_DIRECTORY', '/output')
OUTPUT_DIRECTORY_YEAR_MONTH = bool(os.getenv('OCR_OUTPUT_DIRECTORY_YEAR_MONTH', '')) OUTPUT_DIRECTORY_YEAR_MONTH = getenv_bool('OCR_OUTPUT_DIRECTORY_YEAR_MONTH')
ON_SUCCESS_DELETE = bool(os.getenv('OCR_ON_SUCCESS_DELETE', '')) ON_SUCCESS_DELETE = getenv_bool('OCR_ON_SUCCESS_DELETE')
DESKEW = bool(os.getenv('OCR_DESKEW', '')) DESKEW = getenv_bool('OCR_DESKEW')
OCR_JSON_SETTINGS = json.loads(os.getenv('OCR_JSON_SETTINGS', '{}')) OCR_JSON_SETTINGS = json.loads(os.getenv('OCR_JSON_SETTINGS', '{}'))
POLL_NEW_FILE_SECONDS = int(os.getenv('OCR_POLL_NEW_FILE_SECONDS', '1')) POLL_NEW_FILE_SECONDS = int(os.getenv('OCR_POLL_NEW_FILE_SECONDS', '1'))
USE_POLLING = bool(os.getenv('OCR_USE_POLLING', '')) USE_POLLING = getenv_bool('OCR_USE_POLLING')
LOGLEVEL = os.getenv('OCR_LOGLEVEL', 'INFO') LOGLEVEL = os.getenv('OCR_LOGLEVEL', 'INFO')
PATTERNS = ['*.pdf', '*.PDF'] PATTERNS = ['*.pdf', '*.PDF']
+1 -1
View File
@@ -37,7 +37,7 @@ app.secret_key = "secret"
app.config['MAX_CONTENT_LENGTH'] = 50_000_000 app.config['MAX_CONTENT_LENGTH'] = 50_000_000
app.config.from_envvar("OCRMYPDF_WEBSERVICE_SETTINGS", silent=True) app.config.from_envvar("OCRMYPDF_WEBSERVICE_SETTINGS", silent=True)
ALLOWED_EXTENSIONS = set(["pdf"]) ALLOWED_EXTENSIONS = {"pdf"}
def allowed_file(filename): def allowed_file(filename):
+46
View File
@@ -34,3 +34,49 @@ exclude = '''
| src/ocrmypdf/lib/_leptonica.py | src/ocrmypdf/lib/_leptonica.py
)/ )/
''' '''
[tool.coverage.run]
branch = true
parallel = true
concurrency = ["multiprocessing"]
[tool.coverage.paths]
source = ["src/ocrmypdf"]
[tool.coverage.report]
# Regexes for lines to exclude from consideration
exclude_lines = [
# Have to re-enable the standard pragma
"pragma: no cover",
# Don't complain if tests don't hit defensive assertion code:
"raise AssertionError",
"raise NotImplementedError",
# Don't complain if non-runnable code isn't run:
"if 0:",
"if False:",
"if __name__ == .__main__.:",
"if TYPE_CHECKING:"
]
[tool.isort]
profile = "black"
known_first_party = "ocrmypdf"
known_third_party = ["PIL", "_cffi_backend", "cffi", "flask", "img2pdf", "ocrmypdf", "pdfminer", "pikepdf", "pkg_resources", "pluggy", "pytest", "reportlab", "setuptools", "sphinx_rtd_theme", "tqdm", "watchdog", "werkzeug"]
[tool.pytest.ini_options]
minversion = "6.0"
norecursedirs = ["lib", ".pc", ".git", "venv", "output", "cache", "resources"]
testpaths = ["tests"]
addopts = "-n auto"
markers = ["slow"]
filterwarnings = ["ignore:.*XMLParser.*:DeprecationWarning"]
[tool.mypy]
[[tool.mypy.overrides]]
module = [
'pluggy', 'tqdm', 'coloredlogs', 'img2pdf', 'cffi', '_cffi_backend', 'pdfminer.*', 'reportlab.*'
]
ignore_missing_imports = true
+1 -1
View File
@@ -5,6 +5,6 @@ img2pdf == 0.4.0
pdfminer.six == 20201018 pdfminer.six == 20201018
pikepdf == 2.10.0 pikepdf == 2.10.0
pluggy == 0.13.1 pluggy == 0.13.1
Pillow == 8.2.0 Pillow == 8.3.2
reportlab == 3.5.66 reportlab == 3.5.66
tqdm == 4.59.0 tqdm == 4.59.0
+66 -98
View File
@@ -2,23 +2,15 @@
name = ocrmypdf name = ocrmypdf
description = OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched description = OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
long_description = file: README.md long_description = file: README.md
long_description_content_type = text/markdown; charset=UTF-8 long_description_content_type = text/markdown
url = https://github.com/jbarlow83/OCRmyPDF url = https://github.com/jbarlow83/OCRmyPDF
author = James R. Barlow author = James R. Barlow
author_email = james@purplerock.ca author_email = james@purplerock.ca
license = MPL-2.0
license_file = LICENSE
license_files = license_files =
LICENSE LICENSE
keywords =
PDF
OCR
optical character recognition
PDF/A
scanning
classifiers = classifiers =
Programming Language :: Python :: 3.6
Programming Language :: Python :: 3.7
Programming Language :: Python :: 3.8
Programming Language :: Python :: 3.9
Development Status :: 5 - Production/Stable Development Status :: 5 - Production/Stable
Environment :: Console Environment :: Console
Intended Audience :: End Users/Desktop Intended Audience :: End Users/Desktop
@@ -30,68 +22,82 @@ classifiers =
Operating System :: POSIX Operating System :: POSIX
Operating System :: POSIX :: BSD Operating System :: POSIX :: BSD
Operating System :: POSIX :: Linux Operating System :: POSIX :: Linux
Programming Language :: Python :: 3
Programming Language :: Python :: 3 :: Only
Programming Language :: Python :: 3.6
Programming Language :: Python :: 3.7
Programming Language :: Python :: 3.8
Programming Language :: Python :: 3.9
Topic :: Scientific/Engineering :: Image Recognition Topic :: Scientific/Engineering :: Image Recognition
Topic :: Text Processing :: Indexing Topic :: Text Processing :: Indexing
Topic :: Text Processing :: Linguistic Topic :: Text Processing :: Linguistic
keywords =
PDF
OCR
optical character recognition
PDF/A
scanning
project_urls = project_urls =
Documentation = https://ocrmypdf.readthedocs.io/ Documentation = https://ocrmypdf.readthedocs.io/
Source = https://github.com/jbarlow83/ocrmypdf Source = https://github.com/jbarlow83/ocrmypdf
Tracker = https://github.com/jbarlow83/ocrmypdf/issues Tracker = https://github.com/jbarlow83/ocrmypdf/issues
[options] [options]
zip_safe = False
packages = find: packages = find:
install_requires =
Pillow>=8.2.0
cffi>=1.9.1 # must be a setup and install requirement
coloredlogs>=14.0 # strictly optional
img2pdf>=0.3.0,<0.5 # pure Python
importlib-metadata>=4;python_version<'3.8' # until Python 3.8
importlib-resources>=5;python_version<'3.9' # until Python 3.9
pdfminer.six!=20200720,>=20191110,<=20201018
pikepdf>=2.10.0
pikepdf<3;implementation_name=="pypy" and python_version=='3.6'
pluggy>=0.13.0,<2
reportlab>=3.5.66
tqdm>=4
python_requires = >=3.6
include_package_data = True
package_dir = package_dir =
=src =src
platforms = any platforms = any
include_package_data=True setup_requires =
install_requires = cffi>=1.9.1 # to build the leptonica module
cffi >= 1.9.1 # must be a setup and install requirement setuptools_scm
coloredlogs >= 14.0 # strictly optional setuptools_scm_git_archive
img2pdf >= 0.3.0, < 0.5 # pure Python, so track HEAD closely zip_safe = False
pdfminer.six >= 20191110, != 20200720, <= 20201018
pikepdf >= 2.10.0 [options.packages.find]
Pillow >= 8.2.0 where = src
pluggy >= 0.13.0, < 1.0
reportlab >= 3.5.66 [options.entry_points]
setuptools console_scripts =
tqdm >= 4 ocrmypdf = ocrmypdf.__main__:run
python_requires = >= 3.6
setup_requires = # can be removed whenever we can drop pip 9 support [options.extras_require]
cffi >= 1.9.1 # to build the leptonica module docs =
setuptools_scm # so that version will work sphinx
setuptools_scm_git_archive # enable version from github tarballs sphinx-issues
sphinx-rtd-theme
extended_test =
PyMuPDF==1.13.4
test =
coverage[toml]>=5
pytest>=6.0.0
pytest-cov>=2.11.1
pytest-xdist>=2.2.0
python-xmp-toolkit==2.0.1 # also requires apt-get install libexempi3
watcher =
watchdog>=1.0.2,<3
webservice =
Flask>=1,<3
[options.package_data] [options.package_data]
ocrmypdf = ocrmypdf =
data/sRGB.icc data/sRGB.icc
py.typed py.typed
[options.packages.find]
where = src
[options.extras_require]
test =
pytest >= 6.0.0
pytest-xdist >= 2.2.0
pytest-cov >= 2.11.1
python-xmp-toolkit == 2.0.1 # also requires apt-get install libexempi3
# or brew install exempi
docs =
sphinx
sphinx-rtd-theme
sphinx-issues
extended_test =
PyMuPDF == 1.13.4
watcher =
watchdog >= 1.0.2, < 3
webservice =
Flask >= 1, < 3
[options.entry_points]
console_scripts =
ocrmypdf = ocrmypdf.__main__:run
[bdist_wheel] [bdist_wheel]
python-tag = py36 python-tag = py36
@@ -100,48 +106,10 @@ test = pytest
[check-manifest] [check-manifest]
ignore = ignore =
.github .github
[tool:pytest] [flake8]
norecursedirs = lib .pc .git output cache resources ignore = D203,F401,W503,E501,E203,F841
testpaths = tests exclude = .git,__pycache__,docs/conf.py,build,dist,.venv,.venvpp,.eggs,tmp,src/ocrmypdf/lib/
filterwarnings = max-complexity = 10
ignore:.*XMLParser.*:DeprecationWarning max-line-length = 100
markers =
slow
addopts =
-n auto
[isort]
multi_line_output = 3
include_trailing_comma = True
force_grid_wrap = 0
use_parentheses = True
line_length = 88
known_first_party = ocrmypdf
known_third_party = PIL,_cffi_backend,cffi,flask,img2pdf,pdfminer,pikepdf,pkg_resources,pluggy,pytest,reportlab,setuptools,sphinx_rtd_theme,tqdm,watchdog,werkzeug
[coverage:paths]
source =
src/ocrmypdf
[coverage:run]
branch = true
parallel = true
concurrency = multiprocessing
[coverage:report]
# Regexes for lines to exclude from consideration
exclude_lines =
# Have to re-enable the standard pragma
pragma: no cover
# Don't complain if tests don't hit defensive assertion code:
raise AssertionError
raise NotImplementedError
# Don't complain if non-runnable code isn't run:
if 0:
if False:
if __name__ == .__main__.:
if TYPE_CHECKING:
+5 -3
View File
@@ -117,7 +117,7 @@ def get_languages():
if line.startswith('Error'): if line.startswith('Error'):
raise MissingDependencyError(lang_error(output)) raise MissingDependencyError(lang_error(output))
_header, *rest = output.splitlines() _header, *rest = output.splitlines()
return set(lang.strip() for lang in rest) return {lang.strip() for lang in rest}
def tess_base_args(langs: List[str], engine_mode: Optional[int]) -> List[str]: def tess_base_args(langs: List[str], engine_mode: Optional[int]) -> List[str]:
@@ -250,7 +250,8 @@ def generate_hocr(
# Reminder: test suite tesseract test plugins will break after any changes # Reminder: test suite tesseract test plugins will break after any changes
# to the number of order parameters here # to the number of order parameters here
args_tesseract.extend([input_file, prefix, 'hocr', 'txt'] + tessconfig) args_tesseract.extend([os.fspath(input_file), os.fspath(prefix), 'hocr', 'txt'])
args_tesseract.extend(tessconfig)
try: try:
p = run(args_tesseract, stdout=PIPE, stderr=STDOUT, timeout=timeout, check=True) p = run(args_tesseract, stdout=PIPE, stderr=STDOUT, timeout=timeout, check=True)
stdout = p.stdout stdout = p.stdout
@@ -324,7 +325,8 @@ def generate_pdf(
# Reminder: test suite tesseract test plugins might break after any changes # Reminder: test suite tesseract test plugins might break after any changes
# to the number of order parameters here # to the number of order parameters here
args_tesseract.extend([input_file, prefix, 'pdf', 'txt'] + tessconfig) args_tesseract.extend([os.fspath(input_file), os.fspath(prefix), 'pdf', 'txt'])
args_tesseract.extend(tessconfig)
try: try:
p = run(args_tesseract, stdout=PIPE, stderr=STDOUT, timeout=timeout, check=True) p = run(args_tesseract, stdout=PIPE, stderr=STDOUT, timeout=timeout, check=True)
stdout = p.stdout stdout = p.stdout
+3 -3
View File
@@ -45,7 +45,7 @@ def _setup_unpaper_io(tmpdir: Path, input_file: Path) -> Tuple[Path, Path]:
im = im.convert(mode='1') im = im.convert(mode='1')
else: else:
im = im.convert(mode='RGB') im = im.convert(mode='RGB')
except IOError as e: except OSError as e:
raise MissingDependencyError( raise MissingDependencyError(
"Could not convert image with type " + im.mode "Could not convert image with type " + im.mode
) from e ) from e
@@ -96,12 +96,12 @@ def run(
try: try:
with Image.open(output_pnm) as imout: with Image.open(output_pnm) as imout:
imout.save(output_file, dpi=(dpi, dpi)) imout.save(output_file, dpi=(dpi, dpi))
except (FileNotFoundError, OSError): except OSError as e:
raise SubprocessOutputError( raise SubprocessOutputError(
"unpaper: failed to produce the expected output file. " "unpaper: failed to produce the expected output file. "
+ " Called with: " + " Called with: "
+ str(args_unpaper) + str(args_unpaper)
) from None ) from e
def validate_custom_args(args: str) -> List[str]: def validate_custom_args(args: str) -> List[str]:
+38 -43
View File
@@ -11,8 +11,19 @@ from contextlib import suppress
from pathlib import Path from pathlib import Path
from typing import Optional from typing import Optional
import pikepdf from pikepdf import (
from pikepdf.objects import Dictionary, Name Dictionary,
Name,
Object,
Operator,
Page,
Pdf,
PdfError,
PdfMatrix,
Stream,
parse_content_stream,
unparse_content_stream,
)
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
MAX_REPLACE_PAGES = 100 MAX_REPLACE_PAGES = 100
@@ -47,44 +58,28 @@ def strip_invisible_text(pdf, page):
render_mode = 0 render_mode = 0
text_objects = [] text_objects = []
rich_page = pikepdf.Page(page) rich_page = Page(page)
rich_page.contents_coalesce() rich_page.contents_coalesce()
for operands, operator in pikepdf.parse_content_stream(page, ''): for operands, operator in parse_content_stream(page, ''):
if not in_text_obj: if not in_text_obj:
if operator == pikepdf.Operator('BT'): if operator == Operator('BT'):
in_text_obj = True in_text_obj = True
render_mode = 0 render_mode = 0
text_objects.append((operands, operator)) text_objects.append((operands, operator))
else: else:
stream.append((operands, operator)) stream.append((operands, operator))
else: else:
if operator == pikepdf.Operator('Tr'): if operator == Operator('Tr'):
render_mode = operands[0] render_mode = operands[0]
text_objects.append((operands, operator)) text_objects.append((operands, operator))
if operator == pikepdf.Operator('ET'): if operator == Operator('ET'):
in_text_obj = False in_text_obj = False
if render_mode != 3: if render_mode != 3:
stream.extend(text_objects) stream.extend(text_objects)
text_objects.clear() text_objects.clear()
def convert(op): content_stream = unparse_content_stream(stream)
try: page.Contents = Stream(pdf, content_stream)
return op.unparse()
except AttributeError:
return str(op).encode('ascii')
lines = []
for operands, operator in stream:
if operator == pikepdf.Operator('INLINE IMAGE'):
iim = operands[0]
line = iim.unparse()
else:
line = b' '.join(convert(op) for op in operands) + b' ' + operator.unparse()
lines.append(line)
content_stream = b'\n'.join(lines)
page.Contents = pikepdf.Stream(pdf, content_stream)
class OcrGrafter: class OcrGrafter:
@@ -92,14 +87,14 @@ class OcrGrafter:
self.context = context self.context = context
self.path_base = context.origin self.path_base = context.origin
self.pdf_base = pikepdf.open(self.path_base) self.pdf_base = Pdf.open(self.path_base)
self.font, self.font_key = None, None self.font, self.font_key = None, None
self.pdfinfo = context.pdfinfo self.pdfinfo = context.pdfinfo
self.output_file = context.get_path('graft_layers.pdf') self.output_file = context.get_path('graft_layers.pdf')
self.procset = self.pdf_base.make_indirect( self.procset = self.pdf_base.make_indirect(
pikepdf.Object.parse(b'[ /PDF /Text /ImageB /ImageC /ImageI ]') Object.parse(b'[ /PDF /Text /ImageB /ImageC /ImageI ]')
) )
self.emplacements = 1 self.emplacements = 1
@@ -123,7 +118,7 @@ class OcrGrafter:
# We are updating the old page with a rasterized PDF of the new # We are updating the old page with a rasterized PDF of the new
# page (without changing objgen, to preserve references) # page (without changing objgen, to preserve references)
log.debug("Emplacement update") log.debug("Emplacement update")
with pikepdf.open(image) as pdf_image: with Pdf.open(path_image) as pdf_image:
self.emplacements += 1 self.emplacements += 1
foreign_image_page = pdf_image.pages[0] foreign_image_page = pdf_image.pages[0]
self.pdf_base.pages.append(foreign_image_page) self.pdf_base.pages.append(foreign_image_page)
@@ -196,7 +191,7 @@ class OcrGrafter:
self.pdf_base.save(next_file) self.pdf_base.save(next_file)
self.pdf_base.close() self.pdf_base.close()
self.pdf_base = pikepdf.open(next_file) self.pdf_base = Pdf.open(next_file)
self.procset = self.pdf_base.pages[0].Resources.ProcSet self.procset = self.pdf_base.pages[0].Resources.ProcSet
self.font, self.font_key = None, None # Ensure we reacquire this information self.font, self.font_key = None, None # Ensure we reacquire this information
self.interim_count += 1 self.interim_count += 1
@@ -212,7 +207,7 @@ class OcrGrafter:
font, font_key = None, None font, font_key = None, None
possible_font_names = ('/f-0-0', '/F1') possible_font_names = ('/f-0-0', '/F1')
try: try:
with pikepdf.open(text) as pdf_text: with Pdf.open(text) as pdf_text:
try: try:
pdf_text_fonts = pdf_text.pages[0].Resources.get('/Font', {}) pdf_text_fonts = pdf_text.pages[0].Resources.get('/Font', {})
except (AttributeError, IndexError, KeyError): except (AttributeError, IndexError, KeyError):
@@ -226,7 +221,7 @@ class OcrGrafter:
if pdf_text_font: if pdf_text_font:
font = self.pdf_base.copy_foreign(pdf_text_font) font = self.pdf_base.copy_foreign(pdf_text_font)
return font, font_key return font, font_key
except (FileNotFoundError, pikepdf.PdfError): except (FileNotFoundError, PdfError):
# PdfError occurs if a 0-length file is written e.g. due to OCR timeout # PdfError occurs if a 0-length file is written e.g. due to OCR timeout
return None, None return None, None
@@ -235,9 +230,9 @@ class OcrGrafter:
*, *,
page_num: int, page_num: int,
textpdf: Path, textpdf: Path,
font: pikepdf.Object, font: Object,
font_key: pikepdf.Object, font_key: Object,
procset: pikepdf.Object, procset: Object,
text_rotation: int, text_rotation: int,
strip_old_text: bool, strip_old_text: bool,
): ):
@@ -248,7 +243,7 @@ class OcrGrafter:
return return
# This is a pointer indicating a specific page in the base file # This is a pointer indicating a specific page in the base file
with pikepdf.open(textpdf) as pdf_text: with Pdf.open(textpdf) as pdf_text:
pdf_text_contents = pdf_text.pages[0].Contents.read_bytes() pdf_text_contents = pdf_text.pages[0].Contents.read_bytes()
base_page = self.pdf_base.pages.p(page_num) base_page = self.pdf_base.pages.p(page_num)
@@ -263,13 +258,13 @@ class OcrGrafter:
mediabox = [float(base_page.MediaBox[v]) for v in range(4)] mediabox = [float(base_page.MediaBox[v]) for v in range(4)]
wp, hp = mediabox[2] - mediabox[0], mediabox[3] - mediabox[1] wp, hp = mediabox[2] - mediabox[0], mediabox[3] - mediabox[1]
translate = pikepdf.PdfMatrix().translated(-wt / 2, -ht / 2) translate = PdfMatrix().translated(-wt / 2, -ht / 2)
untranslate = pikepdf.PdfMatrix().translated(wp / 2, hp / 2) untranslate = PdfMatrix().translated(wp / 2, hp / 2)
corner = pikepdf.PdfMatrix().translated(mediabox[0], mediabox[1]) corner = PdfMatrix().translated(mediabox[0], mediabox[1])
# -rotation because the input is a clockwise angle and this formula # -rotation because the input is a clockwise angle and this formula
# uses CCW # uses CCW
text_rotation = -text_rotation % 360 text_rotation = -text_rotation % 360
rotate = pikepdf.PdfMatrix().rotated(text_rotation) rotate = PdfMatrix().rotated(text_rotation)
# Because of rounding of DPI, we might get a text layer that is not # Because of rounding of DPI, we might get a text layer that is not
# identically sized to the target page. Scale to adjust. Normally this # identically sized to the target page. Scale to adjust. Normally this
@@ -280,7 +275,7 @@ class OcrGrafter:
scale_y = hp / ht scale_y = hp / ht
# log.debug('%r', scale_x, scale_y) # log.debug('%r', scale_x, scale_y)
scale = pikepdf.PdfMatrix().scaled(scale_x, scale_y) scale = PdfMatrix().scaled(scale_x, scale_y)
# Translate the text so it is centered at (0, 0), rotate it there, adjust # Translate the text so it is centered at (0, 0), rotate it there, adjust
# for a size different between initial and text PDF, then untranslate, and # for a size different between initial and text PDF, then untranslate, and
@@ -303,14 +298,14 @@ class OcrGrafter:
pdf_draw_xobj = ( pdf_draw_xobj = (
(b'q %s cm\n' % ctm.encode()) + (b'%s Do\n' % text_xobj_name) + b'\nQ\n' (b'q %s cm\n' % ctm.encode()) + (b'%s Do\n' % text_xobj_name) + b'\nQ\n'
) )
new_text_layer = pikepdf.Stream(self.pdf_base, pdf_draw_xobj) new_text_layer = Stream(self.pdf_base, pdf_draw_xobj)
if strip_old_text: if strip_old_text:
strip_invisible_text(self.pdf_base, base_page) strip_invisible_text(self.pdf_base, base_page)
if hasattr(pikepdf.Page, 'contents_add'): if hasattr(Page, 'contents_add'):
# pikepdf >= 2.14 adds this method and deprecates the one below # pikepdf >= 2.14 adds this method and deprecates the one below
pikepdf.Page(base_page).contents_add(new_text_layer, prepend=True) Page(base_page).contents_add(new_text_layer, prepend=True)
else: else:
# pikepdf < 2.14 # pikepdf < 2.14
base_page.page_contents_add( base_page.page_contents_add(
+5 -6
View File
@@ -48,7 +48,7 @@ def triage_image_file(input_file, output_file, options):
log.info("Input file is not a PDF, checking if it is an image...") log.info("Input file is not a PDF, checking if it is an image...")
try: try:
im = Image.open(input_file) im = Image.open(input_file)
except EnvironmentError as e: except OSError as e:
# Recover the original filename # Recover the original filename
log.error(str(e).replace(str(input_file), str(options.input_file))) log.error(str(e).replace(str(input_file), str(options.input_file)))
raise UnsupportedImageFormatError() from e raise UnsupportedImageFormatError() from e
@@ -135,7 +135,7 @@ def triage(original_filename, input_file, output_file, options):
# Origin file is a pdf create a symlink with pdf extension # Origin file is a pdf create a symlink with pdf extension
safe_symlink(input_file, output_file) safe_symlink(input_file, output_file)
return output_file return output_file
except EnvironmentError as e: except OSError as e:
log.debug(f"Temporary file was at: {input_file}") log.debug(f"Temporary file was at: {input_file}")
msg = str(e).replace(str(input_file), original_filename) msg = str(e).replace(str(input_file), original_filename)
raise InputFileError(msg) from e raise InputFileError(msg) from e
@@ -521,13 +521,12 @@ def create_ocr_image(image: Path, page_context: PageContext):
# be None) # be None)
bbox = [float(v) for v in textarea] bbox = [float(v) for v in textarea]
xyscale = tuple(float(coord) / 72.0 for coord in im.info['dpi']) xyscale = tuple(float(coord) / 72.0 for coord in im.info['dpi'])
pixcoords = [ pixcoords = (
bbox[0] * xyscale[0], bbox[0] * xyscale[0],
im.height - bbox[3] * xyscale[1], im.height - bbox[3] * xyscale[1],
bbox[2] * xyscale[0], bbox[2] * xyscale[0],
im.height - bbox[1] * xyscale[1], im.height - bbox[1] * xyscale[1],
] )
pixcoords = [int(round(c)) for c in pixcoords]
log.debug('blanking %r', pixcoords) log.debug('blanking %r', pixcoords)
draw.rectangle(pixcoords, fill=white) draw.rectangle(pixcoords, fill=white)
# draw.rectangle(pixcoords, outline=pink) # draw.rectangle(pixcoords, outline=pink)
@@ -856,7 +855,7 @@ def merge_sidecars(txt_files: Iterable[Optional[Path]], context: PdfContext):
if frm != 1: if frm != 1:
stream.write('\f') # Form feed between pages stream.write('\f') # Form feed between pages
if txt_file: if txt_file:
with open(txt_file, 'r', encoding="utf-8") as in_: with open(txt_file, encoding="utf-8") as in_:
txt = in_.read() txt = in_.read()
# Some OCR engines (e.g. Tesseract v4 alpha) add form feeds # Some OCR engines (e.g. Tesseract v4 alpha) add form feeds
# between pages, and some do not. For consistency, we ignore # between pages, and some do not. For consistency, we ignore
+18 -10
View File
@@ -290,15 +290,16 @@ def exec_concurrent(context: PdfContext, executor: Executor):
# Copy text file to destination # Copy text file to destination
copy_final(text, options.sidecar, context) copy_final(text, options.sidecar, context)
# Merge layers to one single pdf if options.output_type != 'none':
pdf = ocrgraft.finalize() # Merge layers to one single pdf
pdf = ocrgraft.finalize()
# PDF/A and metadata # PDF/A and metadata
log.info("Postprocessing...") log.info("Postprocessing...")
pdf = post_process(pdf, context, executor) pdf = post_process(pdf, context, executor)
# Copy PDF file to destination # Copy PDF file to destination
copy_final(pdf, options.output_file, context) copy_final(pdf, options.output_file, context)
def configure_debug_logging(log_filename: Path, prefix: str = ''): def configure_debug_logging(log_filename: Path, prefix: str = ''):
@@ -399,7 +400,7 @@ def run_pipeline(options, *, plugin_manager, api=False):
return ExitCode.invalid_output_pdf return ExitCode.invalid_output_pdf
report_output_file_size(options, start_input_file, options.output_file) report_output_file_size(options, start_input_file, options.output_file)
except (KeyboardInterrupt if not api else NeverRaise) as e: except (KeyboardInterrupt if not api else NeverRaise):
if options.verbose >= 1: if options.verbose >= 1:
log.exception("KeyboardInterrupt") log.exception("KeyboardInterrupt")
else: else:
@@ -413,7 +414,14 @@ def run_pipeline(options, *, plugin_manager, api=False):
else: else:
log.error(type(e).__name__) log.error(type(e).__name__)
return e.exit_code return e.exit_code
except (Exception if not api else NeverRaise) as e: # pylint: disable=broad-except except (PIL.Image.DecompressionBombError if not api else NeverRaise) as e:
log.exception(
"A decompression bomb error was encountered while executing the "
"pipeline. Use the argument --max-image-mpixels to raise the maximum "
"image pixel limit."
)
return ExitCode.other_error
except (Exception if not api else NeverRaise): # pylint: disable=broad-except
log.exception("An exception occurred while executing the pipeline") log.exception("An exception occurred while executing the pipeline")
return ExitCode.other_error return ExitCode.other_error
finally: finally:
@@ -421,7 +429,7 @@ def run_pipeline(options, *, plugin_manager, api=False):
try: try:
debug_log_handler.close() debug_log_handler.close()
log.removeHandler(debug_log_handler) log.removeHandler(debug_log_handler)
except EnvironmentError as e: except OSError as e:
print(e, file=sys.stderr) print(e, file=sys.stderr)
cleanup_working_files(work_folder, options) cleanup_working_files(work_folder, options)
+19 -18
View File
@@ -13,7 +13,7 @@ import sys
import unicodedata import unicodedata
from pathlib import Path from pathlib import Path
from shutil import copyfileobj from shutil import copyfileobj
from typing import List, Set, Tuple, Union from typing import List, Set, Tuple
import pikepdf import pikepdf
import PIL import PIL
@@ -26,12 +26,7 @@ from ocrmypdf.exceptions import (
MissingDependencyError, MissingDependencyError,
OutputFileAccessError, OutputFileAccessError,
) )
from ocrmypdf.helpers import ( from ocrmypdf.helpers import is_file_writable, monotonic, safe_symlink, samefile
is_file_writable,
is_iterable_notstr,
monotonic,
safe_symlink,
)
from ocrmypdf.hocrtransform import HOCR_OK_LANGS from ocrmypdf.hocrtransform import HOCR_OK_LANGS
from ocrmypdf.subprocess import check_external_program from ocrmypdf.subprocess import check_external_program
@@ -68,7 +63,7 @@ def check_options_languages(options, ocr_engine_languages):
missing_languages = options.languages - ocr_engine_languages missing_languages = options.languages - ocr_engine_languages
if missing_languages: if missing_languages:
msg = ( msg = (
f"OCR engine does not have language data for the following " "OCR engine does not have language data for the following "
"requested languages: \n" "requested languages: \n"
) )
msg += '\n'.join(lang for lang in missing_languages) msg += '\n'.join(lang for lang in missing_languages)
@@ -80,12 +75,18 @@ def check_options_output(options):
is_latin = options.languages.issubset(HOCR_OK_LANGS) is_latin = options.languages.issubset(HOCR_OK_LANGS)
if options.pdf_renderer.startswith('hocr') and not is_latin: if options.pdf_renderer.startswith('hocr') and not is_latin:
msg = ( log.warning(
"The 'hocr' PDF renderer is known to cause problems with one " "The 'hocr' PDF renderer is known to cause problems with one "
"or more of the languages in your document. Use " "or more of the languages in your document. Use "
"--pdf-renderer auto (the default) to avoid this issue." "`--pdf-renderer auto` (the default) to avoid this issue."
)
if options.output_type == 'none' and options.output_file != os.devnull:
raise BadArgsError(
"Since you specified `--pdf-renderer none`, the output file "
f"{options.output_file} cannot be produced. Set the output file to "
f"{os.devnull} to suppress this message."
) )
log.warning(msg)
lossless_reconstruction = False lossless_reconstruction = False
if not any( if not any(
@@ -112,6 +113,10 @@ def check_options_sidecar(options):
raise BadArgsError( raise BadArgsError(
"--sidecar filename must be specified when output file is stdout." "--sidecar filename must be specified when output file is stdout."
) )
elif options.output_file == os.devnull:
raise BadArgsError(
"--sidecar filename must be specified when output file is /dev/null or NUL."
)
options.sidecar = options.output_file + '.txt' options.sidecar = options.output_file + '.txt'
if options.sidecar == options.input_file or options.sidecar == options.output_file: if options.sidecar == options.input_file or options.sidecar == options.output_file:
raise BadArgsError( raise BadArgsError(
@@ -142,8 +147,6 @@ def check_options_preprocessing(options):
def _pages_from_ranges(ranges: str) -> Set[int]: def _pages_from_ranges(ranges: str) -> Set[int]:
if is_iterable_notstr(ranges):
return set(ranges)
pages: List[int] = [] pages: List[int] = []
page_groups = ranges.replace(' ', '').split(',') page_groups = ranges.replace(' ', '').split(',')
for g in page_groups: for g in page_groups:
@@ -182,10 +185,8 @@ def _pages_from_ranges(ranges: str) -> Set[int]:
def check_options_ocr_behavior(options): def check_options_ocr_behavior(options):
exclusive_options = sum( exclusive_options = sum(
[ (1 if opt else 0)
(1 if opt else 0) for opt in (options.force_ocr, options.skip_text, options.redo_ocr)
for opt in (options.force_ocr, options.skip_text, options.redo_ocr)
]
) )
if exclusive_options >= 2: if exclusive_options >= 2:
raise BadArgsError("Choose only one of --force-ocr, --skip-text, --redo-ocr.") raise BadArgsError("Choose only one of --force-ocr, --skip-text, --redo-ocr.")
@@ -302,7 +303,7 @@ def check_closed_streams(options): # pragma: no cover
if options.input_file == '-': if options.input_file == '-':
log.error("Trying to read from stdin but stdin seems closed") log.error("Trying to read from stdin but stdin seems closed")
return False return False
sys.stdin = open(os.devnull, 'r') sys.stdin = open(os.devnull)
if sys.stdout is None: if sys.stdout is None:
if options.output_file == '-': if options.output_file == '-':
+5 -2
View File
@@ -5,9 +5,12 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
import pkg_resources try:
from importlib_metadata import version as _package_version
except ImportError:
from importlib.metadata import version as _package_version
PROGRAM_NAME = 'ocrmypdf' PROGRAM_NAME = 'ocrmypdf'
# Official PEP 396 # Official PEP 396
__version__ = pkg_resources.get_distribution('ocrmypdf').version __version__ = _package_version('ocrmypdf')
+16 -5
View File
@@ -15,10 +15,7 @@ from pathlib import Path
from typing import AnyStr, BinaryIO, Iterable, Optional, Union from typing import AnyStr, BinaryIO, Iterable, Optional, Union
from warnings import warn from warnings import warn
from ocrmypdf._logging import ( # pylint: disable=unused-import from ocrmypdf._logging import PageNumberFilter, TqdmConsole
PageNumberFilter,
TqdmConsole,
)
from ocrmypdf._plugin_manager import get_plugin_manager from ocrmypdf._plugin_manager import get_plugin_manager
from ocrmypdf._sync import run_pipeline from ocrmypdf._sync import run_pipeline
from ocrmypdf._validation import check_options from ocrmypdf._validation import check_options
@@ -31,7 +28,7 @@ except ModuleNotFoundError:
coloredlogs = None coloredlogs = None
StrPath = Union[os.PathLike, AnyStr] StrPath = Union[Path, AnyStr]
PathOrIO = Union[BinaryIO, StrPath] PathOrIO = Union[BinaryIO, StrPath]
_api_lock = threading.Lock() _api_lock = threading.Lock()
@@ -338,3 +335,17 @@ def ocr( # pylint: disable=unused-argument
options = create_options(**create_options_kwargs) options = create_options(**create_options_kwargs)
check_options(options, plugin_manager) check_options(options, plugin_manager)
return run_pipeline(options=options, plugin_manager=plugin_manager, api=True) return run_pipeline(options=options, plugin_manager=plugin_manager, api=True)
__all__ = [
'PageNumberFilter',
'TqdmConsole',
'Verbosity',
'check_options',
'configure_logging',
'create_options',
'get_parser',
'get_plugin_manager',
'ocr',
'run_pipeline',
]
+10 -8
View File
@@ -20,9 +20,8 @@ import signal
import sys import sys
import threading import threading
from contextlib import suppress from contextlib import suppress
from multiprocessing import Pool as ProcessPool from multiprocessing.pool import Pool, ThreadPool
from multiprocessing.pool import ThreadPool from typing import Callable, Iterable, Type, Union
from typing import Callable, Iterable, Union
from tqdm import tqdm from tqdm import tqdm
@@ -31,7 +30,10 @@ from ocrmypdf._logging import TqdmConsole
from ocrmypdf.exceptions import InputFileError from ocrmypdf.exceptions import InputFileError
from ocrmypdf.helpers import remove_all_log_handlers from ocrmypdf.helpers import remove_all_log_handlers
ProcessPool = Pool
Queue = Union[multiprocessing.Queue, queue.Queue] Queue = Union[multiprocessing.Queue, queue.Queue]
UserInit = Callable[[], None]
WorkerInit = Callable[[Queue, UserInit, int], None]
def log_listener(q: Queue): def log_listener(q: Queue):
@@ -62,7 +64,7 @@ def process_sigbus(*args):
raise InputFileError("A worker process lost access to an input file") raise InputFileError("A worker process lost access to an input file")
def process_init(q: Queue, user_init: Callable[[], None], loglevel): def process_init(q: Queue, user_init: UserInit, loglevel) -> None:
"""Initialize a process pool worker""" """Initialize a process pool worker"""
# Ignore SIGINT (our parent process will kill us gracefully) # Ignore SIGINT (our parent process will kill us gracefully)
@@ -85,7 +87,7 @@ def process_init(q: Queue, user_init: Callable[[], None], loglevel):
return return
def thread_init(_queue: Queue, user_init: Callable[[], None], _loglevel): def thread_init(q: Queue, user_init: UserInit, loglevel) -> None:
# As a thread, block SIGBUS so the main thread deals with it... # As a thread, block SIGBUS so the main thread deals with it...
with suppress(AttributeError): with suppress(AttributeError):
signal.pthread_sigmask(signal.SIG_BLOCK, {signal.SIGBUS}) signal.pthread_sigmask(signal.SIG_BLOCK, {signal.SIGBUS})
@@ -107,9 +109,9 @@ class StandardExecutor(Executor):
task_finished: Callable, task_finished: Callable,
): ):
if use_threads: if use_threads:
log_queue = queue.Queue(-1) log_queue: Queue = queue.Queue(-1)
pool_class = ThreadPool pool_class: Type[Pool] = ThreadPool
initializer = thread_init initializer: WorkerInit = thread_init
else: else:
log_queue = multiprocessing.Queue(-1) log_queue = multiprocessing.Queue(-1)
pool_class = ProcessPool pool_class = ProcessPool
@@ -11,7 +11,6 @@ import os
from ocrmypdf import hookimpl from ocrmypdf import hookimpl
from ocrmypdf._exec import tesseract from ocrmypdf._exec import tesseract
from ocrmypdf.cli import numeric from ocrmypdf.cli import numeric
from ocrmypdf.exceptions import MissingDependencyError
from ocrmypdf.helpers import clamp from ocrmypdf.helpers import clamp
from ocrmypdf.pluginspec import OcrEngine from ocrmypdf.pluginspec import OcrEngine
from ocrmypdf.subprocess import check_external_program from ocrmypdf.subprocess import check_external_program
+13 -8
View File
@@ -6,7 +6,7 @@
import argparse import argparse
from typing import Optional, Type, TypeVar from typing import Any, Callable, Optional, TypeVar
from ocrmypdf._version import PROGRAM_NAME as _PROGRAM_NAME from ocrmypdf._version import PROGRAM_NAME as _PROGRAM_NAME
from ocrmypdf._version import __version__ as _VERSION from ocrmypdf._version import __version__ as _VERSION
@@ -14,7 +14,9 @@ from ocrmypdf._version import __version__ as _VERSION
T = TypeVar('T') T = TypeVar('T')
def numeric(basetype: Type[T], min_: Optional[T] = None, max_: Optional[T] = None): def numeric(
basetype: Callable[[Any], T], min_: Optional[T] = None, max_: Optional[T] = None
):
"""Validator for numeric params""" """Validator for numeric params"""
min_ = basetype(min_) if min_ is not None else None min_ = basetype(min_) if min_ is not None else None
max_ = basetype(max_) if max_ is not None else None max_ = basetype(max_) if max_ is not None else None
@@ -22,7 +24,7 @@ def numeric(basetype: Type[T], min_: Optional[T] = None, max_: Optional[T] = Non
def _numeric(string): def _numeric(string):
value = basetype(string) value = basetype(string)
if (min_ is not None and value < min_) or (max_ is not None and value > max_): if (min_ is not None and value < min_) or (max_ is not None and value > max_):
msg = "%r not in valid range %r" % (string, (min_, max_)) msg = f"{string!r} not in valid range {(min_, max_)!r}"
raise argparse.ArgumentTypeError(msg) raise argparse.ArgumentTypeError(msg)
return value return value
@@ -145,7 +147,7 @@ Online documentation is located at:
) )
parser.add_argument( parser.add_argument(
'--output-type', '--output-type',
choices=['pdfa', 'pdf', 'pdfa-1', 'pdfa-2', 'pdfa-3'], choices=['pdfa', 'pdf', 'pdfa-1', 'pdfa-2', 'pdfa-3', 'none'],
default='pdfa', default='pdfa',
help="Choose output type. 'pdfa' creates a PDF/A-2b compliant file for " help="Choose output type. 'pdfa' creates a PDF/A-2b compliant file for "
"long term archiving (default, recommended) but may not suitable " "long term archiving (default, recommended) but may not suitable "
@@ -153,7 +155,8 @@ Online documentation is located at:
"also has problems with full Unicode text. 'pdf' attempts to " "also has problems with full Unicode text. 'pdf' attempts to "
"preserve file contents as much as possible. 'pdf-a1' creates a " "preserve file contents as much as possible. 'pdf-a1' creates a "
"PDF/A1-b file. 'pdf-a2' is equivalent to 'pdfa'. 'pdf-a3' creates a " "PDF/A1-b file. 'pdf-a2' is equivalent to 'pdfa'. 'pdf-a3' creates a "
"PDF/A3-b file.", "PDF/A3-b file. 'none' will produce no output, which may be helpful if "
"only the --sidecar is desired.",
) )
# Use null string '\0' as sentinel to indicate the user supplied no argument, # Use null string '\0' as sentinel to indicate the user supplied no argument,
@@ -338,8 +341,9 @@ Online documentation is located at:
"Control how PDF is optimized after processing:" "Control how PDF is optimized after processing:"
"0 - do not optimize; " "0 - do not optimize; "
"1 - do safe, lossless optimizations (default); " "1 - do safe, lossless optimizations (default); "
"2 - do some lossy optimizations; " "2 - do lossy JPEG and JPEG2000 optimizations; "
"3 - do aggressive lossy optimizations (including lossy JBIG2)" "3 - do more aggressive lossy JPEG and JPEG2000 optimizations. "
"To enable lossy JBIG2, see --jbig2-lossy."
), ),
) )
optimizing.add_argument( optimizing.add_argument(
@@ -377,7 +381,8 @@ Online documentation is located at:
action='store_true', action='store_true',
help=( help=(
"Enable JBIG2 lossy mode (better compression, not suitable for some " "Enable JBIG2 lossy mode (better compression, not suitable for some "
"use cases - see documentation)." "use cases - see documentation). Only takes effect if --optimize 1 or "
"higher is also enabled."
), ),
) )
optimizing.add_argument( optimizing.add_argument(
+8
View File
@@ -0,0 +1,8 @@
# © 2021 James R. Barlow: github.com/jbarlow83
#
# This Source Code Form is subject to the terms of the Mozilla Public
# License, v. 2.0. If a copy of the MPL was not distributed with this
# file, You can obtain one at http://mozilla.org/MPL/2.0/.
"""Data files used to generate certain PDFs."""
+8 -4
View File
@@ -28,7 +28,7 @@ from enum import Enum, auto
from itertools import islice, repeat, takewhile, zip_longest from itertools import islice, repeat, takewhile, zip_longest
from multiprocessing import Pipe, Process from multiprocessing import Pipe, Process
from multiprocessing.connection import Connection, wait from multiprocessing.connection import Connection, wait
from typing import Callable, Iterable, Iterator from typing import Callable, Iterable, Iterator, List
from ocrmypdf import Executor, hookimpl from ocrmypdf import Executor, hookimpl
from ocrmypdf._concurrent import NullProgressBar from ocrmypdf._concurrent import NullProgressBar
@@ -60,7 +60,9 @@ def process_sigbus(*args):
class ConnectionLogHandler(logging.handlers.QueueHandler): class ConnectionLogHandler(logging.handlers.QueueHandler):
def __init__(self, conn: Connection) -> None: def __init__(self, conn: Connection) -> None:
super().__init__(None) # sets the parent's queue to None - parent only touches queue
# in enqueue() which we override
super().__init__(None) # type: ignore
self.conn = conn self.conn = conn
def enqueue(self, record): def enqueue(self, record):
@@ -126,8 +128,8 @@ class LambdaExecutor(Executor):
if not grouped_args: if not grouped_args:
return return
processes = [] processes: List[Process] = []
connections = [] connections: List[Connection] = []
for chunk in grouped_args: for chunk in grouped_args:
parent_conn, child_conn = Pipe() parent_conn, child_conn = Pipe()
@@ -152,6 +154,8 @@ class LambdaExecutor(Executor):
with self.pbar_class(**tqdm_kwargs) as pbar: with self.pbar_class(**tqdm_kwargs) as pbar:
while connections: while connections:
for r in wait(connections): for r in wait(connections):
if not isinstance(r, Connection):
raise NotImplementedError("We only support Connection()")
try: try:
msg_type, msg = r.recv() msg_type, msg = r.recv()
except EOFError: except EOFError:
+3 -3
View File
@@ -189,7 +189,7 @@ def is_file_writable(test_file: os.PathLike) -> bool:
with suppress(OSError): with suppress(OSError):
p.unlink() p.unlink()
return True return True
except (EnvironmentError, RuntimeError) as e: except (OSError, RuntimeError) as e:
log.debug(e) log.debug(e)
log.error(str(e)) log.error(str(e))
return False return False
@@ -226,7 +226,7 @@ def check_pdf(input_file: Path) -> bool:
except ( except (
# Workaround for a problematic pikepdf version # Workaround for a problematic pikepdf version
# pragma: no cover # pragma: no cover
getattr(pikepdf, 'ForeignObjectError') pikepdf.ForeignObjectError
if pikepdf.__version__ == '2.1.0' if pikepdf.__version__ == '2.1.0'
else NeverRaise else NeverRaise
): ):
@@ -273,7 +273,7 @@ def deprecated(func):
def new_func(*args, **kwargs): def new_func(*args, **kwargs):
warnings.simplefilter('always', DeprecationWarning) # turn off filter warnings.simplefilter('always', DeprecationWarning) # turn off filter
warnings.warn( warnings.warn(
"Call to deprecated function {}.".format(func.__name__), f"Call to deprecated function {func.__name__}.",
category=DeprecationWarning, category=DeprecationWarning,
stacklevel=2, stacklevel=2,
) )
+1 -1
View File
@@ -349,7 +349,7 @@ class HocrTransform:
interword_spaces: bool, interword_spaces: bool,
show_bounding_boxes: bool, show_bounding_boxes: bool,
): ):
if not line: if line is None:
return return
pxl_line_coords = self.element_coordinates(line) pxl_line_coords = self.element_coordinates(line)
line_box = self.pt_from_pixel(pxl_line_coords) line_box = self.pt_from_pixel(pxl_line_coords)
+13 -9
View File
@@ -1,5 +1,4 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
# -*- coding: utf-8 -*-
# #
# © 2013-16: jbarlow83 from Github (https://github.com/jbarlow83) # © 2013-16: jbarlow83 from Github (https://github.com/jbarlow83)
# #
@@ -13,6 +12,7 @@
import argparse import argparse
import logging import logging
import os import os
import platform
import sys import sys
import threading import threading
from collections import deque from collections import deque
@@ -23,6 +23,7 @@ from functools import lru_cache
from io import BytesIO, UnsupportedOperation from io import BytesIO, UnsupportedOperation
from os import fspath from os import fspath
from tempfile import TemporaryFile from tempfile import TemporaryFile
from typing import ContextManager, Type
from warnings import warn from warnings import warn
from ocrmypdf.exceptions import MissingDependencyError from ocrmypdf.exceptions import MissingDependencyError
@@ -67,7 +68,7 @@ if os.name == 'nt':
# Loading zlib from other places could cause a version mismatch # Loading zlib from other places could cause a version mismatch
_zlib_path = os.path.join(os.path.dirname(_libpath), 'zlib1.dll') _zlib_path = os.path.join(os.path.dirname(_libpath), 'zlib1.dll')
if not os.path.exists(_zlib_path): if not os.path.exists(_zlib_path):
_zlib_path = find_library('zlib') _zlib_path = find_library('zlib') or ''
try: try:
zlib = ffi.dlopen(_zlib_path) zlib = ffi.dlopen(_zlib_path)
except ffi.error as e: except ffi.error as e:
@@ -86,7 +87,7 @@ except ffi.error as e:
) from e ) from e
class _LeptonicaErrorTrap_Redirect: class _LeptonicaErrorTrap_Redirect(ContextManager):
""" """
Context manager to trap errors reported by Leptonica < 1.79 or on Apple Silicon. Context manager to trap errors reported by Leptonica < 1.79 or on Apple Silicon.
@@ -132,7 +133,7 @@ class _LeptonicaErrorTrap_Redirect:
except Exception: except Exception:
self.leptonica_lock.release() self.leptonica_lock.release()
raise raise
return self return
def __exit__(self, exc_type, exc_value, traceback): def __exit__(self, exc_type, exc_value, traceback):
# Restore old stderr # Restore old stderr
@@ -172,7 +173,7 @@ tls = threading.local()
tls.trap = None tls.trap = None
class _LeptonicaErrorTrap_Queue: class _LeptonicaErrorTrap_Queue(ContextManager):
def __init__(self): def __init__(self):
self.queue = deque() self.queue = deque()
@@ -226,7 +227,7 @@ except (ffi.error, MemoryError):
# Pre-1.79 Leptonica does not have leptSetStderrHandler # Pre-1.79 Leptonica does not have leptSetStderrHandler
# And some platforms, notably Apple ARM 64, do not allow the write+execute # And some platforms, notably Apple ARM 64, do not allow the write+execute
# memory needed to set up the callback function. # memory needed to set up the callback function.
_LeptonicaErrorTrap = _LeptonicaErrorTrap_Redirect _LeptonicaErrorTrap: Type[ContextManager] = _LeptonicaErrorTrap_Redirect
else: else:
# 1.79 have this new symbol # 1.79 have this new symbol
_LeptonicaErrorTrap = _LeptonicaErrorTrap_Queue _LeptonicaErrorTrap = _LeptonicaErrorTrap_Queue
@@ -272,7 +273,7 @@ class LeptonicaObject:
# Leptonica API uses double-pointers for its destroy APIs to prevent # Leptonica API uses double-pointers for its destroy APIs to prevent
# dangling pointers. This means we need to put our single pointer, # dangling pointers. This means we need to put our single pointer,
# cdata, in a temporary CDATA**. # cdata, in a temporary CDATA**.
pp = ffi.new('{} **'.format(cls.LEPTONICA_TYPENAME), cdata) pp = ffi.new(f'{cls.LEPTONICA_TYPENAME} **', cdata)
cls.cdata_destroy(pp) cls.cdata_destroy(pp)
@@ -439,6 +440,9 @@ class Pix(LeptonicaObject):
bio = BytesIO() bio = BytesIO()
pillow_image.save(bio, format='png', compress_level=1) pillow_image.save(bio, format='png', compress_level=1)
py_buffer = bio.getbuffer() py_buffer = bio.getbuffer()
if platform.python_implementation() == 'PyPy':
# PyPy complains that it cannot do from_buffer(memoryview)
py_buffer = bytes(py_buffer)
c_buffer = ffi.from_buffer(py_buffer) c_buffer = ffi.from_buffer(py_buffer)
with _LeptonicaErrorTrap(): with _LeptonicaErrorTrap():
pix = Pix(lept.pixReadMem(c_buffer, len(c_buffer))) pix = Pix(lept.pixReadMem(c_buffer, len(c_buffer)))
@@ -844,7 +848,7 @@ class Box(LeptonicaObject):
def __repr__(self): def __repr__(self):
if self._cdata: if self._cdata:
return '<leptonica.Box x={0} y={1} w={2} h={3}>'.format( return '<leptonica.Box x={} y={} w={} h={}>'.format(
self.x, self.y, self.w, self.h self.x, self.y, self.w, self.h
) )
return '<leptonica.Box NULL>' return '<leptonica.Box NULL>'
@@ -916,7 +920,7 @@ class Sel(LeptonicaObject):
lines = [line.strip() for line in selstr.split('\n') if line.strip()] lines = [line.strip() for line in selstr.split('\n') if line.strip()]
h = len(lines) h = len(lines)
w = len(lines[0]) w = len(lines[0])
lengths = set(len(line) for line in lines) lengths = {len(line) for line in lines}
if len(lengths) != 1: if len(lengths) != 1:
raise ValueError("All lines in selstr must be same length") raise ValueError("All lines in selstr must be same length")
+29 -17
View File
@@ -25,8 +25,16 @@ from typing import (
) )
import img2pdf import img2pdf
import pikepdf from pikepdf import (
from pikepdf import Dictionary, Name, Object, Pdf, PdfImage Dictionary,
Name,
Object,
ObjectStreamMode,
Pdf,
PdfImage,
Stream,
UnsupportedImageTypeError,
)
from PIL import Image from PIL import Image
from ocrmypdf import leptonica from ocrmypdf import leptonica
@@ -63,7 +71,7 @@ def jpg_name(root: Path, xref: Xref) -> Path:
def extract_image_filter( def extract_image_filter(
pike: Pdf, root: Path, image: Object, xref: Xref pike: Pdf, root: Path, image: Stream, xref: Xref
) -> Optional[Tuple[PdfImage, Tuple[Name, Object]]]: ) -> Optional[Tuple[PdfImage, Tuple[Name, Object]]]:
del pike # unused args del pike # unused args
del root del root
@@ -89,7 +97,7 @@ def extract_image_filter(
return None # Don't mess with wide gamut images return None # Don't mess with wide gamut images
if filtdp[0] == Name.JPXDecode: if filtdp[0] == Name.JPXDecode:
log.debug(f"Skipping JPEG2000 iamge, xref {xref}") log.debug(f"Skipping JPEG2000 image, xref {xref}")
return None # Don't do JPEG2000 return None # Don't do JPEG2000
if filtdp[0] == Name.CCITTFaxDecode and filtdp[1].get('/K', 0) >= 0: if filtdp[0] == Name.CCITTFaxDecode and filtdp[1].get('/K', 0) >= 0:
@@ -104,7 +112,7 @@ def extract_image_filter(
def extract_image_jbig2( def extract_image_jbig2(
*, pike: pikepdf.Pdf, root: Path, image: Object, xref: Xref, options *, pike: Pdf, root: Path, image: Stream, xref: Xref, options
) -> Optional[XrefExt]: ) -> Optional[XrefExt]:
del options # unused arg del options # unused arg
@@ -123,16 +131,16 @@ def extract_image_jbig2(
# Showing the palette or ICC to jbig2enc will cause it to perform # Showing the palette or ICC to jbig2enc will cause it to perform
# colorspace transform to 1bpp, which will conflict the palette or # colorspace transform to 1bpp, which will conflict the palette or
# ICC if it exists. # ICC if it exists.
colorspace = pim.obj.get(pikepdf.Name.ColorSpace, None) colorspace = pim.obj.get(Name.ColorSpace, None)
if colorspace is not None or pim.image_mask: if colorspace is not None or pim.image_mask:
try: try:
# Set to DeviceGray temporarily; we already in 1 bpc. # Set to DeviceGray temporarily; we already in 1 bpc.
pim.obj.ColorSpace = pikepdf.Name.DeviceGray pim.obj.ColorSpace = Name.DeviceGray
imgname = root / f'{xref:08d}' imgname = root / f'{xref:08d}'
with imgname.open('wb') as f: with imgname.open('wb') as f:
ext = pim.extract_to(stream=f) ext = pim.extract_to(stream=f)
imgname.rename(imgname.with_suffix(ext)) imgname.rename(imgname.with_suffix(ext))
except pikepdf.UnsupportedImageTypeError: except UnsupportedImageTypeError:
return None return None
finally: finally:
# Restore image colorspace after temporarily setting it to DeviceGray # Restore image colorspace after temporarily setting it to DeviceGray
@@ -145,7 +153,7 @@ def extract_image_jbig2(
def extract_image_generic( def extract_image_generic(
*, pike: Pdf, root: Path, image: PdfImage, xref: Xref, options *, pike: Pdf, root: Path, image: Stream, xref: Xref, options
) -> Optional[XrefExt]: ) -> Optional[XrefExt]:
result = extract_image_filter(pike, root, image, xref) result = extract_image_filter(pike, root, image, xref)
if result is None: if result is None:
@@ -178,7 +186,7 @@ def extract_image_generic(
with imgname.open('wb') as f: with imgname.open('wb') as f:
ext = pim.extract_to(stream=f) ext = pim.extract_to(stream=f)
imgname.rename(imgname.with_suffix(ext)) imgname.rename(imgname.with_suffix(ext))
except pikepdf.UnsupportedImageTypeError: except UnsupportedImageTypeError:
return None return None
return XrefExt(xref, ext) return XrefExt(xref, ext)
elif ( elif (
@@ -365,6 +373,7 @@ def convert_to_jbig2(
When the JBIG2 symbolic coder is not used, each JBIG2 stands on its own When the JBIG2 symbolic coder is not used, each JBIG2 stands on its own
and needs no dictionary. Currently this must be lossless JBIG2. and needs no dictionary. Currently this must be lossless JBIG2.
""" """
jbig2_globals_dict: Optional[Dictionary]
_produce_jbig2_images(jbig2_groups, root, options, executor) _produce_jbig2_images(jbig2_groups, root, options, executor)
@@ -373,7 +382,7 @@ def convert_to_jbig2(
jbig2_symfile = root / (prefix + '.sym') jbig2_symfile = root / (prefix + '.sym')
if jbig2_symfile.exists(): if jbig2_symfile.exists():
jbig2_globals_data = jbig2_symfile.read_bytes() jbig2_globals_data = jbig2_symfile.read_bytes()
jbig2_globals = pikepdf.Stream(pike, jbig2_globals_data) jbig2_globals = Stream(pike, jbig2_globals_data)
jbig2_globals_dict = Dictionary(JBIG2Globals=jbig2_globals) jbig2_globals_dict = Dictionary(JBIG2Globals=jbig2_globals)
elif options.jbig2_page_group_size == 1: elif options.jbig2_page_group_size == 1:
jbig2_globals_dict = None jbig2_globals_dict = None
@@ -444,8 +453,8 @@ def _transcode_png(pike: Pdf, filename: Path, xref: Xref) -> bool:
with output.open('wb') as f: with output.open('wb') as f:
img2pdf.convert(fspath(filename), outputstream=f) img2pdf.convert(fspath(filename), outputstream=f)
with pikepdf.open(output) as pdf_image: with Pdf.open(output) as pdf_image:
foreign_image = next(pdf_image.pages[0].images.values()) foreign_image = next(iter(pdf_image.pages[0].images.values()))
local_image = pike.copy_foreign(foreign_image) local_image = pike.copy_foreign(foreign_image)
im_obj = pike.get_object(xref, 0) im_obj = pike.get_object(xref, 0)
@@ -524,12 +533,15 @@ def transcode_pngs(
_transcode_png(pike, filename, xref) _transcode_png(pike, filename, xref)
DEFAULT_EXECUTOR = SerialExecutor()
def optimize( def optimize(
input_file: Path, input_file: Path,
output_file: Path, output_file: Path,
context, context,
save_settings, save_settings,
executor: Executor = SerialExecutor(), executor: Executor = DEFAULT_EXECUTOR,
) -> None: ) -> None:
options = context.options options = context.options
if options.optimize == 0: if options.optimize == 0:
@@ -543,7 +555,7 @@ def optimize(
if options.jbig2_page_group_size == 0: if options.jbig2_page_group_size == 0:
options.jbig2_page_group_size = 10 if options.jbig2_lossy else 1 options.jbig2_page_group_size = 10 if options.jbig2_lossy else 1
with pikepdf.Pdf.open(input_file) as pike: with Pdf.open(input_file) as pike:
root = output_file.parent / 'images' root = output_file.parent / 'images'
root.mkdir(exist_ok=True) root.mkdir(exist_ok=True)
@@ -575,7 +587,7 @@ def optimize(
if savings < 0: if savings < 0:
log.info("Image optimization did not improve the file - discarded") log.info("Image optimization did not improve the file - discarded")
# We still need to save the file # We still need to save the file
with pikepdf.open(input_file) as pike: with Pdf.open(input_file) as pike:
pike.remove_unreferenced_resources() pike.remove_unreferenced_resources()
pike.save(output_file, **save_settings) pike.save(output_file, **save_settings)
else: else:
@@ -622,7 +634,7 @@ def main(infile, outfile, level, jobs=1):
dict( dict(
compress_streams=True, compress_streams=True,
preserve_pdfa=True, preserve_pdfa=True,
object_stream_mode=pikepdf.ObjectStreamMode.generate, object_stream_mode=ObjectStreamMode.generate,
), ),
) )
copy(fspath(tmpout), fspath(outfile)) copy(fspath(tmpout), fspath(outfile))
+13 -6
View File
@@ -13,13 +13,20 @@ import base64
from pathlib import Path from pathlib import Path
from typing import Dict, Iterator, Union from typing import Dict, Iterator, Union
try:
from importlib_resources import read_binary
except ImportError:
from importlib.resources import read_binary
import pikepdf import pikepdf
import pkg_resources import pkg_resources # deprecated
# Deprecated
ICC_PROFILE_RELPATH = 'data/sRGB.icc' ICC_PROFILE_RELPATH = 'data/sRGB.icc'
# Deprecated
SRGB_ICC_PROFILE = pkg_resources.resource_filename('ocrmypdf', ICC_PROFILE_RELPATH) SRGB_ICC_PROFILE = pkg_resources.resource_filename('ocrmypdf', ICC_PROFILE_RELPATH)
SRGB_ICC_PROFILE_NAME = 'sRGB.icc'
def _postscript_objdef( def _postscript_objdef(
alias: str, alias: str,
@@ -97,12 +104,12 @@ def generate_pdfa_ps(target_filename: Path, icc: str = 'sRGB'):
References: References:
Adobe PDFMARK Reference: https://www.adobe.com/content/dam/acom/en/devnet/acrobat/pdfs/pdfmark_reference.pdf Adobe PDFMARK Reference: https://www.adobe.com/content/dam/acom/en/devnet/acrobat/pdfs/pdfmark_reference.pdf
""" """
if icc == 'sRGB': if icc != 'sRGB':
icc_profile = SRGB_ICC_PROFILE
else:
raise NotImplementedError("Only supporting sRGB") raise NotImplementedError("Only supporting sRGB")
bytes_icc_profile = Path(icc_profile).read_bytes() bytes_icc_profile = read_binary(
'ocrmypdf.data', SRGB_ICC_PROFILE_NAME
)
ps = '\n'.join(_make_postscript(icc, bytes_icc_profile, 3)) ps = '\n'.join(_make_postscript(icc, bytes_icc_profile, 3))
# We should have encoded everything to pure ASCII by this point, and # We should have encoded everything to pure ASCII by this point, and
+84 -41
View File
@@ -9,7 +9,7 @@
import atexit import atexit
import logging import logging
import re import re
from collections import defaultdict, namedtuple from collections import defaultdict
from contextlib import ExitStack from contextlib import ExitStack
from decimal import Decimal from decimal import Decimal
from enum import Enum from enum import Enum
@@ -17,11 +17,27 @@ from functools import partial
from math import hypot, inf, isclose from math import hypot, inf, isclose
from os import PathLike from os import PathLike
from pathlib import Path from pathlib import Path
from typing import Container, Iterator, Optional, Tuple, Union from typing import (
Container,
Dict,
Iterator,
List,
Mapping,
NamedTuple,
Optional,
Tuple,
Union,
)
from warnings import warn from warnings import warn
import pikepdf from pikepdf import (
from pikepdf import Object, Pdf, PdfMatrix Object,
Pdf,
PdfImage,
PdfInlineImage,
PdfMatrix,
parse_content_stream,
)
from ocrmypdf._concurrent import Executor, SerialExecutor from ocrmypdf._concurrent import Executor, SerialExecutor
from ocrmypdf.exceptions import EncryptedPdfError, InputFileError from ocrmypdf.exceptions import EncryptedPdfError, InputFileError
@@ -36,7 +52,7 @@ Encoding = Enum(
'Encoding', 'ccitt jpeg jpeg2000 jbig2 asciihex ascii85 lzw flate runlength' 'Encoding', 'ccitt jpeg jpeg2000 jbig2 asciihex ascii85 lzw flate runlength'
) )
FRIENDLY_COLORSPACE = { FRIENDLY_COLORSPACE: Dict[str, Colorspace] = {
'/DeviceGray': Colorspace.gray, '/DeviceGray': Colorspace.gray,
'/CalGray': Colorspace.gray, '/CalGray': Colorspace.gray,
'/DeviceRGB': Colorspace.rgb, '/DeviceRGB': Colorspace.rgb,
@@ -54,7 +70,7 @@ FRIENDLY_COLORSPACE = {
'/I': Colorspace.index, '/I': Colorspace.index,
} }
FRIENDLY_ENCODING = { FRIENDLY_ENCODING: Dict[str, Encoding] = {
'/CCITTFaxDecode': Encoding.ccitt, '/CCITTFaxDecode': Encoding.ccitt,
'/DCTDecode': Encoding.jpeg, '/DCTDecode': Encoding.jpeg,
'/JPXDecode': Encoding.jpeg2000, '/JPXDecode': Encoding.jpeg2000,
@@ -68,7 +84,7 @@ FRIENDLY_ENCODING = {
'/RL': Encoding.runlength, '/RL': Encoding.runlength,
} }
FRIENDLY_COMP = { FRIENDLY_COMP: Dict[Colorspace, int] = {
Colorspace.gray: 1, Colorspace.gray: 1,
Colorspace.rgb: 3, Colorspace.rgb: 3,
Colorspace.cmyk: 4, Colorspace.cmyk: 4,
@@ -86,16 +102,30 @@ def _is_unit_square(shorthand):
return all(isclose(a, b, rel_tol=1e-3) for a, b in pairwise) return all(isclose(a, b, rel_tol=1e-3) for a, b in pairwise)
XobjectSettings = namedtuple('XobjectSettings', ['name', 'shorthand', 'stack_depth']) class XobjectSettings(NamedTuple):
name: str
shorthand: Tuple[float, float, float, float, float, float]
stack_depth: int
InlineSettings = namedtuple('InlineSettings', ['iimage', 'shorthand', 'stack_depth'])
ContentsInfo = namedtuple( class InlineSettings(NamedTuple):
'ContentsInfo', iimage: PdfInlineImage
['xobject_settings', 'inline_images', 'found_vector', 'found_text', 'name_index'], shorthand: Tuple[float, float, float, float, float, float]
) stack_depth: int
TextboxInfo = namedtuple('TextboxInfo', ['bbox', 'is_visible', 'is_corrupt'])
class ContentsInfo(NamedTuple):
xobject_settings: List[XobjectSettings]
inline_images: List[InlineSettings]
found_vector: bool
found_text: bool
name_index: Mapping[str, List[XobjectSettings]]
class TextboxInfo(NamedTuple):
bbox: Tuple[float, float, float, float]
is_visible: bool
is_corrupt: bool
class VectorMarker: class VectorMarker:
@@ -146,8 +176,8 @@ def _interpret_contents(contentstream: Object, initial_shorthand=UNIT_SQUARE):
stack = [] stack = []
ctm = PdfMatrix(initial_shorthand) ctm = PdfMatrix(initial_shorthand)
xobject_settings = [] xobject_settings: List[XobjectSettings] = []
inline_images = [] inline_images: List[InlineSettings] = []
name_index = defaultdict(lambda: []) name_index = defaultdict(lambda: [])
found_vector = False found_vector = False
found_text = False found_text = False
@@ -157,9 +187,7 @@ def _interpret_contents(contentstream: Object, initial_shorthand=UNIT_SQUARE):
operator_whitelist = ' '.join(vector_ops | text_showing_ops | image_ops) operator_whitelist = ' '.join(vector_ops | text_showing_ops | image_ops)
for n, graphobj in enumerate( for n, graphobj in enumerate(
_normalize_stack( _normalize_stack(parse_content_stream(contentstream, operator_whitelist))
pikepdf.parse_content_stream(contentstream, operator_whitelist)
)
): ):
operands, operator = graphobj operands, operator = graphobj
if operator == 'q': if operator == 'q':
@@ -185,7 +213,7 @@ def _interpret_contents(contentstream: Object, initial_shorthand=UNIT_SQUARE):
name=image_name, shorthand=ctm.shorthand, stack_depth=len(stack) name=image_name, shorthand=ctm.shorthand, stack_depth=len(stack)
) )
xobject_settings.append(settings) xobject_settings.append(settings)
name_index[image_name].append(settings) name_index[str(image_name)].append(settings)
elif operator == 'INLINE IMAGE': # BI/ID/EI are grouped into this elif operator == 'INLINE IMAGE': # BI/ID/EI are grouped into this
iimage = operands[0] iimage = operands[0]
inline = InlineSettings( inline = InlineSettings(
@@ -271,23 +299,28 @@ def _get_dpi(ctm_shorthand, image_size) -> Resolution:
class ImageInfo: class ImageInfo:
DPI_PREC = Decimal('1.000') DPI_PREC = Decimal('1.000')
_comp: Optional[int]
_name: str
def __init__( def __init__(
self, self,
*, *,
name='', name='',
pdfimage: Optional[Object] = None, pdfimage: Optional[Object] = None,
inline: Optional[Object] = None, inline: Optional[PdfInlineImage] = None,
shorthand=None, shorthand=None,
): ):
self._name = str(name) self._name = str(name)
self._shorthand = shorthand self._shorthand = shorthand
pim: Union[PdfInlineImage, PdfImage]
if inline is not None: if inline is not None:
self._origin = 'inline' self._origin = 'inline'
pim = inline.iimage pim = inline
elif pdfimage is not None: elif pdfimage is not None:
self._origin = 'xobject' self._origin = 'xobject'
pim = pikepdf.PdfImage(pdfimage) pim = PdfImage(pdfimage)
else: else:
raise ValueError("Either pdfimage or inline must be set") raise ValueError("Either pdfimage or inline must be set")
self._width = pim.width self._width = pim.width
@@ -303,14 +336,14 @@ class ImageInfo:
self._bpc = int(pim.bits_per_component) self._bpc = int(pim.bits_per_component)
try: try:
self._enc = FRIENDLY_ENCODING.get(pim.filters[0], 'image') self._enc = FRIENDLY_ENCODING.get(pim.filters[0])
except IndexError: except IndexError:
self._enc = '?' self._enc = None
try: try:
self._color = FRIENDLY_COLORSPACE.get(pim.colorspace, '?') self._color = FRIENDLY_COLORSPACE.get(pim.colorspace or '')
except NotImplementedError: except NotImplementedError:
self._color = '?' self._color = None
if self._enc == Encoding.jpeg2000: if self._enc == Encoding.jpeg2000:
self._color = Colorspace.jpeg2000 self._color = Colorspace.jpeg2000
@@ -324,11 +357,14 @@ class ImageInfo:
else: else:
self._comp = 3 self._comp = 3
else: else:
self._comp = FRIENDLY_COMP.get(self._color, '?') if isinstance(self._color, Colorspace):
self._comp = FRIENDLY_COMP.get(self._color)
else:
self._comp = None
# Bit of a hack... infer grayscale if component count is uncertain # Bit of a hack... infer grayscale if component count is uncertain
# but encoding only supports monochrome. # but encoding only supports monochrome.
if self._comp == '?' and self._enc in (Encoding.ccitt, Encoding.jbig2): if self._comp is None and self._enc in (Encoding.ccitt, Encoding.jbig2):
self._comp = FRIENDLY_COMP[Colorspace.gray] self._comp = FRIENDLY_COMP[Colorspace.gray]
@property @property
@@ -353,15 +389,15 @@ class ImageInfo:
@property @property
def color(self): def color(self):
return self._color return self._color if self._color is not None else '?'
@property @property
def comp(self): def comp(self):
return self._comp return self._comp if self._comp is not None else '?'
@property @property
def enc(self): def enc(self):
return self._enc return self._enc if self._enc is not None else 'image'
@property @property
def renderable(self): def renderable(self):
@@ -388,7 +424,7 @@ def _find_inline_images(contentsinfo: ContentsInfo) -> Iterator[ImageInfo]:
for n, inline in enumerate(contentsinfo.inline_images): for n, inline in enumerate(contentsinfo.inline_images):
yield ImageInfo( yield ImageInfo(
name='inline-%02d' % n, shorthand=inline.shorthand, inline=inline name='inline-%02d' % n, shorthand=inline.shorthand, inline=inline.iimage
) )
@@ -413,7 +449,7 @@ def _image_xobjects(container) -> Iterator[Tuple[Object, str]]:
xobjs = resources['/XObject'].as_dict() xobjs = resources['/XObject'].as_dict()
for xobj in xobjs: for xobj in xobjs:
candidate: Object = xobjs[xobj] candidate: Object = xobjs[xobj]
if not '/Subtype' in candidate: if '/Subtype' not in candidate:
continue continue
if candidate['/Subtype'] == '/Image': if candidate['/Subtype'] == '/Image':
pdfimage = candidate pdfimage = candidate
@@ -583,7 +619,7 @@ def _pdf_pageinfo_sync_init(pdf: Pdf, infile: Path, pdfminer_loglevel):
# If the pdf is not opened, open a copy for our worker process to use # If the pdf is not opened, open a copy for our worker process to use
if pdf is None: if pdf is None:
worker_pdf = pikepdf.open(infile) worker_pdf = Pdf.open(infile)
def on_process_close(): def on_process_close():
worker_pdf.close() worker_pdf.close()
@@ -597,7 +633,7 @@ def _pdf_pageinfo_sync(args):
pdf = thread_pdf if thread_pdf is not None else worker_pdf pdf = thread_pdf if thread_pdf is not None else worker_pdf
with ExitStack() as stack: with ExitStack() as stack:
if not pdf: # When called with SerialExecutor if not pdf: # When called with SerialExecutor
pdf = stack.enter_context(pikepdf.open(infile)) pdf = stack.enter_context(Pdf.open(infile))
page = PageInfo(pdf, pageno, infile, check_pages, detailed_analysis) page = PageInfo(pdf, pageno, infile, check_pages, detailed_analysis)
return page return page
@@ -661,6 +697,10 @@ def _pdf_pageinfo_concurrent(
class PageInfo: class PageInfo:
_has_text: Optional[bool]
_has_vector: Optional[bool]
_images: List[ImageInfo]
def __init__( def __init__(
self, self,
pdf: Pdf, pdf: Pdf,
@@ -732,7 +772,7 @@ class PageInfo:
else: else:
self._has_vector = None # i.e. "no information" self._has_vector = None # i.e. "no information"
self._has_text = None self._has_text = None
self._images = None self._images = []
self._dpi = None self._dpi = None
if self._images: if self._images:
@@ -749,7 +789,7 @@ class PageInfo:
@property @property
def has_text(self) -> bool: def has_text(self) -> bool:
return self._has_text return bool(self._has_text)
@property @property
def has_corrupt_text(self) -> bool: def has_corrupt_text(self) -> bool:
@@ -759,7 +799,7 @@ class PageInfo:
@property @property
def has_vector(self) -> bool: def has_vector(self) -> bool:
return self._has_vector return bool(self._has_vector)
@property @property
def width_inches(self) -> Decimal: def width_inches(self) -> Decimal:
@@ -837,6 +877,9 @@ class PageInfo:
) )
DEFAULT_EXECUTOR = SerialExecutor()
class PdfInfo: class PdfInfo:
"""Get summary information about a PDF""" """Get summary information about a PDF"""
@@ -848,13 +891,13 @@ class PdfInfo:
progbar: bool = False, progbar: bool = False,
max_workers: int = None, max_workers: int = None,
check_pages=None, check_pages=None,
executor: Executor = SerialExecutor(), executor: Executor = DEFAULT_EXECUTOR,
): ):
self._infile = infile self._infile = infile
if check_pages is None: if check_pages is None:
check_pages = range(0, 1_000_000_000) check_pages = range(0, 1_000_000_000)
with pikepdf.open(infile) as pdf: with Pdf.open(infile) as pdf:
if pdf.is_encrypted: if pdf.is_encrypted:
raise EncryptedPdfError() # Triggered by encryption with empty passwd raise EncryptedPdfError() # Triggered by encryption with empty passwd
self._pages = _pdf_pageinfo_concurrent( self._pages = _pdf_pageinfo_concurrent(
+1 -1
View File
@@ -135,7 +135,7 @@ class LTStateAwareChar(LTChar):
return self._text return self._text
def __repr__(self): def __repr__(self):
return '<%s %s matrix=%s rendermode=%r font=%r adv=%s text=%r>' % ( return '<{} {} matrix={} rendermode={!r} font={!r} adv={} text={!r}>'.format(
self.__class__.__name__, self.__class__.__name__,
bbox2str(self.bbox), bbox2str(self.bbox),
matrix2str(self.matrix), matrix2str(self.matrix),
+14 -8
View File
@@ -5,7 +5,7 @@
# file, You can obtain one at http://mozilla.org/MPL/2.0/. # file, You can obtain one at http://mozilla.org/MPL/2.0/.
from abc import ABC, abstractmethod, abstractstaticmethod from abc import ABC, abstractmethod
from argparse import ArgumentParser, Namespace from argparse import ArgumentParser, Namespace
from collections import namedtuple from collections import namedtuple
from logging import Handler from logging import Handler
@@ -197,7 +197,7 @@ def rasterize_pdf_page(
@hookspec(firstresult=True) @hookspec(firstresult=True)
def filter_ocr_image(page: 'PageContext', image: 'Image') -> 'Image': def filter_ocr_image(page: 'PageContext', image: 'Image.Image') -> 'Image.Image':
"""Called to filter the image before it is sent to OCR. """Called to filter the image before it is sent to OCR.
This is the image that OCR sees, not what the user sees when they view the This is the image that OCR sees, not what the user sees when they view the
@@ -325,11 +325,13 @@ class OcrEngine(ABC):
Tesseract OCR. Tesseract OCR.
""" """
@abstractstaticmethod @staticmethod
@abstractmethod
def version() -> str: def version() -> str:
"""Returns the version of the OCR engine.""" """Returns the version of the OCR engine."""
@abstractstaticmethod @staticmethod
@abstractmethod
def creator_tag(options: Namespace) -> str: def creator_tag(options: Namespace) -> str:
"""Returns the creator tag to identify this software's role in creating the PDF. """Returns the creator tag to identify this software's role in creating the PDF.
@@ -349,24 +351,28 @@ class OcrEngine(ABC):
to the user, usually in an error message. to the user, usually in an error message.
""" """
@abstractstaticmethod @staticmethod
@abstractmethod
def languages(options: Namespace) -> AbstractSet[str]: def languages(options: Namespace) -> AbstractSet[str]:
"""Returns the set of all languages that are supported by the engine. """Returns the set of all languages that are supported by the engine.
Languages are typically given in 3-letter ISO 3166-1 codes, but actually Languages are typically given in 3-letter ISO 3166-1 codes, but actually
can be any value understood by the OCR engine.""" can be any value understood by the OCR engine."""
@abstractstaticmethod @staticmethod
@abstractmethod
def get_orientation(input_file: Path, options: Namespace) -> OrientationConfidence: def get_orientation(input_file: Path, options: Namespace) -> OrientationConfidence:
"""Returns the orientation of the image.""" """Returns the orientation of the image."""
@abstractstaticmethod @staticmethod
@abstractmethod
def generate_hocr( def generate_hocr(
input_file: Path, output_hocr: Path, output_text: Path, options: Namespace input_file: Path, output_hocr: Path, output_text: Path, options: Namespace
) -> None: ) -> None:
"""Called to produce a hOCR file and sidecar text file.""" """Called to produce a hOCR file and sidecar text file."""
@abstractstaticmethod @staticmethod
@abstractmethod
def generate_pdf( def generate_pdf(
input_file: Path, output_pdf: Path, output_text: Path, options: Namespace input_file: Path, output_pdf: Path, output_text: Path, options: Namespace
) -> None: ) -> None:
+9 -4
View File
@@ -15,7 +15,6 @@ from collections.abc import Mapping
from contextlib import suppress from contextlib import suppress
from distutils.version import LooseVersion, Version from distutils.version import LooseVersion, Version
from functools import lru_cache from functools import lru_cache
from pathlib import Path
from subprocess import PIPE, STDOUT, CalledProcessError, CompletedProcess, Popen from subprocess import PIPE, STDOUT, CalledProcessError, CompletedProcess, Popen
from subprocess import run as subprocess_run from subprocess import run as subprocess_run
from typing import Callable, Optional, Type, Union from typing import Callable, Optional, Type, Union
@@ -27,7 +26,9 @@ from ocrmypdf.exceptions import MissingDependencyError
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
def run(args, *, env=None, logs_errors_to_stdout=False, **kwargs): def run(
args, *, env=None, logs_errors_to_stdout: bool = False, **kwargs
) -> CompletedProcess:
"""Wrapper around :py:func:`subprocess.run` """Wrapper around :py:func:`subprocess.run`
The main purpose of this wrapper is to log subprocess output in an orderly The main purpose of this wrapper is to log subprocess output in an orderly
@@ -65,7 +66,9 @@ def run(args, *, env=None, logs_errors_to_stdout=False, **kwargs):
return proc return proc
def run_polling_stderr(args, *, callback, check=False, env=None, **kwargs): def run_polling_stderr(
args, *, callback: Callable[[str], None], check: bool = False, env=None, **kwargs
) -> CompletedProcess:
"""Run a process like ``ocrmypdf.subprocess.run``, and poll stderr. """Run a process like ``ocrmypdf.subprocess.run``, and poll stderr.
Every line of produced by stderr will be forwarded to the callback function. Every line of produced by stderr will be forwarded to the callback function.
@@ -83,6 +86,8 @@ def run_polling_stderr(args, *, callback, check=False, env=None, **kwargs):
with Popen(args, env=env, **kwargs) as proc: with Popen(args, env=env, **kwargs) as proc:
lines = [] lines = []
while proc.poll() is None: while proc.poll() is None:
if proc.stderr is None:
continue
for msg in iter(proc.stderr.readline, ''): for msg in iter(proc.stderr.readline, ''):
if process_log.isEnabledFor(logging.DEBUG): if process_log.isEnabledFor(logging.DEBUG):
process_log.debug(msg.strip()) process_log.debug(msg.strip())
@@ -102,7 +107,7 @@ def _fix_process_args(args, env, kwargs):
env = os.environ env = os.environ
# Search in spoof path if necessary # Search in spoof path if necessary
program = args[0] program = str(args[0])
if os.name == 'nt': if os.name == 'nt':
from ocrmypdf.subprocess._windows import fix_windows_args from ocrmypdf.subprocess._windows import fix_windows_args
+6 -8
View File
@@ -9,9 +9,9 @@ import os
import shutil import shutil
import sys import sys
from distutils.version import LooseVersion from distutils.version import LooseVersion
from itertools import chain, filterfalse from itertools import chain
from pathlib import Path from pathlib import Path
from typing import Any, Callable, Iterator, Optional, Tuple, TypeVar, cast from typing import Any, Callable, Iterable, Iterator, Set, Tuple, TypeVar
try: try:
import winreg import winreg
@@ -113,7 +113,7 @@ SHIMS = [
] ]
def fix_windows_args(program, args, env): def fix_windows_args(program: str, args, env):
"""Adjust our desired program and command line arguments for use on Windows""" """Adjust our desired program and command line arguments for use on Windows"""
if sys.version_info < (3, 8): if sys.version_info < (3, 8):
@@ -137,14 +137,12 @@ def fix_windows_args(program, args, env):
return args return args
def unique_everseen(iterable, key=None): def unique_everseen(iterable: Iterable[T], key: Callable[[T], T]) -> Iterator[T]:
"List unique elements, preserving order. Remember all elements ever seen." "List unique elements, preserving order."
# unique_everseen('AAAABBBCCDAABBB') --> A B C D # unique_everseen('AAAABBBCCDAABBB') --> A B C D
# unique_everseen('ABBCcAD', str.lower) --> A B C D # unique_everseen('ABBCcAD', str.lower) --> A B C D
seen = set() seen: Set[T] = set()
seen_add = seen.add seen_add = seen.add
if key is None:
key = lambda x: x
for element in iterable: for element in iterable:
k = key(element) k = key(element)
if k not in seen: if k not in seen:
+5
View File
@@ -69,6 +69,11 @@ def outpdf(tmp_path):
return tmp_path / 'out.pdf' return tmp_path / 'out.pdf'
@pytest.fixture(scope="function")
def outtxt(tmp_path):
return tmp_path / 'out.txt'
@pytest.fixture(scope="function") @pytest.fixture(scope="function")
def no_outpdf(tmp_path): def no_outpdf(tmp_path):
"""This just documents the fact that a test is not expected to produce """This just documents the fact that a test is not expected to produce
+1 -1
View File
@@ -87,4 +87,4 @@ def test_jpeg_in_jpeg_out(resources, outpdf):
'tests/plugins/tesseract_noop.py', 'tests/plugins/tesseract_noop.py',
) )
with pikepdf.open(outpdf) as pdf: with pikepdf.open(outpdf) as pdf:
assert next(pdf.pages[0].images.values()).Filter == pikepdf.Name.DCTDecode assert next(iter(pdf.pages[0].images.values())).Filter == pikepdf.Name.DCTDecode
+33 -8
View File
@@ -701,7 +701,7 @@ def test_sidecar_pagecount(resources, outpdf):
pdfinfo = PdfInfo(resources / '3small.pdf') pdfinfo = PdfInfo(resources / '3small.pdf')
num_pages = len(pdfinfo) num_pages = len(pdfinfo)
with open(sidecar, 'r', encoding='utf-8') as f: with open(sidecar, encoding='utf-8') as f:
ocr_text = f.read() ocr_text = f.read()
# There should a formfeed between each pair of pages, so the count of # There should a formfeed between each pair of pages, so the count of
@@ -722,7 +722,7 @@ def test_sidecar_nonempty(resources, outpdf):
'tests/plugins/tesseract_cache.py', 'tests/plugins/tesseract_cache.py',
) )
with open(sidecar, 'r', encoding='utf-8') as f: with open(sidecar, encoding='utf-8') as f:
ocr_text = f.read() ocr_text = f.read()
assert 'the' in ocr_text assert 'the' in ocr_text
@@ -745,14 +745,13 @@ def test_pdfa_n(pdfa_level, resources, outpdf):
assert pdfa_info['conformance'] == f'PDF/A-{pdfa_level}B' assert pdfa_info['conformance'] == f'PDF/A-{pdfa_level}B'
@pytest.mark.skipif( def test_decompression_bomb_error(resources, outpdf):
PIL.__version__ < '5.0.0', reason="Pillow < 5.0.0 doesn't raise the exception"
)
@pytest.mark.slow
def test_decompression_bomb(resources, outpdf):
p, _out, err = run_ocrmypdf(resources / 'hugemono.pdf', outpdf) p, _out, err = run_ocrmypdf(resources / 'hugemono.pdf', outpdf)
assert 'decompression bomb' in err assert 'decompression bomb' in err and '--max-image-mpixels' in err
@pytest.mark.slow
def test_decompression_bomb_succeeds(resources, outpdf):
p, _out, err = run_ocrmypdf( p, _out, err = run_ocrmypdf(
resources / 'hugemono.pdf', outpdf, '--max-image-mpixels', '2000' resources / 'hugemono.pdf', outpdf, '--max-image-mpixels', '2000'
) )
@@ -881,3 +880,29 @@ def test_image_dpi_threshold(resources, outpdf):
'tests/plugins/tesseract_noop.py', 'tests/plugins/tesseract_noop.py',
) )
assert outpdf.exists() assert outpdf.exists()
def test_outputtype_none_bad_setup(resources, outpdf):
p, _out, err = run_ocrmypdf(
resources / 'trivial.pdf',
outpdf,
'--output-type=none',
'--plugin',
'tests/plugins/tesseract_noop.py',
)
assert p.returncode == ExitCode.bad_args
assert 'Set the output file to' in err
def test_outputtype_none(resources, outtxt):
p, _out, err = run_ocrmypdf(
resources / 'trivial.pdf',
os.devnull,
'--output-type=none',
'--sidecar',
outtxt,
'--plugin',
'tests/plugins/tesseract_noop.py',
)
assert p.returncode == ExitCode.ok
assert outtxt.exists()
+1 -1
View File
@@ -287,7 +287,7 @@ def test_srgb_in_unicode_path(tmp_path):
def test_kodak_toc(resources, outpdf): def test_kodak_toc(resources, outpdf):
_output = check_ocrmypdf( check_ocrmypdf(
resources / 'kcs.pdf', resources / 'kcs.pdf',
outpdf, outpdf,
'--output-type', '--output-type',
-4
View File
@@ -50,10 +50,6 @@ def test_nonmonotonic_warning(caplog):
assert 'out of order' in caplog.text assert 'out of order' in caplog.text
def test_list_range():
assert _pages_from_ranges([0, 1, 2]) == {0, 1, 2}
def test_limited_pages(resources, outpdf): def test_limited_pages(resources, outpdf):
multi = resources / 'multipage.pdf' multi = resources / 'multipage.pdf'
ocrmypdf.ocr( ocrmypdf.ocr(
+2 -4
View File
@@ -79,9 +79,7 @@ def test_dpi_needed(image, text, vector, result, rgb_image, outdir):
# Input: # Input:
('', '', '', '', ''), ('', '', '', '', ''),
# Output: # Output:
( (((1, 5), None),),
((1, 5), None),
),
), ),
( (
'no_empty_values', 'no_empty_values',
@@ -147,4 +145,4 @@ def test_dpi_needed(image, text, vector, result, rgb_image, outdir):
), ),
) )
def test_enumerate_compress_ranges(name, input, output): def test_enumerate_compress_ranges(name, input, output):
assert output == tuple(_pipeline.enumerate_compress_ranges(input)) assert output == tuple(_pipeline.enumerate_compress_ranges(input))
+6
View File
@@ -6,6 +6,7 @@
import logging import logging
import os
from unittest.mock import patch from unittest.mock import patch
import pikepdf import pikepdf
@@ -298,3 +299,8 @@ def test_sidecar_equals_output(resources, no_outpdf):
op = no_outpdf op = no_outpdf
with pytest.raises(BadArgsError, match=r'--sidecar'): with pytest.raises(BadArgsError, match=r'--sidecar'):
run_ocrmypdf_api(resources / 'trivial.pdf', op, '--sidecar', op) run_ocrmypdf_api(resources / 'trivial.pdf', op, '--sidecar', op)
def test_devnull_sidecar(resources):
with pytest.raises(BadArgsError, match=r'--sidecar.*NUL'):
run_ocrmypdf_api(resources / 'trivial.pdf', os.devnull, '--sidecar')