Compare commits

..
18 Commits
Author SHA1 Message Date
James R. Barlow 6e49bb3588 v8.2.3 notes 2019-04-03 01:19:12 -07:00
James R. Barlow 427afc0616 Fix LeptonicaErrorTrap when a sys.stderr.fileno() is not available
The LeptonicaErrorTrap was problematic for Celery and other
libraries that mess with stderr.

Closes #359
2019-03-17 14:22:36 -07:00
James R. Barlow 9c7ee2bf23 Better help text for --verbose 2019-03-17 13:29:25 -07:00
James R. Barlow c5cfaa950b readme: tweaks 2019-03-16 14:09:19 -07:00
James R. Barlow 4e2a98ead4 leptonica: fix junkpixt harder 2019-03-16 14:08:58 -07:00
James R. Barlow 210f134b5b Merge branch 'master' of github.com:jbarlow83/OCRmyPDF 2019-03-08 15:38:31 -08:00
James R. Barlow 696c0721a0 docs: fix broken sphinx ref
[ci skip]
2019-03-08 15:38:26 -08:00
James R. Barlow aabab95418 docs: use images folder 2019-03-08 15:38:01 -08:00
James R. Barlow 7d614dd68b docs: explain Automator workflow 2019-03-08 15:37:42 -08:00
jumbliesandjbarlow83 f57dda7939 Update batch.rst (#362)
Added docker instructions for passing "find" filenames into container.  Obviates prior incorrect flag fix.
2019-03-08 12:46:50 -08:00
James R. Barlow 1b4542aa77 Further fixes to external program version testing 2019-03-07 14:27:16 -08:00
James R. Barlow 6c7fca57ec v8.2.1 notes 2019-03-06 22:22:50 -08:00
James R. Barlow 486dc7e22c Fix some test failures missed in prev commit 2019-03-06 13:28:50 -08:00
James R. Barlow dc616bb507 Fix test suite so --clean is not requested when unpaper is not installed 2019-03-05 22:33:13 -08:00
James R. Barlow 902bda43e3 main: fix version testing unnecessarily throwing exception to itself 2019-03-05 22:32:06 -08:00
James R. Barlow f7da63f68b main: fix redundant argument test 2019-03-05 22:29:29 -08:00
James R. Barlow 5da26e4c9c Convert most uses of subprocess.Popen to subprocess.run in test suite 2019-03-05 22:25:22 -08:00
James R. Barlow c19c852705 Fix exception while attempting to print error message for missing program 2019-03-05 16:32:48 -08:00
17 changed files with 225 additions and 195 deletions
+9 -4
View File
@@ -38,9 +38,9 @@ Main features
- If requested deskews and/or cleans the image before performing OCR
- Validates input and output files
- Distributes work across all available CPU cores
- Uses [Tesseract OCR](https://github.com/tesseract-ocr/tesseract) engine
- Supports more than [100 languages](https://github.com/tesseract-ocr/tessdata) recognized by Tesseract
- Battle-tested on thousands of PDFs, a test suite and continuous integration
- Uses [Tesseract OCR](https://github.com/tesseract-ocr/tesseract) engine to recognize more than [100 languages](https://github.com/tesseract-ocr/tessdata)
- Scales properly to handle files with thousands of pages
- Battle-tested on millions of PDFs
For details: please consult the [documentation](https://ocrmypdf.readthedocs.io/en/latest/).
@@ -131,6 +131,11 @@ Press & Media
- [c't 1-2014, page 59](http://heise.de/-2279695): Detailed presentation of OCRmyPDF v1.0 in the leading German IT magazine c't
- [heise Open Source, 09/2014: Texterkennung mit OCRmyPDF](http://heise.de/-2356670)
Business enquiries
------------------
OCRmyPDF would not be the software that it is today is without companies and users choosing to provide support for feature development and consulting enquiries. We are happy to discuss all enquiries, whether for extending the existing feature set, or integrating OCRmyPDF into a larger system.
License
-------
@@ -138,7 +143,7 @@ The OCRmyPDF software is licensed under the GNU GPLv3. Certain files are covered
The license for each test file varies, and is noted in tests/resources/README.rst. The documentation is licensed under Creative Commons Attribution-ShareAlike 4.0 (CC-BY-SA 4.0).
OCRmyPDF versions prior to 6.0 were licensed under the MIT License.
OCRmyPDF versions prior to 6.0 were distributed under the MIT License.
Disclaimer
----------
-1
View File
@@ -85,7 +85,6 @@ For example, if you have a development build of Tesseract don't wish to use the
In this example ``TESSDATA_PREFIX`` is required to redirect Tesseract to an alternate folder for its "tessdata" files.
Overriding other support programs
"""""""""""""""""""""""""""""""""
+20 -7
View File
@@ -27,7 +27,13 @@ This will walk through a directory tree and run OCR on all files in place, print
.. code-block:: bash
find . --printf '%p' -name '*.pdf' -exec ocrmypdf '{}' '{}' \;
find . -printf '%p' -name '*.pdf' -exec ocrmypdf '{}' '{}' \;
Alternatively, with a docker container (mounts a volume to the container where the PDFs are stored):
.. code-block:: bash
find . -printf '%p' -name '*.pdf' -exec docker run --rm -v <host dir>:<container dir> jbarlow83/ocrmypdf-alpine '<container dir>/{}' '<container dir>/{}' \;
This only runs one ``ocrmypdf`` process at a time. This variation uses ``find`` to create a directory list and ``parallel`` to parallelize runs of ``ocrmypdf``, again updating files in place.
@@ -80,9 +86,9 @@ This user contributed script also provides an example of batch processing.
print(full_path)
cmd = ["ocrmypdf", "--deskew", filename, filename]
logging.info(cmd)
proc = subprocess.Popen(
proc = subprocess.run(
cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT)
result = proc.stdout.read()
result = proc.stdout
if proc.returncode == 6:
print("Skipped document because it already contained text")
elif proc.returncode == 0:
@@ -151,7 +157,7 @@ This is only possible for x86-based Synology products. Some Synology products us
# the script is processed as root user via chron
cmd = ['docker', 'run', '--rm', '-v', docker_mount, '-u=1030:65538', 'jbarlow83/ocrmypdf', , '--deskew' , filename, filename_OCR]
logging.info(cmd)
proc = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT)
proc = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT)
result = proc.stdout.read()
logging.info(result)
full_path_OCR = dir_name + '/' + filename_OCR
@@ -163,14 +169,11 @@ This is only possible for x86-based Synology products. Some Synology products us
shutil.move(full_path, full_path_archive)
logging.info('Finished.\n')
Huge batch jobs
"""""""""""""""
If you have thousands of files to work with, contact the author. Consulting work related to OCRmyPDF helps fund this open source project and all inquiries are appreciated.
Hot (watched) folders
---------------------
@@ -210,3 +213,13 @@ Alternatives
""""""""""""
* `Watchman <https://facebook.github.io/watchman/>`_ is a more powerful alternative to ``watchmedo``.
macOS Automator
---------------
You can use the Automator app with macOS, to create a Workflow or Quick Action. Use a *Run Shell Script* action in your workflow. In the context of Automator, the ``PATH`` may be set differently your Terminal's ``PATH``; you may need to explicitly set the PATH to include ``ocrmypdf``. The following example may serve as a starting point:
.. image:: images/macos-workflow.png
:alt: Example macOS Automator script
You may customize the command sent to ocrmypdf.
+1 -1
View File
@@ -104,7 +104,7 @@ Clients must keep their open connection while waiting for OCR to complete. This
Unlike the rest of OCRmyPDF, this web service is licensed under the Affero GPLv3 (AGPLv3) since Ghostscript, a dependency of OCRmyPDF, is also licensed in this way.
In addition to the above, please read our :ref:`general remarks on using OCRmyPDF as a service <ocr-service>`_.
In addition to the above, please read our :ref:`general remarks on using OCRmyPDF as a service <ocr-service>`.
Legacy Ubuntu Docker images
---------------------------

Before

Width:  |  Height:  |  Size: 3.1 KiB

After

Width:  |  Height:  |  Size: 3.1 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 21 KiB

+1 -1
View File
@@ -16,7 +16,7 @@ About PDFs
PDFs are page description files that attempts to preserve a layout exactly. They contain `vector graphics <http://vector-conversions.com/vectorizing/raster_vs_vector.html>`_ that can contain raster objects such as scanned images. Because PDFs can contain multiple pages (unlike many image formats) and can contain fonts and text, it is a good formats for exchanging scanned documents.
.. image:: bitmap_vs_svg.svg
.. image:: images/bitmap_vs_svg.svg
A PDF page might contain multiple images, even if it only appears to have one image. Some scanners or scanning software will segment pages into monochromatic text and color regions for example, to improve the compression ratio and appearance of the page.
+21
View File
@@ -13,6 +13,27 @@ Note that it is licensed under GPLv3, so scripts that ``import ocrmypdf`` and ar
find: [^`]\#([0-9]{1,3})[^0-9]
replace: `#$1 <https://github.com/jbarlow83/OCRmyPDF/issues/$1>`_
v8.2.3
------
- Fixed that ``--mask-barcodes`` would occasionally leave a unwanted temporary file named ``junkpixt`` in the current working folder.
- Fixed (hopefully) handling of Leptonica errors in an environment where a non-standard ``sys.stderr`` is present.
- Improved help text for ``--verbose``.
v8.2.2
------
- Fixed a regression from v8.2.0, an exception that occurred while attempting to report that ``unpaper`` or another optional dependency was unavailable.
- In some cases, ``ocrmypdf [-c|--clean]`` failed to exit with an error when ``unpaper`` is not installed.
v8.2.1
------
- This release was canceled.
v8.2.0
------
+34 -65
View File
@@ -46,6 +46,7 @@ from .exceptions import (
)
from .exec import (
ghostscript,
jbig2enc,
qpdf,
tesseract,
check_external_program,
@@ -230,7 +231,9 @@ jobcontrol.add_argument(
default=[],
nargs='?',
action="append",
help="Print more verbose messages for each additional verbose level",
help="Print more verbose messages for each additional verbose level. Use "
"`-v 1` typically for much more detailed logging. Higher numbers "
"are probably only useful in debugging.",
)
metadata = parser.add_argument_group(
@@ -602,43 +605,19 @@ def check_options_sidecar(options, log):
options.sidecar = options.output_file + '.txt'
def _optional_program_required(name, version_fn, min_version, for_argument):
try:
if version_fn() < min_version:
raise MissingDependencyError(
f"The installed '{name}' is not supported. "
f"Install version {min_version} or newer."
)
except (FileNotFoundError, MissingDependencyError):
raise MissingDependencyError(
f"Install the '{name}' program to use {for_argument}."
)
def _optional_program_recommended(name, version_fn, min_version, for_argument):
try:
if version_fn() < min_version:
raise MissingDependencyError(
f"The installed '{name}' is not supported. "
f"Install version {min_version} or newer."
)
except (FileNotFoundError, MissingDependencyError):
complain(
f"For best results, install the optional program '{name}' to use the "
f"argument {for_argument}."
)
def check_options_preprocessing(options, log):
if options.clean_final:
options.clean = True
if options.unpaper_args and not options.clean:
raise argparse.ArgumentError(None, "--clean is required for --unpaper-args")
if any((options.clean, options.clean_final)):
from .exec import unpaper
_optional_program_required(
'unpaper', unpaper.version, '6.1', '--clean, --clean-final'
if options.clean:
check_external_program(
log=log,
program='unpaper',
package='unpaper',
version_checker=unpaper.version,
need_version='6.1',
required_for=['--clean, --clean-final'],
)
try:
if options.unpaper_args:
@@ -664,19 +643,26 @@ def check_options_ocr_behavior(options, log):
def check_options_optimizing(options, log):
if options.optimize >= 2:
from .exec import pngquant, jbig2enc
_optional_program_required(
'pngquant', pngquant.version, '2.0.1', '--optimize {2,3}'
check_external_program(
log=log,
program='pngquant',
package='pngquant',
version_checker=pngquant.version,
need_version='2.0.1',
required_for='--optimize {2,3}',
)
if options.jbig2_lossy:
_optional_program_required('jbig2', jbig2enc.version, '0.28', '--jbig2-lossy')
elif options.optimize >= 2:
if options.optimize >= 2:
# Although we use JBIG2 for optimize=1, don't nag about it unless the
# user is asking for more optimization
_optional_program_recommended(
'jbig2', jbig2enc.version, '0.28', '--optimize {2,3}'
check_external_program(
log=log,
program='jbig2',
package='jbig2enc',
version_checker=jbig2enc.version,
need_version='0.28',
required_for='--optimize {2,3} | --jbig2-lossy',
recommended=True if not options.jbig2_lossy else False,
)
if options.optimize == 0 and any(
@@ -1027,7 +1013,7 @@ def report_output_file_size(options, _log, input_file, output_file):
)
def check_dependency_versions(log):
def check_dependency_versions(options, log):
check_external_program(
log=log,
program='tesseract',
@@ -1051,27 +1037,10 @@ def check_dependency_versions(log):
return ExitCode.missing_dependency
check_external_program(
log=log,
program='unpaper',
package='unpaper',
version_checker=unpaper.version,
need_version='6.1', # latest sane version
optional=True,
)
if os.environ.get('TRAVIS') != 'true': # Suppress for Ubuntu trusty
check_external_program(
log=log,
program='qpdf',
package='qpdf',
version_checker=qpdf.version,
need_version='8.0.2',
)
check_external_program(
log=log,
program='pngquant',
package='pngquant',
version_checker=pngquant.version,
need_version='2.0.0',
optional=True,
program='qpdf',
package='qpdf',
version_checker=qpdf.version,
need_version='8.0.2',
)
@@ -1091,7 +1060,7 @@ def run_pipeline(args=None):
)
preamble(_log)
check_options(options, _log)
check_dependency_versions(_log)
check_dependency_versions(options, _log)
# Any changes to options will not take effect for options that are already
# bound to function parameters in the pipeline. (For example
+2 -2
View File
@@ -58,9 +58,9 @@ def verify_python3_env(): # pragma: no cover
if os.name == 'posix':
import subprocess
rv = subprocess.Popen(
rv = subprocess.run(
['locale', '-a'], stdout=subprocess.PIPE, stderr=subprocess.PIPE
).communicate()[0]
).stdout
good_locales = set()
has_c_utf8 = False
+57 -44
View File
@@ -21,7 +21,7 @@ import os
import re
import sys
from subprocess import run, STDOUT, PIPE, CalledProcessError
from ..exceptions import MissingDependencyError
from ..exceptions import MissingDependencyError, ExitCode
from collections.abc import Mapping
@@ -43,7 +43,7 @@ def get_version(program, *, version_arg='--version', regex=r'(\d+(\.\d+)*)'):
f"Could not find program '{program}' on the PATH"
) from e
except CalledProcessError as e:
if e.returncode < 0:
if e.returncode != 0:
raise MissingDependencyError(
f"Ran program '{program}' but it exited with an error:\n{e.output}"
) from e
@@ -66,10 +66,17 @@ The program '{program}' could not be executed or was not found on your
system PATH.
'''
unknown_version = '''
OCRmyPDF requires '{program}' {need_version} or higher. Your system has
'{program}' but we cannot tell what version is installed. Contact the
package maintainer.
missing_optional_program = '''
The program '{program}' could not be executed or was not found on your
system PATH. This program is required when you use the
{required_for} arguments. You could try omitting these arguments, or install
the package.
'''
missing_recommend_program = '''
The program '{program}' could not be executed or was not found on your
system PATH. This program is recommended when using the {required_for} arguments,
but not required, so we will proceed. For best results, install the program.
'''
old_version = '''
@@ -77,20 +84,15 @@ OCRmyPDF requires '{program}' {need_version} or higher. Your system appears
to have {found_version}. Please update this program.
'''
okay_its_optional = '''
This program is OPTIONAL, so installation of OCRmyPDF can proceed, but
some functionality may be missing.
'''
not_okay_its_required = '''
This program is REQUIRED for OCRmyPDF to work. Installation will abort.
old_version_required_for = '''
OCRmyPDF requires '{program}' {need_version} or higher when run with the
{required_for} arguments. If you omit these arguments, OCRmyPDF may be able to
proceed. For best results, install the program.
'''
osx_install_advice = '''
If you have homebrew installed, try these command to install the missing
packages:
brew update
brew upgrade
package:
brew install {package}
'''
@@ -105,7 +107,7 @@ installing the RPM for {program}.
'''
def get_platform():
def _get_platform():
if sys.platform.startswith('freebsd'):
return 'freebsd'
elif sys.platform.startswith('linux'):
@@ -113,48 +115,59 @@ def get_platform():
return sys.platform
def _error_trailer(log, program, package, optional, **kwargs):
if optional:
log.error(okay_its_optional.format(**locals()), file=sys.stderr)
else:
log.error(not_okay_its_required.format(**locals()), file=sys.stderr)
def _error_trailer(log, program, package, **kwargs):
if isinstance(package, Mapping):
package = package[get_platform()]
package = package[_get_platform()]
if get_platform() == 'darwin':
log.error(osx_install_advice.format(**locals()), file=sys.stderr)
elif get_platform() == 'linux':
log.error(linux_install_advice.format(**locals()), file=sys.stderr)
if _get_platform() == 'darwin':
log.info(osx_install_advice.format(**locals()))
elif _get_platform() == 'linux':
log.info(linux_install_advice.format(**locals()))
def error_missing_program(log, program, package, optional):
log.error(missing_program.format(**locals()), file=sys.stderr)
_error_trailer(log, **locals())
def _error_missing_program(log, program, package, required_for, recommended):
if required_for:
log.error(missing_optional_program.format(**locals()))
elif recommended:
log.info(missing_recommend_program.format(**locals()))
else:
log.error(missing_program.format(**locals()))
_error_trailer(**locals())
def error_unknown_version(log, program, package, optional, need_version):
log.error(unknown_version.format(**locals()), file=sys.stderr)
_error_trailer(log, **locals())
def error_old_version(log, program, package, optional, need_version, found_version):
log.error(old_version.format(**locals()), file=sys.stderr)
_error_trailer(log, **locals())
def _error_old_version(
log, program, package, need_version, found_version, required_for
):
if required_for:
log.error(old_version_required_for.format(**locals()))
else:
log.error(old_version.format(**locals()))
_error_trailer(**locals())
def check_external_program(
log, program, package, version_checker, need_version, optional=False
*,
log,
program,
package,
version_checker,
need_version,
required_for=None,
recommended=False,
):
try:
found_version = version_checker()
except (CalledProcessError, FileNotFoundError, MissingDependencyError):
error_missing_program(log, program, package, optional)
if not optional:
sys.exit(1)
_error_missing_program(log, program, package, required_for, recommended)
if not recommended:
sys.exit(ExitCode.missing_dependency)
return
if found_version < need_version:
error_old_version(log, program, package, optional, need_version, found_version)
_error_old_version(
log, program, package, need_version, found_version, required_for
)
if not recommended:
sys.exit(ExitCode.missing_dependency)
log.debug(f'Found {program} {found_version}')
+24 -20
View File
@@ -43,11 +43,6 @@ lept = ffi.dlopen(find_library('lept'))
lept.setMsgSeverity(lept.L_SEVERITY_WARNING)
def stderr(*objs):
"""Shorthand print to stderr."""
print("leptonica.py:", *objs, file=sys.stderr)
class _LeptonicaErrorTrap:
"""
Context manager to trap errors reported by Leptonica.
@@ -66,6 +61,7 @@ class _LeptonicaErrorTrap:
def __init__(self):
self.tmpfile = None
self.copy_of_stderr = -1
self.no_stderr = False
def __enter__(self):
from io import UnsupportedOperation
@@ -73,33 +69,40 @@ class _LeptonicaErrorTrap:
self.tmpfile = TemporaryFile()
# Save the old stderr, and redirect stderr to temporary file
sys.stderr.flush()
with suppress(AttributeError):
sys.stderr.flush()
try:
self.copy_of_stderr = os.dup(sys.stderr.fileno())
os.dup2(self.tmpfile.fileno(), sys.stderr.fileno(), inheritable=False)
except AttributeError:
# We are in some unusual context where our Python process does not
# have a sys.stderr. Leptonica still expects to write to file
# descriptor 2, so we are going to ensure it is redirected.
self.copy_of_stderr = None
self.no_stderr = True
os.dup2(self.tmpfile.fileno(), 2, inheritable=False)
except UnsupportedOperation:
self.copy_of_stderr = None
return
def __exit__(self, exc_type, exc_value, traceback):
# Restore old stderr
sys.stderr.flush()
with suppress(AttributeError):
sys.stderr.flush()
if self.copy_of_stderr is not None:
os.dup2(self.copy_of_stderr, sys.stderr.fileno())
os.close(self.copy_of_stderr)
if self.no_stderr:
os.close(2)
# Get data from tmpfile (in with block to ensure it is closed)
with self.tmpfile as tmpfile:
tmpfile.seek(0) # Cursor will be at end, so move back to beginning
leptonica_output = tmpfile.read().decode(errors='replace')
assert self.tmpfile.closed
assert not sys.stderr.closed
# If there are Python errors, let them bubble up
# Get data from tmpfile
self.tmpfile.seek(0) # Cursor will be at end, so move back to beginning
leptonica_output = self.tmpfile.read().decode(errors='replace')
self.tmpfile.close()
# If there are Python errors, record them
if exc_type:
logger.warning(leptonica_output)
return False
# If there are Leptonica errors, wrap them in Python excpetions
if 'Error' in leptonica_output:
@@ -614,9 +617,10 @@ class Pix(LeptonicaObject):
except (LeptonicaError, ValueError, IndexError):
return
finally:
with suppress(FileNotFoundError):
os.unlink('junkpixt.png') # leptonica may produce this
os.unlink('junkpixt')
leptonica_junk = ('junkpixt.png', 'junkpixt')
for junk in leptonica_junk:
with suppress(FileNotFoundError):
os.unlink(junk) # leptonica may produce this
for n, s in enumerate(sarray):
decoded = s.decode()
+21 -13
View File
@@ -19,7 +19,7 @@ import os
import platform
import sys
from pathlib import Path
from subprocess import PIPE, Popen
from subprocess import PIPE, run
import pytest
@@ -72,6 +72,17 @@ def needs_pdfminer(fn):
return fn
@pytest.helpers.register
def have_unpaper():
try:
from ocrmypdf.exec import unpaper
unpaper.version()
except Exception:
return False
return True
TESTS_ROOT = os.path.abspath(os.path.dirname(__file__))
SPOOF_PATH = os.path.join(TESTS_ROOT, 'spoof')
PROJECT_ROOT = os.path.dirname(TESTS_ROOT)
@@ -171,19 +182,16 @@ def run_ocrmypdf(input_file, output_file, *args, env=None, universal_newlines=Tr
if env is None:
env = os.environ
p_args = OCRMYPDF + [str(arg) for arg in args] + [str(input_file), str(output_file)]
p = Popen(
p_args,
close_fds=True,
stdout=PIPE,
stderr=PIPE,
universal_newlines=universal_newlines,
env=env,
p_args = (
OCRMYPDF
+ [str(arg) for arg in args if arg is not None]
+ [str(input_file), str(output_file)]
)
out, err = p.communicate()
# print(err)
return p, out, err
p = run(
p_args, stdout=PIPE, stderr=PIPE, universal_newlines=universal_newlines, env=env
)
# print(p.stderr)
return p, p.stdout, p.stderr
@pytest.helpers.register
+17
View File
@@ -17,7 +17,9 @@
from os import fspath
import os
from pickle import dumps, loads
from unittest.mock import patch
import pytest
from PIL import Image, ImageChops
@@ -85,3 +87,18 @@ def test_leptonica_compile(tmpdir):
# existing compiled library. Also compile in API mode so that we test
# the interfaces, even though we use it ABI mode.
ffibuilder.compile(tmpdir=fspath(tmpdir), target=fspath(tmpdir / 'lepttest.*'))
def test_with_stderr(capsys):
# pytest redirects stderr too; we must disable this for the test to be valid
with capsys.disabled():
with pytest.raises(FileNotFoundError):
lept.Pix.open("does_not_exist1")
def test_without_stderr(capsys):
# pytest redirects stderr too; we must disable this for the test to be valid
with capsys.disabled():
with patch('sys.stderr', new=None):
with pytest.raises(FileNotFoundError):
lept.Pix.open("does_not_exist2")
+15 -26
View File
@@ -21,14 +21,14 @@ import shutil
import sys
from math import isclose
from pathlib import Path
from subprocess import DEVNULL, PIPE, Popen
from subprocess import DEVNULL, PIPE, run, Popen
import PIL
import pytest
from PIL import Image
from ocrmypdf.exceptions import ExitCode, MissingDependencyError
from ocrmypdf.exec import ghostscript, qpdf, tesseract
from ocrmypdf.exec import ghostscript, qpdf, tesseract, unpaper
from ocrmypdf.leptonica import Pix
from ocrmypdf.pdfa import file_claims_pdfa
from ocrmypdf.pdfinfo import Colorspace, Encoding, PdfInfo
@@ -164,7 +164,7 @@ def test_exotic_image(
check_ocrmypdf(
resources / pdf,
outfile,
'-dc',
'-dc' if pytest.helpers.have_unpaper() else '-d',
'-v',
'1',
'--output-type',
@@ -282,8 +282,7 @@ def test_maximum_options(
resources / 'multipage.pdf',
outpdf,
'-d',
'-c',
'-i',
'-ci' if pytest.helpers.have_unpaper() else None,
'-f',
'-k',
'--oversample',
@@ -542,7 +541,6 @@ def test_jbig2_passthrough(spoof_tesseract_cache, resources, outpdf):
'hocr',
env=spoof_tesseract_cache,
)
out_pageinfo = PdfInfo(out)
assert out_pageinfo[0].images[0].enc == Encoding.jbig2
@@ -554,16 +552,13 @@ def test_stdin(spoof_tesseract_noop, ocrmypdf_exec, resources, outpdf):
# Runs: ocrmypdf - output.pdf < testfile.pdf
with open(input_file, 'rb') as input_stream:
p_args = ocrmypdf_exec + ['-', output_file]
p = Popen(
p = run(
p_args,
close_fds=True,
stdout=PIPE,
stderr=PIPE,
stdin=input_stream,
env=spoof_tesseract_noop,
)
out, err = p.communicate()
assert p.returncode == ExitCode.ok
@@ -574,16 +569,13 @@ def test_stdout(spoof_tesseract_noop, ocrmypdf_exec, resources, outpdf):
# Runs: ocrmypdf francais.pdf - > test_stdout.pdf
with open(output_file, 'wb') as output_stream:
p_args = ocrmypdf_exec + [input_file, '-']
p = Popen(
p = run(
p_args,
close_fds=True,
stdout=output_stream,
stderr=PIPE,
stdin=DEVNULL,
env=spoof_tesseract_noop,
)
out, err = p.communicate()
assert p.returncode == ExitCode.ok
assert qpdf.check(output_file, log=None)
@@ -781,10 +773,10 @@ def test_pagesize_consistency(renderer, resources, outpdf):
outpdf,
'--pdf-renderer',
renderer,
'--clean',
'--clean' if pytest.helpers.have_unpaper() else None,
'--deskew',
'--remove-background',
'--clean-final',
'--clean-final' if pytest.helpers.have_unpaper() else None,
)
after_dims = first_page_dimensions(outpdf)
@@ -852,22 +844,21 @@ def test_compression_preserved(
'-',
output_file,
]
p = Popen(
p = run(
p_args,
close_fds=True,
stdout=PIPE,
stderr=PIPE,
stdin=input_stream,
universal_newlines=True,
env=spoof_tesseract_noop,
)
out, err = p.communicate()
if im.mode in ('RGBA', 'LA'):
# If alpha image is input, expect an error
assert p.returncode != ExitCode.ok and b'alpha' in err
assert p.returncode != ExitCode.ok and 'alpha' in p.stderr
return
assert p.returncode == ExitCode.ok, err.decode('utf-8')
assert p.returncode == ExitCode.ok, p.stderr
pdfinfo = PdfInfo(output_file)
@@ -913,17 +904,15 @@ def test_compression_changed(
'-',
output_file,
]
p = Popen(
p = run(
p_args,
close_fds=True,
stdout=PIPE,
stderr=PIPE,
stdin=input_stream,
universal_newlines=True,
env=spoof_tesseract_noop,
)
out, err = p.communicate()
assert p.returncode == ExitCode.ok, err
assert p.returncode == ExitCode.ok, p.stderr
pdfinfo = PdfInfo(output_file)
+2 -2
View File
@@ -117,10 +117,10 @@ def test_pagesize_consistency_tess4(ensure_tess4, resources, outpdf):
outpdf,
'--pdf-renderer',
'sandwich',
'--clean',
'--clean' if pytest.helpers.have_unpaper() else None,
'--deskew',
'--remove-background',
'--clean-final',
'--clean-final' if pytest.helpers.have_unpaper() else None,
env=ensure_tess4,
)
+1 -9
View File
@@ -15,20 +15,12 @@
# You should have received a copy of the GNU General Public License
# along with OCRmyPDF. If not, see <http://www.gnu.org/licenses/>.
import logging
import os
import shutil
from math import isclose
from subprocess import DEVNULL, PIPE, Popen, check_call, check_output
import PyPDF2 as pypdf
import pytest
from ocrmypdf import leptonica
from ocrmypdf.exceptions import ExitCode
from ocrmypdf.exec import ghostscript
from ocrmypdf.pdfa import file_claims_pdfa
from ocrmypdf.pdfinfo import Colorspace, Encoding, PdfInfo
from ocrmypdf.pdfinfo import PdfInfo
check_ocrmypdf = pytest.helpers.check_ocrmypdf
run_ocrmypdf = pytest.helpers.run_ocrmypdf