Compare commits

..
20 Commits
Author SHA1 Message Date
James R. Barlow a2a197ce4c v9.0.2 release notes 2019-09-04 02:34:21 -07:00
James R. Barlow 944d59e5ad Fix --print-parameters issue when chi_sim is not installed 2019-09-04 01:17:52 -07:00
James R. Barlow 1c3e90a892 optimize: solve monochrome by converting to G4 2019-09-04 00:51:47 -07:00
James R. Barlow c728836956 Adjust test requirements 2019-09-04 00:50:48 -07:00
James R. Barlow 0d80fab339 Remove restriction on pytest < 5 2019-09-03 23:47:55 -07:00
James R. Barlow a650caa599 optimize: don't consider 1bpp images for PNG optimization 2019-09-03 23:47:20 -07:00
James R. Barlow c6caff90a1 optimize: only re-insert pngs after pngquant
Previously we attempted to reinsert all PNGs, but it appears to be
unlikely that Leptonica's API is actually capable of optimizing the PNG
before it inserts it.

In any event qpdf has gained image optimization capabilities as well
which we coudld borrow.
2019-09-03 23:46:25 -07:00
James R. Barlow 671c88d3b5 optimize: exclude images with custom Decode tables 2019-09-03 23:37:23 -07:00
James R. Barlow b2cfaedf91 optimize: Don't reinsert 1bpp images
There seems to be version to version inconsistencies between
Leptonica's photometric interpretation of 1bpp images, in
particular commit a0692307 introduces a change to force transcoding
in this situation.

However, I never entirely got to the bottom of where the problem
is, and in any event 1bpp images are probably better optimized
by JBIG2 than pngquant, so we're going to stop running them through
pngquant.
2019-09-03 23:26:13 -07:00
James R. Barlow 19ba3ae011 Allow test_german to xfail if deu language is not installed 2019-09-03 17:38:54 -07:00
James R. Barlow feff1e38bb Use context managers to ensure Pillow images are closed 2019-09-03 17:19:12 -07:00
James R. Barlow c8d6ea6b10 Fix tests broken by --print-parameters change 2019-09-03 17:17:24 -07:00
James R. Barlow b0d9775343 Attempt to resolve black-inversion issue 2019-08-31 01:25:36 -07:00
James R. Barlow 462bfb84fb install: affirm that we now require Tesseract beta 2019-08-31 01:24:31 -07:00
James R. Barlow 11ef78a891 Fix running without eng.traineddata installed raises exception 2019-08-27 14:54:03 -07:00
James R. Barlow 638eb556ef Reactivate user-words test that was always skipped 2019-08-27 14:52:59 -07:00
James R. Barlow fdefcd8af2 travis: Make 3.7 the build leader/deployer 2019-08-26 13:30:07 -07:00
James R. Barlow 09457edad3 alpine: use jbig2enc@community 2019-08-26 12:49:47 -07:00
James R. Barlow 6460a7eb3e docs: leptonica.com -> .org 2019-08-26 12:07:34 -07:00
James R. Barlow 707ebeb151 docs: installation updates 2019-08-11 18:48:56 -07:00
22 changed files with 218 additions and 172 deletions
+1 -1
View File
@@ -14,7 +14,7 @@ RUN \
&& apk add --update \ && apk add --update \
python3-dev \ python3-dev \
py3-setuptools \ py3-setuptools \
jbig2enc@testing \ jbig2enc@community \
ghostscript \ ghostscript \
qpdf@community \ qpdf@community \
qpdf-dev@community \ qpdf-dev@community \
+2 -2
View File
@@ -127,7 +127,7 @@ script:
deploy: deploy:
# release for main pypi # release for main pypi
# 3.6 is considered the build leader and does the deploy, otherwise there is # 3.7 is considered the build leader and does the deploy, otherwise there is
# a race and all versions will try to deploy # a race and all versions will try to deploy
# OTOH if we ever need separate binary wheels then each version needs its # OTOH if we ever need separate binary wheels then each version needs its
# own deploy # own deploy
@@ -139,5 +139,5 @@ deploy:
on: on:
branch: master branch: master
tags: true tags: true
condition: $TRAVIS_PYTHON_VERSION == "3.6" && $TRAVIS_OS_NAME == "linux" condition: $TRAVIS_PYTHON_VERSION == "3.7" && $TRAVIS_OS_NAME == "linux"
skip_upload_docs: true skip_upload_docs: true
+2 -2
View File
@@ -165,8 +165,8 @@ might remove desirable content, especially from poor quality scans.
- ``--deskew`` will correct pages were scanned at a skewed angle by - ``--deskew`` will correct pages were scanned at a skewed angle by
rotating them back into place. Skew determination and correction is rotating them back into place. Skew determination and correction is
performed using `Postl's variance of line performed using `Postl's variance of line
sums <http://www.leptonica.com/skew-measurement.html>`__ algorithm as sums <http://www.leptonica.org/skew-measurement.html>`__ algorithm as
implemented in `Leptonica <http://www.leptonica.com/index.html>`__. implemented in `Leptonica <http://www.leptonica.org/index.html>`__.
- ``--clean`` uses - ``--clean`` uses
`unpaper <https://www.flameeyes.eu/projects/unpaper>`__ to clean up `unpaper <https://www.flameeyes.eu/projects/unpaper>`__ to clean up
pages before OCR, but does not alter the final output. This makes it pages before OCR, but does not alter the final output. This makes it
+30 -24
View File
@@ -21,7 +21,7 @@ installing the Python binary wheels.
Installing on Linux Installing on Linux
=================== ===================
Debian and Ubuntu 16.10 or newer Debian and Ubuntu 18.04 or newer
-------------------------------- --------------------------------
.. |deb-stable| image:: https://repology.org/badge/version-for-repo/debian_stable/ocrmypdf.svg .. |deb-stable| image:: https://repology.org/badge/version-for-repo/debian_stable/ocrmypdf.svg
@@ -33,27 +33,29 @@ Debian and Ubuntu 16.10 or newer
.. |deb-unstable| image:: https://repology.org/badge/version-for-repo/debian_unstable/ocrmypdf.svg .. |deb-unstable| image:: https://repology.org/badge/version-for-repo/debian_unstable/ocrmypdf.svg
:alt: Debian unstable :alt: Debian unstable
.. |ubu-1710| image:: https://repology.org/badge/version-for-repo/ubuntu_17_10/ocrmypdf.svg
:alt: Ubuntu 17.10
.. |ubu-1804| image:: https://repology.org/badge/version-for-repo/ubuntu_18_04/ocrmypdf.svg .. |ubu-1804| image:: https://repology.org/badge/version-for-repo/ubuntu_18_04/ocrmypdf.svg
:alt: Ubuntu 18.04 LTS :alt: Ubuntu 18.04 LTS
.. |ubu-1810| image:: https://repology.org/badge/version-for-repo/ubuntu_18_10/ocrmypdf.svg .. |ubu-1810| image:: https://repology.org/badge/version-for-repo/ubuntu_18_10/ocrmypdf.svg
:alt: Ubuntu 18.10 :alt: Ubuntu 18.10
.. |ubu-1904| image:: https://repology.org/badge/version-for-repo/ubuntu_19_04/ocrmypdf.svg
:alt: Ubuntu 19.04
+-------------------------------------------+ .. |ubu-1910| image:: https://repology.org/badge/version-for-repo/ubuntu_19_10/ocrmypdf.svg
| **OCRmyPDF versions in Debian & Ubuntu** | :alt: Ubuntu 19.10
+-------------------------------------------+
| |latest| |
+-------------------------------------------+
| |deb-stable| |deb-testing| |deb-unstable| |
+-------------------------------------------+
| |ubu-1710| |ubu-1804| |ubu-1810| |
+-------------------------------------------+
Users of Debian 9 ("stretch") or later or Ubuntu 16.10 or later may +-----------------------------------------------+
| **OCRmyPDF versions in Debian & Ubuntu** |
+-----------------------------------------------+
| |latest| |
+-----------------------------------------------+
| |deb-stable| |deb-testing| |deb-unstable| |
+-----------------------------------------------+
| |ubu-1804| |ubu-1810| |ubu-1904| |ubu-1910| |
+-----------------------------------------------+
Users of Debian 9 ("stretch") or later or Ubuntu 18.04 or later may
simply simply
.. code-block:: bash .. code-block:: bash
@@ -64,7 +66,8 @@ As indicated in the table above, Debian and Ubuntu releases may lag
behind the latest version. If the version available for your platform is behind the latest version. If the version available for your platform is
out of date, you could opt to install the latest version from source. out of date, you could opt to install the latest version from source.
See `Installing HEAD revision from See `Installing HEAD revision from
sources <#installing-head-revision-from-sources>`__. sources <#installing-head-revision-from-sources>`__. Ubuntu 16.10 to 17.10
inclusive also had ocrmypdf, but these versions are end of life.
For full details on version availability for your platform, check the For full details on version availability for your platform, check the
`Debian Package Tracker <https://tracker.debian.org/pkg/ocrmypdf>`__ or `Debian Package Tracker <https://tracker.debian.org/pkg/ocrmypdf>`__ or
@@ -81,19 +84,22 @@ For full details on version availability for your platform, check the
Fedora 29 or newer Fedora 29 or newer
------------------ ------------------
.. |fedora-29| image:: https://repology.org/badge/version-for-repo/fedora29/ocrmypdf.svg .. |fedora-29| image:: https://repology.org/badge/version-for-repo/fedora_29/ocrmypdf.svg
:alt: Fedora 29 :alt: Fedora 29
.. |fedora-30| image:: https://repology.org/badge/version-for-repo/fedora_30/ocrmypdf.svg
:alt: Fedora 30
.. |fedora-rawhide| image:: https://repology.org/badge/version-for-repo/fedora_rawhide/ocrmypdf.svg .. |fedora-rawhide| image:: https://repology.org/badge/version-for-repo/fedora_rawhide/ocrmypdf.svg
:alt: Fedore Rawhide :alt: Fedore Rawhide
+------------------------------+ +-----------------------------------------------+
| **OCRmyPDF version** | | **OCRmyPDF version** |
+------------------------------+ +-----------------------------------------------+
| |latest| | | |latest| |
+------------------------------+ +-----------------------------------------------+
| |fedora-29| |fedora-rawhide| | | |fedora-29| |fedora-30| |fedora-rawhide| |
+------------------------------+ +-----------------------------------------------+
Users of Fedora 29 later may simply Users of Fedora 29 later may simply
@@ -507,7 +513,7 @@ manager. ``pip`` cannot provide them.
- Python 3.6 or newer - Python 3.6 or newer
- Ghostscript 9.15 or newer - Ghostscript 9.15 or newer
- qpdf 8.1.0 or newer - qpdf 8.1.0 or newer
- Tesseract 4.0.0-alpha or newer - Tesseract 4.0.0-beta or newer
As of ocrmypdf 7.2.1, the following versions are recommended: As of ocrmypdf 7.2.1, the following versions are recommended:
+19 -2
View File
@@ -13,14 +13,31 @@ Note that it is licensed under GPLv3, so scripts that
``import ocrmypdf`` and are released publicly should probably also be ``import ocrmypdf`` and are released publicly should probably also be
licensed under GPLv3. licensed under GPLv3.
v9.0.2
======
- The image optimizer now skips optimizing flate (PNG) encoded images in some
situations where the optimization effort was likely wasted.
- The image optimizer now ignores images that specify arbitrary decode arrays,
since these are rare.
- Fixed an issue that caused inversion of black and white in monochrome images.
We are not certain but the problem seems to be linked to Leptonica 1.76.0 and
older.
- Fixed some cases where the test suite failed or produced unexpected if
English or German Tesseract language packs were not installed.
- Fixed a runtime error if the Tesseract English language is not installed.
- Improved explicit closing of Pillow images after use.
- Actually fixed of Alpine Docker image build.
- Changed to pikepdf 1.6.3.
v9.0.1 v9.0.1
====== ======
- Fixed test suite failing when either of optional dependencies unpaper and - Fixed test suite failing when either of optional dependencies unpaper and
pngquant were missing. pngquant were missing.
- Fixed Alpine Docker image build. - Attempted fix of Alpine Docker image build.
- Documented that FreeBSD ports are now available. - Documented that FreeBSD ports are now available.
- Changed to pikepdf 1.6.1 (also for Alpine Docker). - Changed to pikepdf 1.6.1.
v9.0.0 v9.0.0
====== ======
+2 -2
View File
@@ -1,6 +1,6 @@
pytest >= 4.4.1, < 5 pytest >= 5.0.0
pytest-helpers-namespace >= 2019.1.8 pytest-helpers-namespace >= 2019.1.8
pytest-xdist == 1.28.0 pytest-xdist >= 1.29.0 # For DumpError fix
pytest-cov >= 2.6.1 pytest-cov >= 2.6.1
python-xmp-toolkit # requires apt-get install libexempi3 python-xmp-toolkit # requires apt-get install libexempi3
# or brew install exempi # or brew install exempi
+2 -3
View File
@@ -54,9 +54,9 @@ def triage_image_file(input_file, output_file, options, log):
# Recover the original filename # Recover the original filename
log.error(str(e).replace(input_file, options.input_file)) log.error(str(e).replace(input_file, options.input_file))
raise UnsupportedImageFormatError() from e raise UnsupportedImageFormatError() from e
else:
log.info("Input file is an image")
with im:
log.info("Input file is an image")
if 'dpi' in im.info: if 'dpi' in im.info:
if im.info['dpi'] <= (96, 96) and not options.image_dpi: if im.info['dpi'] <= (96, 96) and not options.image_dpi:
log.info("Image size: (%d, %d)" % im.size) log.info("Image size: (%d, %d)" % im.size)
@@ -89,7 +89,6 @@ def triage_image_file(input_file, output_file, options, log):
elif im.mode == 'CMYK': elif im.mode == 'CMYK':
log.info('Input CMYK image has no ICC profile, not usable') log.info('Input CMYK image has no ICC profile, not usable')
raise UnsupportedImageFormatError() raise UnsupportedImageFormatError()
im.close()
try: try:
log.info("Image seems valid. Try converting to PDF...") log.info("Image seems valid. Try converting to PDF...")
+1 -1
View File
@@ -107,7 +107,7 @@ def check_options_output(options):
options.pdf_renderer = 'sandwich' options.pdf_renderer = 'sandwich'
if options.pdf_renderer == 'sandwich' and not tesseract.has_textonly_pdf( if options.pdf_renderer == 'sandwich' and not tesseract.has_textonly_pdf(
options.tesseract_env options.tesseract_env, languages
): ):
raise MissingDependencyError( raise MissingDependencyError(
"You are using an alpha version of Tesseract 4.0 that does not support " "You are using an alpha version of Tesseract 4.0 that does not support "
+1 -2
View File
@@ -40,8 +40,7 @@ def available():
def quantize(input_file, output_file, quality_min, quality_max): def quantize(input_file, output_file, quality_min, quality_max):
if input_file.endswith('.jpg'): if input_file.endswith('.jpg'):
im = Image.open(input_file) with Image.open(input_file) as im, NamedTemporaryFile(suffix='.png') as tmp:
with NamedTemporaryFile(suffix='.png') as tmp:
im.save(tmp) im.save(tmp)
args = [ args = [
'pngquant', 'pngquant',
+6 -6
View File
@@ -61,13 +61,13 @@ def v4(tesseract_env=None):
return version(tesseract_env) >= '4' return version(tesseract_env) >= '4'
def has_textonly_pdf(tesseract_env=None): def has_textonly_pdf(tesseract_env=None, langs=None):
"""Does Tesseract have textonly_pdf capability? """Does Tesseract have textonly_pdf capability?
Available in v4.00.00alpha since January 2017. Best to Available in v4.00.00alpha since January 2017. Best to
parse the parameter list parse the parameter list.
""" """
args_tess = ['tesseract', '--print-parameters', 'pdf'] args_tess = tess_base_args(langs, engine_mode=None) + ['--print-parameters', 'pdf']
params = '' params = ''
try: try:
proc = run( proc = run(
@@ -233,8 +233,8 @@ def _generate_null_hocr(output_hocr, output_sidecar, image):
the same size as the input image.""" the same size as the input image."""
from PIL import Image from PIL import Image
im = Image.open(image) with Image.open(image) as im:
w, h = im.size w, h = im.size
with open(output_hocr, 'w', encoding="utf-8") as f: with open(output_hocr, 'w', encoding="utf-8") as f:
f.write(HOCR_TEMPLATE.format(w, h)) f.write(HOCR_TEMPLATE.format(w, h))
@@ -358,7 +358,7 @@ def generate_pdf(
if pagesegmode is not None: if pagesegmode is not None:
args_tesseract.extend(['--psm', str(pagesegmode)]) args_tesseract.extend(['--psm', str(pagesegmode)])
if text_only and has_textonly_pdf(tesseract_env): if text_only and has_textonly_pdf(tesseract_env, language):
args_tesseract.extend(['-c', 'textonly_pdf=1']) args_tesseract.extend(['-c', 'textonly_pdf=1'])
if user_words: if user_words:
+21 -22
View File
@@ -42,33 +42,30 @@ def run(input_file, output_file, dpi, log, mode_args):
SUFFIXES = {'1': '.pbm', 'L': '.pgm', 'RGB': '.ppm'} SUFFIXES = {'1': '.pbm', 'L': '.pgm', 'RGB': '.ppm'}
im = Image.open(input_file) with TemporaryDirectory() as tmpdir, Image.open(input_file) as im:
if im.mode not in SUFFIXES.keys(): if im.mode not in SUFFIXES.keys():
log.info("Converting image to other colorspace") log.info("Converting image to other colorspace")
try:
if im.mode == 'P' and len(im.getcolors()) == 2:
im = im.convert(mode='1')
else:
im = im.convert(mode='RGB')
except IOError as e:
im.close()
raise MissingDependencyError(
"Could not convert image with type " + im.mode
) from e
try: try:
if im.mode == 'P' and len(im.getcolors()) == 2: suffix = SUFFIXES[im.mode]
im = im.convert(mode='1') except KeyError:
else:
im = im.convert(mode='RGB')
except IOError as e:
im.close()
raise MissingDependencyError( raise MissingDependencyError(
"Could not convert image with type " + im.mode "Failed to convert image to a supported format."
) from e ) from e
try:
suffix = SUFFIXES[im.mode]
except KeyError:
im.close()
raise MissingDependencyError(
"Failed to convert image to a supported format."
) from e
with TemporaryDirectory() as tmpdir:
input_pnm = os.path.join(tmpdir, f'input{suffix}') input_pnm = os.path.join(tmpdir, f'input{suffix}')
output_pnm = os.path.join(tmpdir, f'output{suffix}') output_pnm = os.path.join(tmpdir, f'output{suffix}')
im.save(input_pnm, format='PPM') im.save(input_pnm, format='PPM')
im.close()
# To prevent any shenanigans from accepting arbitrary parameters in # To prevent any shenanigans from accepting arbitrary parameters in
# --unpaper-args, we: # --unpaper-args, we:
@@ -95,10 +92,12 @@ def run(input_file, output_file, dpi, log, mode_args):
log.debug(proc.stdout) log.debug(proc.stdout)
# unpaper sets dpi to 72; fix this # unpaper sets dpi to 72; fix this
try: try:
Image.open(output_pnm).save(output_file, dpi=(dpi, dpi)) with Image.open(output_pnm) as imout:
imout.save(output_file, dpi=(dpi, dpi))
except (FileNotFoundError, OSError): except (FileNotFoundError, OSError):
raise SubprocessOutputError( raise SubprocessOutputError(
"unpaper: failed to produce the expected output file. Called with: " "unpaper: failed to produce the expected output file. "
+ " Called with: "
+ str(args_unpaper) + str(args_unpaper)
) from None ) from None
+89 -55
View File
@@ -73,6 +73,9 @@ def extract_image_filter(pike, root, log, image, xref):
if filtdp[0] == Name.JPXDecode: if filtdp[0] == Name.JPXDecode:
return None # Don't do JPEG2000 return None # Don't do JPEG2000
if Name.Decode in image:
return None # Don't mess with custom Decode tables
return pim, filtdp return pim, filtdp
@@ -104,6 +107,10 @@ def extract_image_generic(*, pike, root, log, image, xref, options):
return None return None
pim, filtdp = result pim, filtdp = result
# Don't try to PNG-optimize 1bpp images, since JBIG2 does it better.
if pim.bits_per_component == 1:
return None
if filtdp[0] == Name.DCTDecode and options.optimize >= 2: if filtdp[0] == Name.DCTDecode and options.optimize >= 2:
# This is a simple heuristic derived from some training data, that has # This is a simple heuristic derived from some training data, that has
# about a 70% chance of guessing whether the JPEG is high quality, # about a 70% chance of guessing whether the JPEG is high quality,
@@ -343,6 +350,7 @@ def transcode_jpegs(pike, jpegs, root, log, options):
def transcode_pngs(pike, images, image_name_fn, root, log, options): def transcode_pngs(pike, images, image_name_fn, root, log, options):
modified = set()
if options.optimize >= 2: if options.optimize >= 2:
png_quality = ( png_quality = (
max(10, options.png_quality - 10), max(10, options.png_quality - 10),
@@ -363,6 +371,7 @@ def transcode_pngs(pike, images, image_name_fn, root, log, options):
png_quality[1], png_quality[1],
) )
) )
modified.add(xref)
with tqdm( with tqdm(
desc="PNGs", desc="PNGs",
total=len(futures), total=len(futures),
@@ -372,10 +381,14 @@ def transcode_pngs(pike, images, image_name_fn, root, log, options):
for _future in concurrent.futures.as_completed(futures): for _future in concurrent.futures.as_completed(futures):
pbar.update() pbar.update()
for xref in images: for xref in modified:
im_obj = pike.get_object(xref, 0) im_obj = pike.get_object(xref, 0)
try: try:
compdata = leptonica.CompressedData.open(png_name(root, xref)) pix = leptonica.Pix.open(png_name(root, xref))
if pix.mode == '1':
compdata = pix.generate_pdf_ci_data(leptonica.lept.L_G4_ENCODE, 0)
else:
compdata = leptonica.CompressedData.open(png_name(root, xref))
except leptonica.LeptonicaError as e: except leptonica.LeptonicaError as e:
# Most likely this means file not found, i.e. quantize did not # Most likely this means file not found, i.e. quantize did not
# produce an improved version # produce an improved version
@@ -391,62 +404,83 @@ def transcode_pngs(pike, images, image_name_fn, root, log, options):
f"{len(compdata)} > {int(im_obj.stream_dict.Length)}" f"{len(compdata)} > {int(im_obj.stream_dict.Length)}"
) )
continue continue
if compdata.type == leptonica.lept.L_FLATE_ENCODE:
return rewrite_png(pike, im_obj, compdata, log)
elif compdata.type == leptonica.lept.L_G4_ENCODE:
return rewrite_png_as_g4(pike, im_obj, compdata, log)
# When a PNG is inserted into a PDF, we more or less copy the IDAT section from
# the PDF and transfer the rest of the PNG headers to PDF image metadata.
# One thing we have to do is tell the PDF reader whether a predictor was used
# on the image before Flate encoding. (Typically one is.)
# According to Leptonica source, PDF readers don't actually need us
# to specify the correct predictor, they just need a value of either:
# 1 - no predictor
# 10-14 - there is a predictor
# Leptonica's compdata->predictor only tells TRUE or FALSE
# From there the PNG decoder can infer the rest from the file.
# In practice the predictor should be Paeth, 14, so we'll use that.
# See:
# - PDF RM 7.4.4.4 Table 10
# - https://github.com/DanBloomberg/leptonica/blob/master/src/pdfio2.c#L757
predictor = 14 if compdata.predictor > 0 else 1
dparms = Dictionary(Predictor=predictor)
if predictor > 1:
dparms.BitsPerComponent = compdata.bps # Yes, this is redundant
dparms.Colors = compdata.spp
dparms.Columns = compdata.w
im_obj.BitsPerComponent = compdata.bps def rewrite_png_as_g4(pike, im_obj, compdata, log):
im_obj.Width = compdata.w im_obj.BitsPerComponent = 1
im_obj.Height = compdata.h im_obj.Width = compdata.w
im_obj.Height = compdata.h
if compdata.ncolors > 0: im_obj.write(compdata.read())
# .ncolors is the number of colors in the palette, not the number of
# colors used in a true color image log.debug(f"PNG to G4 {im_obj.objgen}")
palette_pdf_string = compdata.get_palette_pdf_string() if Name.Predictor in im_obj:
palette_data = pikepdf.Object.parse(palette_pdf_string) del im_obj.Predictor
palette_stream = pikepdf.Stream(pike, bytes(palette_data)) if Name.DecodeParms in im_obj:
palette = [ del im_obj.DecodeParms
Name.Indexed, im_obj.DecodeParms = Dictionary(
Name.DeviceRGB, K=-1, BlackIs1=bool(compdata.minisblack), Columns=compdata.w
compdata.ncolors - 1, )
palette_stream,
] im_obj.Filter = Name.CCITTFaxDecode
cs = palette return
else:
if compdata.spp == 1:
# PDF interprets binary-1 as black in 1bpp, but PNG sets def rewrite_png(pike, im_obj, compdata, log):
# black to 0 for 1bpp. Create a palette that informs the PDF # When a PNG is inserted into a PDF, we more or less copy the IDAT section from
# of the mapping - seems cleaner to go this way but pikepdf # the PDF and transfer the rest of the PNG headers to PDF image metadata.
# needs to be patched to support it. # One thing we have to do is tell the PDF reader whether a predictor was used
# palette = [Name.Indexed, Name.DeviceGray, 1, b"\xff\x00"] # on the image before Flate encoding. (Typically one is.)
# cs = palette # According to Leptonica source, PDF readers don't actually need us
cs = Name.DeviceGray # to specify the correct predictor, they just need a value of either:
elif compdata.spp == 3: # 1 - no predictor
cs = Name.DeviceRGB # 10-14 - there is a predictor
elif compdata.spp == 4: # Leptonica's compdata->predictor only tells TRUE or FALSE
cs = Name.DeviceCMYK # 10-14 means the actual predictor is specified in the data, so for any
if compdata.bps == 1: # number >= 10 the PDF reader will use whatever the PNG data specifies.
im_obj.Decode = [1, 0] # Bit of a kludge but this inverts photometric too # In practice Leptonica should use Paeth, 14, but 15 seems to be the
im_obj.ColorSpace = cs # designated value for "optimal". So we will use 15.
im_obj.write(compdata.read(), filter=Name.FlateDecode, decode_parms=dparms) # See:
# - PDF RM 7.4.4.4 Table 10
# - https://github.com/DanBloomberg/leptonica/blob/master/src/pdfio2.c#L757
predictor = 15 if compdata.predictor > 0 else 1
dparms = Dictionary(Predictor=predictor)
if predictor > 1:
dparms.BitsPerComponent = compdata.bps # Yes, this is redundant
dparms.Colors = compdata.spp
dparms.Columns = compdata.w
im_obj.BitsPerComponent = compdata.bps
im_obj.Width = compdata.w
im_obj.Height = compdata.h
log.debug(
f"PNG {im_obj.objgen}: palette={compdata.ncolors} spp={compdata.spp} bps={compdata.bps}"
)
if compdata.ncolors > 0:
# .ncolors is the number of colors in the palette, not the number of
# colors used in a true color image. The palette string is always
# given as RGB tuples even when the image is grayscale; see
# https://github.com/DanBloomberg/leptonica/blob/master/src/colormap.c#L2067
palette_pdf_string = compdata.get_palette_pdf_string()
palette_data = pikepdf.Object.parse(palette_pdf_string)
palette_stream = pikepdf.Stream(pike, bytes(palette_data))
palette = [Name.Indexed, Name.DeviceRGB, compdata.ncolors - 1, palette_stream]
cs = palette
else:
# ncolors == 0 means we are using a colorspace without a palette
if compdata.spp == 1:
cs = Name.DeviceGray
elif compdata.spp == 3:
cs = Name.DeviceRGB
elif compdata.spp == 4:
cs = Name.DeviceCMYK
im_obj.ColorSpace = cs
im_obj.write(compdata.read(), filter=Name.FlateDecode, decode_parms=dparms)
def optimize(input_file, output_file, context, save_settings): def optimize(input_file, output_file, context, save_settings):
+1 -1
View File
@@ -52,7 +52,7 @@ def main():
elif sys.argv[1] == '--list-langs': elif sys.argv[1] == '--list-langs':
print('List of available languages (1):\neng', file=sys.stderr) print('List of available languages (1):\neng', file=sys.stderr)
sys.exit(0) sys.exit(0)
elif sys.argv[1] == '--print-parameters': elif sys.argv[-2] == '--print-parameters':
print("Some parameters", file=sys.stderr) print("Some parameters", file=sys.stderr)
print("textonly_pdf\t1\tSome help text") print("textonly_pdf\t1\tSome help text")
sys.exit(0) sys.exit(0)
+1 -1
View File
@@ -44,7 +44,7 @@ def main():
elif sys.argv[1] == '--list-langs': elif sys.argv[1] == '--list-langs':
print('List of available languages (1):\neng\n', file=sys.stderr) print('List of available languages (1):\neng\n', file=sys.stderr)
sys.exit(0) sys.exit(0)
elif sys.argv[1] == '--print-parameters': elif sys.argv[-2] == '--print-parameters':
print('A parameter list would go here\ntextonly_pdf 0\n', file=sys.stderr) print('A parameter list would go here\ntextonly_pdf 0\n', file=sys.stderr)
sys.exit(0) sys.exit(0)
elif sys.argv[-2] == 'hocr': elif sys.argv[-2] == 'hocr':
+2
View File
@@ -100,6 +100,8 @@ def main():
# Convert non-standard but supported -psm to --psm # Convert non-standard but supported -psm to --psm
sys.argv = ['--psm' if arg == '-psm' else arg for arg in sys.argv] sys.argv = ['--psm' if arg == '-psm' else arg for arg in sys.argv]
if '_OCRMYPDF_TEST_INFILE' not in os.environ:
real_tesseract() # test not properly set up
source = os.environ['_OCRMYPDF_TEST_INFILE'] # required source = os.environ['_OCRMYPDF_TEST_INFILE'] # required
args = parser.parse_args() args = parser.parse_args()
+1 -1
View File
@@ -50,7 +50,7 @@ def main():
elif sys.argv[1] == '--list-langs': elif sys.argv[1] == '--list-langs':
print('List of available languages (1):\neng', file=sys.stderr) print('List of available languages (1):\neng', file=sys.stderr)
sys.exit(0) sys.exit(0)
elif sys.argv[1] == '--print-parameters': elif sys.argv[-2] == '--print-parameters':
print('A parameter list would go here\ntextonly_pdf 0\n', file=sys.stderr) print('A parameter list would go here\ntextonly_pdf 0\n', file=sys.stderr)
sys.exit(0) sys.exit(0)
elif sys.argv[-2] == 'hocr': elif sys.argv[-2] == 'hocr':
+1 -1
View File
@@ -76,7 +76,7 @@ def main():
elif sys.argv[1] == '--list-langs': elif sys.argv[1] == '--list-langs':
print('List of available languages (1):\neng', file=sys.stderr) print('List of available languages (1):\neng', file=sys.stderr)
sys.exit(0) sys.exit(0)
elif sys.argv[1] == '--print-parameters': elif sys.argv[-2] == '--print-parameters':
print("Some parameters", file=sys.stderr) print("Some parameters", file=sys.stderr)
print("textonly_pdf\t1\tSome help text") print("textonly_pdf\t1\tSome help text")
sys.exit(0) sys.exit(0)
+2 -1
View File
@@ -38,7 +38,8 @@ def test_colormap_backgroundnorm(resources):
def crom_pix(resources): def crom_pix(resources):
pix = lept.Pix.open(resources / 'crom.png') pix = lept.Pix.open(resources / 'crom.png')
im = Image.open(resources / 'crom.png') im = Image.open(resources / 'crom.png')
return pix, im yield pix, im
im.close()
def test_pix_basic(crom_pix): def test_pix_basic(crom_pix):
+13 -30
View File
@@ -118,8 +118,8 @@ def test_deskew(spoof_tesseract_noop, resources, outdir):
def test_remove_background(spoof_tesseract_noop, resources, outdir): def test_remove_background(spoof_tesseract_noop, resources, outdir):
# Ensure the input image does not contain pure white/black # Ensure the input image does not contain pure white/black
im = Image.open(resources / 'congress.jpg') with Image.open(resources / 'congress.jpg') as im:
assert im.getextrema() != ((0, 255), (0, 255), (0, 255)) assert im.getextrema() != ((0, 255), (0, 255), (0, 255))
output_pdf = check_ocrmypdf( output_pdf = check_ocrmypdf(
resources / 'congress.jpg', resources / 'congress.jpg',
@@ -145,8 +145,8 @@ def test_remove_background(spoof_tesseract_noop, resources, outdir):
) )
# The output image should contain pure white and black # The output image should contain pure white and black
im = Image.open(output_png) with Image.open(output_png) as im:
assert im.getextrema() == ((0, 255), (0, 255), (0, 255)) assert im.getextrema() == ((0, 255), (0, 255), (0, 255))
# This will run 5 * 2 * 2 = 20 test cases # This will run 5 * 2 * 2 = 20 test cases
@@ -349,10 +349,9 @@ def test_german(spoof_tesseract_cache, resources, outdir):
sidecar, sidecar,
env=spoof_tesseract_cache, env=spoof_tesseract_cache,
) )
print(os.environ) if 'deu' not in tesseract.languages():
assert ( pytest.xfail(reason="tesseract-deu language pack not installed")
p.returncode == ExitCode.ok assert p.returncode == ExitCode.ok, "Requires tesseract deu language pack"
), "This test may fail if Tesseract language packs are missing"
def test_klingon(resources, outpdf): def test_klingon(resources, outpdf):
@@ -652,26 +651,11 @@ THIS FILE IS INVALID
assert p.returncode == ExitCode.invalid_config assert p.returncode == ExitCode.invalid_config
@pytest.mark.skipif(tesseract.v4(), reason='arg has no effect in 4.0-beta1') @pytest.mark.skipif(not tesseract.has_user_words(), reason='not functional until 4.1.0')
def test_user_words(resources, outdir): def test_user_words_ocr(resources, outdir):
# Does not actually test if --user-words causes output to differ
word_list = outdir / 'wordlist.txt' word_list = outdir / 'wordlist.txt'
sidecar_before = outdir / 'sidecar_before.txt' sidecar_after = outdir / 'sidecar.txt'
sidecar_after = outdir / 'sidecar_after.txt'
# Don't know how to make this test pass on various versions and platforms
# so weaken to merely testing that the argument is accepted
consistent = False
if consistent:
check_ocrmypdf(
resources / 'crom.png',
outdir / 'out.pdf',
'--image-dpi',
150,
'--sidecar',
sidecar_before,
)
assert 'cromulent' not in sidecar_before.open().read()
with word_list.open('w') as f: with word_list.open('w') as f:
f.write('cromulent\n') # a perfectly cromulent word f.write('cromulent\n') # a perfectly cromulent word
@@ -687,9 +671,6 @@ def test_user_words(resources, outdir):
word_list, word_list,
) )
if consistent:
assert 'cromulent' in sidecar_after.open().read()
def test_form_xobject(spoof_tesseract_noop, resources, outpdf): def test_form_xobject(spoof_tesseract_noop, resources, outpdf):
check_ocrmypdf( check_ocrmypdf(
@@ -810,6 +791,7 @@ def test_compression_preserved(
assert pdfimage.color == Colorspace.rgb, "Colorspace changed" assert pdfimage.color == Colorspace.rgb, "Colorspace changed"
elif im.mode.startswith('L'): elif im.mode.startswith('L'):
assert pdfimage.color == Colorspace.gray, "Colorspace changed" assert pdfimage.color == Colorspace.gray, "Colorspace changed"
im.close()
@pytest.mark.parametrize( @pytest.mark.parametrize(
@@ -871,6 +853,7 @@ def test_compression_changed(
assert pdfimage.color == Colorspace.rgb, "Colorspace changed" assert pdfimage.color == Colorspace.rgb, "Colorspace changed"
elif im.mode.startswith('L'): elif im.mode.startswith('L'):
assert pdfimage.color == Colorspace.gray, "Colorspace changed" assert pdfimage.color == Colorspace.gray, "Colorspace changed"
im.close()
def test_sidecar_pagecount(spoof_tesseract_cache, resources, outpdf): def test_sidecar_pagecount(spoof_tesseract_cache, resources, outpdf):
+7 -7
View File
@@ -48,11 +48,11 @@ def test_mono_not_inverted(resources, outdir):
xres=10, xres=10,
yres=10, yres=10,
raster_device='pnggray', raster_device='pnggray',
log=logging.getLogger(name='test_mono_flip'), log=logging.getLogger(name='test_mono_not_inverted'),
) )
im = Image.open(fspath(outdir / 'im.png')) with Image.open(fspath(outdir / 'im.png')) as im:
assert im.getpixel((0, 0)) == 255, "Expected white background" assert im.getpixel((0, 0)) == 255, "Expected white background"
@pytest.mark.skipif(not pngquant.available(), reason='need pngquant') @pytest.mark.skipif(not pngquant.available(), reason='need pngquant')
@@ -110,10 +110,10 @@ def test_flate_to_jbig2(resources, outdir, spoof_tesseract_noop):
# This test requires an image that pngquant is capable of converting to # This test requires an image that pngquant is capable of converting to
# to 1bpp - so use an existing 1bpp image, convert up, confirm it can # to 1bpp - so use an existing 1bpp image, convert up, confirm it can
# convert down # convert down
im = Image.open(fspath(resources / 'typewriter.png')) with Image.open(fspath(resources / 'typewriter.png')) as im:
assert im.mode in ('1', 'P') assert im.mode in ('1', 'P')
im = im.convert('L') im = im.convert('L')
im.save(fspath(outdir / 'type8.png')) im.save(fspath(outdir / 'type8.png'))
check_ocrmypdf( check_ocrmypdf(
outdir / 'type8.png', outdir / 'type8.png',
+5 -5
View File
@@ -224,12 +224,12 @@ def test_rotate_deskew_timeout(resources, outdir):
@pytest.mark.parametrize('image_angle', (0, 90, 180, 270)) @pytest.mark.parametrize('image_angle', (0, 90, 180, 270))
def test_rotate_page_level(image_angle, page_angle, resources, outdir): def test_rotate_page_level(image_angle, page_angle, resources, outdir):
def make_rotate_test(prefix, image_angle, page_angle): def make_rotate_test(prefix, image_angle, page_angle):
im = Image.open(fspath(resources / 'typewriter.png'))
if image_angle != 0:
ccw_angle = -image_angle % 360
im = im.transpose(getattr(Image, f'ROTATE_{ccw_angle}'))
memimg = BytesIO() memimg = BytesIO()
im.save(memimg, format='PNG') with Image.open(fspath(resources / 'typewriter.png')) as im:
if image_angle != 0:
ccw_angle = -image_angle % 360
im = im.transpose(getattr(Image, f'ROTATE_{ccw_angle}'))
im.save(memimg, format='PNG')
memimg.seek(0) memimg.seek(0)
mempdf = BytesIO() mempdf = BytesIO()
img2pdf.convert( img2pdf.convert(
+9 -3
View File
@@ -37,15 +37,21 @@ def test_hocr_notlatin_warning(caplog):
def test_old_ghostscript(caplog): def test_old_ghostscript(caplog):
with patch('ocrmypdf.exec.ghostscript.version', return_value='9.19'): with patch('ocrmypdf.exec.ghostscript.version', return_value='9.19'), patch(
'ocrmypdf.exec.tesseract.has_textonly_pdf', return_value=True
):
vd.check_options_output(make_opts(language='chi_sim', output_type='pdfa')) vd.check_options_output(make_opts(language='chi_sim', output_type='pdfa'))
assert 'Ghostscript does not work correctly' in caplog.text assert 'Ghostscript does not work correctly' in caplog.text
with patch('ocrmypdf.exec.ghostscript.version', return_value='9.18'): with patch('ocrmypdf.exec.ghostscript.version', return_value='9.18'), patch(
'ocrmypdf.exec.tesseract.has_textonly_pdf', return_value=True
):
with pytest.raises(MissingDependencyError): with pytest.raises(MissingDependencyError):
vd.check_options_output(make_opts(output_type='pdfa-3')) vd.check_options_output(make_opts(output_type='pdfa-3'))
with patch('ocrmypdf.exec.ghostscript.version', return_value='9.24'): with patch('ocrmypdf.exec.ghostscript.version', return_value='9.24'), patch(
'ocrmypdf.exec.tesseract.has_textonly_pdf', return_value=True
):
with pytest.raises(MissingDependencyError): with pytest.raises(MissingDependencyError):
vd.check_dependency_versions(make_opts()) vd.check_dependency_versions(make_opts())