Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
d8d9c41abb | ||
|
|
a8cad72f72 | ||
|
|
86c04305f4 | ||
|
|
0a110fac55 | ||
|
|
f8970ad862 | ||
|
|
fcfc78b7ee | ||
|
|
8bb244df24 | ||
|
|
8a1cb70479 | ||
|
|
2c579700d6 | ||
|
|
87ff6c8301 | ||
|
|
5915259bee | ||
|
|
969e54f0e3 | ||
|
|
b923612323 | ||
|
|
dc2b161306 | ||
|
|
ae49e3b6db |
@@ -18,8 +18,20 @@ jobs:
|
|||||||
runs-on: ${{ matrix.os }}
|
runs-on: ${{ matrix.os }}
|
||||||
strategy:
|
strategy:
|
||||||
matrix:
|
matrix:
|
||||||
os: [ubuntu-18.04] #, ubuntu-20.04]
|
include:
|
||||||
python: ["3.6"] #, "3.7", "3.8", "3.9"]
|
- os: ubuntu-18.04
|
||||||
|
python: 3.6
|
||||||
|
- os: ubuntu-18.04
|
||||||
|
python: 3.7
|
||||||
|
- os: ubuntu-20.04
|
||||||
|
python: 3.8
|
||||||
|
- os: ubuntu-20.04
|
||||||
|
python: 3.9
|
||||||
|
- os: ubuntu-latest
|
||||||
|
python: 3.9
|
||||||
|
- os: ubuntu-latest
|
||||||
|
python: 3.9
|
||||||
|
tesseract5: true
|
||||||
|
|
||||||
env:
|
env:
|
||||||
OS: ${{ matrix.os }}
|
OS: ${{ matrix.os }}
|
||||||
@@ -35,6 +47,11 @@ jobs:
|
|||||||
with:
|
with:
|
||||||
python-version: ${{ matrix.python }}
|
python-version: ${{ matrix.python }}
|
||||||
|
|
||||||
|
- name: Install Tesseract 5
|
||||||
|
if: matrix.tesseract5
|
||||||
|
run: |
|
||||||
|
sudo add-apt-repository ppa:alex-p/tesseract-ocr-devel
|
||||||
|
|
||||||
- name: Install common packages
|
- name: Install common packages
|
||||||
run: |
|
run: |
|
||||||
sudo apt-get update
|
sudo apt-get update
|
||||||
@@ -232,6 +249,7 @@ jobs:
|
|||||||
name: Build Docker images
|
name: Build Docker images
|
||||||
needs: [wheel_sdist_linux, test_linux, test_macos, test_windows]
|
needs: [wheel_sdist_linux, test_linux, test_macos, test_windows]
|
||||||
runs-on: ubuntu-latest
|
runs-on: ubuntu-latest
|
||||||
|
if: github.event_name != 'pull_request'
|
||||||
steps:
|
steps:
|
||||||
- name: Set image tag to release or branch
|
- name: Set image tag to release or branch
|
||||||
run: echo "DOCKER_IMAGE_TAG=${GITHUB_REF##*/}" >> $GITHUB_ENV
|
run: echo "DOCKER_IMAGE_TAG=${GITHUB_REF##*/}" >> $GITHUB_ENV
|
||||||
|
|||||||
@@ -91,6 +91,11 @@ brew install tesseract-lang
|
|||||||
|
|
||||||
You can then pass the `-l LANG` argument to OCRmyPDF to give a hint as to what languages it should search for. Multiple languages can be requested.
|
You can then pass the `-l LANG` argument to OCRmyPDF to give a hint as to what languages it should search for. Multiple languages can be requested.
|
||||||
|
|
||||||
|
OCRmyPDF supports Tesseract 4.0 and the beta versions of Tesseract 5.0. It will
|
||||||
|
automatically use whichever version it finds first on the `PATH` environment
|
||||||
|
variable. On Windows, if `PATH` does not provide a Tesseract binary, we use
|
||||||
|
the highest version number that is installed according to the Windows Registry.
|
||||||
|
|
||||||
## Documentation and support
|
## Documentation and support
|
||||||
|
|
||||||
Once OCRmyPDF is installed, the built-in help which explains the command syntax and options can be accessed via:
|
Once OCRmyPDF is installed, the built-in help which explains the command syntax and options can be accessed via:
|
||||||
@@ -115,7 +120,7 @@ In addition to the required Python version (3.6+), OCRmyPDF requires external pr
|
|||||||
- [heise Open Source, 09/2014: Texterkennung mit OCRmyPDF](https://heise.de/-2356670)
|
- [heise Open Source, 09/2014: Texterkennung mit OCRmyPDF](https://heise.de/-2356670)
|
||||||
- [heise Durchsuchbare PDF-Dokumente mit OCRmyPDF erstellen](https://www.heise.de/ratgeber/Durchsuchbare-PDF-Dokumente-mit-OCRmyPDF-erstellen-4607592.html)
|
- [heise Durchsuchbare PDF-Dokumente mit OCRmyPDF erstellen](https://www.heise.de/ratgeber/Durchsuchbare-PDF-Dokumente-mit-OCRmyPDF-erstellen-4607592.html)
|
||||||
- [Excellent Utilities: OCRmyPDF](https://www.linuxlinks.com/excellent-utilities-ocrmypdf-add-ocr-text-layer-scanned-pdfs/)
|
- [Excellent Utilities: OCRmyPDF](https://www.linuxlinks.com/excellent-utilities-ocrmypdf-add-ocr-text-layer-scanned-pdfs/)
|
||||||
- [LinuxUser Texterkennung mit OCRmyPDF und Scanbd automatisieren](https://www.linux-community.de/ausgaben/linuxuser/2021/06/texterkennung-mit-ocrmypdf-und-scanbd-automatisieren/)
|
- [LinuxUser Texterkennung mit OCRmyPDF und Scanbd automatisieren](https://www.linux-community.de/ausgaben/linuxuser/2021/06/texterkennung-mit-ocrmypdf-und-scanbd-automatisieren/)
|
||||||
|
|
||||||
## Business enquiries
|
## Business enquiries
|
||||||
|
|
||||||
|
|||||||
+3
-3
@@ -9,11 +9,11 @@ encoding was patented for a long time. All known JBIG2 US patents have
|
|||||||
expired as of 2017, but it is possible that unknown patents exist.
|
expired as of 2017, but it is possible that unknown patents exist.
|
||||||
|
|
||||||
JBIG2 encoding is recommended for OCRmyPDF and is used to losslessly
|
JBIG2 encoding is recommended for OCRmyPDF and is used to losslessly
|
||||||
create smaller PDFs. If JBIG2 encoding not available, lower quality
|
create smaller PDFs. If JBIG2 encoding is not available, lower quality
|
||||||
encodings will be used.
|
encodings will be used.
|
||||||
|
|
||||||
JBIG2 decoding is not patented and is performed automatically by most
|
JBIG2 decoding is not patented and is performed automatically by most
|
||||||
PDF viewers. It is widely supported has been part of the PDF
|
PDF viewers. It is widely supported and has been part of the PDF
|
||||||
specification since 2001.
|
specification since 2001.
|
||||||
|
|
||||||
On macOS, Homebrew packages jbig2enc and OCRmyPDF includes it by
|
On macOS, Homebrew packages jbig2enc and OCRmyPDF includes it by
|
||||||
@@ -37,7 +37,7 @@ Lossy mode JBIG2
|
|||||||
|
|
||||||
OCRmyPDF provides lossy mode JBIG2 as an advanced feature. Users should
|
OCRmyPDF provides lossy mode JBIG2 as an advanced feature. Users should
|
||||||
`review the technical concerns with JBIG2 in lossy
|
`review the technical concerns with JBIG2 in lossy
|
||||||
mode <https://abbyy.technology/en:kb:tip:jbig2_compression_and_ocr>`__
|
mode <https://en.wikipedia.org/wiki/JBIG2#Disadvantages>`__
|
||||||
and decide if this feature is acceptable for their use case.
|
and decide if this feature is acceptable for their use case.
|
||||||
|
|
||||||
JBIG2 lossy mode does achieve higher compression ratios than any other
|
JBIG2 lossy mode does achieve higher compression ratios than any other
|
||||||
|
|||||||
@@ -41,8 +41,8 @@ layer. First, it runs all PDFs through
|
|||||||
`pikepdf <https://github.com/pikepdf/pikepdf>`__, a library based on
|
`pikepdf <https://github.com/pikepdf/pikepdf>`__, a library based on
|
||||||
`qpdf <https://github.com/qpdf/qpdf>`__, a program that repairs PDFs
|
`qpdf <https://github.com/qpdf/qpdf>`__, a program that repairs PDFs
|
||||||
with syntax errors. This is done because, in the author's experience, a
|
with syntax errors. This is done because, in the author's experience, a
|
||||||
significant number of PDFs in the wild especially those created by
|
significant number of PDFs in the wild, especially those created by
|
||||||
scanners are not well-formed files. qpdf makes it more likely that
|
scanners, are not well-formed files. qpdf makes it more likely that
|
||||||
OCRmyPDF will succeed, but offers no security guarantees. qpdf is also
|
OCRmyPDF will succeed, but offers no security guarantees. qpdf is also
|
||||||
used to split the PDF into single page PDFs.
|
used to split the PDF into single page PDFs.
|
||||||
|
|
||||||
|
|||||||
+13
-2
@@ -12,11 +12,22 @@ may be unreliable. Use the API to depend on precise behavior.
|
|||||||
The public API may be useful in scripts that launch OCRmyPDF processes or that
|
The public API may be useful in scripts that launch OCRmyPDF processes or that
|
||||||
wish to use some of its features for working with PDFs.
|
wish to use some of its features for working with PDFs.
|
||||||
|
|
||||||
|
v12.3.3
|
||||||
|
=======
|
||||||
|
|
||||||
|
- watcher.py: fixed interpretation of boolean env vars (:issue:`821`).
|
||||||
|
- Adjust CI scripts to test Tesseract 5 betas.
|
||||||
|
- Document our support for the Tesseract 5 betas.
|
||||||
|
|
||||||
|
v12.3.2
|
||||||
|
=======
|
||||||
|
|
||||||
|
- Indicate support for flask 2.x, watcher 2.x (:issue:`815, 816`).
|
||||||
|
|
||||||
v12.3.1
|
v12.3.1
|
||||||
=======
|
=======
|
||||||
|
|
||||||
- Fixed issue with selection of text when using the hOCR renderer. (:issue:`813`)
|
- Fixed issue with selection of text when using the hOCR renderer (:issue:`813`).
|
||||||
- Fixed build errors with the Docker image by upgrading to a newer Ubuntu.
|
- Fixed build errors with the Docker image by upgrading to a newer Ubuntu.
|
||||||
Also set the timezone of this image to UTC.
|
Also set the timezone of this image to UTC.
|
||||||
|
|
||||||
@@ -24,7 +35,7 @@ v12.3.0
|
|||||||
=======
|
=======
|
||||||
|
|
||||||
- Fixed a regression introduced in Pillow 8.3.0. Pillow no longer rounds DPI
|
- Fixed a regression introduced in Pillow 8.3.0. Pillow no longer rounds DPI
|
||||||
for image resolutions. We now account for this. (:issue:`802`)
|
for image resolutions. We now account for this (:issue:`802`).
|
||||||
- We no longer use some API calls that are deprecated in the latest versions of
|
- We no longer use some API calls that are deprecated in the latest versions of
|
||||||
pikepdf.
|
pikepdf.
|
||||||
- Improved error message when a language is requested that doesn't look like a
|
- Improved error message when a language is requested that doesn't look like a
|
||||||
|
|||||||
+10
-4
@@ -1,3 +1,4 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
# Copyright (C) 2019 Ian Alexander: https://github.com/ianalexander
|
# Copyright (C) 2019 Ian Alexander: https://github.com/ianalexander
|
||||||
# Copyright (C) 2020 James R Barlow: https://github.com/jbarlow83
|
# Copyright (C) 2020 James R Barlow: https://github.com/jbarlow83
|
||||||
#
|
#
|
||||||
@@ -36,14 +37,19 @@ import ocrmypdf
|
|||||||
|
|
||||||
# pylint: disable=logging-format-interpolation
|
# pylint: disable=logging-format-interpolation
|
||||||
|
|
||||||
|
|
||||||
|
def getenv_bool(name: str, default: str = 'False'):
|
||||||
|
return os.getenv(name, default).lower() in ('true', 'yes', 'y', '1')
|
||||||
|
|
||||||
|
|
||||||
INPUT_DIRECTORY = os.getenv('OCR_INPUT_DIRECTORY', '/input')
|
INPUT_DIRECTORY = os.getenv('OCR_INPUT_DIRECTORY', '/input')
|
||||||
OUTPUT_DIRECTORY = os.getenv('OCR_OUTPUT_DIRECTORY', '/output')
|
OUTPUT_DIRECTORY = os.getenv('OCR_OUTPUT_DIRECTORY', '/output')
|
||||||
OUTPUT_DIRECTORY_YEAR_MONTH = bool(os.getenv('OCR_OUTPUT_DIRECTORY_YEAR_MONTH', ''))
|
OUTPUT_DIRECTORY_YEAR_MONTH = getenv_bool('OCR_OUTPUT_DIRECTORY_YEAR_MONTH')
|
||||||
ON_SUCCESS_DELETE = bool(os.getenv('OCR_ON_SUCCESS_DELETE', ''))
|
ON_SUCCESS_DELETE = getenv_bool('OCR_ON_SUCCESS_DELETE')
|
||||||
DESKEW = bool(os.getenv('OCR_DESKEW', ''))
|
DESKEW = getenv_bool('OCR_DESKEW')
|
||||||
OCR_JSON_SETTINGS = json.loads(os.getenv('OCR_JSON_SETTINGS', '{}'))
|
OCR_JSON_SETTINGS = json.loads(os.getenv('OCR_JSON_SETTINGS', '{}'))
|
||||||
POLL_NEW_FILE_SECONDS = int(os.getenv('OCR_POLL_NEW_FILE_SECONDS', '1'))
|
POLL_NEW_FILE_SECONDS = int(os.getenv('OCR_POLL_NEW_FILE_SECONDS', '1'))
|
||||||
USE_POLLING = bool(os.getenv('OCR_USE_POLLING', ''))
|
USE_POLLING = getenv_bool('OCR_USE_POLLING')
|
||||||
LOGLEVEL = os.getenv('OCR_LOGLEVEL', 'INFO')
|
LOGLEVEL = os.getenv('OCR_LOGLEVEL', 'INFO')
|
||||||
PATTERNS = ['*.pdf', '*.PDF']
|
PATTERNS = ['*.pdf', '*.PDF']
|
||||||
|
|
||||||
|
|||||||
+2
-10
@@ -1,3 +1,4 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
# webservice.py wrapper for OCRmyPDF
|
# webservice.py wrapper for OCRmyPDF
|
||||||
# Copyright (C) 2019 James R. Barlow: github.com/jbarlow83
|
# Copyright (C) 2019 James R. Barlow: github.com/jbarlow83
|
||||||
#
|
#
|
||||||
@@ -28,16 +29,7 @@ import shlex
|
|||||||
from subprocess import PIPE, run
|
from subprocess import PIPE, run
|
||||||
from tempfile import TemporaryDirectory
|
from tempfile import TemporaryDirectory
|
||||||
|
|
||||||
from flask import (
|
from flask import Flask, Response, request, send_from_directory
|
||||||
Flask,
|
|
||||||
Response,
|
|
||||||
abort,
|
|
||||||
flash,
|
|
||||||
redirect,
|
|
||||||
request,
|
|
||||||
send_from_directory,
|
|
||||||
url_for,
|
|
||||||
)
|
|
||||||
from werkzeug.utils import secure_filename
|
from werkzeug.utils import secure_filename
|
||||||
|
|
||||||
app = Flask(__name__)
|
app = Flask(__name__)
|
||||||
|
|||||||
@@ -34,3 +34,41 @@ exclude = '''
|
|||||||
| src/ocrmypdf/lib/_leptonica.py
|
| src/ocrmypdf/lib/_leptonica.py
|
||||||
)/
|
)/
|
||||||
'''
|
'''
|
||||||
|
|
||||||
|
[tool.coverage.run]
|
||||||
|
branch = true
|
||||||
|
parallel = true
|
||||||
|
concurrency = ["multiprocessing"]
|
||||||
|
|
||||||
|
[tool.coverage.paths]
|
||||||
|
source = ["src/ocrmypdf"]
|
||||||
|
|
||||||
|
[tool.coverage.report]
|
||||||
|
# Regexes for lines to exclude from consideration
|
||||||
|
exclude_lines = [
|
||||||
|
# Have to re-enable the standard pragma
|
||||||
|
"pragma: no cover",
|
||||||
|
|
||||||
|
# Don't complain if tests don't hit defensive assertion code:
|
||||||
|
"raise AssertionError",
|
||||||
|
"raise NotImplementedError",
|
||||||
|
|
||||||
|
# Don't complain if non-runnable code isn't run:
|
||||||
|
"if 0:",
|
||||||
|
"if False:",
|
||||||
|
"if __name__ == .__main__.:",
|
||||||
|
"if TYPE_CHECKING:"
|
||||||
|
]
|
||||||
|
|
||||||
|
[tool.isort]
|
||||||
|
profile = "black"
|
||||||
|
known_first_party = "ocrmypdf"
|
||||||
|
known_third_party = ["PIL","_cffi_backend","cffi","flask","img2pdf","pdfminer","pikepdf","pkg_resources","pluggy","pytest","reportlab","setuptools","sphinx_rtd_theme","tqdm","watchdog","werkzeug"]
|
||||||
|
|
||||||
|
[tool.pytest.ini_options]
|
||||||
|
minversion = "6.0"
|
||||||
|
norecursedirs = ["lib", ".pc", ".git", "venv", "output", "cache", "resources"]
|
||||||
|
testpaths = ["tests"]
|
||||||
|
addopts = "-n auto"
|
||||||
|
markers = ["slow"]
|
||||||
|
filterwarnings = ["ignore:.*XMLParser.*:DeprecationWarning"]
|
||||||
@@ -72,6 +72,7 @@ where = src
|
|||||||
|
|
||||||
[options.extras_require]
|
[options.extras_require]
|
||||||
test =
|
test =
|
||||||
|
coverage[toml] >= 5
|
||||||
pytest >= 6.0.0
|
pytest >= 6.0.0
|
||||||
pytest-xdist >= 2.2.0
|
pytest-xdist >= 2.2.0
|
||||||
pytest-cov >= 2.11.1
|
pytest-cov >= 2.11.1
|
||||||
@@ -84,9 +85,9 @@ docs =
|
|||||||
extended_test =
|
extended_test =
|
||||||
PyMuPDF == 1.13.4
|
PyMuPDF == 1.13.4
|
||||||
watcher =
|
watcher =
|
||||||
watchdog >= 1.0.2, < 2
|
watchdog >= 1.0.2, < 3
|
||||||
webservice =
|
webservice =
|
||||||
Flask >= 1, < 2
|
Flask >= 1, < 3
|
||||||
|
|
||||||
[options.entry_points]
|
[options.entry_points]
|
||||||
console_scripts =
|
console_scripts =
|
||||||
@@ -101,47 +102,3 @@ test = pytest
|
|||||||
[check-manifest]
|
[check-manifest]
|
||||||
ignore =
|
ignore =
|
||||||
.github
|
.github
|
||||||
|
|
||||||
[tool:pytest]
|
|
||||||
norecursedirs = lib .pc .git output cache resources
|
|
||||||
testpaths = tests
|
|
||||||
filterwarnings =
|
|
||||||
ignore:.*XMLParser.*:DeprecationWarning
|
|
||||||
markers =
|
|
||||||
slow
|
|
||||||
addopts =
|
|
||||||
-n auto
|
|
||||||
|
|
||||||
[isort]
|
|
||||||
multi_line_output = 3
|
|
||||||
include_trailing_comma = True
|
|
||||||
force_grid_wrap = 0
|
|
||||||
use_parentheses = True
|
|
||||||
line_length = 88
|
|
||||||
known_first_party = ocrmypdf
|
|
||||||
known_third_party = PIL,_cffi_backend,cffi,flask,img2pdf,pdfminer,pikepdf,pkg_resources,pluggy,pytest,reportlab,setuptools,sphinx_rtd_theme,tqdm,watchdog,werkzeug
|
|
||||||
|
|
||||||
[coverage:paths]
|
|
||||||
source =
|
|
||||||
src/ocrmypdf
|
|
||||||
|
|
||||||
[coverage:run]
|
|
||||||
branch = true
|
|
||||||
parallel = true
|
|
||||||
concurrency = multiprocessing
|
|
||||||
|
|
||||||
[coverage:report]
|
|
||||||
# Regexes for lines to exclude from consideration
|
|
||||||
exclude_lines =
|
|
||||||
# Have to re-enable the standard pragma
|
|
||||||
pragma: no cover
|
|
||||||
|
|
||||||
# Don't complain if tests don't hit defensive assertion code:
|
|
||||||
raise AssertionError
|
|
||||||
raise NotImplementedError
|
|
||||||
|
|
||||||
# Don't complain if non-runnable code isn't run:
|
|
||||||
if 0:
|
|
||||||
if False:
|
|
||||||
if __name__ == .__main__.:
|
|
||||||
if TYPE_CHECKING:
|
|
||||||
|
|||||||
Reference in New Issue
Block a user