v10 release notes and dependencies

This commit is contained in:
James R. Barlow
2020-06-10 14:27:47 -07:00
parent 7caf1e85ff
commit 17a4831745
4 changed files with 33 additions and 17 deletions
+27 -13
View File
@@ -13,32 +13,46 @@ Note that it is licensed under GPLv3, so scripts that
``import ocrmypdf`` and are released publicly should probably also be
licensed under GPLv3.
v10.0.0 (not yet released)
==========================
v10.0.0
=======
**Breaking changes**
- Support for pdfminer.six version 20181108 has been dropped, along with a
monkeypatch that made this version work.
- Ghostscript is no longer used for finding the location of text in PDFs, and
APIs related to this feature have been removed.
- Output messages are now displayed in color (when supported by the terminal)
and prefixes describing the severity of the message are removed. As such
programs that parse OCRmyPDF's log message will need to be revised. (Please
consider using OCRmyPDF as a library instead.)
- The minimum version for certain dependencies has increased.
- Many API changes; see developer changes.
- The Python libraries pluggy and coloredlogs are now required.
**New features and improvements**
- PDF page scanning is now parallelized across CPUs, speeding up this phase
for files with a high page count.
- PDF page scanning is optimized, addressing some performance regressions.
- A plugin architecture has been added, currently allowing one to more easily
use a different OCR engine or PDF renderer from Tesseract and Ghostscript,
respectively. A plugin can also override some decisions, such changing
the OCR settings after initial scanning.
- Colored log messages.
**Developer changes**
- The test spoofing mechanism, used to correct handling of failures in
Tesseract and Ghostscript, has been removed in favor of using plugins for
testing. The spoofing mechanism was fairly complex and required many special
hacks for Windows.
- Code describing the resolution in DPI of images was refactored into a
``ocrmypdf.helpers.Resolution`` class.
- A deprecated parameter in ``ocrmypdf.exec.ghostscript.generate_pdfa`` was
removed.
- The deprecated module ``ocrmypdf.exec.qpdf`` was removed.
- The module ``ocrmypdf._exec`` is now private to OCRmyPDF.
- The ``ocrmypdf.hocrtransform`` module has been updated to follow PEP8 naming
conventions.
**New features**
- PDF page scanning is now parallelized across CPUs, speeding up the "Scan"
phase for files with a high page count.
- Colored log messages.
- Ghostscript is no longer used for finding the location of text in PDFs, and
APIs related to this feature have been removed.
- Lots of internal reorganization to support plugins.
v9.8.2
======
+2 -1
View File
@@ -5,7 +5,8 @@ cffi == 1.14.0
coloredlogs == 14.0 # technically optional
img2pdf == 0.3.4
pdfminer.six == 20200517
pikepdf == 1.11.1
pikepdf == 1.14.0
pluggy == 0.13.1
Pillow == 7.1.1
reportlab == 3.5.34
tqdm == 4.45.0
+1 -1
View File
@@ -1,7 +1,7 @@
pytest >= 5.0.0
pytest-helpers-namespace >= 2019.1.8
pytest-xdist >= 1.31.0
pytest-cov >= 2.8.0
pytest-cov >= 2.9.0
python-xmp-toolkit == 2.0.1 # requires apt-get install libexempi3
# or brew install exempi
#PyMuPDF == 1.13.4 # optional
+3 -2
View File
@@ -83,8 +83,9 @@ setup(
'coloredlogs >= 14.0', # strictly optional
'img2pdf >= 0.3.0, < 0.4', # pure Python, so track HEAD closely
'pdfminer.six >= 20191110, <= 20200517',
'pikepdf >= 1.8.1, < 2',
'Pillow >= 6.2.0',
'pikepdf >= 1.14.0, < 2',
'Pillow >= 7.0.0',
'pluggy >= 0.13.0',
'reportlab >= 3.3.0', # oldest released version with sane image handling
'tqdm >= 4',
],