Compare commits
15
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
b8fcd7cf83 | ||
|
|
0bcc5dcabb | ||
|
|
9f74f306e1 | ||
|
|
65bf748d59 | ||
|
|
65dba03703 | ||
|
|
cb1126b743 | ||
|
|
8ef928fcdd | ||
|
|
67fa41c39a | ||
|
|
1e14c2101f | ||
|
|
7f12775032 | ||
|
|
d3bee93ed5 | ||
|
|
7409cbed29 | ||
|
|
55a44f823a | ||
|
|
8cd87bd26f | ||
|
|
03932e877a |
@@ -1,8 +0,0 @@
|
|||||||
# Always use Unix convention for new lines
|
|
||||||
* text eol=lf
|
|
||||||
|
|
||||||
# These files are binary and should be left untouched
|
|
||||||
# (binary is a macro for -text -diff)
|
|
||||||
*.jar binary
|
|
||||||
*.pdf binary
|
|
||||||
*.PDF binary
|
|
||||||
-17
@@ -1,17 +0,0 @@
|
|||||||
tmp/
|
|
||||||
log/
|
|
||||||
*.pyc
|
|
||||||
tests/output/
|
|
||||||
.ruffus_history.sqlite
|
|
||||||
*.sublime-*
|
|
||||||
/*.pdf
|
|
||||||
build/
|
|
||||||
dist/
|
|
||||||
*.egg-info/
|
|
||||||
venv/
|
|
||||||
*/test/output
|
|
||||||
bin/
|
|
||||||
include/
|
|
||||||
lib/
|
|
||||||
pip-selfcheck.json
|
|
||||||
pyvenv.cfg
|
|
||||||
+29
@@ -0,0 +1,29 @@
|
|||||||
|
|
||||||
|
language: generic
|
||||||
|
sudo: required
|
||||||
|
dist: trusty
|
||||||
|
|
||||||
|
before_install:
|
||||||
|
- sudo add-apt-repository ppa:heyarje/libav-11 -y
|
||||||
|
- sudo add-apt-repository ppa:alex-p/tesseract-ocr -y
|
||||||
|
- sudo add-apt-repository ppa:evl.ms/evil -y # for a newer version (2.5.0) of pngquant
|
||||||
|
- sudo apt-get update -qq
|
||||||
|
- sudo apt-get install libleptonica-dev -y # required to build jbig2enc
|
||||||
|
- sudo apt-get install zlib1g-dev -y # required to build jbig2enc
|
||||||
|
# - sudo apt-get install imagemagick -y # required to convert logo to desktop icon
|
||||||
|
|
||||||
|
script:
|
||||||
|
- export OCRMYPDF_VERSION=8.3.2
|
||||||
|
- bash build-appimage.sh
|
||||||
|
# remove previous installed libraries to ensure that the tests use the libraries of the AppImage
|
||||||
|
- sudo apt-get remove libleptonica-dev zlib1g-dev -y
|
||||||
|
- bash test/test-appimage.sh
|
||||||
|
- wget https://github.com/probonopd/uploadtool/raw/master/upload.sh
|
||||||
|
|
||||||
|
after_success:
|
||||||
|
- bash upload.sh OCRmyPDF*.AppImage
|
||||||
|
|
||||||
|
branches:
|
||||||
|
except:
|
||||||
|
# Do not build tags that we create when we upload to GitHub Releases
|
||||||
|
- /^(?i:continuous)/
|
||||||
-20
@@ -1,20 +0,0 @@
|
|||||||
Copyright (c) 2013-2015, The OCRmyPDF Authors
|
|
||||||
|
|
||||||
Permission is hereby granted, free of charge, to any person obtaining a
|
|
||||||
copy of this software and associated documentation files (the
|
|
||||||
"Software"), to deal in the Software without restriction, including
|
|
||||||
without limitation the rights to use, copy, modify, merge, publish,
|
|
||||||
distribute, sublicense, and/or sell copies of the Software, and to
|
|
||||||
permit persons to whom the Software is furnished to do so, subject to
|
|
||||||
the following conditions:
|
|
||||||
|
|
||||||
The above copyright notice and this permission notice shall be included
|
|
||||||
in all copies or substantial portions of the Software.
|
|
||||||
|
|
||||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS
|
|
||||||
OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
|
|
||||||
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
|
|
||||||
IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY
|
|
||||||
CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT,
|
|
||||||
TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE
|
|
||||||
SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
|
|
||||||
@@ -1,4 +0,0 @@
|
|||||||
recursive-include ocrmypdf/jhove/bin *.jar
|
|
||||||
recursive-include ocrmypdf/jhove/lib *.jar
|
|
||||||
recursive-include ocrmypdf/jhove/conf *.conf
|
|
||||||
recursive-exclude tests/output *
|
|
||||||
@@ -1,6 +0,0 @@
|
|||||||
#!/bin/sh
|
|
||||||
##############################################################################
|
|
||||||
# Copyright (c) 2013-14: fritz-hh from Github (https://github.com/fritz-hh)
|
|
||||||
##############################################################################
|
|
||||||
|
|
||||||
python3 -m ocrmypdf.main "$@"
|
|
||||||
@@ -0,0 +1,40 @@
|
|||||||
|
# OCRmyPDF-AppImage [](https://travis-ci.com/FPille/OCRmyPDF-AppImage)
|
||||||
|
[AppImage][APPIMAGE] for [OCRmyPDF][OCRMYPDF]
|
||||||
|
|
||||||
|
## Usage
|
||||||
|
Download OCRmyPDF*.AppImage, make it executable and run it.
|
||||||
|
```
|
||||||
|
wget https://github.com/FPille/OCRmyPDF-AppImage/releases/download/continuous/OCRmyPDF-8.3.2-x86_64.AppImage
|
||||||
|
chmod +x OCRmyPDF*.AppImage
|
||||||
|
./OCRmyPDF*.AppImage --help
|
||||||
|
```
|
||||||
|
|
||||||
|
Beside OCRmyPDF additional command line programs can be run with this AppImage like:
|
||||||
|
* ghostscript
|
||||||
|
* img2pdf
|
||||||
|
* pngquant
|
||||||
|
* python3.6
|
||||||
|
* qpdf
|
||||||
|
* tesseract
|
||||||
|
* unpaper
|
||||||
|
|
||||||
|
Just use the program name as first parameter plus options:
|
||||||
|
```
|
||||||
|
./OCRmyPDF*.AppImage tesseract -v
|
||||||
|
tesseract 4.1.0
|
||||||
|
leptonica-1.76.0
|
||||||
|
libjpeg 8d (libjpeg-turbo 1.3.0) : libpng 1.2.50 : libtiff 4.0.3 : zlib 1.2.11 : libwebp 0.4.0 : libopenjp2 2.3.0
|
||||||
|
Found AVX2
|
||||||
|
Found AVX
|
||||||
|
Found SSE
|
||||||
|
```
|
||||||
|
Or create a symlink for the corresponding program:
|
||||||
|
```
|
||||||
|
ln -s OCRmyPDF*.AppImage tesseract
|
||||||
|
./tesseract --list-langs
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
|
[APPIMAGE]: https://appimage.org
|
||||||
|
[OCRMYPDF]: https://github.com/jbarlow83/OCRmyPDF
|
||||||
|
|
||||||
-95
@@ -1,95 +0,0 @@
|
|||||||
OCRmyPDF
|
|
||||||
========
|
|
||||||
|
|
||||||
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to
|
|
||||||
be searched.
|
|
||||||
|
|
||||||
Main features
|
|
||||||
-------------
|
|
||||||
|
|
||||||
- Generates a searchable
|
|
||||||
`PDF/A <https://en.wikipedia.org/?title=PDF/A>`__ file from a PDF
|
|
||||||
file containing only images
|
|
||||||
- Places OCRed text accurately below the image to ease copy / paste
|
|
||||||
- Keeps the exact resolution of the original embedded images
|
|
||||||
|
|
||||||
- or if requested oversamples the images before OCRing so as to get
|
|
||||||
better results
|
|
||||||
|
|
||||||
- If requested deskews and/or cleans the image before performing OCR
|
|
||||||
- Validates the generated file against the PDF/A-1b specification using
|
|
||||||
`JHOVE <http://jhove.sourceforge.net/>`__
|
|
||||||
- Provides debug mode to enable easy verification of the OCR results
|
|
||||||
- Processes several pages in parallel when more than one CPU core is
|
|
||||||
available
|
|
||||||
- Uses Tesseract OCR engine
|
|
||||||
|
|
||||||
For details: please consult the release notes
|
|
||||||
|
|
||||||
Motivation
|
|
||||||
----------
|
|
||||||
|
|
||||||
I searched the web for a free command line tool to OCR PDF files on
|
|
||||||
Linux/UNIX: I found many, but none of them were really satisfying. -
|
|
||||||
Either they produced PDF files with misplaced text under the image
|
|
||||||
(making copy/paste impossible) - Or they did not display correctly some
|
|
||||||
escaped HTML characters located in the hocr file produced by the OCR
|
|
||||||
engine - Or they changed the resolution of the embedded images - Or they
|
|
||||||
generated PDF file having a ridiculous big size - Or they crashed when
|
|
||||||
trying to OCR some of my PDF files - Or they did not produce valid PDF
|
|
||||||
files (even though they were readable with my current PDF reader) - On
|
|
||||||
top of that none of them produced PDF/A files (format dedicated for long
|
|
||||||
time storage / archiving)
|
|
||||||
|
|
||||||
... so I decided to develop my own tool (using various existing scripts
|
|
||||||
as an inspiration)
|
|
||||||
|
|
||||||
Install
|
|
||||||
-------
|
|
||||||
|
|
||||||
Download OCRmyPDF here: https://github.com/fritz-hh/OCRmyPDF/releases
|
|
||||||
|
|
||||||
To install, download the tarball and run:
|
|
||||||
|
|
||||||
pip install ocrmypdf-3.0rc2.tar.gz
|
|
||||||
|
|
||||||
You can install it within a virtual environment or system-wide.
|
|
||||||
|
|
||||||
To run the program::
|
|
||||||
|
|
||||||
ocrmypdf --help
|
|
||||||
|
|
||||||
If not yet installed, the script will notify you about dependencies that
|
|
||||||
need to be installed. The script requires specific versions of the
|
|
||||||
dependencies. Older version than the ones mentioned in the release notes
|
|
||||||
are likely not to be compatible to OCRmyPDF.
|
|
||||||
|
|
||||||
Support
|
|
||||||
-------
|
|
||||||
|
|
||||||
In case you detect an issue, please:
|
|
||||||
|
|
||||||
- Check if your issue is already known
|
|
||||||
- If no problem report exists on github, please create one here:
|
|
||||||
https://github.com/fritz-hh/OCRmyPDF/issues
|
|
||||||
- Describe your problem thoroughly
|
|
||||||
- Append the console output of the script when running the debug mode
|
|
||||||
(-v 1 option)
|
|
||||||
- If possible provide your input PDF file as well as the content of the
|
|
||||||
temporary folder (using a file sharing service like
|
|
||||||
www.file-upload.net)
|
|
||||||
|
|
||||||
Press & Media
|
|
||||||
-------------
|
|
||||||
|
|
||||||
- `c't 1-2014, page 59 <http://www.heise.de/ct/inhalt/2014/1/58/>`__:
|
|
||||||
Detailed presentation of OCRmyPDF v1.0 in the leading German IT
|
|
||||||
magazine c't
|
|
||||||
- `heise Open Source, 09/2014: Texterkennung mit
|
|
||||||
OCRmyPDF <http://www.heise.de/-2356670>`__
|
|
||||||
|
|
||||||
Disclaimer
|
|
||||||
----------
|
|
||||||
|
|
||||||
The software is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR
|
|
||||||
CONDITIONS OF ANY KIND, either express or implied.
|
|
||||||
@@ -1,452 +0,0 @@
|
|||||||
RELEASE NOTES
|
|
||||||
=============
|
|
||||||
|
|
||||||
Please always read this file before installing the package
|
|
||||||
|
|
||||||
Download software here: https://github.com/fritz-hh/OCRmyPDF/tags
|
|
||||||
|
|
||||||
v3.0-rc2:
|
|
||||||
=========
|
|
||||||
|
|
||||||
New features
|
|
||||||
------------
|
|
||||||
|
|
||||||
- Easier installation with Python's package manager
|
|
||||||
- Now installs ``ocrmypdf`` to ``/usr/local/bin`` or equivalent for system-wide
|
|
||||||
access
|
|
||||||
- Tesseract 3.03 PDF page can be used instead for better positioning
|
|
||||||
of recognized text (``--pdf-renderer tesseract``)
|
|
||||||
- Improved command line syntax and usage help (``--help``)
|
|
||||||
- PDF metadata (title, author, keywords) are now transferred to the
|
|
||||||
output PDF
|
|
||||||
- PDF metadata can also be set from the command line (``--title``, etc.)
|
|
||||||
- Added test cases to confirm everything is working
|
|
||||||
- Added option to skip extremely large pages that take too long to OCR and are
|
|
||||||
often not OCRable (e.g. large scanned maps or diagrams); other pages are still
|
|
||||||
processed (``--skip-big``)
|
|
||||||
- Added option to kill Tesseract OCR process if it seems to be taking too long on
|
|
||||||
a page, while still processing other pages (``--tesseract-timeout``)
|
|
||||||
|
|
||||||
Changes
|
|
||||||
-------
|
|
||||||
|
|
||||||
- New, robust rewrite in Python 3.4+ with ruffus_ pipelines
|
|
||||||
- Now uses Ghostscript 9.14's improved color conversion model
|
|
||||||
- All "tasks" in the pipeline can be executed in parallel on any
|
|
||||||
available CPUs, increasing performance
|
|
||||||
- The ``-o DPI`` argument has been phased out, in favor of ``--oversample DPI``
|
|
||||||
- Removed several dependencies, so it's easier to install. We no
|
|
||||||
longer use:
|
|
||||||
|
|
||||||
- GNU parallel_
|
|
||||||
- ImageMagick_
|
|
||||||
- Python 2.7
|
|
||||||
- shell scripts
|
|
||||||
|
|
||||||
- Some new external dependencies are required:
|
|
||||||
|
|
||||||
- MuPDF_ tools
|
|
||||||
- Ghostscript 9.14+
|
|
||||||
- Unpaper_ 6.1 (optional)
|
|
||||||
- some automatically managed Python dependencies
|
|
||||||
|
|
||||||
.. _ruffus: http://www.ruffus.org.uk/index.html
|
|
||||||
.. _parallel: https://www.gnu.org/software/parallel/
|
|
||||||
.. _ImageMagick: http://www.imagemagick.org/script/index.php
|
|
||||||
.. _MuPDF: http://mupdf.com/docs/
|
|
||||||
.. _Unpaper: https://github.com/Flameeyes/unpaper
|
|
||||||
|
|
||||||
Compatibility notes
|
|
||||||
-------------------
|
|
||||||
|
|
||||||
- ``./OCRmyPDF.sh`` script is still available for now
|
|
||||||
- Stacking the verbosity option like ``-vvv`` is no longer supported
|
|
||||||
|
|
||||||
- The configuration file ``config.sh`` has been removed. Instead, you can
|
|
||||||
feed a file to the arguments for common settings:
|
|
||||||
|
|
||||||
::
|
|
||||||
|
|
||||||
ocrmypdf input.pdf output.pdf @settings.txt
|
|
||||||
|
|
||||||
where ``settings.txt`` contains, for example:
|
|
||||||
|
|
||||||
::
|
|
||||||
|
|
||||||
-l deu --author 'A. Merkel' --pdf-renderer tesseract
|
|
||||||
|
|
||||||
|
|
||||||
Fixes
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Handling of filenames containing spaces: fixed
|
|
||||||
|
|
||||||
Notes
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Some dependencies may work with lower versions than tested, so try
|
|
||||||
overriding dependencies if they are "in the way" to see if they work.
|
|
||||||
|
|
||||||
|
|
||||||
v2.2-stable (2014-09-29):
|
|
||||||
=======
|
|
||||||
|
|
||||||
New features
|
|
||||||
------------
|
|
||||||
|
|
||||||
- None
|
|
||||||
|
|
||||||
Changes
|
|
||||||
-------
|
|
||||||
|
|
||||||
- Update to jhove v1.11
|
|
||||||
- Request the python library reportlab v3.0 or newer (So that we could remove a patch to the previous version of reportlab leading to issues for some users)
|
|
||||||
|
|
||||||
Fixes
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Fix bug on Mac OS X (resolution of simlink to OCRmyPDF.sh script) (thanks to jbarlow83)
|
|
||||||
- Check if the input pdf file exists before to continue
|
|
||||||
|
|
||||||
Tested with
|
|
||||||
-----------
|
|
||||||
|
|
||||||
- Operating system: FreeBSD 9.2
|
|
||||||
- Dependencies:
|
|
||||||
- parallel 20140822
|
|
||||||
- poppler-utils 0.24.5
|
|
||||||
- ImageMagick 6.8.9-4 2014-09-17
|
|
||||||
- Unpaper 0.3
|
|
||||||
- tesseract 3.02.02
|
|
||||||
- Python 2.7.8
|
|
||||||
- ghostcript (gs): 9.06
|
|
||||||
- java: openjdk version "1.7.0_65"
|
|
||||||
|
|
||||||
|
|
||||||
v2.1-stable (2014-09-20):
|
|
||||||
=========================
|
|
||||||
|
|
||||||
New features
|
|
||||||
------------
|
|
||||||
|
|
||||||
- None
|
|
||||||
|
|
||||||
Changes
|
|
||||||
-------
|
|
||||||
|
|
||||||
- None
|
|
||||||
|
|
||||||
Fixes
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Allow execution via simlink
|
|
||||||
- Add support for tesseract 3.03
|
|
||||||
- Add support for newer version of reportlab
|
|
||||||
- Lowered minimum version of gnu parallel
|
|
||||||
- Various typo
|
|
||||||
|
|
||||||
Tested with
|
|
||||||
-----------
|
|
||||||
|
|
||||||
- Operating system: FreeBSD 9.1
|
|
||||||
- Dependencies:
|
|
||||||
- parallel 20130222
|
|
||||||
- poppler-utils 0.22.2
|
|
||||||
- ImageMagick 6.8.0-7 2013-03-30
|
|
||||||
- Unpaper 0.3
|
|
||||||
- tesseract 3.02.02
|
|
||||||
- Python 2.7.3
|
|
||||||
- ghoscript (gs): 9.06
|
|
||||||
- java: openjdk version "1.7.0\_17"
|
|
||||||
|
|
||||||
v2.0-stable (2014-01-25):
|
|
||||||
=========================
|
|
||||||
|
|
||||||
New features
|
|
||||||
------------
|
|
||||||
|
|
||||||
- Check if the language(s) passed using the -l option is supported by
|
|
||||||
tesseract (fixes #60)
|
|
||||||
|
|
||||||
Changes
|
|
||||||
-------
|
|
||||||
|
|
||||||
- Allow OCRmyPDF to be used with tesseract 3.02.01, even though OCR
|
|
||||||
might fail for few PDF file (see issue #28). Rationale: For some
|
|
||||||
linux distribution, no newer version than tesseract 3.02.01 is
|
|
||||||
available
|
|
||||||
|
|
||||||
Fixes
|
|
||||||
-----
|
|
||||||
|
|
||||||
- More robust algorithm for checking the version of the installed
|
|
||||||
tesseract package
|
|
||||||
|
|
||||||
Tested with
|
|
||||||
-----------
|
|
||||||
|
|
||||||
- Operating system: FreeBSD 9.1
|
|
||||||
- Dependencies:
|
|
||||||
- parallel 20130222
|
|
||||||
- poppler-utils 0.22.2
|
|
||||||
- ImageMagick 6.8.0-7 2013-03-30
|
|
||||||
- Unpaper 0.3
|
|
||||||
- tesseract 3.02.02
|
|
||||||
- Python 2.7.3
|
|
||||||
- ghoscript (gs): 9.06
|
|
||||||
- java: openjdk version "1.7.0\_17"
|
|
||||||
|
|
||||||
v2.0-rc2 (2014-01-16):
|
|
||||||
======================
|
|
||||||
|
|
||||||
New features
|
|
||||||
------------
|
|
||||||
|
|
||||||
- None
|
|
||||||
|
|
||||||
Changes
|
|
||||||
-------
|
|
||||||
|
|
||||||
- Size reduction of final PDF file: (fixes #50)
|
|
||||||
- Support for monochrome (Black&White) images (massive size reduction
|
|
||||||
in final PDF: >80%)
|
|
||||||
- Reduced size of grayscale images (by 13% on test PDF file)
|
|
||||||
- Preventing fi, fl ligatures does not require anymore to pass an
|
|
||||||
additional config file to tesseract using the -C option (fixes #58)
|
|
||||||
- Location of temporary folder according to content of environment
|
|
||||||
variable TMPDIR.
|
|
||||||
- Dependency to pdftk removed
|
|
||||||
- Check for compatible versions of dependencies: (fixes #51)
|
|
||||||
- parallel and tesseract
|
|
||||||
- python libraries reportlab and lxml
|
|
||||||
|
|
||||||
Fixes
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Improved portability with various shells (dash, bash, tcsh) and OS
|
|
||||||
(FreeBSD, MAC OSX, Linux) (fixes #59)
|
|
||||||
- Corrected bug in case the input PDF file contains a space character
|
|
||||||
(fixes #48)
|
|
||||||
- Prevent spurious error message in case there is no image in a PDF
|
|
||||||
page
|
|
||||||
- Prevent collision of temporary folder names (fixes #57)
|
|
||||||
|
|
||||||
Tested with
|
|
||||||
-----------
|
|
||||||
|
|
||||||
- Operating system: FreeBSD 9.1
|
|
||||||
- Dependencies:
|
|
||||||
- parallel 20130222
|
|
||||||
- poppler-utils 0.22.2
|
|
||||||
- ImageMagick 6.8.0-7 2013-03-30
|
|
||||||
- Unpaper 0.3
|
|
||||||
- tesseract 3.02.02
|
|
||||||
- Python 2.7.3
|
|
||||||
- ghoscript (gs): 9.06
|
|
||||||
- java: openjdk version "1.7.0\_17"
|
|
||||||
|
|
||||||
v2.0-rc1 (2014-01-07):
|
|
||||||
======================
|
|
||||||
|
|
||||||
New features
|
|
||||||
------------
|
|
||||||
|
|
||||||
- Huge performance improvement on machines having multiple CPU/cores
|
|
||||||
(processing of several pages concurrently) (fixes #18)
|
|
||||||
- By default prevent from processing a PDF file already containing
|
|
||||||
fonts (i.e. text)(it can be overridden with the -f flag) (fixes #16)
|
|
||||||
- Warn if the resolution is too low to get reasonable OCR results
|
|
||||||
(fixes #37)
|
|
||||||
- New option (-o) to perform automatic oversampling if the image
|
|
||||||
resolution is too low. This can improve OCR results.
|
|
||||||
- Warn if using a tesseract version older than v3.02.02 (as older
|
|
||||||
versions are known to produce invalid output) (fixes #41)
|
|
||||||
- Echo version of the installed dependencies (e.g. tesseract) in debug
|
|
||||||
mode in order to ease support (fixes #35)
|
|
||||||
- Echo the arguments passed to the script in debug mode to ease support
|
|
||||||
|
|
||||||
Changes
|
|
||||||
-------
|
|
||||||
|
|
||||||
- In debug mode: The debug page is now placed after the respective
|
|
||||||
"normal" page
|
|
||||||
- Reduced disk space usage in temporary folder if -d (deskew) or -c
|
|
||||||
(cleanup) options are not selected
|
|
||||||
- New file src/config.sh containing various configuration parameters
|
|
||||||
- Documentation of the tesseract config file "tess-cfg/no\_ligature"
|
|
||||||
improved
|
|
||||||
- Improved consistency of the temporary file names
|
|
||||||
|
|
||||||
Fixes
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Improved robustness:
|
|
||||||
- in case vertical resolution differs from horizontal resolution (fixes
|
|
||||||
#38)
|
|
||||||
- in case a PDF page contains more than one image (fixes #36)
|
|
||||||
- Fix a problem occurring if python 3 is the standard interpreter
|
|
||||||
(fixes #33)
|
|
||||||
- Fix a problem occurring if the input PDF file contains special
|
|
||||||
characters like "#" (fixes #34)
|
|
||||||
|
|
||||||
Tested with
|
|
||||||
-----------
|
|
||||||
|
|
||||||
- Operating system: FreeBSD 9.1
|
|
||||||
- Dependencies:
|
|
||||||
- parallel 20130222
|
|
||||||
- poppler-utils 0.22.2
|
|
||||||
- ImageMagick 6.8.0-7 2013-03-30
|
|
||||||
- Unpaper 0.3
|
|
||||||
- tesseract 3.02.02
|
|
||||||
- Python 2.7.3
|
|
||||||
- pdftk 1.45
|
|
||||||
- ghoscript (gs): 9.06
|
|
||||||
- java: openjdk version "1.7.0\_17"
|
|
||||||
|
|
||||||
v1.1-stable (2014-01-06):
|
|
||||||
=========================
|
|
||||||
|
|
||||||
New features
|
|
||||||
------------
|
|
||||||
|
|
||||||
- N/A
|
|
||||||
|
|
||||||
Changes
|
|
||||||
-------
|
|
||||||
|
|
||||||
- N/A
|
|
||||||
|
|
||||||
Fixes
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Fixed syntax error (bashism) leading to an error message on certain
|
|
||||||
systems (fixes #42)
|
|
||||||
|
|
||||||
Tested with
|
|
||||||
-----------
|
|
||||||
|
|
||||||
- Operating system: FreeBSD 9.1
|
|
||||||
- Dependencies:
|
|
||||||
- poppler-utils 0.22.2
|
|
||||||
- ImageMagick 6.8.0-7 2013-03-30
|
|
||||||
- Unpaper 0.3
|
|
||||||
- tesseract 3.02.02
|
|
||||||
- Python 2.7.3
|
|
||||||
- pdftk 1.45
|
|
||||||
- ghoscript (gs): 9.06
|
|
||||||
- java: openjdk version "1.7.0\_17"
|
|
||||||
|
|
||||||
v1.0-stable (2013-05-06):
|
|
||||||
=========================
|
|
||||||
|
|
||||||
New features
|
|
||||||
------------
|
|
||||||
|
|
||||||
- In debug mode: compute and echo time required for processing (fixes
|
|
||||||
#26)
|
|
||||||
|
|
||||||
Changes
|
|
||||||
-------
|
|
||||||
|
|
||||||
- Removed feature to add metadata in final pdf file (because it lead to
|
|
||||||
to final PDF file that does not comply to the PDF/A-1 format)
|
|
||||||
- Removed feature to set same owner & permissions in final PDF file
|
|
||||||
than in input file
|
|
||||||
- Removed many unused jhove files (e.g. documentation, \*.java and
|
|
||||||
\*.class files)
|
|
||||||
|
|
||||||
Fixes
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Correction to handle correctly path and input PDF files having spaces
|
|
||||||
(fixes #31)
|
|
||||||
- Resolutions (x/y) that are nearly equal are now supported (fixes #25)
|
|
||||||
- Fix compatibility issue with Ubuntu server 12.04 / Ubuntu server
|
|
||||||
10.04 / Linux Mint 13 Maya and probably other Linux distributions
|
|
||||||
(fixes #27)
|
|
||||||
- Commit missing jhove files (\*.jar mainly) due to wrong .gitignore
|
|
||||||
|
|
||||||
Tested with
|
|
||||||
-----------
|
|
||||||
|
|
||||||
- Operating system: FreeBSD 9.1
|
|
||||||
- Dependencies:
|
|
||||||
- poppler-utils 0.22.2
|
|
||||||
- ImageMagick 6.8.0-7 2013-03-30
|
|
||||||
- Unpaper 0.3
|
|
||||||
- tesseract 3.02.02
|
|
||||||
- Python 2.7.3
|
|
||||||
- pdftk 1.45
|
|
||||||
- ghoscript (gs): 9.06
|
|
||||||
- java: openjdk version "1.7.0\_17"
|
|
||||||
|
|
||||||
v1.0-rc2 (2013-04-29):
|
|
||||||
======================
|
|
||||||
|
|
||||||
New features
|
|
||||||
------------
|
|
||||||
|
|
||||||
- Keep temporary files if debug mode is set (fixes #22)
|
|
||||||
- Set same owner & permissions in final PDF file than in input file
|
|
||||||
(fixes #9)
|
|
||||||
- Added metadata in final pdf file (fixes #4)
|
|
||||||
|
|
||||||
Changes
|
|
||||||
-------
|
|
||||||
|
|
||||||
- N/A
|
|
||||||
|
|
||||||
Fixes
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Fixed wrong image cropping when deskew option is activated
|
|
||||||
- Exit with error message if page size is not found in hocr file (fixes
|
|
||||||
#21)
|
|
||||||
- Various minor fixes in log messages
|
|
||||||
|
|
||||||
Tested with
|
|
||||||
-----------
|
|
||||||
|
|
||||||
- Operating system: FreeBSD 9.1
|
|
||||||
- Dependencies:
|
|
||||||
- poppler-utils 0.22.2
|
|
||||||
- ImageMagick 6.8.0-7 2013-03-30
|
|
||||||
- Unpaper 0.3
|
|
||||||
- tesseract 3.02.02
|
|
||||||
- Python 2.7.3
|
|
||||||
- pdftk 1.45
|
|
||||||
- ghoscript (gs): 9.06
|
|
||||||
- java: openjdk version "1.7.0\_17"
|
|
||||||
|
|
||||||
v1.0-rc1 (2013-04-26):
|
|
||||||
======================
|
|
||||||
|
|
||||||
New features
|
|
||||||
------------
|
|
||||||
|
|
||||||
- First release candidate
|
|
||||||
|
|
||||||
Changes
|
|
||||||
-------
|
|
||||||
|
|
||||||
- N/A
|
|
||||||
|
|
||||||
Fixes
|
|
||||||
-----
|
|
||||||
|
|
||||||
- N/A
|
|
||||||
|
|
||||||
Tested with
|
|
||||||
-----------
|
|
||||||
|
|
||||||
- Operating system: FreeBSD 9.1
|
|
||||||
- Dependencies:
|
|
||||||
- poppler-utils 0.22.2
|
|
||||||
- ImageMagick 6.8.0-7 2013-03-30
|
|
||||||
- Unpaper 0.3
|
|
||||||
- tesseract 3.02.02
|
|
||||||
- Python 2.7.3
|
|
||||||
- pdftk 1.45
|
|
||||||
- ghoscript (gs): 9.06
|
|
||||||
- java: openjdk version "1.7.0\_17"
|
|
||||||
-96
@@ -1,96 +0,0 @@
|
|||||||
Recoding in 5 python modules
|
|
||||||
==================
|
|
||||||
|
|
||||||
- Less platform dependent implementation
|
|
||||||
- Higher versality (wrt addition of new intput / output file types)
|
|
||||||
|
|
||||||
The functionality of each module is described below:
|
|
||||||
|
|
||||||
Normalize inputs (inputs can be a pdf file, an image, a folder containing images)
|
|
||||||
----------------
|
|
||||||
|
|
||||||
- For pdf:
|
|
||||||
- Identify if page needs to be ocred (see -s and -f parameters)
|
|
||||||
- If the page needs to be OCRed:
|
|
||||||
- Extract the image corresponding to the page and save it in a tmp folder. 3 approaches to extract images:
|
|
||||||
- extract raw image from pdf and rotate it according to pdf page rotation
|
|
||||||
- if not possible: identify resolution and rasterize
|
|
||||||
- if not possible: use default resolution and rasterize
|
|
||||||
- If not:
|
|
||||||
- Save the page AS-IS in the tmp folder that should contained the final page
|
|
||||||
- For image(s):
|
|
||||||
- Just copy the images with standardized name into the tmp folder containing pages to be OCRed
|
|
||||||
|
|
||||||
Preprocess normalized inputs (perform jobs in parallel)
|
|
||||||
----------------------------
|
|
||||||
|
|
||||||
- Orientation (if requested by user)
|
|
||||||
- Correct orientation
|
|
||||||
- Skew angle (if requested by user)
|
|
||||||
- Correct skew angle
|
|
||||||
- Cleaning (if requested by user)
|
|
||||||
- Clean image
|
|
||||||
|
|
||||||
Perform OCR (perform jobs in parallel)
|
|
||||||
-----------
|
|
||||||
|
|
||||||
- Perform OCR and save resulting hocr file for each respective page (perfom jobs in parallel)
|
|
||||||
|
|
||||||
Generate output for each page
|
|
||||||
-----------------------------
|
|
||||||
|
|
||||||
- For pdf (if output file has a "pdf" extension):
|
|
||||||
- Generate pdf pages from hocr files (note: pdf pages can already exist if OCR has been skipped for them)
|
|
||||||
- For txt
|
|
||||||
- generate txt file for each page (containing txt located into hocr file)
|
|
||||||
|
|
||||||
Build final output
|
|
||||||
------------------
|
|
||||||
|
|
||||||
- For pdf:
|
|
||||||
- Concatenate pdf pages
|
|
||||||
- Convert to pdf/1-a
|
|
||||||
- Verify conformity to pdf/1-a
|
|
||||||
- For txt
|
|
||||||
- Concatenate all txt files into the final output txt file
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
Tmp folder structure
|
|
||||||
=========================
|
|
||||||
|
|
||||||
- tmp_xxxxx/
|
|
||||||
- a_raw_images (either from images or extracted from pdf file)
|
|
||||||
- b_preprocessed_images (after deswing anf cleaning)
|
|
||||||
- c_ocr_out
|
|
||||||
- d_output_pages (one file per page (at first 1 pdf file per page. Later on other formats might be supported)
|
|
||||||
- e_output_final (concatenate output pages and conversion into PDF/1-a standard, Later on support other formats)
|
|
||||||
|
|
||||||
ocrmypdf arguments
|
|
||||||
==================
|
|
||||||
|
|
||||||
ocrmypdf [-h] [-v] [-k] [-g] [-o dpi] [-f|-s] [-r] [-d] [-c] [-i] [-l lan1[+lan2...]] [-C] inputpath outputfile1 [outputfile2...]
|
|
||||||
|
|
||||||
- Overall parameters
|
|
||||||
- [-h] : Display this help message
|
|
||||||
- [-v] : Increase the verbosity (this option can be used more than once) (e.g. -vvv)
|
|
||||||
- [-k] : Do not delete the temporary files
|
|
||||||
- [-g] : Activate debug mode (max verbosity, keep tmp files, generate debug pages)
|
|
||||||
- Normalization parameters:
|
|
||||||
- [-o dpi] : If page resolution is lower x dpi, provide OCR engine with an oversampled image. (Can improve OCR results)
|
|
||||||
- [-f] : Force to OCR the whole document, even if some page already contain font data (only for pdf inputs)
|
|
||||||
- [-s] : If pages contain font data, do not OCR that page, but include the page (as is) in the final output (only for pdf inputs)
|
|
||||||
- Prepocessing parameters:
|
|
||||||
- [-r] : Correct orientation
|
|
||||||
- [-d] : Deskew each page
|
|
||||||
- [-c] : Clean each page
|
|
||||||
- [-i] : Incorporate cleaned image in final output
|
|
||||||
- OCR parameters:
|
|
||||||
- [-l lan1[+lan2...]] : Document language(s). Multiple languages may be specified, separated by '+' characters.
|
|
||||||
- [-C cfg] : Pass an additional cofg file to the tesseract OCR engine. (this option can be used more than once)
|
|
||||||
- output generation parameters:
|
|
||||||
- None by now
|
|
||||||
- input files:
|
|
||||||
- inputpath : path to image, pdf file or folder to be processed
|
|
||||||
- output files:
|
|
||||||
- outputfile1 [outputfile2 ...] : *.pdf file or *.txt file to be generated (argumenst can be repeated if both pdf and txt file should be generated
|
|
||||||
Executable
+105
@@ -0,0 +1,105 @@
|
|||||||
|
#! /bin/bash
|
||||||
|
|
||||||
|
HERE="$(dirname "$(readlink -f "${0}")")"
|
||||||
|
|
||||||
|
export PATH="$HERE/usr/bin:$HERE/usr/local/bin:$HERE/usr/python/bin:$PATH"
|
||||||
|
export LD_PRELOAD="$HERE/usr/lib/liblept.so.5"
|
||||||
|
export LD_LIBRARY_PATH="$HERE/usr/lib:$HERE/usr/lib/x86_64-linux-gnu:$LD_LIBRARY_PATH"
|
||||||
|
export TESSDATA_PREFIX="$HERE/usr/share/tesseract-ocr/4.00/tessdata"
|
||||||
|
export GS_LIB="$HERE/usr/share/ghostscript/9.26/lib:$HERE/usr/share/ghostscript/9.26/Resource:$HERE/usr/share/ghostscript/9.26/Resource/Init"
|
||||||
|
|
||||||
|
# Allow the AppImage to be symlinked to e.g., /usr/bin/commandname
|
||||||
|
# or called with ./Some*.AppImage commandname ...
|
||||||
|
# refer to https://github.com/AppImage/AppImageKit/wiki/Bundling-command-line-tools
|
||||||
|
|
||||||
|
if [ ! -z "$APPIMAGE" ] ; then
|
||||||
|
BINARY_NAME=$(basename "$ARGV0")
|
||||||
|
else
|
||||||
|
BINARY_NAME=$(basename "$0")
|
||||||
|
export APPDIR="$HERE" # required for the wrapper scripts of linuxdeploy-plugin-python
|
||||||
|
fi
|
||||||
|
|
||||||
|
usage() {
|
||||||
|
echo "
|
||||||
|
==============================================================================
|
||||||
|
AppImage for OCRmyPDF
|
||||||
|
==============================================================================
|
||||||
|
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be
|
||||||
|
searched or copy-pasted.
|
||||||
|
|
||||||
|
usage:
|
||||||
|
$ARGV0 [ocrmypdf] [--help] [--list-programs]
|
||||||
|
[--list-licenses] [--show-license]
|
||||||
|
|
||||||
|
ocrmypdf execute OCRmyPDF
|
||||||
|
|
||||||
|
--help show this help message
|
||||||
|
|
||||||
|
--list-programs list all programs contained in this AppImage
|
||||||
|
|
||||||
|
--list-licenses list all licenses contained in this AppImage
|
||||||
|
|
||||||
|
--show-license [LICENSE] show content of license file
|
||||||
|
"
|
||||||
|
}
|
||||||
|
|
||||||
|
if [ "$1" == "--help" ] ; then
|
||||||
|
usage
|
||||||
|
exit $?
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [ "$1" == "--list-programs" ] ; then
|
||||||
|
pushd "$HERE"
|
||||||
|
echo ""
|
||||||
|
echo "Run \"$ARGV0\" with one of the following arguments to run the respective program."
|
||||||
|
echo ""
|
||||||
|
find . -type f -perm /111 ! -path '*/lib/*' -execdir basename {} ";" | sort -u | column
|
||||||
|
echo ""
|
||||||
|
exit $?
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [ "$1" == "--list-licenses" ] ; then
|
||||||
|
pushd "$HERE"
|
||||||
|
echo ""
|
||||||
|
echo "Run \"$ARGV0\" with one of the following arguments to display the respective license file."
|
||||||
|
echo ""
|
||||||
|
find . -type f \( ! -path '*/tesseract-ocr-*' -o -path '*/tesseract-ocr-eng/*' \) \
|
||||||
|
\( -iname "license*" -o -iname "*copyright*" -o -iname "*copying*" \) -printf "--show-license %P\n" | sort | column
|
||||||
|
echo ""
|
||||||
|
exit $?
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [ "$1" == "--show-license" ] ; then
|
||||||
|
pushd "$HERE"
|
||||||
|
shift
|
||||||
|
if [ -f "$1" ] ; then
|
||||||
|
less -N "$1"
|
||||||
|
exit $?
|
||||||
|
else
|
||||||
|
echo "\"$1\" is not a valid license file path."
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [ ! -z "$1" ] && [ -e "$HERE/bin/$1" ] ; then
|
||||||
|
MAIN="$HERE/bin/$1" ; shift
|
||||||
|
elif [ ! -z "$1" ] && [ -e "$HERE/usr/bin/$1" ] ; then
|
||||||
|
MAIN="$HERE/usr/bin/$1" ; shift
|
||||||
|
elif [ ! -z "$1" ] && [ -e "$HERE/usr/python/bin/$1" ] ; then
|
||||||
|
MAIN="$HERE/usr/python/bin/$1" ; shift
|
||||||
|
elif [ ! -z "$1" ] && [ -e "$HERE/usr/local/bin/$1" ] ; then
|
||||||
|
MAIN="$HERE/usr/local/bin/$1" ; shift
|
||||||
|
elif [ -e "$HERE/bin/$BINARY_NAME" ] ; then
|
||||||
|
MAIN="$HERE/bin/$BINARY_NAME"
|
||||||
|
elif [ -e "$HERE/usr/bin/$BINARY_NAME" ] ; then
|
||||||
|
MAIN="$HERE/usr/bin/$BINARY_NAME"
|
||||||
|
elif [ -e "$HERE/usr/python/bin/$BINARY_NAME" ] ; then
|
||||||
|
MAIN="$HERE/usr/python/bin/$BINARY_NAME"
|
||||||
|
elif [ -e "$HERE/usr/local/bin/$BINARY_NAME" ] ; then
|
||||||
|
MAIN="$HERE/usr/local/bin/$BINARY_NAME"
|
||||||
|
else
|
||||||
|
usage
|
||||||
|
exit $?
|
||||||
|
fi
|
||||||
|
|
||||||
|
exec "${MAIN}" "$@"
|
||||||
@@ -0,0 +1,120 @@
|
|||||||
|
#! /bin/bash
|
||||||
|
|
||||||
|
set -x
|
||||||
|
set -e
|
||||||
|
|
||||||
|
# use RAM disk if possible
|
||||||
|
if [ "$CI" == "" ] && [ -d /dev/shm ]; then
|
||||||
|
TEMP_BASE=/dev/shm
|
||||||
|
else
|
||||||
|
TEMP_BASE=/tmp
|
||||||
|
fi
|
||||||
|
|
||||||
|
BUILD_DIR=$(mktemp -d -p "$TEMP_BASE" OCRmyPDF-AppImage-build-XXXXXX)
|
||||||
|
|
||||||
|
cleanup () {
|
||||||
|
if [ -d "$BUILD_DIR" ]; then
|
||||||
|
rm -rf "$BUILD_DIR"
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
|
trap cleanup EXIT
|
||||||
|
|
||||||
|
# store repo root as variable
|
||||||
|
REPO_ROOT=$(readlink -f "$(dirname "$(dirname "$0")")")
|
||||||
|
OLD_CWD=$(readlink -f .)
|
||||||
|
|
||||||
|
pushd "$BUILD_DIR"
|
||||||
|
|
||||||
|
mkdir -p AppDir
|
||||||
|
mkdir -p PackageDir
|
||||||
|
mkdir -p jbig2
|
||||||
|
|
||||||
|
# download linuxdeploy AppImage and linuxdeploy-plugin-python AppImage
|
||||||
|
wget https://github.com/TheAssassin/linuxdeploy/releases/download/continuous/linuxdeploy-x86_64.AppImage
|
||||||
|
# wget https://github.com/niess/linuxdeploy-plugin-python/releases/download/continuous/linuxdeploy-plugin-python-x86_64.AppImage
|
||||||
|
|
||||||
|
# use adapted linuxdeploy-plugin-python instead of the original one (otherwise OCRmyPDF breaks)
|
||||||
|
wget https://github.com/FPille/linuxdeploy-plugin-python/releases/download/continuous/linuxdeploy-plugin-python-x86_64.AppImage
|
||||||
|
|
||||||
|
chmod +x linuxdeploy*.AppImage
|
||||||
|
|
||||||
|
|
||||||
|
ARCH=$(uname -i)
|
||||||
|
export ARCH
|
||||||
|
|
||||||
|
|
||||||
|
# .desktop file
|
||||||
|
cat > ocrmypdf.desktop <<\EOF
|
||||||
|
[Desktop Entry]
|
||||||
|
Name=ocrmypdf
|
||||||
|
Type=Application
|
||||||
|
Exec=ocrmypdf
|
||||||
|
Icon=ocrmypdf
|
||||||
|
Terminal=true
|
||||||
|
Comment=OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
|
||||||
|
Categories=Graphics;Scanning;OCR;
|
||||||
|
EOF
|
||||||
|
|
||||||
|
|
||||||
|
# download logo and convert it to desktop icon
|
||||||
|
# requires Imagemagick (convert)
|
||||||
|
wget https://raw.githubusercontent.com/jbarlow83/OCRmyPDF/master/docs/images/logo-social.png
|
||||||
|
convert logo-social.png -resize 512x512\> -size 512x512 xc:white +swap -gravity center -composite ocrmypdf.png
|
||||||
|
|
||||||
|
|
||||||
|
# download and intsall packages required by OCRmyPDF
|
||||||
|
pushd PackageDir
|
||||||
|
packages=(tesseract-ocr tesseract-ocr-all libavformat56 ghostscript qpdf pngquant)
|
||||||
|
|
||||||
|
for i in "${packages[@]}"
|
||||||
|
do
|
||||||
|
apt-get -d -o dir::cache="$PWD" -o Debug::NoLocking=1 --reinstall install "$i" -y
|
||||||
|
done
|
||||||
|
|
||||||
|
wget -q 'https://www.dropbox.com/s/vaq0kbwi6e6au80/unpaper_6.1-1.deb?raw=1' -O unpaper_6.1-1.deb
|
||||||
|
|
||||||
|
find . -type f -name \*.deb -exec dpkg-deb -X {} "$BUILD_DIR"/AppDir \;
|
||||||
|
popd
|
||||||
|
|
||||||
|
|
||||||
|
# compile and install jbig2
|
||||||
|
# requires libleptonica-dev, zlib1g-dev
|
||||||
|
wget -q https://github.com/agl/jbig2enc/archive/0.29.tar.gz -O - | \
|
||||||
|
tar xz -C jbig2 --strip-components=1
|
||||||
|
pushd jbig2
|
||||||
|
./autogen.sh
|
||||||
|
./configure --prefix="$BUILD_DIR"/AppDir/usr
|
||||||
|
make && make install
|
||||||
|
popd
|
||||||
|
|
||||||
|
pushd "$BUILD_DIR"/AppDir
|
||||||
|
# add some tools to AppDir
|
||||||
|
#cp -f /usr/bin/column ./usr/bin/
|
||||||
|
#cp -f /bin/less ./usr/bin/
|
||||||
|
|
||||||
|
# remove unnecessary data from AppDir
|
||||||
|
[ -d bin ] && rm -rf ./bin
|
||||||
|
[ -d etc ] && rm -rf ./etc
|
||||||
|
[ -d var ] && rm -rf ./var
|
||||||
|
popd
|
||||||
|
|
||||||
|
|
||||||
|
# export LD_LIBRARY_PATH so that dependencies of shared libraries can be deployed by linuxdeploy-x86_64.AppImage
|
||||||
|
export LD_LIBRARY_PATH="$BUILD_DIR/AppDir/usr/lib:$BUILD_DIR/AppDir/usr/lib/x86_64-linux-gnu:$LD_LIBRARY_PATH"
|
||||||
|
|
||||||
|
#OCRMYPDF_VERSION=8.3.2 # exported in .travis.yml file
|
||||||
|
export PIP_REQUIREMENTS="ocrmypdf==$OCRMYPDF_VERSION"
|
||||||
|
export VERSION="$OCRMYPDF_VERSION"
|
||||||
|
export OUTPUT=OCRmyPDF-"$VERSION"-"$ARCH".AppImage
|
||||||
|
export PYTHON_SOURCE=https://www.python.org/ftp/python/3.6.8/Python-3.6.8.tgz
|
||||||
|
|
||||||
|
./linuxdeploy-x86_64.AppImage --appdir AppDir --plugin python \
|
||||||
|
-d ocrmypdf.desktop -i ocrmypdf.png \
|
||||||
|
--custom-apprun "$REPO_ROOT"/appimage/AppRun.sh --output appimage
|
||||||
|
|
||||||
|
|
||||||
|
# move AppImage back to old CWD
|
||||||
|
mv "$OUTPUT" "$OLD_CWD"/
|
||||||
|
|
||||||
|
popd
|
||||||
-45
@@ -1,45 +0,0 @@
|
|||||||
#!/usr/local/bin/python2
|
|
||||||
##############################################################################
|
|
||||||
# Copyright (c) 2013-14: fritz-hh from Github (https://github.com/fritz-hh)
|
|
||||||
##############################################################################
|
|
||||||
|
|
||||||
import argparse
|
|
||||||
from argparse import RawTextHelpFormatter
|
|
||||||
|
|
||||||
if __name__ == '__main__':
|
|
||||||
|
|
||||||
parser = argparse.ArgumentParser(prog="ocrmypdf", description="OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched", formatter_class=RawTextHelpFormatter)
|
|
||||||
parser.add_argument('-v', '--verbove', action='count', help="Increase the verbosity (this option can be used more than once) (e.g. -vvv)")
|
|
||||||
parser.add_argument('-k', '--keep-tmp', action='store_true', help="Do not delete the temporary files")
|
|
||||||
parser.add_argument('-g', '--debug', action='store_true',
|
|
||||||
help="Activate debug mode:\n" +
|
|
||||||
"- Generates a PDF file containing each page twice (once with the image, once without the image\n" +
|
|
||||||
" but with the OCRed text as well as the detected bounding boxes)\n" +
|
|
||||||
"- Set the verbosity to the highest possible\n" +
|
|
||||||
"- Do not delete the temporary files")
|
|
||||||
parser.add_argument('-d', '--deskew', action='store_true', help="Deskew each page before performing OCR")
|
|
||||||
parser.add_argument('-c', '--clean', action='store_true', help="Clean each page before performing OCR")
|
|
||||||
parser.add_argument('-i', action='store_true', help="Incorporate the cleaned image in the final PDF file (by default the original image, or the deskewed image if the -d option is set)")
|
|
||||||
parser.add_argument('-o', '--oversample', metavar='dpi', type=int,
|
|
||||||
help="If the resolution of an image is lower than dpi value provided as argument, provide the OCR engine with\n" +
|
|
||||||
"an oversampled image having the latter dpi value. This can improve the OCR results but can lead to a larger output PDF file.\n" +
|
|
||||||
"(default: no oversampling performed)")
|
|
||||||
group_sf = parser.add_mutually_exclusive_group()
|
|
||||||
group_sf.add_argument('-f', action='store_true',
|
|
||||||
help="Force to OCR the whole document, even if some page already contain font data.\n" +
|
|
||||||
"(which should not be the case for PDF files built from scnanned images)\n" +
|
|
||||||
"Any text data will be rendered to raster format and then fed through OCR.")
|
|
||||||
group_sf.add_argument('-s', action='store_true', help="If pages contain font data, do not OCR that page, but include the page (as is) in the final output.")
|
|
||||||
parser.add_argument('-l', '--language', metavar='lan',
|
|
||||||
help="Language(s) of the PDF file. The language should be set correctly in order to get good OCR results.\n" +
|
|
||||||
"Any language supported by tesseract is supported (Tesseract uses 3-character ISO 639-2 language codes)\n" +
|
|
||||||
"Multiple languages may be specified, separated by '+' characters.\n" +
|
|
||||||
"(The default language is defined in the config file)")
|
|
||||||
parser.add_argument('-C', '--tess-config', metavar='file',
|
|
||||||
help="Pass an additional configuration file to the tesseract OCR engine.\n" +
|
|
||||||
"(this option can be used more than once)\n" +
|
|
||||||
"Note 1: The configuration file must be available in the \"tessdata/configs\" folder of your tesseract installation")
|
|
||||||
parser.add_argument('inputfile', help="PDF file to be OCRed")
|
|
||||||
parser.add_argument('outputfile', help="The PDF/A file that will be generated")
|
|
||||||
|
|
||||||
args = parser.parse_args()
|
|
||||||
@@ -1,51 +0,0 @@
|
|||||||
#!/usr/bin/env python3
|
|
||||||
# © 2015 James R. Barlow: github.com/jbarlow83
|
|
||||||
|
|
||||||
from tempfile import NamedTemporaryFile
|
|
||||||
from subprocess import Popen, PIPE, check_call
|
|
||||||
from shutil import copy
|
|
||||||
|
|
||||||
|
|
||||||
def rasterize_pdf(input_file, output_file, xres, yres, raster_device, log):
|
|
||||||
with NamedTemporaryFile(delete=True) as tmp:
|
|
||||||
args_gs = [
|
|
||||||
'gs',
|
|
||||||
'-dBATCH', '-dNOPAUSE',
|
|
||||||
'-sDEVICE=%s' % raster_device,
|
|
||||||
'-o', tmp.name,
|
|
||||||
'-r{0}x{1}'.format(str(xres), str(yres)),
|
|
||||||
input_file
|
|
||||||
]
|
|
||||||
|
|
||||||
p = Popen(args_gs, close_fds=True, stdout=PIPE, stderr=PIPE,
|
|
||||||
universal_newlines=True)
|
|
||||||
stdout, stderr = p.communicate()
|
|
||||||
if stdout:
|
|
||||||
log.debug(stdout)
|
|
||||||
if stderr:
|
|
||||||
log.error(stderr)
|
|
||||||
|
|
||||||
if p.returncode == 0:
|
|
||||||
copy(tmp.name, output_file)
|
|
||||||
else:
|
|
||||||
log.error('Ghostscript rendering failed')
|
|
||||||
|
|
||||||
|
|
||||||
def generate_pdfa(pdf_pages, output_file):
|
|
||||||
with NamedTemporaryFile(delete=True) as gs_pdf:
|
|
||||||
args_gs = [
|
|
||||||
"gs",
|
|
||||||
"-dQUIET",
|
|
||||||
"-dBATCH",
|
|
||||||
"-dNOPAUSE",
|
|
||||||
"-sDEVICE=pdfwrite",
|
|
||||||
"-sColorConversionStrategy=/RGB",
|
|
||||||
"-sProcessColorModel=DeviceRGB",
|
|
||||||
"-dPDFA",
|
|
||||||
"-sPDFACompatibilityPolicy=2",
|
|
||||||
"-sOutputICCProfile=srgb.icc",
|
|
||||||
"-sOutputFile=" + gs_pdf.name,
|
|
||||||
]
|
|
||||||
args_gs.extend(pdf_pages)
|
|
||||||
check_call(args_gs)
|
|
||||||
copy(gs_pdf.name, output_file)
|
|
||||||
@@ -1,231 +0,0 @@
|
|||||||
#!/usr/local/bin/python3
|
|
||||||
##############################################################################
|
|
||||||
# Copyright (c) 2013-14: fritz-hh from Github
|
|
||||||
# (https://github.com/fritz-hh)
|
|
||||||
#
|
|
||||||
# Copyright (c) 2010: Jonathan Brinley from Github
|
|
||||||
# (https://github.com/jbrinley/HocrConverter)
|
|
||||||
# Initial version by Jonathan Brinley, jonathanbrinley@gmail.com
|
|
||||||
##############################################################################
|
|
||||||
from reportlab.pdfgen.canvas import Canvas
|
|
||||||
from reportlab.lib.units import inch
|
|
||||||
from lxml import etree as ElementTree
|
|
||||||
from PIL import Image
|
|
||||||
from collections import namedtuple
|
|
||||||
import re
|
|
||||||
import argparse
|
|
||||||
|
|
||||||
|
|
||||||
Rect = namedtuple('Rect', ['x1', 'y1', 'x2', 'y2'])
|
|
||||||
|
|
||||||
|
|
||||||
class HocrTransformError(Exception):
|
|
||||||
pass
|
|
||||||
|
|
||||||
|
|
||||||
class HocrTransform():
|
|
||||||
|
|
||||||
"""
|
|
||||||
A class for converting documents from the hOCR format.
|
|
||||||
For details of the hOCR format, see:
|
|
||||||
http://docs.google.com/View?docid=dfxcv4vc_67g844kf
|
|
||||||
"""
|
|
||||||
|
|
||||||
def __init__(self, hocrFileName, dpi):
|
|
||||||
self.dpi = dpi
|
|
||||||
self.boxPattern = re.compile(r'bbox((\s+\d+){4})')
|
|
||||||
|
|
||||||
self.hocr = ElementTree.ElementTree()
|
|
||||||
self.hocr.parse(hocrFileName)
|
|
||||||
|
|
||||||
# if the hOCR file has a namespace, ElementTree requires its use to
|
|
||||||
# find elements
|
|
||||||
matches = re.match(r'({.*})html', self.hocr.getroot().tag)
|
|
||||||
self.xmlns = ''
|
|
||||||
if matches:
|
|
||||||
self.xmlns = matches.group(1)
|
|
||||||
|
|
||||||
# get dimension in pt (not pixel!!!!) of the OCRed image
|
|
||||||
self.width, self.height = None, None
|
|
||||||
for div in self.hocr.findall(
|
|
||||||
".//%sdiv[@class='ocr_page']" % (self.xmlns)):
|
|
||||||
coords = self.element_coordinates(div)
|
|
||||||
pt_coords = self.pt_from_pixel(coords)
|
|
||||||
self.width = pt_coords.x2 - pt_coords.x1
|
|
||||||
self.height = pt_coords.y2 - pt_coords.y1
|
|
||||||
# there shouldn't be more than one, and if there is, we don't want
|
|
||||||
# it
|
|
||||||
break
|
|
||||||
if self.width is None or self.height is None:
|
|
||||||
raise HocrTransformError("hocr file is missing page dimensions")
|
|
||||||
|
|
||||||
def __str__(self):
|
|
||||||
"""
|
|
||||||
Return the textual content of the HTML body
|
|
||||||
"""
|
|
||||||
if self.hocr is None:
|
|
||||||
return ''
|
|
||||||
body = self.hocr.find(".//%sbody" % (self.xmlns))
|
|
||||||
if body:
|
|
||||||
return self._get_element_text(body)
|
|
||||||
else:
|
|
||||||
return ''
|
|
||||||
|
|
||||||
def _get_element_text(self, element):
|
|
||||||
"""
|
|
||||||
Return the textual content of the element and its children
|
|
||||||
"""
|
|
||||||
text = ''
|
|
||||||
if element.text is not None:
|
|
||||||
text += element.text
|
|
||||||
for child in element.getchildren():
|
|
||||||
text += self._get_element_text(child)
|
|
||||||
if element.tail is not None:
|
|
||||||
text += element.tail
|
|
||||||
return text
|
|
||||||
|
|
||||||
def element_coordinates(self, element):
|
|
||||||
"""
|
|
||||||
Returns a tuple containing the coordinates of the bounding box around
|
|
||||||
an element
|
|
||||||
"""
|
|
||||||
out = (0, 0, 0, 0)
|
|
||||||
if 'title' in element.attrib:
|
|
||||||
matches = self.boxPattern.search(element.attrib['title'])
|
|
||||||
if matches:
|
|
||||||
coords = matches.group(1).split()
|
|
||||||
out = Rect._make(int(coords[n]) for n in range(4))
|
|
||||||
return out
|
|
||||||
|
|
||||||
def pt_from_pixel(self, pxl):
|
|
||||||
"""
|
|
||||||
Returns the quantity in PDF units (pt) given quantity in pixels
|
|
||||||
"""
|
|
||||||
return Rect._make(
|
|
||||||
(c / self.dpi * inch) for c in pxl)
|
|
||||||
|
|
||||||
def replace_unsupported_chars(self, s):
|
|
||||||
"""
|
|
||||||
Given an input string, returns the corresponding string that:
|
|
||||||
- is available in the helvetica facetype
|
|
||||||
- does not contain any ligature (to allow easy search in the PDF file)
|
|
||||||
"""
|
|
||||||
# The 'u' before the character to replace indicates that it is a
|
|
||||||
# unicode character
|
|
||||||
s = s.replace(u"fl", "fl")
|
|
||||||
s = s.replace(u"fi", "fi")
|
|
||||||
return s
|
|
||||||
|
|
||||||
def to_pdf(self, outFileName, imageFileName=None, showBoundingboxes=False,
|
|
||||||
fontname="Helvetica", invisibleText=False):
|
|
||||||
"""
|
|
||||||
Creates a PDF file with an image superimposed on top of the text.
|
|
||||||
Text is positioned according to the bounding box of the lines in
|
|
||||||
the hOCR file.
|
|
||||||
The image need not be identical to the image used to create the hOCR
|
|
||||||
file.
|
|
||||||
It can have a lower resolution, different color mode, etc.
|
|
||||||
"""
|
|
||||||
# create the PDF file
|
|
||||||
# page size in points (1/72 in.)
|
|
||||||
pdf = Canvas(
|
|
||||||
outFileName, pagesize=(self.width, self.height), pageCompression=1)
|
|
||||||
|
|
||||||
# draw bounding box for each paragraph
|
|
||||||
# light blue for bounding box of paragraph
|
|
||||||
pdf.setStrokeColorRGB(0, 1, 1)
|
|
||||||
# light blue for bounding box of paragraph
|
|
||||||
pdf.setFillColorRGB(0, 1, 1)
|
|
||||||
pdf.setLineWidth(0) # no line for bounding box
|
|
||||||
for elem in self.hocr.findall(
|
|
||||||
".//%sp[@class='%s']" % (self.xmlns, "ocr_par")):
|
|
||||||
|
|
||||||
elemtxt = self._get_element_text(elem).rstrip()
|
|
||||||
if len(elemtxt) == 0:
|
|
||||||
continue
|
|
||||||
|
|
||||||
pxl_coords = self.element_coordinates(elem)
|
|
||||||
pt = self.pt_from_pixel(pxl_coords)
|
|
||||||
|
|
||||||
# draw the bbox border
|
|
||||||
if showBoundingboxes:
|
|
||||||
pdf.rect(
|
|
||||||
pt.x1, self.height - pt.y2, pt.x2 - pt.x1, pt.y2 - pt.y1,
|
|
||||||
fill=1)
|
|
||||||
|
|
||||||
# check if element with class 'ocrx_word' are available
|
|
||||||
# otherwise use 'ocr_line' as fallback
|
|
||||||
elemclass = "ocr_line"
|
|
||||||
if self.hocr.find(
|
|
||||||
".//%sspan[@class='ocrx_word']" % (self.xmlns)) is not None:
|
|
||||||
elemclass = "ocrx_word"
|
|
||||||
|
|
||||||
# itterate all text elements
|
|
||||||
# light green for bounding box of word/line
|
|
||||||
pdf.setStrokeColorRGB(1, 0, 0)
|
|
||||||
pdf.setLineWidth(0.5) # bounding box line width
|
|
||||||
pdf.setDash(6, 3) # bounding box is dashed
|
|
||||||
pdf.setFillColorRGB(0, 0, 0) # text in black
|
|
||||||
for elem in self.hocr.findall(
|
|
||||||
".//%sspan[@class='%s']" % (self.xmlns, elemclass)):
|
|
||||||
|
|
||||||
elemtxt = self._get_element_text(elem).rstrip()
|
|
||||||
|
|
||||||
elemtxt = self.replace_unsupported_chars(elemtxt)
|
|
||||||
|
|
||||||
if len(elemtxt) == 0:
|
|
||||||
continue
|
|
||||||
|
|
||||||
pxl_coords = self.element_coordinates(elem)
|
|
||||||
pt = self.pt_from_pixel(pxl_coords)
|
|
||||||
|
|
||||||
# draw the bbox border
|
|
||||||
if showBoundingboxes:
|
|
||||||
pdf.rect(
|
|
||||||
pt.x1, self.height - pt.y2, pt.x2 - pt.x1, pt.y2 - pt.y1,
|
|
||||||
fill=0)
|
|
||||||
|
|
||||||
text = pdf.beginText()
|
|
||||||
fontsize = pt.y2 - pt.y1
|
|
||||||
text.setFont(fontname, fontsize)
|
|
||||||
if invisibleText:
|
|
||||||
text.setTextRenderMode(3) # Invisible (indicates OCR text)
|
|
||||||
|
|
||||||
# set cursor to bottom left corner of bbox (adjust for dpi)
|
|
||||||
text.setTextOrigin(pt.x1, self.height - pt.y2)
|
|
||||||
|
|
||||||
# scale the width of the text to fill the width of the bbox
|
|
||||||
text.setHorizScale(
|
|
||||||
100 * (pt.x2 - pt.x1) / pdf.stringWidth(
|
|
||||||
elemtxt, fontname, fontsize))
|
|
||||||
|
|
||||||
# write the text to the page
|
|
||||||
text.textLine(elemtxt)
|
|
||||||
pdf.drawText(text)
|
|
||||||
|
|
||||||
# put the image on the page, scaled to fill the page
|
|
||||||
if imageFileName is not None:
|
|
||||||
pdf.drawImage(imageFileName, 0, 0,
|
|
||||||
width=self.width, height=self.height)
|
|
||||||
|
|
||||||
# finish up the page and save it
|
|
||||||
pdf.showPage()
|
|
||||||
pdf.save()
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
parser = argparse.ArgumentParser(description='Convert hocr file to PDF')
|
|
||||||
parser.add_argument('-b', '--boundingboxes', action="store_true",
|
|
||||||
default=False, help='Show bounding boxes borders')
|
|
||||||
parser.add_argument('-r', '--resolution', type=int,
|
|
||||||
default=300,
|
|
||||||
help='Resolution of the image that was OCRed')
|
|
||||||
parser.add_argument('-i', '--image', default=None,
|
|
||||||
help='Path to the image to be placed above the text')
|
|
||||||
parser.add_argument('hocrfile', help='Path to the hocr file to be parsed')
|
|
||||||
parser.add_argument(
|
|
||||||
'outputfile', help='Path to the PDF file to be generated')
|
|
||||||
args = parser.parse_args()
|
|
||||||
|
|
||||||
hocr = HocrTransform(args.hocrfile, args.resolution)
|
|
||||||
hocr.to_pdf(args.outputfile, args.image, args.boundingboxes)
|
|
||||||
@@ -1,502 +0,0 @@
|
|||||||
GNU LESSER GENERAL PUBLIC LICENSE
|
|
||||||
Version 2.1, February 1999
|
|
||||||
|
|
||||||
Copyright (C) 1991, 1999 Free Software Foundation, Inc.
|
|
||||||
59 Temple Place, Suite 330, Boston, MA 02111-1307 USA
|
|
||||||
Everyone is permitted to copy and distribute verbatim copies
|
|
||||||
of this license document, but changing it is not allowed.
|
|
||||||
|
|
||||||
[This is the first released version of the Lesser GPL. It also counts
|
|
||||||
as the successor of the GNU Library Public License, version 2, hence
|
|
||||||
the version number 2.1.]
|
|
||||||
|
|
||||||
Preamble
|
|
||||||
|
|
||||||
The licenses for most software are designed to take away your
|
|
||||||
freedom to share and change it. By contrast, the GNU General Public
|
|
||||||
Licenses are intended to guarantee your freedom to share and change
|
|
||||||
free software--to make sure the software is free for all its users.
|
|
||||||
|
|
||||||
This license, the Lesser General Public License, applies to some
|
|
||||||
specially designated software packages--typically libraries--of the
|
|
||||||
Free Software Foundation and other authors who decide to use it. You
|
|
||||||
can use it too, but we suggest you first think carefully about whether
|
|
||||||
this license or the ordinary General Public License is the better
|
|
||||||
strategy to use in any particular case, based on the explanations below.
|
|
||||||
|
|
||||||
When we speak of free software, we are referring to freedom of use,
|
|
||||||
not price. Our General Public Licenses are designed to make sure that
|
|
||||||
you have the freedom to distribute copies of free software (and charge
|
|
||||||
for this service if you wish); that you receive source code or can get
|
|
||||||
it if you want it; that you can change the software and use pieces of
|
|
||||||
it in new free programs; and that you are informed that you can do
|
|
||||||
these things.
|
|
||||||
|
|
||||||
To protect your rights, we need to make restrictions that forbid
|
|
||||||
distributors to deny you these rights or to ask you to surrender these
|
|
||||||
rights. These restrictions translate to certain responsibilities for
|
|
||||||
you if you distribute copies of the library or if you modify it.
|
|
||||||
|
|
||||||
For example, if you distribute copies of the library, whether gratis
|
|
||||||
or for a fee, you must give the recipients all the rights that we gave
|
|
||||||
you. You must make sure that they, too, receive or can get the source
|
|
||||||
code. If you link other code with the library, you must provide
|
|
||||||
complete object files to the recipients, so that they can relink them
|
|
||||||
with the library after making changes to the library and recompiling
|
|
||||||
it. And you must show them these terms so they know their rights.
|
|
||||||
|
|
||||||
We protect your rights with a two-step method: (1) we copyright the
|
|
||||||
library, and (2) we offer you this license, which gives you legal
|
|
||||||
permission to copy, distribute and/or modify the library.
|
|
||||||
|
|
||||||
To protect each distributor, we want to make it very clear that
|
|
||||||
there is no warranty for the free library. Also, if the library is
|
|
||||||
modified by someone else and passed on, the recipients should know
|
|
||||||
that what they have is not the original version, so that the original
|
|
||||||
author's reputation will not be affected by problems that might be
|
|
||||||
introduced by others.
|
|
||||||
|
|
||||||
Finally, software patents pose a constant threat to the existence of
|
|
||||||
any free program. We wish to make sure that a company cannot
|
|
||||||
effectively restrict the users of a free program by obtaining a
|
|
||||||
restrictive license from a patent holder. Therefore, we insist that
|
|
||||||
any patent license obtained for a version of the library must be
|
|
||||||
consistent with the full freedom of use specified in this license.
|
|
||||||
|
|
||||||
Most GNU software, including some libraries, is covered by the
|
|
||||||
ordinary GNU General Public License. This license, the GNU Lesser
|
|
||||||
General Public License, applies to certain designated libraries, and
|
|
||||||
is quite different from the ordinary General Public License. We use
|
|
||||||
this license for certain libraries in order to permit linking those
|
|
||||||
libraries into non-free programs.
|
|
||||||
|
|
||||||
When a program is linked with a library, whether statically or using
|
|
||||||
a shared library, the combination of the two is legally speaking a
|
|
||||||
combined work, a derivative of the original library. The ordinary
|
|
||||||
General Public License therefore permits such linking only if the
|
|
||||||
entire combination fits its criteria of freedom. The Lesser General
|
|
||||||
Public License permits more lax criteria for linking other code with
|
|
||||||
the library.
|
|
||||||
|
|
||||||
We call this license the "Lesser" General Public License because it
|
|
||||||
does Less to protect the user's freedom than the ordinary General
|
|
||||||
Public License. It also provides other free software developers Less
|
|
||||||
of an advantage over competing non-free programs. These disadvantages
|
|
||||||
are the reason we use the ordinary General Public License for many
|
|
||||||
libraries. However, the Lesser license provides advantages in certain
|
|
||||||
special circumstances.
|
|
||||||
|
|
||||||
For example, on rare occasions, there may be a special need to
|
|
||||||
encourage the widest possible use of a certain library, so that it becomes
|
|
||||||
a de-facto standard. To achieve this, non-free programs must be
|
|
||||||
allowed to use the library. A more frequent case is that a free
|
|
||||||
library does the same job as widely used non-free libraries. In this
|
|
||||||
case, there is little to gain by limiting the free library to free
|
|
||||||
software only, so we use the Lesser General Public License.
|
|
||||||
|
|
||||||
In other cases, permission to use a particular library in non-free
|
|
||||||
programs enables a greater number of people to use a large body of
|
|
||||||
free software. For example, permission to use the GNU C Library in
|
|
||||||
non-free programs enables many more people to use the whole GNU
|
|
||||||
operating system, as well as its variant, the GNU/Linux operating
|
|
||||||
system.
|
|
||||||
|
|
||||||
Although the Lesser General Public License is Less protective of the
|
|
||||||
users' freedom, it does ensure that the user of a program that is
|
|
||||||
linked with the Library has the freedom and the wherewithal to run
|
|
||||||
that program using a modified version of the Library.
|
|
||||||
|
|
||||||
The precise terms and conditions for copying, distribution and
|
|
||||||
modification follow. Pay close attention to the difference between a
|
|
||||||
"work based on the library" and a "work that uses the library". The
|
|
||||||
former contains code derived from the library, whereas the latter must
|
|
||||||
be combined with the library in order to run.
|
|
||||||
|
|
||||||
GNU LESSER GENERAL PUBLIC LICENSE
|
|
||||||
TERMS AND CONDITIONS FOR COPYING, DISTRIBUTION AND MODIFICATION
|
|
||||||
|
|
||||||
0. This License Agreement applies to any software library or other
|
|
||||||
program which contains a notice placed by the copyright holder or
|
|
||||||
other authorized party saying it may be distributed under the terms of
|
|
||||||
this Lesser General Public License (also called "this License").
|
|
||||||
Each licensee is addressed as "you".
|
|
||||||
|
|
||||||
A "library" means a collection of software functions and/or data
|
|
||||||
prepared so as to be conveniently linked with application programs
|
|
||||||
(which use some of those functions and data) to form executables.
|
|
||||||
|
|
||||||
The "Library", below, refers to any such software library or work
|
|
||||||
which has been distributed under these terms. A "work based on the
|
|
||||||
Library" means either the Library or any derivative work under
|
|
||||||
copyright law: that is to say, a work containing the Library or a
|
|
||||||
portion of it, either verbatim or with modifications and/or translated
|
|
||||||
straightforwardly into another language. (Hereinafter, translation is
|
|
||||||
included without limitation in the term "modification".)
|
|
||||||
|
|
||||||
"Source code" for a work means the preferred form of the work for
|
|
||||||
making modifications to it. For a library, complete source code means
|
|
||||||
all the source code for all modules it contains, plus any associated
|
|
||||||
interface definition files, plus the scripts used to control compilation
|
|
||||||
and installation of the library.
|
|
||||||
|
|
||||||
Activities other than copying, distribution and modification are not
|
|
||||||
covered by this License; they are outside its scope. The act of
|
|
||||||
running a program using the Library is not restricted, and output from
|
|
||||||
such a program is covered only if its contents constitute a work based
|
|
||||||
on the Library (independent of the use of the Library in a tool for
|
|
||||||
writing it). Whether that is true depends on what the Library does
|
|
||||||
and what the program that uses the Library does.
|
|
||||||
|
|
||||||
1. You may copy and distribute verbatim copies of the Library's
|
|
||||||
complete source code as you receive it, in any medium, provided that
|
|
||||||
you conspicuously and appropriately publish on each copy an
|
|
||||||
appropriate copyright notice and disclaimer of warranty; keep intact
|
|
||||||
all the notices that refer to this License and to the absence of any
|
|
||||||
warranty; and distribute a copy of this License along with the
|
|
||||||
Library.
|
|
||||||
|
|
||||||
You may charge a fee for the physical act of transferring a copy,
|
|
||||||
and you may at your option offer warranty protection in exchange for a
|
|
||||||
fee.
|
|
||||||
|
|
||||||
2. You may modify your copy or copies of the Library or any portion
|
|
||||||
of it, thus forming a work based on the Library, and copy and
|
|
||||||
distribute such modifications or work under the terms of Section 1
|
|
||||||
above, provided that you also meet all of these conditions:
|
|
||||||
|
|
||||||
a) The modified work must itself be a software library.
|
|
||||||
|
|
||||||
b) You must cause the files modified to carry prominent notices
|
|
||||||
stating that you changed the files and the date of any change.
|
|
||||||
|
|
||||||
c) You must cause the whole of the work to be licensed at no
|
|
||||||
charge to all third parties under the terms of this License.
|
|
||||||
|
|
||||||
d) If a facility in the modified Library refers to a function or a
|
|
||||||
table of data to be supplied by an application program that uses
|
|
||||||
the facility, other than as an argument passed when the facility
|
|
||||||
is invoked, then you must make a good faith effort to ensure that,
|
|
||||||
in the event an application does not supply such function or
|
|
||||||
table, the facility still operates, and performs whatever part of
|
|
||||||
its purpose remains meaningful.
|
|
||||||
|
|
||||||
(For example, a function in a library to compute square roots has
|
|
||||||
a purpose that is entirely well-defined independent of the
|
|
||||||
application. Therefore, Subsection 2d requires that any
|
|
||||||
application-supplied function or table used by this function must
|
|
||||||
be optional: if the application does not supply it, the square
|
|
||||||
root function must still compute square roots.)
|
|
||||||
|
|
||||||
These requirements apply to the modified work as a whole. If
|
|
||||||
identifiable sections of that work are not derived from the Library,
|
|
||||||
and can be reasonably considered independent and separate works in
|
|
||||||
themselves, then this License, and its terms, do not apply to those
|
|
||||||
sections when you distribute them as separate works. But when you
|
|
||||||
distribute the same sections as part of a whole which is a work based
|
|
||||||
on the Library, the distribution of the whole must be on the terms of
|
|
||||||
this License, whose permissions for other licensees extend to the
|
|
||||||
entire whole, and thus to each and every part regardless of who wrote
|
|
||||||
it.
|
|
||||||
|
|
||||||
Thus, it is not the intent of this section to claim rights or contest
|
|
||||||
your rights to work written entirely by you; rather, the intent is to
|
|
||||||
exercise the right to control the distribution of derivative or
|
|
||||||
collective works based on the Library.
|
|
||||||
|
|
||||||
In addition, mere aggregation of another work not based on the Library
|
|
||||||
with the Library (or with a work based on the Library) on a volume of
|
|
||||||
a storage or distribution medium does not bring the other work under
|
|
||||||
the scope of this License.
|
|
||||||
|
|
||||||
3. You may opt to apply the terms of the ordinary GNU General Public
|
|
||||||
License instead of this License to a given copy of the Library. To do
|
|
||||||
this, you must alter all the notices that refer to this License, so
|
|
||||||
that they refer to the ordinary GNU General Public License, version 2,
|
|
||||||
instead of to this License. (If a newer version than version 2 of the
|
|
||||||
ordinary GNU General Public License has appeared, then you can specify
|
|
||||||
that version instead if you wish.) Do not make any other change in
|
|
||||||
these notices.
|
|
||||||
|
|
||||||
Once this change is made in a given copy, it is irreversible for
|
|
||||||
that copy, so the ordinary GNU General Public License applies to all
|
|
||||||
subsequent copies and derivative works made from that copy.
|
|
||||||
|
|
||||||
This option is useful when you wish to copy part of the code of
|
|
||||||
the Library into a program that is not a library.
|
|
||||||
|
|
||||||
4. You may copy and distribute the Library (or a portion or
|
|
||||||
derivative of it, under Section 2) in object code or executable form
|
|
||||||
under the terms of Sections 1 and 2 above provided that you accompany
|
|
||||||
it with the complete corresponding machine-readable source code, which
|
|
||||||
must be distributed under the terms of Sections 1 and 2 above on a
|
|
||||||
medium customarily used for software interchange.
|
|
||||||
|
|
||||||
If distribution of object code is made by offering access to copy
|
|
||||||
from a designated place, then offering equivalent access to copy the
|
|
||||||
source code from the same place satisfies the requirement to
|
|
||||||
distribute the source code, even though third parties are not
|
|
||||||
compelled to copy the source along with the object code.
|
|
||||||
|
|
||||||
5. A program that contains no derivative of any portion of the
|
|
||||||
Library, but is designed to work with the Library by being compiled or
|
|
||||||
linked with it, is called a "work that uses the Library". Such a
|
|
||||||
work, in isolation, is not a derivative work of the Library, and
|
|
||||||
therefore falls outside the scope of this License.
|
|
||||||
|
|
||||||
However, linking a "work that uses the Library" with the Library
|
|
||||||
creates an executable that is a derivative of the Library (because it
|
|
||||||
contains portions of the Library), rather than a "work that uses the
|
|
||||||
library". The executable is therefore covered by this License.
|
|
||||||
Section 6 states terms for distribution of such executables.
|
|
||||||
|
|
||||||
When a "work that uses the Library" uses material from a header file
|
|
||||||
that is part of the Library, the object code for the work may be a
|
|
||||||
derivative work of the Library even though the source code is not.
|
|
||||||
Whether this is true is especially significant if the work can be
|
|
||||||
linked without the Library, or if the work is itself a library. The
|
|
||||||
threshold for this to be true is not precisely defined by law.
|
|
||||||
|
|
||||||
If such an object file uses only numerical parameters, data
|
|
||||||
structure layouts and accessors, and small macros and small inline
|
|
||||||
functions (ten lines or less in length), then the use of the object
|
|
||||||
file is unrestricted, regardless of whether it is legally a derivative
|
|
||||||
work. (Executables containing this object code plus portions of the
|
|
||||||
Library will still fall under Section 6.)
|
|
||||||
|
|
||||||
Otherwise, if the work is a derivative of the Library, you may
|
|
||||||
distribute the object code for the work under the terms of Section 6.
|
|
||||||
Any executables containing that work also fall under Section 6,
|
|
||||||
whether or not they are linked directly with the Library itself.
|
|
||||||
|
|
||||||
6. As an exception to the Sections above, you may also combine or
|
|
||||||
link a "work that uses the Library" with the Library to produce a
|
|
||||||
work containing portions of the Library, and distribute that work
|
|
||||||
under terms of your choice, provided that the terms permit
|
|
||||||
modification of the work for the customer's own use and reverse
|
|
||||||
engineering for debugging such modifications.
|
|
||||||
|
|
||||||
You must give prominent notice with each copy of the work that the
|
|
||||||
Library is used in it and that the Library and its use are covered by
|
|
||||||
this License. You must supply a copy of this License. If the work
|
|
||||||
during execution displays copyright notices, you must include the
|
|
||||||
copyright notice for the Library among them, as well as a reference
|
|
||||||
directing the user to the copy of this License. Also, you must do one
|
|
||||||
of these things:
|
|
||||||
|
|
||||||
a) Accompany the work with the complete corresponding
|
|
||||||
machine-readable source code for the Library including whatever
|
|
||||||
changes were used in the work (which must be distributed under
|
|
||||||
Sections 1 and 2 above); and, if the work is an executable linked
|
|
||||||
with the Library, with the complete machine-readable "work that
|
|
||||||
uses the Library", as object code and/or source code, so that the
|
|
||||||
user can modify the Library and then relink to produce a modified
|
|
||||||
executable containing the modified Library. (It is understood
|
|
||||||
that the user who changes the contents of definitions files in the
|
|
||||||
Library will not necessarily be able to recompile the application
|
|
||||||
to use the modified definitions.)
|
|
||||||
|
|
||||||
b) Use a suitable shared library mechanism for linking with the
|
|
||||||
Library. A suitable mechanism is one that (1) uses at run time a
|
|
||||||
copy of the library already present on the user's computer system,
|
|
||||||
rather than copying library functions into the executable, and (2)
|
|
||||||
will operate properly with a modified version of the library, if
|
|
||||||
the user installs one, as long as the modified version is
|
|
||||||
interface-compatible with the version that the work was made with.
|
|
||||||
|
|
||||||
c) Accompany the work with a written offer, valid for at
|
|
||||||
least three years, to give the same user the materials
|
|
||||||
specified in Subsection 6a, above, for a charge no more
|
|
||||||
than the cost of performing this distribution.
|
|
||||||
|
|
||||||
d) If distribution of the work is made by offering access to copy
|
|
||||||
from a designated place, offer equivalent access to copy the above
|
|
||||||
specified materials from the same place.
|
|
||||||
|
|
||||||
e) Verify that the user has already received a copy of these
|
|
||||||
materials or that you have already sent this user a copy.
|
|
||||||
|
|
||||||
For an executable, the required form of the "work that uses the
|
|
||||||
Library" must include any data and utility programs needed for
|
|
||||||
reproducing the executable from it. However, as a special exception,
|
|
||||||
the materials to be distributed need not include anything that is
|
|
||||||
normally distributed (in either source or binary form) with the major
|
|
||||||
components (compiler, kernel, and so on) of the operating system on
|
|
||||||
which the executable runs, unless that component itself accompanies
|
|
||||||
the executable.
|
|
||||||
|
|
||||||
It may happen that this requirement contradicts the license
|
|
||||||
restrictions of other proprietary libraries that do not normally
|
|
||||||
accompany the operating system. Such a contradiction means you cannot
|
|
||||||
use both them and the Library together in an executable that you
|
|
||||||
distribute.
|
|
||||||
|
|
||||||
7. You may place library facilities that are a work based on the
|
|
||||||
Library side-by-side in a single library together with other library
|
|
||||||
facilities not covered by this License, and distribute such a combined
|
|
||||||
library, provided that the separate distribution of the work based on
|
|
||||||
the Library and of the other library facilities is otherwise
|
|
||||||
permitted, and provided that you do these two things:
|
|
||||||
|
|
||||||
a) Accompany the combined library with a copy of the same work
|
|
||||||
based on the Library, uncombined with any other library
|
|
||||||
facilities. This must be distributed under the terms of the
|
|
||||||
Sections above.
|
|
||||||
|
|
||||||
b) Give prominent notice with the combined library of the fact
|
|
||||||
that part of it is a work based on the Library, and explaining
|
|
||||||
where to find the accompanying uncombined form of the same work.
|
|
||||||
|
|
||||||
8. You may not copy, modify, sublicense, link with, or distribute
|
|
||||||
the Library except as expressly provided under this License. Any
|
|
||||||
attempt otherwise to copy, modify, sublicense, link with, or
|
|
||||||
distribute the Library is void, and will automatically terminate your
|
|
||||||
rights under this License. However, parties who have received copies,
|
|
||||||
or rights, from you under this License will not have their licenses
|
|
||||||
terminated so long as such parties remain in full compliance.
|
|
||||||
|
|
||||||
9. You are not required to accept this License, since you have not
|
|
||||||
signed it. However, nothing else grants you permission to modify or
|
|
||||||
distribute the Library or its derivative works. These actions are
|
|
||||||
prohibited by law if you do not accept this License. Therefore, by
|
|
||||||
modifying or distributing the Library (or any work based on the
|
|
||||||
Library), you indicate your acceptance of this License to do so, and
|
|
||||||
all its terms and conditions for copying, distributing or modifying
|
|
||||||
the Library or works based on it.
|
|
||||||
|
|
||||||
10. Each time you redistribute the Library (or any work based on the
|
|
||||||
Library), the recipient automatically receives a license from the
|
|
||||||
original licensor to copy, distribute, link with or modify the Library
|
|
||||||
subject to these terms and conditions. You may not impose any further
|
|
||||||
restrictions on the recipients' exercise of the rights granted herein.
|
|
||||||
You are not responsible for enforcing compliance by third parties with
|
|
||||||
this License.
|
|
||||||
|
|
||||||
11. If, as a consequence of a court judgment or allegation of patent
|
|
||||||
infringement or for any other reason (not limited to patent issues),
|
|
||||||
conditions are imposed on you (whether by court order, agreement or
|
|
||||||
otherwise) that contradict the conditions of this License, they do not
|
|
||||||
excuse you from the conditions of this License. If you cannot
|
|
||||||
distribute so as to satisfy simultaneously your obligations under this
|
|
||||||
License and any other pertinent obligations, then as a consequence you
|
|
||||||
may not distribute the Library at all. For example, if a patent
|
|
||||||
license would not permit royalty-free redistribution of the Library by
|
|
||||||
all those who receive copies directly or indirectly through you, then
|
|
||||||
the only way you could satisfy both it and this License would be to
|
|
||||||
refrain entirely from distribution of the Library.
|
|
||||||
|
|
||||||
If any portion of this section is held invalid or unenforceable under any
|
|
||||||
particular circumstance, the balance of the section is intended to apply,
|
|
||||||
and the section as a whole is intended to apply in other circumstances.
|
|
||||||
|
|
||||||
It is not the purpose of this section to induce you to infringe any
|
|
||||||
patents or other property right claims or to contest validity of any
|
|
||||||
such claims; this section has the sole purpose of protecting the
|
|
||||||
integrity of the free software distribution system which is
|
|
||||||
implemented by public license practices. Many people have made
|
|
||||||
generous contributions to the wide range of software distributed
|
|
||||||
through that system in reliance on consistent application of that
|
|
||||||
system; it is up to the author/donor to decide if he or she is willing
|
|
||||||
to distribute software through any other system and a licensee cannot
|
|
||||||
impose that choice.
|
|
||||||
|
|
||||||
This section is intended to make thoroughly clear what is believed to
|
|
||||||
be a consequence of the rest of this License.
|
|
||||||
|
|
||||||
12. If the distribution and/or use of the Library is restricted in
|
|
||||||
certain countries either by patents or by copyrighted interfaces, the
|
|
||||||
original copyright holder who places the Library under this License may add
|
|
||||||
an explicit geographical distribution limitation excluding those countries,
|
|
||||||
so that distribution is permitted only in or among countries not thus
|
|
||||||
excluded. In such case, this License incorporates the limitation as if
|
|
||||||
written in the body of this License.
|
|
||||||
|
|
||||||
13. The Free Software Foundation may publish revised and/or new
|
|
||||||
versions of the Lesser General Public License from time to time.
|
|
||||||
Such new versions will be similar in spirit to the present version,
|
|
||||||
but may differ in detail to address new problems or concerns.
|
|
||||||
|
|
||||||
Each version is given a distinguishing version number. If the Library
|
|
||||||
specifies a version number of this License which applies to it and
|
|
||||||
"any later version", you have the option of following the terms and
|
|
||||||
conditions either of that version or of any later version published by
|
|
||||||
the Free Software Foundation. If the Library does not specify a
|
|
||||||
license version number, you may choose any version ever published by
|
|
||||||
the Free Software Foundation.
|
|
||||||
|
|
||||||
14. If you wish to incorporate parts of the Library into other free
|
|
||||||
programs whose distribution conditions are incompatible with these,
|
|
||||||
write to the author to ask for permission. For software which is
|
|
||||||
copyrighted by the Free Software Foundation, write to the Free
|
|
||||||
Software Foundation; we sometimes make exceptions for this. Our
|
|
||||||
decision will be guided by the two goals of preserving the free status
|
|
||||||
of all derivatives of our free software and of promoting the sharing
|
|
||||||
and reuse of software generally.
|
|
||||||
|
|
||||||
NO WARRANTY
|
|
||||||
|
|
||||||
15. BECAUSE THE LIBRARY IS LICENSED FREE OF CHARGE, THERE IS NO
|
|
||||||
WARRANTY FOR THE LIBRARY, TO THE EXTENT PERMITTED BY APPLICABLE LAW.
|
|
||||||
EXCEPT WHEN OTHERWISE STATED IN WRITING THE COPYRIGHT HOLDERS AND/OR
|
|
||||||
OTHER PARTIES PROVIDE THE LIBRARY "AS IS" WITHOUT WARRANTY OF ANY
|
|
||||||
KIND, EITHER EXPRESSED OR IMPLIED, INCLUDING, BUT NOT LIMITED TO, THE
|
|
||||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR
|
|
||||||
PURPOSE. THE ENTIRE RISK AS TO THE QUALITY AND PERFORMANCE OF THE
|
|
||||||
LIBRARY IS WITH YOU. SHOULD THE LIBRARY PROVE DEFECTIVE, YOU ASSUME
|
|
||||||
THE COST OF ALL NECESSARY SERVICING, REPAIR OR CORRECTION.
|
|
||||||
|
|
||||||
16. IN NO EVENT UNLESS REQUIRED BY APPLICABLE LAW OR AGREED TO IN
|
|
||||||
WRITING WILL ANY COPYRIGHT HOLDER, OR ANY OTHER PARTY WHO MAY MODIFY
|
|
||||||
AND/OR REDISTRIBUTE THE LIBRARY AS PERMITTED ABOVE, BE LIABLE TO YOU
|
|
||||||
FOR DAMAGES, INCLUDING ANY GENERAL, SPECIAL, INCIDENTAL OR
|
|
||||||
CONSEQUENTIAL DAMAGES ARISING OUT OF THE USE OR INABILITY TO USE THE
|
|
||||||
LIBRARY (INCLUDING BUT NOT LIMITED TO LOSS OF DATA OR DATA BEING
|
|
||||||
RENDERED INACCURATE OR LOSSES SUSTAINED BY YOU OR THIRD PARTIES OR A
|
|
||||||
FAILURE OF THE LIBRARY TO OPERATE WITH ANY OTHER SOFTWARE), EVEN IF
|
|
||||||
SUCH HOLDER OR OTHER PARTY HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH
|
|
||||||
DAMAGES.
|
|
||||||
|
|
||||||
END OF TERMS AND CONDITIONS
|
|
||||||
|
|
||||||
How to Apply These Terms to Your New Libraries
|
|
||||||
|
|
||||||
If you develop a new library, and you want it to be of the greatest
|
|
||||||
possible use to the public, we recommend making it free software that
|
|
||||||
everyone can redistribute and change. You can do so by permitting
|
|
||||||
redistribution under these terms (or, alternatively, under the terms of the
|
|
||||||
ordinary General Public License).
|
|
||||||
|
|
||||||
To apply these terms, attach the following notices to the library. It is
|
|
||||||
safest to attach them to the start of each source file to most effectively
|
|
||||||
convey the exclusion of warranty; and each file should have at least the
|
|
||||||
"copyright" line and a pointer to where the full notice is found.
|
|
||||||
|
|
||||||
<one line to give the library's name and a brief idea of what it does.>
|
|
||||||
Copyright (C) <year> <name of author>
|
|
||||||
|
|
||||||
This library is free software; you can redistribute it and/or
|
|
||||||
modify it under the terms of the GNU Lesser General Public
|
|
||||||
License as published by the Free Software Foundation; either
|
|
||||||
version 2.1 of the License, or (at your option) any later version.
|
|
||||||
|
|
||||||
This library is distributed in the hope that it will be useful,
|
|
||||||
but WITHOUT ANY WARRANTY; without even the implied warranty of
|
|
||||||
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU
|
|
||||||
Lesser General Public License for more details.
|
|
||||||
|
|
||||||
You should have received a copy of the GNU Lesser General Public
|
|
||||||
License along with this library; if not, write to the Free Software
|
|
||||||
Foundation, Inc., 59 Temple Place, Suite 330, Boston, MA 02111-1307 USA
|
|
||||||
|
|
||||||
Also add information on how to contact you by electronic and paper mail.
|
|
||||||
|
|
||||||
You should also get your employer (if you work as a programmer) or your
|
|
||||||
school, if any, to sign a "copyright disclaimer" for the library, if
|
|
||||||
necessary. Here is a sample; alter the names:
|
|
||||||
|
|
||||||
Yoyodyne, Inc., hereby disclaims all copyright interest in the
|
|
||||||
library `Frob' (a library for tweaking knobs) written by James Random Hacker.
|
|
||||||
|
|
||||||
<signature of Ty Coon>, 1 April 1990
|
|
||||||
Ty Coon, President of Vice
|
|
||||||
|
|
||||||
That's all there is to it!
|
|
||||||
@@ -1,17 +0,0 @@
|
|||||||
JHOVE - JSTOR/Harvard Object Validation Environment
|
|
||||||
Copyright 2003-2008 by JSTOR and the President and Fellows of Harvard College
|
|
||||||
|
|
||||||
This program is free software; you can redistribute it and/or modify
|
|
||||||
it under the terms of the GNU Lessor General Public License as
|
|
||||||
published by the Free Software Foundation; either version 2.1 of the
|
|
||||||
License, or (at your option) any later version.
|
|
||||||
|
|
||||||
This program is distributed in the hope that it will be useful, but
|
|
||||||
WITHOUT ANY WARRANTY; without even the implied warranty of
|
|
||||||
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU
|
|
||||||
Lesser General Public License for more details.
|
|
||||||
|
|
||||||
You should have received a copy of the GNU Lesser General Public
|
|
||||||
License along with this program; if not, write to the Free Software
|
|
||||||
Foundation, Inc., 59 Temple Place, Suite 330, Boston, MA 02111-1307
|
|
||||||
USA
|
|
||||||
@@ -1,227 +0,0 @@
|
|||||||
JHOVE - JSTOR/Harvard Object Validation Environment
|
|
||||||
Copyright 2003-2012 by JSTOR and the President and Fellows of Harvard College
|
|
||||||
JHOVE is made available under the GNU Lesser General Public License (LGPL;
|
|
||||||
see the file LICENSE for details)
|
|
||||||
|
|
||||||
Rev. 1.11, 2013-09-29
|
|
||||||
|
|
||||||
JHOVE (the JSTOR/Harvard Object Validation Environment, pronounced "jhove")
|
|
||||||
is an extensible software framework for performing format identification,
|
|
||||||
validation, and characterization of digital objects.
|
|
||||||
|
|
||||||
o Format identification is the process of determining the format to which a
|
|
||||||
digital object conforms: "I have a digital object; what format is it?"
|
|
||||||
o Format validation is the process of determining the level of compliance of a
|
|
||||||
digital object to the specification for its purported format: "I have an
|
|
||||||
object purportedly of format F; is it?"
|
|
||||||
o Format characterization is the process of determing the format-specific
|
|
||||||
significant properties of an object of a given format: "I have an object of
|
|
||||||
format F; what are its salient properties?"
|
|
||||||
|
|
||||||
These actions are frequently necessary during routine operation of digital
|
|
||||||
repositories and for digital preservation activities.
|
|
||||||
|
|
||||||
The output from JHOVE is controlled by output handlers. JHOVE uses an
|
|
||||||
extensible plug-in architecture; it can be configured at the time of its
|
|
||||||
invocation to include whatever specific format modules and output handlers
|
|
||||||
that are desired. The initial release of JHOVE includes modules for
|
|
||||||
arbitrary byte streams, ASCII and UTF-8 encoded text, AIFF and WAVE audio,
|
|
||||||
GIF, JPEG, JPEG 2000, TIFF, and PDF; and text and XML output handlers.
|
|
||||||
|
|
||||||
The JHOVE project is a collaboration of JSTOR and the Harvard University
|
|
||||||
Library. Development of JHOVE was funded in part by the Andrew W. Mellon
|
|
||||||
Foundation. JHOVE is made available under the GNU Lesser General Public
|
|
||||||
License (LGPL; see the file LICENSE for details).
|
|
||||||
|
|
||||||
JHOVE is currently being maintained by indpendent developers.
|
|
||||||
|
|
||||||
REQUIREMENTS
|
|
||||||
|
|
||||||
1. Java J2SE 1.5
|
|
||||||
(JHOVE was originally implemented using the Sun J2SE SDK 1.4.1 and has
|
|
||||||
been tested to work with 1.5)
|
|
||||||
|
|
||||||
2. If you would like to compile the JHOVE source code, then
|
|
||||||
Apache Ant, a Java-based build tool <http://ant.apache.org/> is necessary.
|
|
||||||
Note that the JAVA_HOME environment variable must be set appropriately for
|
|
||||||
Ant to work properly.
|
|
||||||
(JHOVE was implemented and tested using Ant 1.5.1.)
|
|
||||||
|
|
||||||
DISTRIBUTION
|
|
||||||
|
|
||||||
The JHOVE distribution package includes:
|
|
||||||
|
|
||||||
jhove/ # JHOVE home directory
|
|
||||||
COPYING # GNU Lesser General Public License
|
|
||||||
LICENSE # JHOVE license information
|
|
||||||
README
|
|
||||||
RELEASENOTES # JHOVE release notes
|
|
||||||
bin/
|
|
||||||
jhove.jar # JHOVE API package
|
|
||||||
jhove-handler.jar # Standard output handler package
|
|
||||||
jhove-module.jar # Standard module package
|
|
||||||
JhoveApp.jar # JHOVE command line application
|
|
||||||
JhoveView.jar # JHOVE with Swing GUI front-end
|
|
||||||
build.xml # Ant configuration file
|
|
||||||
classes/
|
|
||||||
build.xml # Ant configuration file
|
|
||||||
edu/ ... # JHOVE API packages
|
|
||||||
ADump.* # AIFF dump utility class
|
|
||||||
GDump.* # GIF dump utility class
|
|
||||||
Jhove.* # JHOVE main class
|
|
||||||
JDump.* # JPEG dump utility class
|
|
||||||
J2Dump.* # JPEG 2000 dump utility class
|
|
||||||
PDump.* # PDF dump utility class
|
|
||||||
TDump.* # TIFF dump utility class
|
|
||||||
UserHome.* # user.home property utility class
|
|
||||||
WDump.* # WAVE dump utility class
|
|
||||||
conf/
|
|
||||||
jhove.conf # JHOVE configuration file
|
|
||||||
jhove.xsd # JHOVE output schema
|
|
||||||
jhoveConfig.xsd # JHOVE configuration file schema
|
|
||||||
doc/
|
|
||||||
*.html # API documentation
|
|
||||||
...
|
|
||||||
examples/ # Sample files
|
|
||||||
ascii/ ...
|
|
||||||
gif/ ...
|
|
||||||
jpeg/ ...
|
|
||||||
jpeg2000/ ...
|
|
||||||
pdf/ ...
|
|
||||||
tiff/ ...
|
|
||||||
utf-8/ ...
|
|
||||||
adump* # AIFF dump Bourne shell driver
|
|
||||||
adump.bat* # AIFF dump DOS shell driver script
|
|
||||||
gdump* # GIF dump Bourne shell driver
|
|
||||||
gdump.bat* # GIF dump DOS shell driver script
|
|
||||||
jdump* # JPEG dump Bourne shell driver
|
|
||||||
jdump.bat* # JPEG dump DOS shell driver script
|
|
||||||
j2dump* # JPEG 2000 dump Bourne shell driver
|
|
||||||
j2dump.bat* # JPEG 2000 dump DOS shell driver
|
|
||||||
jhove.tmpl* # Template for JHOVE Bourne shell driver script
|
|
||||||
jhove_bat.tmpl* # Template for JHOVE DOS shell driver script
|
|
||||||
pdump* # PDF dump Bourne shell driver
|
|
||||||
pdump.bat* # PDF dump DOS shell driver script
|
|
||||||
tdump* # TIFF dump Bourne shell driver
|
|
||||||
tdump.bat* # TIFF dump DOS shell driver script
|
|
||||||
userhome* # user.home Bourne shell driver
|
|
||||||
userhome.bat* # user.home DOS shell driver script
|
|
||||||
wdump* # WAVE dump Bourne shell driver
|
|
||||||
wdump.bat* # WAVE dump DOS shell driver script
|
|
||||||
|
|
||||||
INSTALLATION
|
|
||||||
|
|
||||||
Edit the configuration file, jhove/conf/jhove.conf, and set the absolute
|
|
||||||
pathname of the JHOVE home directory and the temporary directory (in which
|
|
||||||
temporary files are created):
|
|
||||||
|
|
||||||
<jhoveHome>jhove-home-directory</jhoveHome>
|
|
||||||
<tempDirectory>temporary-directory</tempDirectory>
|
|
||||||
|
|
||||||
The JHOVE home directory is the top-most directory in the distribution TAR
|
|
||||||
or ZIP file. On Unix systems, "/var/tmp" is an appropriate temporary
|
|
||||||
directory; on Windows, "C:\Temp". For example, if the distribution TAR
|
|
||||||
file is disaggregated on a Unix system in the directory "/users/stephen/
|
|
||||||
projects", then the configuration file should read:
|
|
||||||
|
|
||||||
<jhoveHome>/users/stephen/projects/jhove</jhoveHome>
|
|
||||||
<tempDirectory>/var/tmp</jhoveHome>
|
|
||||||
|
|
||||||
In the JHOVE home directory, copy the JHOVE Bourne shell driver script
|
|
||||||
template, "jhove.tmpl", to "jhove" (or the equivalent Windows shell
|
|
||||||
script, "jhove_bat.tmpl" to "jhove.bat"), and set the
|
|
||||||
JHOVE home directory, Java home directory, and Java interpreter:
|
|
||||||
|
|
||||||
JHOVE_HOME=jhove-home-directory
|
|
||||||
JAVA_HOME=java-home-directory
|
|
||||||
JAVA=java-interpreter
|
|
||||||
|
|
||||||
The JAVA_HOME property should provide the absolute pathname of the Java
|
|
||||||
runtime or SDK installation; JAVA should provide the absolute pathname of the
|
|
||||||
Java interpreter. For example:
|
|
||||||
|
|
||||||
JHOVE_HOME=/users/stephen/projects/jhove
|
|
||||||
JAVA_HOME=/usr/local/j2re1.4.1_02
|
|
||||||
JAVA=$JAVA_HOME/bin/java
|
|
||||||
|
|
||||||
In the DOS shell driver script, jhove.bat, the equivalent three
|
|
||||||
variables are:
|
|
||||||
|
|
||||||
SET JHOVE_HOME=jhove-home-directory
|
|
||||||
SET JAVA_HOME=java-home-directory
|
|
||||||
SET JAVA=%JAVA_HOME%\bin\java
|
|
||||||
|
|
||||||
For example:
|
|
||||||
|
|
||||||
SET JHOVE_HOME="C:\Program Files\jhove"
|
|
||||||
SET JAVA_HOME="C:\Program Files\java\j2re1.4.1_02"
|
|
||||||
SET JAVA=%JAVA_HOME%\bin\java
|
|
||||||
|
|
||||||
The quotation marks are necessary because of the embedded space characters.
|
|
||||||
On Windows platforms it may also be necessary to add the Java bin subdirectory
|
|
||||||
to the System PATH environment variable:
|
|
||||||
|
|
||||||
PATH=C:\Program Files\java\j2re1.4.1_02\bin;...
|
|
||||||
|
|
||||||
(For information on setting a Windows environment variable, consult your local
|
|
||||||
documentation or system administrator.)
|
|
||||||
|
|
||||||
USAGE
|
|
||||||
|
|
||||||
java Jhove [-c config] [-m module] [-h handler] [-e encoding] [-H handler]
|
|
||||||
[-o output] [-x saxclass] [-t tempdir] [-b bufsize]
|
|
||||||
[-l loglevel] [[-krs] dir-file-or-uri [...]]
|
|
||||||
|
|
||||||
where -c config Configuration file pathname
|
|
||||||
-m module Module name
|
|
||||||
-h handler Output handler name (defaults to TEXT)
|
|
||||||
-e encoding Character encoding used by output handler (defaults to UTF-8)
|
|
||||||
-H handler About handler name
|
|
||||||
-o output Output file pathname (defaults to standard output)
|
|
||||||
-x saxclass SAX parser class (defaults to J2SE default)
|
|
||||||
-t tempdir Temporary directory in which to create temporary files
|
|
||||||
-b bufsize Buffer size for buffered I/O (defaults to J2SE 1.4 default)
|
|
||||||
-l loglevel Logging level
|
|
||||||
-k Calculate CRC32, MD5, and SHA-1 checksums
|
|
||||||
-r Display raw data flags, not textual equivalents
|
|
||||||
-s Format identification based on internal signatures only
|
|
||||||
dir-file-or-uri Directory or file pathname or URI of formated content
|
|
||||||
stream
|
|
||||||
|
|
||||||
All named modules and output handlers must be found on the Java CLASSPATH at
|
|
||||||
the time of invocation. The JHOVE driver script, jhove/jhove, automatically
|
|
||||||
sets the CLASSPATH and invokes the Jhove main class:
|
|
||||||
|
|
||||||
jhove [-c config] [-m module] [-h handler] [-e encoding] [-H handler]
|
|
||||||
[-o output] [-x saxclass] [-t tempdir] [-b bufsize] [-l loglevel]
|
|
||||||
[[-krs] dir-file-or-uri [...]]
|
|
||||||
|
|
||||||
The following additional programs are available, primarily for testing
|
|
||||||
and debugging purposes. They display a minimally processed, human-readable
|
|
||||||
version of the contents of AIFF, GIF, JPEG, JPEG 2000, PDF, TIFF, and WAVE
|
|
||||||
files:
|
|
||||||
|
|
||||||
java ADump aiff-file
|
|
||||||
java GDump gif-file
|
|
||||||
java JDump jpeg-file
|
|
||||||
java J2Dump jpeg2000-file
|
|
||||||
java PDump pdf-file
|
|
||||||
java TDump tiff-file
|
|
||||||
java WDump wave-file
|
|
||||||
|
|
||||||
For convenience, the following driver scripts are also available:
|
|
||||||
|
|
||||||
adump aiff-file
|
|
||||||
gdump gif-file
|
|
||||||
jdump jpeg-file
|
|
||||||
j2dump jpeg2000-file
|
|
||||||
pdump pdf-file
|
|
||||||
tdump tiff-file
|
|
||||||
wdump wave-file
|
|
||||||
|
|
||||||
The JHOVE Swing-based GUI interface can be invoked from a command shell from
|
|
||||||
the jhove/bin sub-directory:
|
|
||||||
|
|
||||||
java -jar JhoveView.jar -c <configFile>
|
|
||||||
|
|
||||||
where <configFile> is the pathname of the JHOVE configuration file.
|
|
||||||
File diff suppressed because it is too large
Load Diff
Binary file not shown.
Binary file not shown.
@@ -1,25 +0,0 @@
|
|||||||
JHOVE - JSTOR/Harvard Object Validation Environment
|
|
||||||
Copyright 2003 by JSTOR and the President and Fellows of Harvard College
|
|
||||||
JHOVE is made available under the GNU General Public License (see the file
|
|
||||||
LICENSE for details)
|
|
||||||
|
|
||||||
Rev. 2003-11-25
|
|
||||||
|
|
||||||
The following jar files are meant to be used for embedding JHOVE functionality
|
|
||||||
into new applications or systems.
|
|
||||||
|
|
||||||
jhove.jar Contains the JHOVE API interfaces and classes
|
|
||||||
jhove-module.jar Contains the standard JHOVE modules ()
|
|
||||||
jhove-handler.jar Contains the standard JHOVE output handlers (TEXT and XML)
|
|
||||||
|
|
||||||
The following jar file is meant to be used with the stand-alone JHOVE
|
|
||||||
application using a command-line interface. It contains the main Jhove class
|
|
||||||
and the contents of jhove.jar, jhove-module.jar, and jhove-handler.jar.
|
|
||||||
|
|
||||||
JhoveApp.jar
|
|
||||||
|
|
||||||
The following jar file is meant to be used with the stand-alone JHOVE
|
|
||||||
application using a Swing GUI interface. It contains the main JhoveView class
|
|
||||||
and the contents of jhove.jar, jhove-module.jar, and jhove-handler.jar.
|
|
||||||
|
|
||||||
JhoveView.jar
|
|
||||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -1,78 +0,0 @@
|
|||||||
<project name="Jhove" default="dist" basedir=".">
|
|
||||||
<description>Project build file
|
|
||||||
Jhove - JSTOR/Harvard Object Validation Environment
|
|
||||||
Version 1.0 2004-09-10
|
|
||||||
Copyright 2004 by JSTOR and the President and Fellows of Harvard College
|
|
||||||
</description>
|
|
||||||
|
|
||||||
<!-- ant (or ant dist) Build everything
|
|
||||||
ant debug Build everything with debug enabled
|
|
||||||
ant clean Delete backup files
|
|
||||||
ant cleanclass Delete backup and class files
|
|
||||||
ant cleandist Delete backup, class, and jar files
|
|
||||||
ant javadoc Build javadocs
|
|
||||||
-->
|
|
||||||
|
|
||||||
<!-- set global properties for this build -->
|
|
||||||
<property name="bin" location="bin"/>
|
|
||||||
<property name="classes" location="classes"/>
|
|
||||||
<property name="doc" location="doc"/>
|
|
||||||
|
|
||||||
<target name="dist" description="Create distribution">
|
|
||||||
<ant dir="${classes}" inheritAll="false">
|
|
||||||
<property name="dbg" value="off"/>
|
|
||||||
</ant>
|
|
||||||
<chmod file="jhove" perm="ugo+x"/>
|
|
||||||
<chmod file="${bin}/JhoveView.jar" perm="ugo+x"/>
|
|
||||||
</target>
|
|
||||||
|
|
||||||
<target name="debug" description="Create distribution with debug enabled">
|
|
||||||
<ant dir="${classes}" inheritAll="false">
|
|
||||||
<property name="dbg" value="on"/>
|
|
||||||
</ant>
|
|
||||||
</target>
|
|
||||||
|
|
||||||
<target name="view" description="Create JhoveView application">
|
|
||||||
<ant dir="${classes}" target="view" inheritAll="false">
|
|
||||||
<property name="dbs" value="on"/>
|
|
||||||
</ant>
|
|
||||||
</target>
|
|
||||||
|
|
||||||
<target name="clean" description="Delete backup files">
|
|
||||||
<ant dir="${classes}" target="main-clean" inheritAll="false"/>
|
|
||||||
</target>
|
|
||||||
|
|
||||||
<target name="cleanclass" depends="clean">
|
|
||||||
<ant dir="${classes}" target="main-cleanclass" inheritAll="false"/>
|
|
||||||
</target>
|
|
||||||
|
|
||||||
<target name="cleandist" depends="cleanclass">
|
|
||||||
<delete file="${bin}/jhove.jar"/>
|
|
||||||
<delete file="${bin}/jhove-handler.jar"/>
|
|
||||||
<delete file="${bin}/jhove-module.jar"/>
|
|
||||||
<delete file="${bin}/JhoveApp.jar"/>
|
|
||||||
<delete file="${bin}/JhoveView.jar"/>
|
|
||||||
</target>
|
|
||||||
|
|
||||||
<target name="javadoc">
|
|
||||||
<javadoc sourcepath="${classes}" destdir="${doc}"
|
|
||||||
windowtitle="JHOVE Documentation"
|
|
||||||
Overview="${classes}/overview.html">
|
|
||||||
<package name="edu.harvard.hul.ois.jhove"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.handler"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.handler.audit"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.module"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.module.aiff"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.module.gif"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.module.html"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.module.iff"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.module.jpeg"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.module.jpeg2000"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.module.pdf"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.module.tiff"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.module.wave"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.module.xml"/>
|
|
||||||
<package name="edu.harvard.hul.ois.jhove.viewer"/>
|
|
||||||
</javadoc>
|
|
||||||
</target>
|
|
||||||
</project>
|
|
||||||
@@ -1,63 +0,0 @@
|
|||||||
JHOVE - JSTOR/Harvard Object Validation Environment
|
|
||||||
Copyright 2003-2007 by JSTOR and the President and Fellows of Harvard College
|
|
||||||
JHOVE is made available under the GNU General Public License (see the file
|
|
||||||
LICENSE for details)
|
|
||||||
|
|
||||||
Rev. 2007-08-30
|
|
||||||
|
|
||||||
Edit the configuration file, jhove.conf, and set the JHOVE home
|
|
||||||
directory:
|
|
||||||
|
|
||||||
<jhoveHome>jhove-home-directory</jhoveHome>
|
|
||||||
|
|
||||||
and temporary directory:
|
|
||||||
|
|
||||||
<tempDirectory>temporary-directory</tempDirectory>
|
|
||||||
|
|
||||||
On most Unix systems, a reasonable temporary directory is "/var/tmp";
|
|
||||||
on Windows, "C:\temp".
|
|
||||||
|
|
||||||
The optional
|
|
||||||
|
|
||||||
<bufferSize>buffer-size</bufferSize>
|
|
||||||
|
|
||||||
element defines the buffer size used for buffer I/O operations.
|
|
||||||
|
|
||||||
The optional
|
|
||||||
|
|
||||||
<mixVersion>1.0</mixVersion>
|
|
||||||
|
|
||||||
element specifies that the XML output handler should conform to the
|
|
||||||
MIX 1.0 schema. The default behavior is for handler output to conform
|
|
||||||
to the MIX 0.2 schema.
|
|
||||||
|
|
||||||
The optional
|
|
||||||
|
|
||||||
<sigBytes>n</sigBytes>
|
|
||||||
|
|
||||||
element specifies that JHOVE modules will look for format signatures
|
|
||||||
in the first <n> bytes of the file. The default value is 1024.
|
|
||||||
|
|
||||||
All class names must be fully qualified with their package name:
|
|
||||||
|
|
||||||
<module>
|
|
||||||
<class>fully-package-qualified-class-name</class>
|
|
||||||
<init>optional-initialization-argument</init>
|
|
||||||
<param>optional-invocation-argument</param>
|
|
||||||
</module>
|
|
||||||
|
|
||||||
The optional <init> argument is passed to the module once at the time
|
|
||||||
its class is instantiated. See module-specific documentation for a
|
|
||||||
description of any initialization options.
|
|
||||||
|
|
||||||
The optional <param> argument is passed to the module every time it is
|
|
||||||
invoked. See module-specific documentation for a description of any
|
|
||||||
invocation options.
|
|
||||||
|
|
||||||
The order in which format modules are defined is important; when
|
|
||||||
performing a format identification operation, JHOVE will search for a
|
|
||||||
matching module in the order in which the modules are defined in the
|
|
||||||
configuration file. In general, the modules for more generic formats
|
|
||||||
should come later in the list. For example, the standard module ASCII
|
|
||||||
should be defined before the UTF-8 module, since all ASCII objects
|
|
||||||
are, by definition, UTF-8 objects, but not vice versa.
|
|
||||||
@@ -1,47 +0,0 @@
|
|||||||
<?xml version="1.0" encoding="UTF-8"?>
|
|
||||||
<jhoveConfig version="1.0"
|
|
||||||
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
|
|
||||||
xmlns="http://hul.harvard.edu/ois/xml/ns/jhove/jhoveConfig"
|
|
||||||
xsi:schemaLocation="http://hul.harvard.edu/ois/xml/ns/jhove/jhoveConfig
|
|
||||||
http://hul.harvard.edu/ois/xml/xsd/jhove/1.6/jhoveConfig.xsd">
|
|
||||||
<jhoveHome>/users/stephen/projects/jhove</jhoveHome>
|
|
||||||
<defaultEncoding>utf-8</defaultEncoding>
|
|
||||||
<tempDirectory>/var/tmp</tempDirectory>
|
|
||||||
<bufferSize>131072</bufferSize>
|
|
||||||
<mixVersion>1.0</mixVersion>
|
|
||||||
<sigBytes>1024</sigBytes>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.AiffModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.WaveModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.PdfModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.Jpeg2000Module</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.JpegModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.GifModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.TiffModule</class>
|
|
||||||
<param>byteoffset=true</param>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.XmlModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.HtmlModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.AsciiModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.Utf8Module</class>
|
|
||||||
</module>
|
|
||||||
</jhoveConfig>
|
|
||||||
@@ -1,51 +0,0 @@
|
|||||||
<?xml version="1.0" encoding="UTF-8"?>
|
|
||||||
<jhoveConfig version="1.0"
|
|
||||||
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
|
|
||||||
xmlns="http://hul.harvard.edu/ois/xml/ns/jhove/jhoveConfig"
|
|
||||||
xsi:schemaLocation="http://hul.harvard.edu/ois/xml/ns/jhove/jhoveConfig
|
|
||||||
http://hul.harvard.edu/ois/xml/xsd/jhove/1.6/jhoveConfig.xsd">
|
|
||||||
<jhoveHome>/users/stephen/projects/jhove</jhoveHome>
|
|
||||||
<defaultEncoding>utf-8</defaultEncoding>
|
|
||||||
<tempDirectory>/var/tmp</tempDirectory>
|
|
||||||
<bufferSize>131072</bufferSize>
|
|
||||||
<mixVersion>1.0</mixVersion>
|
|
||||||
<sigBytes>1024</sigBytes>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.AiffModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.WaveModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.PdfModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.Jpeg2000Module</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.JpegModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.GifModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.TiffModule</class>
|
|
||||||
<param>byteoffset=true</param>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.XmlModule</class>
|
|
||||||
<param>withTextMD=true</param>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.HtmlModule</class>
|
|
||||||
<param>withTextMD=true</param>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.AsciiModule</class>
|
|
||||||
<param>withTextMD=true</param>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.Utf8Module</class>
|
|
||||||
<param>withTextMD=true</param>
|
|
||||||
</module>
|
|
||||||
</jhoveConfig>
|
|
||||||
@@ -1,45 +0,0 @@
|
|||||||
<?xml version="1.0" encoding="UTF-8"?>
|
|
||||||
<jhoveConfig version="1.0"
|
|
||||||
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
|
|
||||||
xmlns="http://hul.harvard.edu/ois/xml/ns/jhove/jhoveConfig"
|
|
||||||
xsi:schemaLocation="http://hul.harvard.edu/ois/xml/ns/jhove/jhoveConfig
|
|
||||||
http://hul.harvard.edu/ois/xml/xsd/jhove/1.6/jhoveConfig.xsd">
|
|
||||||
<jhoveHome>./jhove/</jhoveHome>
|
|
||||||
<defaultEncoding>utf-8</defaultEncoding>
|
|
||||||
<tempDirectory>/var/tmp</tempDirectory>
|
|
||||||
<bufferSize>131072</bufferSize>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.AiffModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.WaveModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.PdfModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.Jpeg2000Module</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.JpegModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.GifModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.TiffModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.XmlModule</class>
|
|
||||||
<param>schema=http://www.example.com/schema;/home/schemas/exampleschema.xsd</param>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.HtmlModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.AsciiModule</class>
|
|
||||||
</module>
|
|
||||||
<module>
|
|
||||||
<class>edu.harvard.hul.ois.jhove.module.Utf8Module</class>
|
|
||||||
</module>
|
|
||||||
</jhoveConfig>
|
|
||||||
@@ -1,88 +0,0 @@
|
|||||||
#!/usr/bin/perl
|
|
||||||
|
|
||||||
########################################################################
|
|
||||||
# Jhove - JSTOR/Harvard Object Validation Environment
|
|
||||||
# Copyright 2004 by JSTOR and the President and Fellows of Harvard College
|
|
||||||
#
|
|
||||||
# A Perl script for plugging local path information into the
|
|
||||||
# various script files of JHOVE, as well as conf/jhove.conf.
|
|
||||||
#
|
|
||||||
# This is configured only for Unix (including OS X).
|
|
||||||
#
|
|
||||||
# Usage: configure.pl jhove_home_directory [java_home_directory [java_runtime_directory]]
|
|
||||||
#
|
|
||||||
# If invoked with no arguments, it will output a usage message.
|
|
||||||
#
|
|
||||||
########################################################################
|
|
||||||
use File::Copy;
|
|
||||||
|
|
||||||
sub mung {
|
|
||||||
my $f = $_[0];
|
|
||||||
my $bak = $f . "~";
|
|
||||||
#If there is no backup file, copy the file to the
|
|
||||||
#backup. Otherwise work from the backup.
|
|
||||||
if (!(-e $bak)) {
|
|
||||||
rename ($f, $bak);
|
|
||||||
}
|
|
||||||
open (INFILE, $bak);
|
|
||||||
open (OUTFILE, ">" . $f);
|
|
||||||
|
|
||||||
#Walks through each line of file, making substitutions.
|
|
||||||
#Remember that the JAVA_HOME and JAVA arguments are optional.
|
|
||||||
while (<INFILE>) {
|
|
||||||
s/^JHOVE_HOME=.*/JHOVE_HOME=$ARGV[0]/;
|
|
||||||
if ($narg >= 2) {
|
|
||||||
s/^JAVA_HOME=.*/JAVA_HOME=$ARGV[1]/;
|
|
||||||
}
|
|
||||||
if ($narg >= 3) {
|
|
||||||
s/^JAVA=.*/JAVA=$ARGV[2]/;
|
|
||||||
}
|
|
||||||
print OUTFILE;
|
|
||||||
}
|
|
||||||
close (INFILE);
|
|
||||||
close (OUTFILE);
|
|
||||||
if (-e $f) {
|
|
||||||
print ("Fixed " . $f . "\n");
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|
||||||
$narg = $#ARGV + 1;
|
|
||||||
if ($narg <= 0) {
|
|
||||||
print "Usage: configure.pl jhove_home_directory [java_home_directory [java_runtime_directory]]\n";
|
|
||||||
exit;
|
|
||||||
}
|
|
||||||
print "JHOVE_HOME will be set to " . $ARGV[0] . "\n";
|
|
||||||
if ($narg >= 2) {
|
|
||||||
print "JAVA_HOME will be set to " . $ARGV[1] . "\n";
|
|
||||||
}
|
|
||||||
if ($narg >= 3) {
|
|
||||||
print "JAVA will be set to " . $ARGV[2] . "\n";
|
|
||||||
}
|
|
||||||
mung ("jhove");
|
|
||||||
mung ("adump");
|
|
||||||
mung ("gdump");
|
|
||||||
mung ("jdump");
|
|
||||||
mung ("j2dump");
|
|
||||||
mung ("pdump");
|
|
||||||
mung ("tdump");
|
|
||||||
mung ("wdump");
|
|
||||||
|
|
||||||
#Fix up the config file. We assume that the <jhoveHome>
|
|
||||||
#element is all on one line.
|
|
||||||
if (!(-e "conf/jhove.conf~")) {
|
|
||||||
rename ("conf/jhove.conf", "conf/jhove.conf~");
|
|
||||||
}
|
|
||||||
open (INFILE, "conf/jhove.conf~");
|
|
||||||
open (OUTFILE, ">conf/jhove.conf");
|
|
||||||
while (<INFILE>) {
|
|
||||||
s!<jhoveHome>.*</jhoveHome>!<jhoveHome>$ARGV[0]</jhoveHome>!;
|
|
||||||
print OUTFILE;
|
|
||||||
}
|
|
||||||
close (INFILE);
|
|
||||||
close (OUTFILE);
|
|
||||||
if (-e "conf/jhove.conf") {
|
|
||||||
print "Fixed conf/jhove.conf\n";
|
|
||||||
}
|
|
||||||
exit;
|
|
||||||
|
|
||||||
@@ -1,37 +0,0 @@
|
|||||||
#!/bin/sh
|
|
||||||
|
|
||||||
########################################################################
|
|
||||||
# gdump - JSTOR/Harvard Object Validation Environment
|
|
||||||
# Copyright 2004-2005 by the President and Fellows of Harvard College
|
|
||||||
# JHOVE is made available under the GNU General Public License (see the
|
|
||||||
# file LICENSE for details)
|
|
||||||
#
|
|
||||||
# Driver script for the GIF dump utility
|
|
||||||
#
|
|
||||||
# Usage: gdump file
|
|
||||||
#
|
|
||||||
# where file is a GIF file
|
|
||||||
#
|
|
||||||
# Configuration constants:
|
|
||||||
|
|
||||||
JHOVE_HOME=/users/stephen/projects/jhove
|
|
||||||
|
|
||||||
JAVA_HOME=/usr/java # Java JRE directory
|
|
||||||
JAVA=$JAVA_HOME/bin/java # Java interpreter
|
|
||||||
|
|
||||||
EXTRA_JARS= # Extra .jar files to add to CLASSPATH
|
|
||||||
|
|
||||||
# NOTE: Nothing below this line should be edited
|
|
||||||
########################################################################
|
|
||||||
|
|
||||||
CP=${JHOVE_HOME}/bin/JhoveApp.jar:${EXTRA_JARS}
|
|
||||||
|
|
||||||
# Retrieve a copy of all command line arguments to pass to the application.
|
|
||||||
|
|
||||||
ARGS=""
|
|
||||||
for ARG do
|
|
||||||
ARGS="$ARGS $ARG"
|
|
||||||
done
|
|
||||||
|
|
||||||
# Set the CLASSPATH and invoke the Java loader.
|
|
||||||
${JAVA} -classpath $CP GDump $ARGS
|
|
||||||
@@ -1,37 +0,0 @@
|
|||||||
#!/bin/sh
|
|
||||||
|
|
||||||
########################################################################
|
|
||||||
# j2dump - JSTOR/Harvard Object Validation Environment
|
|
||||||
# Copyright 2004-2005 by the President and Fellows of Harvard College
|
|
||||||
# JHOVE is made available under the GNU General Public License (see the
|
|
||||||
# file LICENSE for details)
|
|
||||||
#
|
|
||||||
# Driver script for the JPEG 2000 dump utility
|
|
||||||
#
|
|
||||||
# Usage: j2dump file
|
|
||||||
#
|
|
||||||
# where file is a JPEG file
|
|
||||||
#
|
|
||||||
# Configuration constants:
|
|
||||||
|
|
||||||
JHOVE_HOME=/users/stephen/projects/jhove
|
|
||||||
|
|
||||||
JAVA_HOME=/usr/java # Java JRE directory
|
|
||||||
JAVA=$JAVA_HOME/bin/java # Java interpreter
|
|
||||||
|
|
||||||
EXTRA_JARS= # Extra .jar files to add to CLASSPATH
|
|
||||||
|
|
||||||
# NOTE: Nothing below this line should be edited
|
|
||||||
########################################################################
|
|
||||||
|
|
||||||
CP=${JHOVE_HOME}/bin/JhoveApp.jar:${EXTRA_JARS}
|
|
||||||
|
|
||||||
# Retrieve a copy of all command line arguments to pass to the application.
|
|
||||||
|
|
||||||
ARGS=""
|
|
||||||
for ARG do
|
|
||||||
ARGS="$ARGS $ARG"
|
|
||||||
done
|
|
||||||
|
|
||||||
# Set the CLASSPATH and invoke the Java loader.
|
|
||||||
${JAVA} -classpath $CP J2Dump $ARGS
|
|
||||||
@@ -1,37 +0,0 @@
|
|||||||
#!/bin/sh
|
|
||||||
|
|
||||||
########################################################################
|
|
||||||
# jdump - JSTOR/Harvard Object Validation Environment
|
|
||||||
# Copyright 2004-2005 by the President and Fellows of Harvard College
|
|
||||||
# JHOVE is made available under the GNU General Public License (see the
|
|
||||||
# file LICENSE for details)
|
|
||||||
#
|
|
||||||
# Driver script for the JPEG dump utility
|
|
||||||
#
|
|
||||||
# Usage: jdump file
|
|
||||||
#
|
|
||||||
# where file is a JPEG file
|
|
||||||
#
|
|
||||||
# Configuration constants:
|
|
||||||
|
|
||||||
JHOVE_HOME=/users/stephen/projects/jhove
|
|
||||||
|
|
||||||
JAVA_HOME=/usr/java # Java JRE directory
|
|
||||||
JAVA=$JAVA_HOME/bin/java # Java interpreter
|
|
||||||
|
|
||||||
EXTRA_JARS= # Extra .jar files to add to CLASSPATH
|
|
||||||
|
|
||||||
# NOTE: Nothing below this line should be edited
|
|
||||||
########################################################################
|
|
||||||
|
|
||||||
CP=${JHOVE_HOME}/bin/JhoveApp.jar:${EXTRA_JARS}
|
|
||||||
|
|
||||||
# Retrieve a copy of all command line arguments to pass to the application.
|
|
||||||
|
|
||||||
ARGS=""
|
|
||||||
for ARG do
|
|
||||||
ARGS="$ARGS $ARG"
|
|
||||||
done
|
|
||||||
|
|
||||||
# Set the CLASSPATH and invoke the Java loader.
|
|
||||||
${JAVA} -classpath $CP JDump $ARGS
|
|
||||||
@@ -1,58 +0,0 @@
|
|||||||
#!/bin/sh
|
|
||||||
|
|
||||||
########################################################################
|
|
||||||
# JHOVE - JSTOR/Harvard Object Validation Environment
|
|
||||||
# Copyright 2003-2005 by JSTOR and the President and Fellows of Harvard College
|
|
||||||
# JHOVE is made available under the GNU General Public License (see the
|
|
||||||
# file LICENSE for details)
|
|
||||||
#
|
|
||||||
# Usage: jhove [-c config] [-m module] [-h handler] [-e encoding] [-H handler]
|
|
||||||
# [-o output] [-x saxclass] [-t tempdir] [-b bufsize]
|
|
||||||
# [-l loglevel] [[-krs] dir-file-or-uri [...]]
|
|
||||||
#
|
|
||||||
# where -c config Configuration file pathname
|
|
||||||
# -m module Module name
|
|
||||||
# -h handler Output handler name (defaults to TEXT)
|
|
||||||
# -e encoding Character encoding of output handler (defaults to UTF-8)
|
|
||||||
# -H handler About handler name
|
|
||||||
# -o output Output file pathname (defaults to standard output)
|
|
||||||
# -x saxclass SAX parser class (defaults to J2SE 1.4 default)
|
|
||||||
# -t tempdir Temporary directory in which to create temporary files
|
|
||||||
# -b bufsize Buffer size for buffered I/O (defaults to J2SE 1.4 default)
|
|
||||||
# -k Calculate CRC32, MD5, and SHA-1 checksums
|
|
||||||
# -r Display raw data flags, not textual equivalents
|
|
||||||
# -s Format identification based on internal signatures only
|
|
||||||
# dir-file-or-uri Directory, file pathname or URI of formatted content
|
|
||||||
#
|
|
||||||
# CHANGE for JHOVE 1.8:
|
|
||||||
# You no longer have to figure out where JAVA_HOME is; that's the
|
|
||||||
# operating system's job. If the OS tells you it can't find Java,
|
|
||||||
# adjust your shell's path or revert to the old way (commented out).
|
|
||||||
# Configuration constants:
|
|
||||||
|
|
||||||
#JHOVE_HOME=/users/gary/dev/jhove
|
|
||||||
JHOVE_HOME=[fill in path to jhove directory]
|
|
||||||
|
|
||||||
JAVA_HOME=/usr/java # Java JRE directory -- change to your local java home
|
|
||||||
JAVA=$JAVA_HOME/bin/java # Java interpreter -- usually won't need change
|
|
||||||
|
|
||||||
#XTRA_JARS=/users/stephen/xercesImpl.jar
|
|
||||||
EXTRA_JARS= # Extra .jar files to add to CLASSPATH
|
|
||||||
|
|
||||||
# NOTE: Nothing below this line should be edited
|
|
||||||
########################################################################
|
|
||||||
|
|
||||||
CP=${JHOVE_HOME}/bin/JhoveApp.jar:${EXTRA_JARS}
|
|
||||||
|
|
||||||
# Retrieve a copy of all command line arguments to pass to the application.
|
|
||||||
|
|
||||||
ARGS=""
|
|
||||||
for ARG do
|
|
||||||
ARGS="$ARGS $ARG"
|
|
||||||
done
|
|
||||||
|
|
||||||
# Set the CLASSPATH and invoke the Java loader.
|
|
||||||
#{JAVA} -classpath $CP Jhove $ARGS -x org.apache.xerces.parsers.SAXParser
|
|
||||||
#${JAVA} -classpath $CP Jhove $ARGS
|
|
||||||
# New way, doesn't require you to use JAVA_HOME.
|
|
||||||
java -classpath $CP Jhove $ARGS
|
|
||||||
@@ -1,54 +0,0 @@
|
|||||||
#!/bin/sh
|
|
||||||
|
|
||||||
########################################################################
|
|
||||||
# JHOVE - JSTOR/Harvard Object Validation Environment
|
|
||||||
# Copyright 2003-2004 by JSTOR and the President and Fellows of Harvard College
|
|
||||||
# JHOVE is made available under the GNU General Public License (see the
|
|
||||||
# file LICENSE for details)
|
|
||||||
#
|
|
||||||
# Copy jhove.tmpl to jhove, and replace the value of JHOVE_HOME with
|
|
||||||
# the path to your jhove directory.
|
|
||||||
#
|
|
||||||
# Usage: jhove [-c config] [-m module [-p param]] [-h handler [-P param]]
|
|
||||||
# [-e encoding] [-H handler] [-o output] [-x saxclass]
|
|
||||||
# [-t tempdir] [-b bufsize] [[-krs] dir-file-or-uri [...]]
|
|
||||||
#
|
|
||||||
# where -c config Configuration file pathname
|
|
||||||
# -m module Module name
|
|
||||||
# -p param Module-specific parameter
|
|
||||||
# -h handler Output handler name (defaults to TEXT)
|
|
||||||
# -P param Handler-specific parameter
|
|
||||||
# -o output Output file pathname (defaults to standard output)
|
|
||||||
# -x saxclass SAX parser class (defaults to J2SE 1.4 default)
|
|
||||||
# -t tempdir Temporary directory in which to create temporary files
|
|
||||||
# -b bufsize Buffer size for buffered I/O (defaults to J2SE 1.4 default)
|
|
||||||
# -k Calculate CRC32, MD5, and SHA-1 checksums
|
|
||||||
# -r Display raw data flags, not textual equivalents
|
|
||||||
# -s Format identification based on internal signatures only
|
|
||||||
# dir-file-or-uri Directory, file pathname or URI of formatted content
|
|
||||||
#
|
|
||||||
# Configuration constants:
|
|
||||||
|
|
||||||
JHOVE_HOME=[your directory path]/jhove
|
|
||||||
|
|
||||||
JAVA_HOME=/usr/java
|
|
||||||
JAVA=/usr/bin/java
|
|
||||||
|
|
||||||
#XTRA_JARS=/users/stephen/xercesImpl.jar
|
|
||||||
EXTRA_JARS= # Extra .jar files to add to CLASSPATH
|
|
||||||
|
|
||||||
# NOTE: Nothing below this line should be edited
|
|
||||||
########################################################################
|
|
||||||
|
|
||||||
CP=${JHOVE_HOME}/bin/JhoveApp.jar:${EXTRA_JARS}
|
|
||||||
|
|
||||||
# Retrieve a copy of all command line arguments to pass to the application.
|
|
||||||
|
|
||||||
ARGS=""
|
|
||||||
for ARG do
|
|
||||||
ARGS="$ARGS $ARG"
|
|
||||||
done
|
|
||||||
|
|
||||||
# Set the CLASSPATH and invoke the Java loader.
|
|
||||||
#{JAVA} -classpath $CP Jhove $ARGS -x org.apache.xerces.parsers.SAXParser
|
|
||||||
${JAVA} -classpath $CP Jhove $ARGS
|
|
||||||
@@ -1,60 +0,0 @@
|
|||||||
@ECHO OFF
|
|
||||||
REM JHOVE - JSTOR/Harvard Object Validation Environment
|
|
||||||
REM Copyright 2003-2005 by JSTOR and the President and Fellows of Harvard College
|
|
||||||
REM JHOVE is made available under the GNU General Public License (see the
|
|
||||||
REM file LICENSE for details)
|
|
||||||
REM
|
|
||||||
REM Usage: jhove [-c config] [-m module] [-h handler] [-e encoding]
|
|
||||||
REM [-H handler] [-o output] [-x saxclass] [-t tempdir]
|
|
||||||
REM [-b bufsize] [-l loglevel] [[-krs] dir-file-or-uri [...]]
|
|
||||||
REM
|
|
||||||
REM For Windows systems, copy jhove_bat.tmpl to jhove.bat and change
|
|
||||||
REM the value of JHOVE_HOME to the path to your jhove directory.
|
|
||||||
REM
|
|
||||||
REM where -c config Configuration file pathname
|
|
||||||
REM -m module Module name
|
|
||||||
REM -h handler Output handler name (defaults to TEXT)
|
|
||||||
REM -e encoding Character encoding of output handler (defaults to UTF-8)
|
|
||||||
REM -H handler About handler name
|
|
||||||
REM -o output Output file pathname (defaults to standard output)
|
|
||||||
REM -x saxclass SAX parser class (defaults to J2SE 1.4 default)
|
|
||||||
REM -t tempdir Temporary directory in which to create temporary files
|
|
||||||
REM -b bufsize Buffer size for buffered I/O (defaults to J2SE default)
|
|
||||||
REM -l loglevel Logging level
|
|
||||||
REM -k Calculate CRC32, MD5, and SHA-1 checksums
|
|
||||||
REM -r Display raw data flags, not textual equivalents
|
|
||||||
REM -s Format identification based on internal signatures only
|
|
||||||
REM dir-file-or-uri Directory, file pathname, or URI of formatted content
|
|
||||||
REM
|
|
||||||
REM Configuration constants:
|
|
||||||
REM JHOVE_HOME Jhove installation directory
|
|
||||||
REM JAVA_HOME Java JRE directory
|
|
||||||
REM JAVA Java interpreter
|
|
||||||
REM EXTRA_JARS Extra jar files to add to CLASSPATH
|
|
||||||
|
|
||||||
REM SET JHOVE_HOME="C:\Program Files\jhove"
|
|
||||||
SET JHOVE_HOME="[your directory path]\jhove"
|
|
||||||
|
|
||||||
SET EXTRA_JARS=
|
|
||||||
|
|
||||||
REM NOTE: Nothing below this line should be edited
|
|
||||||
REM #########################################################################
|
|
||||||
|
|
||||||
|
|
||||||
SET CP=%JHOVE_HOME%\bin\JhoveApp.jar
|
|
||||||
IF "%EXTRA_JARS%"=="" GOTO FI
|
|
||||||
SET CP=%CP%:%EXTRA_JARS
|
|
||||||
:FI
|
|
||||||
|
|
||||||
REM Retrieve a copy of all command line arguments to pass to the application
|
|
||||||
|
|
||||||
SET ARGS=
|
|
||||||
:WHILE
|
|
||||||
IF "%1"=="" GOTO LOOP
|
|
||||||
SET ARGS=%ARGS% %1
|
|
||||||
SHIFT
|
|
||||||
GOTO WHILE
|
|
||||||
:LOOP
|
|
||||||
|
|
||||||
REM Set the CLASSPATH and invoke the Java loader
|
|
||||||
JAVA -classpath %CP% Jhove %ARGS%
|
|
||||||
Binary file not shown.
Binary file not shown.
@@ -1,14 +0,0 @@
|
|||||||
#!/usr/bin/perl -w
|
|
||||||
|
|
||||||
# Generate MD5 checksum
|
|
||||||
|
|
||||||
use Digest::MD5;
|
|
||||||
|
|
||||||
die "usage: md5.pl file\n" if $#ARGV < 0;
|
|
||||||
|
|
||||||
open (FILE, "<$ARGV[0]") or die "can't open file\"$ARGV[0]\"!\n";
|
|
||||||
$ctx = Digest::MD5->new->addfile (*FILE);
|
|
||||||
close (FILE);
|
|
||||||
$digest = $ctx->hexdigest;
|
|
||||||
|
|
||||||
print "$digest\n";
|
|
||||||
@@ -1,44 +0,0 @@
|
|||||||
#!/bin/sh
|
|
||||||
#DO NOT RUN THIS ON A DEVELOPEMENT DIRECTORY, ONLY ON A
|
|
||||||
#CHECKED-OUT COPY TO BE PACKAGED!
|
|
||||||
|
|
||||||
if [ "$1" = "" ]; then
|
|
||||||
echo "Usage: packagejhove.sh [version]"
|
|
||||||
echo "e.g., packagejhove.sh 1_8"
|
|
||||||
exit 1
|
|
||||||
fi
|
|
||||||
|
|
||||||
echo "This script will prepare your directory for uploading."
|
|
||||||
echo "DO NOT RUN IT unless you're a developer and know what"
|
|
||||||
echo "you are doing. "
|
|
||||||
|
|
||||||
echo "Start in the top-level directory of the JHOVE checkout."
|
|
||||||
echo "Run ant and ant javadoc and do any necessary testing and"
|
|
||||||
echo "committing before running this script."
|
|
||||||
echo
|
|
||||||
echo
|
|
||||||
|
|
||||||
echo "To continue, enter the secret phrase."
|
|
||||||
read OATH
|
|
||||||
if [ "$OATH" != "I solemnly swear that I am up to no good" ]; then
|
|
||||||
exit 1
|
|
||||||
fi
|
|
||||||
|
|
||||||
cd ..
|
|
||||||
cat >>CVSS << EOF
|
|
||||||
CVS
|
|
||||||
.cvsignore
|
|
||||||
EOF
|
|
||||||
tar cvfX jhove-$1.tar CVSS jhove
|
|
||||||
gzip jhove-$1.tar
|
|
||||||
cp -r jhove jhove-zip
|
|
||||||
cd jhove-zip
|
|
||||||
find . \( -name CVS -o -name .cvsignore \) -exec rm -r {} \;
|
|
||||||
cd ..
|
|
||||||
mv jhove jhove-ok
|
|
||||||
mv jhove-zip jhove
|
|
||||||
zip -r jhove-$1.zip jhove
|
|
||||||
rm -r jhove
|
|
||||||
mv jhove-ok jhove
|
|
||||||
jhove/md5.pl jhove-$1.tar.gz >jhove-$1.tar.gz.md5
|
|
||||||
jhove/md5.pl jhove-$1.zip >jhove-$1.zip.md5
|
|
||||||
@@ -1,37 +0,0 @@
|
|||||||
#!/bin/sh
|
|
||||||
|
|
||||||
########################################################################
|
|
||||||
# pdump - JSTOR/Harvard Object Validation Environment
|
|
||||||
# Copyright 2003-2005 by JSTOR and the President and Fellows of Harvard College
|
|
||||||
# JHOVE is made available under the GNU General Public License (see the
|
|
||||||
# file LICENSE for details)
|
|
||||||
#
|
|
||||||
# Driver script for the PDF dump utility
|
|
||||||
#
|
|
||||||
# Usage: pdump file
|
|
||||||
#
|
|
||||||
# where file is a PDF file
|
|
||||||
#
|
|
||||||
# Configuration constants:
|
|
||||||
|
|
||||||
JHOVE_HOME=/users/stephen/projects/jhove
|
|
||||||
|
|
||||||
JAVA_HOME=/usr/java # Java JRE directory
|
|
||||||
JAVA=$JAVA_HOME/bin/java # Java interpreter
|
|
||||||
|
|
||||||
EXTRA_JARS= # Extra .jar files to add to CLASSPATH
|
|
||||||
|
|
||||||
# NOTE: Nothing below this line should be edited
|
|
||||||
########################################################################
|
|
||||||
|
|
||||||
CP=${JHOVE_HOME}/bin/JhoveApp.jar:${EXTRA_JARS}
|
|
||||||
|
|
||||||
# Retrieve a copy of all command line arguments to pass to the application.
|
|
||||||
|
|
||||||
ARGS=""
|
|
||||||
for ARG do
|
|
||||||
ARGS="$ARGS $ARG"
|
|
||||||
done
|
|
||||||
|
|
||||||
# Set the CLASSPATH and invoke the Java loader.
|
|
||||||
${JAVA} -classpath $CP PDump $ARGS
|
|
||||||
@@ -1,24 +0,0 @@
|
|||||||
#!/bin/sh
|
|
||||||
|
|
||||||
########################################################################
|
|
||||||
# userhome - JSTOR/Harvard Object Validation Environment
|
|
||||||
# Copyright 2004-2006 by the President and Fellows of Harvard College
|
|
||||||
# JHOVE is made available under the GNU General Public License (see the
|
|
||||||
# file LICENSE for details)
|
|
||||||
#
|
|
||||||
# Driver script to display the default Java user.home property
|
|
||||||
#
|
|
||||||
# Usage: userhome
|
|
||||||
#
|
|
||||||
# Configuration constants:
|
|
||||||
|
|
||||||
JHOVE_HOME=/users/stephen/projects/jhove
|
|
||||||
|
|
||||||
JAVA_HOME=/usr/java # Java JRE directory
|
|
||||||
JAVA=$JAVA_HOME/bin/java # Java interpreter
|
|
||||||
|
|
||||||
# NOTE: Nothing below this line should be edited
|
|
||||||
########################################################################
|
|
||||||
|
|
||||||
# Set the CLASSPATH and invoke the Java loader.
|
|
||||||
${JAVA} -classpath ${JHOVE_HOME}/classes UserHome
|
|
||||||
@@ -1,331 +0,0 @@
|
|||||||
#!/usr/bin/env python2
|
|
||||||
# -*- coding: utf-8 -*-
|
|
||||||
#
|
|
||||||
# © 2013-15: jbarlow83 from Github (https://github.com/jbarlow83)
|
|
||||||
#
|
|
||||||
#
|
|
||||||
# Use Leptonica to detect find and remove page skew. Leptonica uses the method
|
|
||||||
# of differential square sums, which its author claim is faster and more robust
|
|
||||||
# than the Hough transform used by ImageMagick.
|
|
||||||
|
|
||||||
from __future__ import print_function, absolute_import, division
|
|
||||||
import argparse
|
|
||||||
import ctypes as C
|
|
||||||
import sys
|
|
||||||
import os
|
|
||||||
import logging
|
|
||||||
from tempfile import TemporaryFile
|
|
||||||
|
|
||||||
logger = logging.getLogger(__name__)
|
|
||||||
|
|
||||||
|
|
||||||
def stderr(*objs):
|
|
||||||
"""Python 2/3 compatible print to stderr.
|
|
||||||
"""
|
|
||||||
print("leptonica.py:", *objs, file=sys.stderr)
|
|
||||||
|
|
||||||
|
|
||||||
from ctypes.util import find_library
|
|
||||||
lept_lib = find_library('lept')
|
|
||||||
if not lept_lib:
|
|
||||||
stderr("Could not find the Leptonica library")
|
|
||||||
sys.exit(3)
|
|
||||||
try:
|
|
||||||
lept = C.cdll.LoadLibrary(lept_lib)
|
|
||||||
except Exception:
|
|
||||||
stderr("Could not load the Leptonica library from %s", lept_lib)
|
|
||||||
sys.exit(3)
|
|
||||||
|
|
||||||
|
|
||||||
class _PIXCOLORMAP(C.Structure):
|
|
||||||
"""struct PixColormap from Leptonica src/pix.h
|
|
||||||
"""
|
|
||||||
|
|
||||||
_fields_ = [
|
|
||||||
("array", C.c_void_p),
|
|
||||||
("depth", C.c_int32),
|
|
||||||
("nalloc", C.c_int32),
|
|
||||||
("n", C.c_int32)
|
|
||||||
]
|
|
||||||
|
|
||||||
|
|
||||||
class _PIX(C.Structure):
|
|
||||||
"""struct Pix from Leptonica src/pix.h
|
|
||||||
"""
|
|
||||||
|
|
||||||
_fields_ = [
|
|
||||||
("w", C.c_uint32),
|
|
||||||
("h", C.c_uint32),
|
|
||||||
("d", C.c_uint32),
|
|
||||||
("wpl", C.c_uint32),
|
|
||||||
("refcount", C.c_uint32),
|
|
||||||
("xres", C.c_int32),
|
|
||||||
("yres", C.c_int32),
|
|
||||||
("informat", C.c_int32),
|
|
||||||
("text", C.POINTER(C.c_char)),
|
|
||||||
("colormap", C.POINTER(_PIXCOLORMAP)),
|
|
||||||
("data", C.POINTER(C.c_uint32))
|
|
||||||
]
|
|
||||||
|
|
||||||
|
|
||||||
PIX = C.POINTER(_PIX)
|
|
||||||
|
|
||||||
lept.pixRead.argtypes = [C.c_char_p]
|
|
||||||
lept.pixRead.restype = PIX
|
|
||||||
lept.pixScale.argtypes = [PIX, C.c_float, C.c_float]
|
|
||||||
lept.pixScale.restype = PIX
|
|
||||||
lept.pixDeskew.argtypes = [PIX, C.c_int32]
|
|
||||||
lept.pixDeskew.restype = PIX
|
|
||||||
lept.pixFindSkew.argtypes = [PIX, C.POINTER(C.c_float), C.POINTER(C.c_float)]
|
|
||||||
lept.pixFindSkew.restype = C.c_int32
|
|
||||||
lept.pixWriteImpliedFormat.argtypes = [C.c_char_p, PIX, C.c_int32, C.c_int32]
|
|
||||||
lept.pixWriteImpliedFormat.restype = C.c_int32
|
|
||||||
lept.pixDestroy.argtypes = [C.POINTER(PIX)]
|
|
||||||
lept.pixDestroy.restype = None
|
|
||||||
lept.getLeptonicaVersion.argtypes = []
|
|
||||||
lept.getLeptonicaVersion.restype = C.c_char_p
|
|
||||||
|
|
||||||
|
|
||||||
class LeptonicaErrorTrap(object):
|
|
||||||
"""Context manager to trap errors reported by Leptonica.
|
|
||||||
|
|
||||||
Leptonica's error return codes are unreliable to the point of being
|
|
||||||
almost useless. It does, however, write errors to stderr provided that is
|
|
||||||
not disabled at its compile time. Fortunately this is done using error
|
|
||||||
macros so it is very self-consistent.
|
|
||||||
|
|
||||||
This context manager redirects stderr to a temporary file which is then
|
|
||||||
read and parsed for error messages. As a side benefit, debug messages
|
|
||||||
from Leptonica are also suppressed.
|
|
||||||
|
|
||||||
"""
|
|
||||||
def __enter__(self):
|
|
||||||
self.tmpfile = TemporaryFile()
|
|
||||||
|
|
||||||
# Save the old stderr, and redirect stderr to temporary file
|
|
||||||
self.old_stderr_fileno = os.dup(sys.stderr.fileno())
|
|
||||||
os.dup2(self.tmpfile.fileno(), sys.stderr.fileno())
|
|
||||||
return
|
|
||||||
|
|
||||||
def __exit__(self, exc_type, exc_value, traceback):
|
|
||||||
# Restore old stderr
|
|
||||||
os.dup2(self.old_stderr_fileno, sys.stderr.fileno())
|
|
||||||
|
|
||||||
# Get data from tmpfile (in with block to ensure it is closed)
|
|
||||||
with self.tmpfile as tmpfile:
|
|
||||||
tmpfile.seek(0) # Cursor will be at end, so move back to beginning
|
|
||||||
leptonica_output = tmpfile.read().decode(errors='replace')
|
|
||||||
|
|
||||||
# If there are Python errors, let them bubble up
|
|
||||||
if exc_type:
|
|
||||||
logger.warning(leptonica_output)
|
|
||||||
return False
|
|
||||||
|
|
||||||
# If there are Leptonica errors, wrap them in Python excpetions
|
|
||||||
if 'Error' in leptonica_output:
|
|
||||||
if 'image file not found' in leptonica_output:
|
|
||||||
raise FileNotFoundError()
|
|
||||||
if 'pixWrite: stream not opened' in leptonica_output:
|
|
||||||
raise LeptonicaIOError()
|
|
||||||
raise LeptonicaError(leptonica_output)
|
|
||||||
|
|
||||||
return False
|
|
||||||
|
|
||||||
|
|
||||||
class LeptonicaError(Exception):
|
|
||||||
pass
|
|
||||||
|
|
||||||
|
|
||||||
class LeptonicaIOError(LeptonicaError):
|
|
||||||
pass
|
|
||||||
|
|
||||||
|
|
||||||
def pixRead(filename):
|
|
||||||
"""Load an image file into a PIX object.
|
|
||||||
|
|
||||||
Leptonica can load TIFF, PNM (PBM, PGM, PPM), PNG, and JPEG. If loading
|
|
||||||
fails then the object will wrap a C null pointer.
|
|
||||||
|
|
||||||
"""
|
|
||||||
with LeptonicaErrorTrap():
|
|
||||||
return lept.pixRead(filename.encode(sys.getfilesystemencoding()))
|
|
||||||
|
|
||||||
|
|
||||||
def pixScale(pix, scalex, scaley):
|
|
||||||
"""Returns the pix object rescaled according to the proportions given."""
|
|
||||||
with LeptonicaErrorTrap():
|
|
||||||
return lept.pixScale(pix, scalex, scaley)
|
|
||||||
|
|
||||||
|
|
||||||
def pixDeskew(pix, reduction_factor=0):
|
|
||||||
"""Returns the deskewed pix object.
|
|
||||||
|
|
||||||
A clone of the original is returned when the algorithm cannot find a skew
|
|
||||||
angle with sufficient confidence.
|
|
||||||
|
|
||||||
reduction_factor -- amount to downsample (0 for default) when searching
|
|
||||||
for skew angle
|
|
||||||
|
|
||||||
"""
|
|
||||||
with LeptonicaErrorTrap():
|
|
||||||
return lept.pixDeskew(pix, reduction_factor)
|
|
||||||
|
|
||||||
|
|
||||||
def pixFindSkew(pix):
|
|
||||||
"""Returns a tuple (deskew angle in degrees, confidence value).
|
|
||||||
|
|
||||||
Returns (None, None) if no angle is available.
|
|
||||||
|
|
||||||
"""
|
|
||||||
with LeptonicaErrorTrap():
|
|
||||||
angle = C.c_float(0.0)
|
|
||||||
confidence = C.c_float(0.0)
|
|
||||||
result = lept.pixFindSkew(pix, C.byref(angle), C.byref(confidence))
|
|
||||||
if result == 0:
|
|
||||||
return (angle.value, confidence.value)
|
|
||||||
else:
|
|
||||||
return (None, None)
|
|
||||||
|
|
||||||
|
|
||||||
def pixWriteImpliedFormat(filename, pix, jpeg_quality=0, jpeg_progressive=0):
|
|
||||||
"""Write pix to the filename, with the extension indicating format.
|
|
||||||
|
|
||||||
jpeg_quality -- quality (iff JPEG; 1 - 100, 0 for default)
|
|
||||||
jpeg_progressive -- (iff JPEG; 0 for baseline seq., 1 for progressive)
|
|
||||||
|
|
||||||
"""
|
|
||||||
fileroot, extension = os.path.splitext(filename)
|
|
||||||
fix_pnm = False
|
|
||||||
if extension.lower() in ('.pbm', '.pgm', '.ppm'):
|
|
||||||
# Leptonica does not process handle these extensions correctly, but
|
|
||||||
# does handle .pnm correctly. Add another .pnm suffix.
|
|
||||||
filename += '.pnm'
|
|
||||||
fix_pnm = True
|
|
||||||
|
|
||||||
with LeptonicaErrorTrap():
|
|
||||||
lept.pixWriteImpliedFormat(
|
|
||||||
filename.encode(sys.getfilesystemencoding()),
|
|
||||||
pix, jpeg_quality, jpeg_progressive)
|
|
||||||
|
|
||||||
if fix_pnm:
|
|
||||||
from shutil import move
|
|
||||||
move(filename, filename[:-4]) # Remove .pnm suffix
|
|
||||||
|
|
||||||
|
|
||||||
def pixDestroy(pix):
|
|
||||||
"""Destroy the pix object.
|
|
||||||
|
|
||||||
Function signature is pixDestroy(struct Pix **), hence C.byref() to pass
|
|
||||||
the address of the pointer.
|
|
||||||
|
|
||||||
"""
|
|
||||||
with LeptonicaErrorTrap():
|
|
||||||
lept.pixDestroy(C.byref(pix))
|
|
||||||
|
|
||||||
|
|
||||||
def getLeptonicaVersion():
|
|
||||||
"""Get Leptonica version string.
|
|
||||||
|
|
||||||
Caveat: Leptonica expects the caller to free this memory. We don't,
|
|
||||||
since that would involve binding to libc to access libc.free(),
|
|
||||||
a pointless effort to reclaim 100 bytes of memory.
|
|
||||||
|
|
||||||
"""
|
|
||||||
return lept.getLeptonicaVersion().decode()
|
|
||||||
|
|
||||||
|
|
||||||
def deskew(infile, outfile, dpi):
|
|
||||||
try:
|
|
||||||
pix_source = pixRead(infile)
|
|
||||||
except LeptonicaIOError:
|
|
||||||
raise LeptonicaIOError("Failed to open file: %s" % infile)
|
|
||||||
|
|
||||||
if dpi < 150:
|
|
||||||
reduction_factor = 1 # Don't downsample too much if DPI is already low
|
|
||||||
else:
|
|
||||||
reduction_factor = 0 # Use default
|
|
||||||
pix_deskewed = pixDeskew(pix_source, reduction_factor)
|
|
||||||
|
|
||||||
try:
|
|
||||||
pixWriteImpliedFormat(outfile, pix_deskewed)
|
|
||||||
except LeptonicaIOError:
|
|
||||||
raise LeptonicaIOError("Failed to open destination file: %s" % outfile)
|
|
||||||
pixDestroy(pix_source)
|
|
||||||
pixDestroy(pix_deskewed)
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == '__main__':
|
|
||||||
parser = argparse.ArgumentParser(
|
|
||||||
description="Python wrapper to access Leptonica")
|
|
||||||
|
|
||||||
subparsers = parser.add_subparsers(title='commands',
|
|
||||||
description='supported operations')
|
|
||||||
|
|
||||||
parser_deskew = subparsers.add_parser('deskew')
|
|
||||||
parser_deskew.add_argument('-r', '--dpi', dest='dpi', action='store',
|
|
||||||
type=int, default=300, help='input resolution')
|
|
||||||
parser_deskew.add_argument('infile', help='image to deskew')
|
|
||||||
parser_deskew.add_argument('outfile', help='deskewed output image')
|
|
||||||
parser_deskew.set_defaults(func=deskew)
|
|
||||||
|
|
||||||
args = parser.parse_args()
|
|
||||||
|
|
||||||
if getLeptonicaVersion() != u'leptonica-1.69':
|
|
||||||
print("Unexpected leptonica version: %s" % getLeptonicaVersion())
|
|
||||||
|
|
||||||
args.func(args)
|
|
||||||
|
|
||||||
|
|
||||||
def _test_output(mode, extension, im_format):
|
|
||||||
from PIL import Image
|
|
||||||
from tempfile import NamedTemporaryFile
|
|
||||||
|
|
||||||
with NamedTemporaryFile(prefix='test-lept-pnm', suffix=extension, delete=True) as tmpfile:
|
|
||||||
im = Image.new(mode=mode, size=(100, 100))
|
|
||||||
im.save(tmpfile)
|
|
||||||
|
|
||||||
pix = pixRead(tmpfile.name)
|
|
||||||
pixWriteImpliedFormat(tmpfile.name, pix)
|
|
||||||
pixDestroy(pix)
|
|
||||||
|
|
||||||
im_roundtrip = Image.open(tmpfile.name)
|
|
||||||
assert im_roundtrip.mode == im.mode, "leptonica mode differs"
|
|
||||||
assert im_roundtrip.format == im_format, \
|
|
||||||
"{0}: leptonica produced a {1}".format(
|
|
||||||
extension,
|
|
||||||
im_roundtrip.format)
|
|
||||||
|
|
||||||
|
|
||||||
def test_pnm_output():
|
|
||||||
params = [['1', '.pbm', 'PPM'], ['L', '.pgm', 'PPM'],
|
|
||||||
['RGB', '.ppm', 'PPM']]
|
|
||||||
for param in params:
|
|
||||||
_test_output(*param)
|
|
||||||
|
|
||||||
|
|
||||||
def test_skew_angle():
|
|
||||||
from PIL import Image, ImageDraw
|
|
||||||
from tempfile import NamedTemporaryFile
|
|
||||||
|
|
||||||
im = Image.new(mode='1', size=(1000, 1000), color=1)
|
|
||||||
|
|
||||||
draw = ImageDraw.Draw(im)
|
|
||||||
for n in range(20):
|
|
||||||
draw.line([(50, 25 + 50*n), (950, 25 + 50*n)], width=1)
|
|
||||||
del draw
|
|
||||||
|
|
||||||
test_angles = [0.1 * ang for ang in range(1, 10)] + \
|
|
||||||
[float(ang) for ang in range(1, 7)]
|
|
||||||
test_angles += [-ang for ang in test_angles]
|
|
||||||
test_angles = sorted(test_angles)
|
|
||||||
|
|
||||||
for rotate_angle in test_angles:
|
|
||||||
rotated_im = im.rotate(rotate_angle)
|
|
||||||
with NamedTemporaryFile(prefix='lept-skew', suffix='.png', delete=True) as tmpfile:
|
|
||||||
rotated_im.save(tmpfile)
|
|
||||||
pix = pixRead(tmpfile.name)
|
|
||||||
angle, confidence = pixFindSkew(pix)
|
|
||||||
pixDestroy(pix)
|
|
||||||
print('{0} {1} {2}'.format(rotate_angle, angle, confidence), file=sys.stderr)
|
|
||||||
|
|
||||||
|
|
||||||
@@ -1,842 +0,0 @@
|
|||||||
#!/usr/bin/env python3
|
|
||||||
# © 2015 James R. Barlow: github.com/jbarlow83
|
|
||||||
|
|
||||||
from contextlib import suppress
|
|
||||||
from tempfile import NamedTemporaryFile, mkdtemp
|
|
||||||
import sys
|
|
||||||
import os
|
|
||||||
import fileinput
|
|
||||||
import re
|
|
||||||
import shutil
|
|
||||||
import warnings
|
|
||||||
import multiprocessing
|
|
||||||
import atexit
|
|
||||||
import textwrap
|
|
||||||
|
|
||||||
import PyPDF2 as pypdf
|
|
||||||
from PIL import Image
|
|
||||||
|
|
||||||
from subprocess import Popen, check_call, PIPE, CalledProcessError, \
|
|
||||||
TimeoutExpired
|
|
||||||
try:
|
|
||||||
from subprocess import DEVNULL
|
|
||||||
except ImportError:
|
|
||||||
DEVNULL = open(os.devnull, 'wb')
|
|
||||||
|
|
||||||
|
|
||||||
from ruffus import transform, suffix, merge, active_if, regex, jobs_limit, \
|
|
||||||
formatter, follows, split, collate, check_if_uptodate
|
|
||||||
import ruffus.cmdline as cmdline
|
|
||||||
|
|
||||||
from .hocrtransform import HocrTransform
|
|
||||||
from .pageinfo import pdf_get_all_pageinfo
|
|
||||||
from .pdfa import generate_pdfa_def
|
|
||||||
from . import ghostscript
|
|
||||||
from . import tesseract
|
|
||||||
|
|
||||||
|
|
||||||
warnings.simplefilter('ignore', pypdf.utils.PdfReadWarning)
|
|
||||||
|
|
||||||
|
|
||||||
BASEDIR = os.path.dirname(os.path.realpath(__file__))
|
|
||||||
JHOVE_PATH = os.path.realpath(os.path.join(BASEDIR, 'jhove'))
|
|
||||||
JHOVE_JAR = os.path.join(JHOVE_PATH, 'bin', 'JhoveApp.jar')
|
|
||||||
JHOVE_CFG = os.path.join(JHOVE_PATH, 'conf', 'jhove.conf')
|
|
||||||
|
|
||||||
EXIT_BAD_ARGS = 1
|
|
||||||
EXIT_BAD_INPUT_FILE = 2
|
|
||||||
EXIT_MISSING_DEPENDENCY = 3
|
|
||||||
EXIT_INVALID_OUTPUT_PDFA = 4
|
|
||||||
EXIT_FILE_ACCESS_ERROR = 5
|
|
||||||
EXIT_ALREADY_DONE_OCR = 6
|
|
||||||
EXIT_OTHER_ERROR = 15
|
|
||||||
|
|
||||||
# -------------
|
|
||||||
# External dependencies
|
|
||||||
|
|
||||||
MINIMUM_TESS_VERSION = '3.02.02'
|
|
||||||
|
|
||||||
|
|
||||||
def complain(message):
|
|
||||||
print(textwrap.wrap(message), file=sys.stderr)
|
|
||||||
|
|
||||||
|
|
||||||
if tesseract.version() < MINIMUM_TESS_VERSION:
|
|
||||||
complain(
|
|
||||||
"Please install tesseract {0} or newer "
|
|
||||||
"(currently installed version is {1})".format(
|
|
||||||
MINIMUM_TESS_VERSION, tesseract.version()))
|
|
||||||
sys.exit(EXIT_MISSING_DEPENDENCY)
|
|
||||||
|
|
||||||
|
|
||||||
# -------------
|
|
||||||
# Parser
|
|
||||||
|
|
||||||
parser = cmdline.get_argparse(
|
|
||||||
prog="ocrmypdf",
|
|
||||||
description="Generate searchable PDF file from an image-only PDF file.",
|
|
||||||
version='3.0rc2',
|
|
||||||
fromfile_prefix_chars='@',
|
|
||||||
ignored_args=[
|
|
||||||
'touch_files_only', 'recreate_database', 'checksum_file_name',
|
|
||||||
'key_legend_in_graph', 'draw_graph_horizontally', 'flowchart_format',
|
|
||||||
'forced_tasks', 'target_tasks'])
|
|
||||||
|
|
||||||
parser.add_argument(
|
|
||||||
'input_file',
|
|
||||||
help="PDF file containing the images to be OCRed")
|
|
||||||
parser.add_argument(
|
|
||||||
'output_file',
|
|
||||||
help="output searchable PDF file")
|
|
||||||
parser.add_argument(
|
|
||||||
'-l', '--language', action='append',
|
|
||||||
help="language of the file to be OCRed")
|
|
||||||
|
|
||||||
metadata = parser.add_argument_group(
|
|
||||||
"Metadata options",
|
|
||||||
"Set output PDF/A metadata (default: use input document's title)")
|
|
||||||
metadata.add_argument(
|
|
||||||
'--title', type=str,
|
|
||||||
help="set document title (place multiple words in quotes)")
|
|
||||||
metadata.add_argument(
|
|
||||||
'--author', type=str,
|
|
||||||
help="set document author")
|
|
||||||
metadata.add_argument(
|
|
||||||
'--subject', type=str,
|
|
||||||
help="set document")
|
|
||||||
metadata.add_argument(
|
|
||||||
'--keywords', type=str,
|
|
||||||
help="set document keywords")
|
|
||||||
|
|
||||||
|
|
||||||
preprocessing = parser.add_argument_group(
|
|
||||||
"Preprocessing options",
|
|
||||||
"Improve OCR quality and final image")
|
|
||||||
preprocessing.add_argument(
|
|
||||||
'-d', '--deskew', action='store_true',
|
|
||||||
help="deskew each page before performing OCR")
|
|
||||||
preprocessing.add_argument(
|
|
||||||
'-c', '--clean', action='store_true',
|
|
||||||
help="clean pages with unpaper before performing OCR")
|
|
||||||
preprocessing.add_argument(
|
|
||||||
'-i', '--clean-final', action='store_true',
|
|
||||||
help="incorporate the cleaned image in the final PDF file")
|
|
||||||
preprocessing.add_argument(
|
|
||||||
'--oversample', metavar='DPI', type=int, default=0,
|
|
||||||
help="oversample images to improve OCR results slightly")
|
|
||||||
|
|
||||||
parser.add_argument(
|
|
||||||
'-f', '--force-ocr', action='store_true',
|
|
||||||
help="force image into OCR, even if the page already contains text")
|
|
||||||
parser.add_argument(
|
|
||||||
'-s', '--skip-text', action='store_true',
|
|
||||||
help="skip OCR on any pages that already contain text")
|
|
||||||
parser.add_argument(
|
|
||||||
'--skip-big', type=float, metavar='MPixels',
|
|
||||||
help="skip OCR on pages larger than the specified amount of megapixels")
|
|
||||||
# parser.add_argument(
|
|
||||||
# '--exact-image', action='store_true',
|
|
||||||
# help="Use original page from PDF without re-rendering")
|
|
||||||
|
|
||||||
advanced = parser.add_argument_group(
|
|
||||||
"Advanced",
|
|
||||||
"Advanced options for power users")
|
|
||||||
advanced.add_argument(
|
|
||||||
'--tesseract-config', default=[], type=list, action='append',
|
|
||||||
help="additional Tesseract configuration files")
|
|
||||||
advanced.add_argument(
|
|
||||||
'--pdf-renderer', choices=['tesseract', 'hocr'], default='hocr',
|
|
||||||
help='choose OCR PDF renderer')
|
|
||||||
advanced.add_argument(
|
|
||||||
'--tesseract-timeout', default=180.0, type=float,
|
|
||||||
help='give up on OCR after timeout')
|
|
||||||
|
|
||||||
debugging = parser.add_argument_group(
|
|
||||||
"Debugging",
|
|
||||||
"Arguments to help with troubleshooting and debugging")
|
|
||||||
debugging.add_argument(
|
|
||||||
'-k', '--keep-temporary-files', action='store_true',
|
|
||||||
help="keep temporary files (helpful for debugging)")
|
|
||||||
debugging.add_argument(
|
|
||||||
'-g', '--debug-rendering', action='store_true',
|
|
||||||
help="render each page twice with debug information on second page")
|
|
||||||
|
|
||||||
options = parser.parse_args()
|
|
||||||
|
|
||||||
|
|
||||||
# ----------
|
|
||||||
# Languages
|
|
||||||
|
|
||||||
if not options.language:
|
|
||||||
options.language = ['eng'] # Enforce English hegemony
|
|
||||||
|
|
||||||
# Support v2.x "eng+deu" language syntax
|
|
||||||
if '+' in options.language[0]:
|
|
||||||
options.language = options.language[0].split('+')
|
|
||||||
|
|
||||||
if not set(options.language).issubset(tesseract.languages()):
|
|
||||||
complain(
|
|
||||||
"The installed version of tesseract does not have language "
|
|
||||||
"data for the following requested languages: ")
|
|
||||||
for lang in (set(options.language) - tesseract.languages()):
|
|
||||||
complain(lang, file=sys.stderr)
|
|
||||||
sys.exit(EXIT_BAD_ARGS)
|
|
||||||
|
|
||||||
|
|
||||||
# ----------
|
|
||||||
# Arguments
|
|
||||||
|
|
||||||
|
|
||||||
if any((options.deskew, options.clean, options.clean_final)):
|
|
||||||
try:
|
|
||||||
from . import unpaper
|
|
||||||
except ImportError:
|
|
||||||
complain(
|
|
||||||
"Install the 'unpaper' program to use --deskew or --clean.")
|
|
||||||
sys.exit(EXIT_BAD_ARGS)
|
|
||||||
else:
|
|
||||||
unpaper = None
|
|
||||||
|
|
||||||
if options.debug_rendering and options.pdf_renderer == 'tesseract':
|
|
||||||
complain(
|
|
||||||
"Ignoring --debug-rendering because it is not supported with"
|
|
||||||
"--pdf-renderer=tesseract.")
|
|
||||||
|
|
||||||
if options.force_ocr and options.skip_text:
|
|
||||||
complain(
|
|
||||||
"Error: --force-ocr and --skip-text are mutually incompatible.")
|
|
||||||
sys.exit(EXIT_BAD_ARGS)
|
|
||||||
|
|
||||||
if options.clean and not options.clean_final \
|
|
||||||
and options.pdf_renderer == 'tesseract':
|
|
||||||
complain(
|
|
||||||
"Tesseract PDF renderer cannot render --clean pages without "
|
|
||||||
"also performing --clean-final, so --clean-final is assumed.")
|
|
||||||
|
|
||||||
|
|
||||||
# ----------
|
|
||||||
# Logging
|
|
||||||
|
|
||||||
|
|
||||||
_logger, _logger_mutex = cmdline.setup_logging(__name__, options.log_file,
|
|
||||||
options.verbose)
|
|
||||||
|
|
||||||
|
|
||||||
class WrappedLogger:
|
|
||||||
|
|
||||||
def __init__(self, my_logger, my_mutex):
|
|
||||||
self.logger = my_logger
|
|
||||||
self.mutex = my_mutex
|
|
||||||
|
|
||||||
def log(self, *args, **kwargs):
|
|
||||||
with self.mutex:
|
|
||||||
self.logger.log(*args, **kwargs)
|
|
||||||
|
|
||||||
def debug(self, *args, **kwargs):
|
|
||||||
with self.mutex:
|
|
||||||
self.logger.debug(*args, **kwargs)
|
|
||||||
|
|
||||||
def info(self, *args, **kwargs):
|
|
||||||
with self.mutex:
|
|
||||||
self.logger.info(*args, **kwargs)
|
|
||||||
|
|
||||||
def warning(self, *args, **kwargs):
|
|
||||||
with self.mutex:
|
|
||||||
self.logger.warning(*args, **kwargs)
|
|
||||||
|
|
||||||
def error(self, *args, **kwargs):
|
|
||||||
with self.mutex:
|
|
||||||
self.logger.error(*args, **kwargs)
|
|
||||||
|
|
||||||
def critical(self, *args, **kwargs):
|
|
||||||
with self.mutex:
|
|
||||||
self.logger.critical(*args, **kwargs)
|
|
||||||
|
|
||||||
_log = WrappedLogger(_logger, _logger_mutex)
|
|
||||||
|
|
||||||
|
|
||||||
def re_symlink(input_file, soft_link_name, log=_log):
|
|
||||||
"""
|
|
||||||
Helper function: relinks soft symbolic link if necessary
|
|
||||||
"""
|
|
||||||
# Guard against soft linking to oneself
|
|
||||||
if input_file == soft_link_name:
|
|
||||||
log.debug("Warning: No symbolic link made. You are using " +
|
|
||||||
"the original data directory as the working directory.")
|
|
||||||
return
|
|
||||||
|
|
||||||
# Soft link already exists: delete for relink?
|
|
||||||
if os.path.lexists(soft_link_name):
|
|
||||||
# do not delete or overwrite real (non-soft link) file
|
|
||||||
if not os.path.islink(soft_link_name):
|
|
||||||
raise Exception("%s exists and is not a link" % soft_link_name)
|
|
||||||
try:
|
|
||||||
os.unlink(soft_link_name)
|
|
||||||
except:
|
|
||||||
log.debug("Can't unlink %s" % (soft_link_name))
|
|
||||||
|
|
||||||
if not os.path.exists(input_file):
|
|
||||||
raise Exception("trying to create a broken symlink to %s" % input_file)
|
|
||||||
|
|
||||||
log.debug("os.symlink(%s, %s)" % (input_file, soft_link_name))
|
|
||||||
|
|
||||||
# Create symbolic link using absolute path
|
|
||||||
os.symlink(
|
|
||||||
os.path.abspath(input_file),
|
|
||||||
soft_link_name
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
# -------------
|
|
||||||
# The Pipeline
|
|
||||||
|
|
||||||
manager = multiprocessing.Manager()
|
|
||||||
_pdfinfo = manager.list()
|
|
||||||
_pdfinfo_lock = manager.Lock()
|
|
||||||
|
|
||||||
work_folder = mkdtemp(prefix="com.github.ocrmypdf.")
|
|
||||||
|
|
||||||
|
|
||||||
@atexit.register
|
|
||||||
def cleanup_working_files(*args):
|
|
||||||
if options.keep_temporary_files:
|
|
||||||
print("Temporary working files saved at:")
|
|
||||||
print(work_folder)
|
|
||||||
else:
|
|
||||||
with suppress(FileNotFoundError):
|
|
||||||
shutil.rmtree(work_folder)
|
|
||||||
|
|
||||||
|
|
||||||
@transform(
|
|
||||||
input=options.input_file,
|
|
||||||
filter=suffix('.pdf'),
|
|
||||||
output='.repaired.pdf',
|
|
||||||
output_dir=work_folder,
|
|
||||||
extras=[_log, _pdfinfo, _pdfinfo_lock])
|
|
||||||
def repair_pdf(
|
|
||||||
input_file,
|
|
||||||
output_file,
|
|
||||||
log,
|
|
||||||
pdfinfo,
|
|
||||||
pdfinfo_lock):
|
|
||||||
args_mutool = [
|
|
||||||
'mutool', 'clean',
|
|
||||||
input_file, output_file
|
|
||||||
]
|
|
||||||
check_call(args_mutool)
|
|
||||||
|
|
||||||
with pdfinfo_lock:
|
|
||||||
pdfinfo.extend(pdf_get_all_pageinfo(output_file))
|
|
||||||
log.info(pdfinfo)
|
|
||||||
|
|
||||||
|
|
||||||
def get_pageinfo(input_file, pdfinfo, pdfinfo_lock):
|
|
||||||
pageno = int(os.path.basename(input_file)[0:6]) - 1
|
|
||||||
with pdfinfo_lock:
|
|
||||||
pageinfo = pdfinfo[pageno].copy()
|
|
||||||
return pageinfo
|
|
||||||
|
|
||||||
|
|
||||||
def is_ocr_required(pageinfo, log):
|
|
||||||
page = pageinfo['pageno'] + 1
|
|
||||||
ocr_required = True
|
|
||||||
if not pageinfo['images']:
|
|
||||||
# If the page has no images, then it contains vector content or text
|
|
||||||
# or both. It seems quite unlikely that one would find meaningful text
|
|
||||||
# from rasterizing vector content. So skip the page.
|
|
||||||
log.info(
|
|
||||||
"Page {0} has no images - skipping OCR".format(page)
|
|
||||||
)
|
|
||||||
ocr_required = False
|
|
||||||
elif pageinfo['has_text']:
|
|
||||||
s = "Page {0} already has text! – {1}"
|
|
||||||
|
|
||||||
if not options.force_ocr and not options.skip_text:
|
|
||||||
log.error(s.format(page,
|
|
||||||
"aborting (use --force-ocr to force OCR)"))
|
|
||||||
sys.exit(EXIT_ALREADY_DONE_OCR)
|
|
||||||
elif options.force_ocr:
|
|
||||||
log.info(s.format(page,
|
|
||||||
"rasterizing text and running OCR anyway"))
|
|
||||||
ocr_required = True
|
|
||||||
elif options.skip_text:
|
|
||||||
log.info(s.format(page,
|
|
||||||
"skipping all processing on this page"))
|
|
||||||
ocr_required = False
|
|
||||||
|
|
||||||
if ocr_required and options.skip_big:
|
|
||||||
pixel_count = pageinfo['width_pixels'] * pageinfo['height_pixels']
|
|
||||||
if pixel_count > (options.skip_big * 1000000):
|
|
||||||
ocr_required = False
|
|
||||||
log.info(
|
|
||||||
"Page {0} is very large; skipping due to -b".format(page))
|
|
||||||
|
|
||||||
return ocr_required
|
|
||||||
|
|
||||||
|
|
||||||
@split(
|
|
||||||
repair_pdf,
|
|
||||||
os.path.join(work_folder, '*.page.pdf'),
|
|
||||||
extras=[_log, _pdfinfo, _pdfinfo_lock])
|
|
||||||
def split_pages(
|
|
||||||
input_file,
|
|
||||||
output_files,
|
|
||||||
log,
|
|
||||||
pdfinfo,
|
|
||||||
pdfinfo_lock):
|
|
||||||
|
|
||||||
for oo in output_files:
|
|
||||||
with suppress(FileNotFoundError):
|
|
||||||
os.unlink(oo)
|
|
||||||
args_pdfseparate = [
|
|
||||||
'pdfseparate',
|
|
||||||
input_file,
|
|
||||||
os.path.join(work_folder, '%06d.page.pdf')
|
|
||||||
]
|
|
||||||
check_call(args_pdfseparate)
|
|
||||||
|
|
||||||
from glob import glob
|
|
||||||
for filename in glob(os.path.join(work_folder, '*.page.pdf')):
|
|
||||||
pageinfo = get_pageinfo(filename, pdfinfo, pdfinfo_lock)
|
|
||||||
|
|
||||||
alt_suffix = '.ocr.page.pdf' if is_ocr_required(pageinfo, log) \
|
|
||||||
else '.skip.page.pdf'
|
|
||||||
re_symlink(
|
|
||||||
filename,
|
|
||||||
os.path.join(
|
|
||||||
work_folder,
|
|
||||||
os.path.basename(filename)[0:6] + alt_suffix))
|
|
||||||
|
|
||||||
|
|
||||||
@transform(
|
|
||||||
input=split_pages,
|
|
||||||
filter=suffix('.ocr.page.pdf'),
|
|
||||||
output='.page.png',
|
|
||||||
output_dir=work_folder,
|
|
||||||
extras=[_log, _pdfinfo, _pdfinfo_lock])
|
|
||||||
def rasterize_with_ghostscript(
|
|
||||||
input_file,
|
|
||||||
output_file,
|
|
||||||
log,
|
|
||||||
pdfinfo,
|
|
||||||
pdfinfo_lock):
|
|
||||||
|
|
||||||
pageinfo = get_pageinfo(input_file, pdfinfo, pdfinfo_lock)
|
|
||||||
|
|
||||||
device = 'png16m' # 24-bit
|
|
||||||
if all(image['comp'] == 1 for image in pageinfo['images']):
|
|
||||||
if all(image['bpc'] == 1 for image in pageinfo['images']):
|
|
||||||
device = 'pngmono'
|
|
||||||
elif not any(image['color'] == 'color'
|
|
||||||
for image in pageinfo['images']):
|
|
||||||
device = 'pnggray'
|
|
||||||
|
|
||||||
xres = max(pageinfo['xres'], options.oversample or 0)
|
|
||||||
yres = max(pageinfo['yres'], options.oversample or 0)
|
|
||||||
|
|
||||||
ghostscript.rasterize_pdf(input_file, output_file, xres, yres, device, log)
|
|
||||||
|
|
||||||
|
|
||||||
@transform(
|
|
||||||
input=rasterize_with_ghostscript,
|
|
||||||
filter=suffix(".page.png"),
|
|
||||||
output=".pp-deskew.png",
|
|
||||||
extras=[_log, _pdfinfo, _pdfinfo_lock])
|
|
||||||
def preprocess_deskew(
|
|
||||||
input_file,
|
|
||||||
output_file,
|
|
||||||
log,
|
|
||||||
pdfinfo,
|
|
||||||
pdfinfo_lock):
|
|
||||||
|
|
||||||
if not options.deskew:
|
|
||||||
re_symlink(input_file, output_file, log)
|
|
||||||
return
|
|
||||||
|
|
||||||
pageinfo = get_pageinfo(input_file, pdfinfo, pdfinfo_lock)
|
|
||||||
dpi = int(pageinfo['xres'])
|
|
||||||
|
|
||||||
unpaper.deskew(input_file, output_file, dpi, log)
|
|
||||||
|
|
||||||
|
|
||||||
@transform(
|
|
||||||
input=preprocess_deskew,
|
|
||||||
filter=suffix(".pp-deskew.png"),
|
|
||||||
output=".pp-clean.png",
|
|
||||||
extras=[_log, _pdfinfo, _pdfinfo_lock])
|
|
||||||
def preprocess_clean(
|
|
||||||
input_file,
|
|
||||||
output_file,
|
|
||||||
log,
|
|
||||||
pdfinfo,
|
|
||||||
pdfinfo_lock):
|
|
||||||
|
|
||||||
if not options.clean:
|
|
||||||
re_symlink(input_file, output_file, log)
|
|
||||||
return
|
|
||||||
|
|
||||||
pageinfo = get_pageinfo(input_file, pdfinfo, pdfinfo_lock)
|
|
||||||
dpi = int(pageinfo['xres'])
|
|
||||||
|
|
||||||
unpaper.clean(input_file, output_file, dpi, log)
|
|
||||||
|
|
||||||
|
|
||||||
@active_if(options.pdf_renderer == 'hocr')
|
|
||||||
@transform(
|
|
||||||
input=preprocess_clean,
|
|
||||||
filter=suffix(".pp-clean.png"),
|
|
||||||
output=".hocr",
|
|
||||||
extras=[_log, _pdfinfo, _pdfinfo_lock])
|
|
||||||
def ocr_tesseract_hocr(
|
|
||||||
input_file,
|
|
||||||
output_file,
|
|
||||||
log,
|
|
||||||
pdfinfo,
|
|
||||||
pdfinfo_lock):
|
|
||||||
|
|
||||||
pageinfo = get_pageinfo(input_file, pdfinfo, pdfinfo_lock)
|
|
||||||
|
|
||||||
args_tesseract = [
|
|
||||||
'tesseract',
|
|
||||||
'-l', '+'.join(options.language),
|
|
||||||
input_file,
|
|
||||||
output_file,
|
|
||||||
'hocr'
|
|
||||||
] + options.tesseract_config
|
|
||||||
p = Popen(args_tesseract, close_fds=True, stdout=PIPE, stderr=PIPE,
|
|
||||||
universal_newlines=True)
|
|
||||||
try:
|
|
||||||
stdout, stderr = p.communicate(timeout=options.tesseract_timeout)
|
|
||||||
except TimeoutExpired:
|
|
||||||
p.kill()
|
|
||||||
stdout, stderr = p.communicate()
|
|
||||||
# Generate a HOCR file with no recognized text if tesseract times out
|
|
||||||
# Temporary workaround to hocrTransform not being able to function if
|
|
||||||
# it does not have a valid hOCR file.
|
|
||||||
with open(output_file, 'w', encoding="utf-8") as f:
|
|
||||||
f.write(tesseract.HOCR_TEMPLATE.format(
|
|
||||||
pageinfo['width_pixels'],
|
|
||||||
pageinfo['height_pixels']))
|
|
||||||
else:
|
|
||||||
if stdout:
|
|
||||||
log.info(stdout)
|
|
||||||
if stderr:
|
|
||||||
log.error(stderr)
|
|
||||||
|
|
||||||
if p.returncode != 0:
|
|
||||||
raise CalledProcessError(p.returncode, args_tesseract)
|
|
||||||
|
|
||||||
if os.path.exists(output_file + '.html'):
|
|
||||||
# Tesseract 3.02 appends suffix ".html" on its own (.hocr.html)
|
|
||||||
shutil.move(output_file + '.html', output_file)
|
|
||||||
elif os.path.exists(output_file + '.hocr'):
|
|
||||||
# Tesseract 3.03 appends suffix ".hocr" on its own (.hocr.hocr)
|
|
||||||
shutil.move(output_file + '.hocr', output_file)
|
|
||||||
|
|
||||||
# Tesseract 3.03 inserts source filename into hocr file without
|
|
||||||
# escaping it, creating invalid XML and breaking the parser.
|
|
||||||
# As a workaround, rewrite the hocr file, replacing the filename
|
|
||||||
# with a space.
|
|
||||||
regex_nested_single_quotes = re.compile(
|
|
||||||
r"""title='image "([^"]*)";""")
|
|
||||||
with fileinput.input(files=(output_file,), inplace=True) as f:
|
|
||||||
for line in f:
|
|
||||||
line = regex_nested_single_quotes.sub(
|
|
||||||
r"""title='image " ";""", line)
|
|
||||||
print(line, end='') # fileinput.input redirects stdout
|
|
||||||
|
|
||||||
|
|
||||||
@active_if(options.pdf_renderer == 'hocr')
|
|
||||||
@collate(
|
|
||||||
input=[rasterize_with_ghostscript, preprocess_deskew, preprocess_clean],
|
|
||||||
filter=regex(r".*/(\d{6})(?:\.page|\.pp-deskew|\.pp-clean)\.png"),
|
|
||||||
output=os.path.join(work_folder, r'\1.image'),
|
|
||||||
extras=[_log, _pdfinfo, _pdfinfo_lock])
|
|
||||||
def select_image_for_pdf(
|
|
||||||
infiles,
|
|
||||||
output_file,
|
|
||||||
log,
|
|
||||||
pdfinfo,
|
|
||||||
pdfinfo_lock):
|
|
||||||
if options.clean_final:
|
|
||||||
image_suffix = '.pp-clean.png'
|
|
||||||
elif options.deskew:
|
|
||||||
image_suffix = '.pp-deskew.png'
|
|
||||||
else:
|
|
||||||
image_suffix = '.page.png'
|
|
||||||
image = next(ii for ii in infiles if ii.endswith(image_suffix))
|
|
||||||
|
|
||||||
pageinfo = get_pageinfo(image, pdfinfo, pdfinfo_lock)
|
|
||||||
if all(image['enc'] == 'jpeg' for image in pageinfo['images']):
|
|
||||||
# If all images were JPEGs originally, produce a JPEG as output
|
|
||||||
Image.open(image).save(output_file, format='JPEG')
|
|
||||||
else:
|
|
||||||
re_symlink(image, output_file)
|
|
||||||
|
|
||||||
|
|
||||||
@active_if(options.pdf_renderer == 'hocr')
|
|
||||||
@collate(
|
|
||||||
input=[select_image_for_pdf, ocr_tesseract_hocr],
|
|
||||||
filter=regex(r".*/(\d{6})(?:\.image|\.hocr)"),
|
|
||||||
output=os.path.join(work_folder, r'\1.rendered.pdf'),
|
|
||||||
extras=[_log, _pdfinfo, _pdfinfo_lock])
|
|
||||||
def render_hocr_page(
|
|
||||||
infiles,
|
|
||||||
output_file,
|
|
||||||
log,
|
|
||||||
pdfinfo,
|
|
||||||
pdfinfo_lock):
|
|
||||||
hocr = next(ii for ii in infiles if ii.endswith('.hocr'))
|
|
||||||
image = next(ii for ii in infiles if ii.endswith('.image'))
|
|
||||||
|
|
||||||
pageinfo = get_pageinfo(image, pdfinfo, pdfinfo_lock)
|
|
||||||
dpi = round(max(pageinfo['xres'], pageinfo['yres'], options.oversample))
|
|
||||||
|
|
||||||
hocrtransform = HocrTransform(hocr, dpi)
|
|
||||||
hocrtransform.to_pdf(output_file, imageFileName=image,
|
|
||||||
showBoundingboxes=False, invisibleText=True)
|
|
||||||
|
|
||||||
|
|
||||||
@active_if(options.pdf_renderer == 'hocr')
|
|
||||||
@active_if(options.debug_rendering)
|
|
||||||
@collate(
|
|
||||||
input=[select_image_for_pdf, ocr_tesseract_hocr],
|
|
||||||
filter=regex(r".*/(\d{6})(?:\.image|\.hocr)"),
|
|
||||||
output=os.path.join(work_folder, r'\1.debug.pdf'),
|
|
||||||
extras=[_log, _pdfinfo, _pdfinfo_lock])
|
|
||||||
def render_hocr_debug_page(
|
|
||||||
infiles,
|
|
||||||
output_file,
|
|
||||||
log,
|
|
||||||
pdfinfo,
|
|
||||||
pdfinfo_lock):
|
|
||||||
hocr = next(ii for ii in infiles if ii.endswith('.hocr'))
|
|
||||||
image = next(ii for ii in infiles if ii.endswith('.image'))
|
|
||||||
|
|
||||||
pageinfo = get_pageinfo(image, pdfinfo, pdfinfo_lock)
|
|
||||||
dpi = round(max(pageinfo['xres'], pageinfo['yres'], options.oversample))
|
|
||||||
|
|
||||||
hocrtransform = HocrTransform(hocr, dpi)
|
|
||||||
hocrtransform.to_pdf(output_file, imageFileName=None,
|
|
||||||
showBoundingboxes=True, invisibleText=False)
|
|
||||||
|
|
||||||
|
|
||||||
@active_if(options.pdf_renderer == 'tesseract')
|
|
||||||
@collate(
|
|
||||||
input=[preprocess_clean, split_pages],
|
|
||||||
filter=regex(r".*/(\d{6})(?:\.pp-clean\.png|\.page\.pdf)"),
|
|
||||||
output=os.path.join(work_folder, r'\1.rendered.pdf'),
|
|
||||||
extras=[_log, _pdfinfo, _pdfinfo_lock])
|
|
||||||
def tesseract_ocr_and_render_pdf(
|
|
||||||
input_files,
|
|
||||||
output_file,
|
|
||||||
log,
|
|
||||||
pdfinfo,
|
|
||||||
pdfinfo_lock):
|
|
||||||
|
|
||||||
input_image = next((ii for ii in input_files if ii.endswith('.png')), '')
|
|
||||||
input_pdf = next((ii for ii in input_files if ii.endswith('.pdf')))
|
|
||||||
if not input_image:
|
|
||||||
# Skipping this page
|
|
||||||
re_symlink(input_pdf, output_file)
|
|
||||||
return
|
|
||||||
|
|
||||||
args_tesseract = [
|
|
||||||
'tesseract',
|
|
||||||
'-l', '+'.join(options.language),
|
|
||||||
input_image,
|
|
||||||
os.path.splitext(output_file)[0], # Tesseract appends suffix
|
|
||||||
'pdf'
|
|
||||||
] + options.tesseract_config
|
|
||||||
p = Popen(args_tesseract, close_fds=True, stdout=PIPE, stderr=PIPE,
|
|
||||||
universal_newlines=True)
|
|
||||||
|
|
||||||
try:
|
|
||||||
stdout, stderr = p.communicate(timeout=options.tesseract_timeout)
|
|
||||||
if stdout:
|
|
||||||
log.info(stdout)
|
|
||||||
if stderr:
|
|
||||||
log.error(stderr)
|
|
||||||
except TimeoutError:
|
|
||||||
p.kill()
|
|
||||||
log.info("Tesseract - page timed out")
|
|
||||||
re_symlink(input_pdf, output_file)
|
|
||||||
|
|
||||||
|
|
||||||
@transform(
|
|
||||||
input=repair_pdf,
|
|
||||||
filter=suffix('.repaired.pdf'),
|
|
||||||
output='.pdfa_def.ps',
|
|
||||||
output_dir=work_folder,
|
|
||||||
extras=[_log])
|
|
||||||
def generate_postscript_stub(
|
|
||||||
input_file,
|
|
||||||
output_file,
|
|
||||||
log):
|
|
||||||
|
|
||||||
pdf = pypdf.PdfFileReader(input_file)
|
|
||||||
|
|
||||||
def from_document_info(key):
|
|
||||||
# pdf.documentInfo.get() DOES NOT work as expected
|
|
||||||
try:
|
|
||||||
s = pdf.documentInfo[key]
|
|
||||||
return str(s)
|
|
||||||
except KeyError:
|
|
||||||
return ''
|
|
||||||
|
|
||||||
pdfmark = {
|
|
||||||
'title': from_document_info('/Title'),
|
|
||||||
'author': from_document_info('/Author'),
|
|
||||||
'keywords': from_document_info('/Keywords'),
|
|
||||||
'subject': from_document_info('/Subject'),
|
|
||||||
}
|
|
||||||
if options.title:
|
|
||||||
pdfmark['title'] = options.title
|
|
||||||
if options.author:
|
|
||||||
pdfmark['author'] = options.author
|
|
||||||
if options.keywords:
|
|
||||||
pdfmark['keywords'] = options.keywords
|
|
||||||
if options.subject:
|
|
||||||
pdfmark['subject'] = options.subject
|
|
||||||
|
|
||||||
generate_pdfa_def(output_file, pdfmark)
|
|
||||||
|
|
||||||
|
|
||||||
@transform(
|
|
||||||
input=split_pages,
|
|
||||||
filter=suffix('.skip.page.pdf'),
|
|
||||||
output='.done.pdf',
|
|
||||||
output_dir=work_folder,
|
|
||||||
extras=[_log])
|
|
||||||
def skip_page(
|
|
||||||
input_file,
|
|
||||||
output_file,
|
|
||||||
log):
|
|
||||||
re_symlink(input_file, output_file, log)
|
|
||||||
|
|
||||||
|
|
||||||
@merge(
|
|
||||||
input=[render_hocr_page, render_hocr_debug_page, skip_page,
|
|
||||||
tesseract_ocr_and_render_pdf, generate_postscript_stub],
|
|
||||||
output=os.path.join(work_folder, 'merged.pdf'),
|
|
||||||
extras=[_log, _pdfinfo, _pdfinfo_lock])
|
|
||||||
def merge_pages(
|
|
||||||
input_files,
|
|
||||||
output_file,
|
|
||||||
log,
|
|
||||||
pdfinfo,
|
|
||||||
pdfinfo_lock):
|
|
||||||
|
|
||||||
def input_file_order(s):
|
|
||||||
'''Sort order: All rendered pages followed
|
|
||||||
by their debug page, if any, followed by Postscript stub.
|
|
||||||
Ghostscript documentation has the Postscript stub at the
|
|
||||||
beginning, but it works at the end and also gets document info
|
|
||||||
right that way.'''
|
|
||||||
if s.endswith('.ps'):
|
|
||||||
return 99999999
|
|
||||||
key = int(os.path.basename(s)[0:6]) * 10
|
|
||||||
if 'debug' in os.path.basename(s):
|
|
||||||
key += 1
|
|
||||||
return key
|
|
||||||
|
|
||||||
pdf_pages = sorted(input_files, key=input_file_order)
|
|
||||||
log.info(pdf_pages)
|
|
||||||
ghostscript.generate_pdfa(pdf_pages, output_file)
|
|
||||||
|
|
||||||
|
|
||||||
@transform(
|
|
||||||
input=merge_pages,
|
|
||||||
filter=formatter(),
|
|
||||||
output=options.output_file,
|
|
||||||
extras=[_log, _pdfinfo, _pdfinfo_lock])
|
|
||||||
def validate_pdfa(
|
|
||||||
input_file,
|
|
||||||
output_file,
|
|
||||||
log,
|
|
||||||
pdfinfo,
|
|
||||||
pdfinfo_lock):
|
|
||||||
|
|
||||||
args_jhove = [
|
|
||||||
'java',
|
|
||||||
'-jar', JHOVE_JAR,
|
|
||||||
'-c', JHOVE_CFG,
|
|
||||||
'-m', 'PDF-hul',
|
|
||||||
input_file
|
|
||||||
]
|
|
||||||
p_jhove = Popen(args_jhove, close_fds=True, universal_newlines=True,
|
|
||||||
stdout=PIPE, stderr=DEVNULL)
|
|
||||||
stdout, _ = p_jhove.communicate()
|
|
||||||
|
|
||||||
log.debug(stdout)
|
|
||||||
if p_jhove.returncode != 0:
|
|
||||||
log.error(stdout)
|
|
||||||
raise RuntimeError(
|
|
||||||
"Unexpected error while checking compliance to PDF/A file.")
|
|
||||||
|
|
||||||
pdf_is_valid = True
|
|
||||||
if re.search(r'ErrorMessage', stdout,
|
|
||||||
re.IGNORECASE | re.MULTILINE):
|
|
||||||
pdf_is_valid = False
|
|
||||||
if re.search(r'^\s+Status.*not valid', stdout,
|
|
||||||
re.IGNORECASE | re.MULTILINE):
|
|
||||||
pdf_is_valid = False
|
|
||||||
if re.search(r'^\s+Status.*Not well-formed', stdout,
|
|
||||||
re.IGNORECASE | re.MULTILINE):
|
|
||||||
pdf_is_valid = False
|
|
||||||
|
|
||||||
pdf_is_pdfa = False
|
|
||||||
if re.search(r'^\s+Profile:.*PDF/A-1', stdout,
|
|
||||||
re.IGNORECASE | re.MULTILINE):
|
|
||||||
pdf_is_pdfa = True
|
|
||||||
|
|
||||||
if not pdf_is_valid:
|
|
||||||
log.warning('Output file: The generated PDF/A file is INVALID')
|
|
||||||
elif pdf_is_valid and not pdf_is_pdfa:
|
|
||||||
log.warning('Output file: Generated file is a VALID PDF but not PDF/A')
|
|
||||||
elif pdf_is_valid and pdf_is_pdfa:
|
|
||||||
log.info('Output file: The generated PDF/A file is VALID')
|
|
||||||
shutil.copy(input_file, output_file)
|
|
||||||
|
|
||||||
|
|
||||||
# @active_if(ocr_required and options.exact_image)
|
|
||||||
# @merge([render_hocr_blank_page, extract_single_page],
|
|
||||||
# os.path.join(work_folder, "%04i.merged.pdf") % pageno)
|
|
||||||
# def merge_hocr_with_original_page(infiles, output_file):
|
|
||||||
# with open(infiles[0], 'rb') as hocr_input, \
|
|
||||||
# open(infiles[1], 'rb') as page_input, \
|
|
||||||
# open(output_file, 'wb') as output:
|
|
||||||
# hocr_reader = pypdf.PdfFileReader(hocr_input)
|
|
||||||
# page_reader = pypdf.PdfFileReader(page_input)
|
|
||||||
# writer = pypdf.PdfFileWriter()
|
|
||||||
|
|
||||||
# the_page = hocr_reader.getPage(0)
|
|
||||||
# the_page.mergePage(page_reader.getPage(0))
|
|
||||||
# writer.addPage(the_page)
|
|
||||||
# writer.write(output)
|
|
||||||
|
|
||||||
|
|
||||||
def available_cpu_count():
|
|
||||||
try:
|
|
||||||
return multiprocessing.cpu_count()
|
|
||||||
except NotImplementedError:
|
|
||||||
pass
|
|
||||||
|
|
||||||
try:
|
|
||||||
import psutil
|
|
||||||
return psutil.cpu_count()
|
|
||||||
except (ImportError, AttributeError):
|
|
||||||
pass
|
|
||||||
|
|
||||||
complain(
|
|
||||||
"Could not get CPU count. Assuming one (1) CPU."
|
|
||||||
"Use -j N to set manually.")
|
|
||||||
return 1
|
|
||||||
|
|
||||||
|
|
||||||
def run_pipeline():
|
|
||||||
cmdline.run(options, multiprocess=available_cpu_count())
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == '__main__':
|
|
||||||
run_pipeline()
|
|
||||||
@@ -1,137 +0,0 @@
|
|||||||
#!/usr/bin/env python3
|
|
||||||
# © 2015 James R. Barlow: github.com/jbarlow83
|
|
||||||
|
|
||||||
from subprocess import Popen, PIPE
|
|
||||||
from decimal import Decimal, getcontext
|
|
||||||
import re
|
|
||||||
import sys
|
|
||||||
import PyPDF2 as pypdf
|
|
||||||
|
|
||||||
|
|
||||||
FRIENDLY_COLORSPACE = {
|
|
||||||
'/DeviceGray': 'gray',
|
|
||||||
'/CalGray': 'gray',
|
|
||||||
'/DeviceRGB': 'rgb',
|
|
||||||
'/CalRGB': 'rgb',
|
|
||||||
'/DeviceCMYK': 'cmyk',
|
|
||||||
'/Lab': 'lab',
|
|
||||||
'/ICCBased': 'icc',
|
|
||||||
'/Indexed': 'index',
|
|
||||||
'/Separation': 'sep',
|
|
||||||
'/DeviceN': 'devn',
|
|
||||||
'/Pattern': '-'
|
|
||||||
}
|
|
||||||
|
|
||||||
FRIENDLY_ENCODING = {
|
|
||||||
'/CCITTFaxDecode': 'ccitt',
|
|
||||||
'/DCTDecode': 'jpeg',
|
|
||||||
'/JPXDecode': 'jpx',
|
|
||||||
'/JBIG2Decode': 'jbig2',
|
|
||||||
}
|
|
||||||
|
|
||||||
FRIENDLY_COMP = {
|
|
||||||
'gray': 1,
|
|
||||||
'rgb': 3,
|
|
||||||
'cmyk': 4,
|
|
||||||
'lab': 3,
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|
||||||
def _page_has_inline_images(page):
|
|
||||||
# PDF always uses \r\n for separator regardless of platform
|
|
||||||
# Really basic heuristic that might trigger the odd false positive
|
|
||||||
# This is only finds the first image and is not quite spec compliant
|
|
||||||
contents = page.getContents()
|
|
||||||
data = contents.getData()
|
|
||||||
begin_image, image_data, end_image = False, False, False
|
|
||||||
for data in re.split(b'\s+', data):
|
|
||||||
if data == b'BI':
|
|
||||||
begin_image = True
|
|
||||||
elif data == b'ID':
|
|
||||||
image_data = True
|
|
||||||
elif data == b'EI':
|
|
||||||
end_image = True
|
|
||||||
if all((begin_image, image_data, end_image)):
|
|
||||||
return True
|
|
||||||
|
|
||||||
|
|
||||||
def _find_page_images(page, pageinfo):
|
|
||||||
try:
|
|
||||||
page['/Resources']['/XObject']
|
|
||||||
except KeyError:
|
|
||||||
return
|
|
||||||
|
|
||||||
# Look for XObject (out of line images)
|
|
||||||
for xobj in page['/Resources']['/XObject']:
|
|
||||||
# PyPDF2 returns the keys as an iterator
|
|
||||||
pdfimage = page['/Resources']['/XObject'][xobj]
|
|
||||||
if pdfimage['/Subtype'] != '/Image':
|
|
||||||
continue
|
|
||||||
if '/ImageMask' in pdfimage:
|
|
||||||
if pdfimage['/ImageMask']:
|
|
||||||
continue
|
|
||||||
image = {}
|
|
||||||
image['width'] = pdfimage['/Width']
|
|
||||||
image['height'] = pdfimage['/Height']
|
|
||||||
image['bpc'] = pdfimage['/BitsPerComponent']
|
|
||||||
if '/Filter' in pdfimage:
|
|
||||||
filter_ = pdfimage['/Filter']
|
|
||||||
if isinstance(filter_, pypdf.generic.ArrayObject):
|
|
||||||
filter_ = filter_[0]
|
|
||||||
image['enc'] = FRIENDLY_ENCODING.get(filter_, 'image')
|
|
||||||
else:
|
|
||||||
image['enc'] = 'image'
|
|
||||||
if '/ColorSpace' in pdfimage:
|
|
||||||
cs = pdfimage['/ColorSpace']
|
|
||||||
if isinstance(cs, pypdf.generic.ArrayObject):
|
|
||||||
cs = cs[0]
|
|
||||||
image['color'] = FRIENDLY_COLORSPACE.get(cs, '-')
|
|
||||||
else:
|
|
||||||
image['color'] = 'jpx' if image['enc'] == 'jpx' else '?'
|
|
||||||
|
|
||||||
image['comp'] = FRIENDLY_COMP.get(image['color'], '?')
|
|
||||||
image['dpi_w'] = image['width'] / pageinfo['width_inches']
|
|
||||||
image['dpi_h'] = image['height'] / pageinfo['height_inches']
|
|
||||||
image['dpi'] = (image['dpi_w'] * image['dpi_h']) ** Decimal(0.5)
|
|
||||||
yield image
|
|
||||||
|
|
||||||
|
|
||||||
def _pdf_get_pageinfo(infile, page: int):
|
|
||||||
pageinfo = {}
|
|
||||||
pageinfo['pageno'] = page
|
|
||||||
pageinfo['images'] = []
|
|
||||||
|
|
||||||
pdf = pypdf.PdfFileReader(infile)
|
|
||||||
page = pdf.pages[page - 1]
|
|
||||||
|
|
||||||
text = page.extractText()
|
|
||||||
pageinfo['has_text'] = (text.strip() != '')
|
|
||||||
|
|
||||||
width_pt = page['/MediaBox'][2] - page['/MediaBox'][0]
|
|
||||||
height_pt = page['/MediaBox'][3] - page['/MediaBox'][1]
|
|
||||||
pageinfo['width_inches'] = width_pt / Decimal(72.0)
|
|
||||||
pageinfo['height_inches'] = height_pt / Decimal(72.0)
|
|
||||||
|
|
||||||
pageinfo['images'] = [im for im in _find_page_images(page, pageinfo)]
|
|
||||||
|
|
||||||
# Look for inline images
|
|
||||||
if _page_has_inline_images(page):
|
|
||||||
raise NotImplementedError(
|
|
||||||
"Warning: input PDF contains inline images - not supported")
|
|
||||||
|
|
||||||
if pageinfo['images']:
|
|
||||||
xres = max(image['dpi_w'] for image in pageinfo['images'])
|
|
||||||
yres = max(image['dpi_h'] for image in pageinfo['images'])
|
|
||||||
pageinfo['xres'], pageinfo['yres'] = xres, yres
|
|
||||||
pageinfo['width_pixels'] = \
|
|
||||||
int(round(xres * pageinfo['width_inches']))
|
|
||||||
pageinfo['height_pixels'] = \
|
|
||||||
int(round(yres * pageinfo['height_inches']))
|
|
||||||
|
|
||||||
return pageinfo
|
|
||||||
|
|
||||||
|
|
||||||
def pdf_get_all_pageinfo(infile):
|
|
||||||
pdf = pypdf.PdfFileReader(infile)
|
|
||||||
getcontext().prec = 6
|
|
||||||
return [_pdf_get_pageinfo(infile, n) for n in range(pdf.numPages)]
|
|
||||||
@@ -1,130 +0,0 @@
|
|||||||
#!/usr/bin/env python3
|
|
||||||
# © 2015 James R. Barlow: github.com/jbarlow83
|
|
||||||
#
|
|
||||||
# Generate a PDFA_def.ps file for Ghostscript >= 9.14
|
|
||||||
|
|
||||||
from __future__ import print_function, absolute_import, division
|
|
||||||
from string import Template
|
|
||||||
from subprocess import Popen, PIPE
|
|
||||||
import os
|
|
||||||
import codecs
|
|
||||||
|
|
||||||
|
|
||||||
# This is a template written in PostScript which is needed to create PDF/A
|
|
||||||
# files, from the Ghostscript documentation. Lines beginning with % are
|
|
||||||
# comments. Python substitution variables have a '$' prefix.
|
|
||||||
pdfa_def_template = u"""%!
|
|
||||||
% This is a sample prefix file for creating a PDF/A document.
|
|
||||||
% Feel free to modify entries marked with "Customize".
|
|
||||||
% This assumes an ICC profile to reside in the file (ISO Coated sb.icc),
|
|
||||||
% unless the user modifies the corresponding line below.
|
|
||||||
|
|
||||||
% Define entries in the document Info dictionary :
|
|
||||||
/ICCProfile ($icc_profile)
|
|
||||||
def
|
|
||||||
|
|
||||||
[ /Title <$title>
|
|
||||||
/Author <$author>
|
|
||||||
/Subject <$subject>
|
|
||||||
/Keywords <$keywords>
|
|
||||||
/DOCINFO pdfmark
|
|
||||||
|
|
||||||
% Define an ICC profile :
|
|
||||||
|
|
||||||
[/_objdef {icc_PDFA} /type /stream /OBJ pdfmark
|
|
||||||
[{icc_PDFA}
|
|
||||||
<<
|
|
||||||
/N currentpagedevice /ProcessColorModel known {
|
|
||||||
currentpagedevice /ProcessColorModel get dup /DeviceGray eq
|
|
||||||
{pop 1} {
|
|
||||||
/DeviceRGB eq
|
|
||||||
{3}{4} ifelse
|
|
||||||
} ifelse
|
|
||||||
} {
|
|
||||||
(ERROR, unable to determine ProcessColorModel) == flush
|
|
||||||
} ifelse
|
|
||||||
>> /PUT pdfmark
|
|
||||||
[{icc_PDFA} ICCProfile (r) file /PUT pdfmark
|
|
||||||
|
|
||||||
% Define the output intent dictionary :
|
|
||||||
|
|
||||||
[/_objdef {OutputIntent_PDFA} /type /dict /OBJ pdfmark
|
|
||||||
[{OutputIntent_PDFA} <<
|
|
||||||
/Type /OutputIntent % Must be so (the standard requires).
|
|
||||||
/S /GTS_PDFA1 % Must be so (the standard requires).
|
|
||||||
/DestOutputProfile {icc_PDFA} % Must be so (see above).
|
|
||||||
/OutputConditionIdentifier ($icc_identifier)
|
|
||||||
>> /PUT pdfmark
|
|
||||||
[{Catalog} <</OutputIntents [ {OutputIntent_PDFA} ]>> /PUT pdfmark
|
|
||||||
"""
|
|
||||||
|
|
||||||
|
|
||||||
def encode_text_string(s: str) -> str:
|
|
||||||
'''Encode text string to hex string for use in a PDF
|
|
||||||
|
|
||||||
From PDF 32000-1:2008 a string object may be included in hexademical form
|
|
||||||
if it is enclosed in angle brackets. For general Unicode the string should
|
|
||||||
be UTF-16 (big endian) with byte order marks. A non-hexademical
|
|
||||||
representation is doable but this is preferable since it allows the output
|
|
||||||
Postscript file to be completely ASCII and no escaping of Postscript
|
|
||||||
characters is necessary.
|
|
||||||
'''
|
|
||||||
if s == '':
|
|
||||||
return ''
|
|
||||||
utf16_bytes = s.encode('utf-16be')
|
|
||||||
ascii_hex_bytes = codecs.encode(b'\xfe\xff' + utf16_bytes, 'hex')
|
|
||||||
ascii_hex_str = ascii_hex_bytes.decode('ascii').lower()
|
|
||||||
return ascii_hex_str
|
|
||||||
|
|
||||||
|
|
||||||
def _get_pdfa_def(icc_profile, icc_identifier, pdfmark):
|
|
||||||
pdfmark_utf16 = {k: encode_text_string(v) for k, v in pdfmark.items()}
|
|
||||||
|
|
||||||
t = Template(pdfa_def_template)
|
|
||||||
result = t.substitute(icc_profile=icc_profile,
|
|
||||||
icc_identifier=icc_identifier,
|
|
||||||
title=pdfmark_utf16.get('title', ''),
|
|
||||||
author=pdfmark_utf16.get('author', ''),
|
|
||||||
subject=pdfmark_utf16.get('subject', ''),
|
|
||||||
keywords=pdfmark_utf16.get('keywords', ''))
|
|
||||||
return result
|
|
||||||
|
|
||||||
|
|
||||||
def _get_postscript_icc_path():
|
|
||||||
"Parse Ghostscript's help message to find where iccprofiles are stored"
|
|
||||||
|
|
||||||
p_gs = Popen(['gs', '--help'], close_fds=True, universal_newlines=True,
|
|
||||||
stdout=PIPE, stderr=PIPE)
|
|
||||||
out, _ = p_gs.communicate()
|
|
||||||
lines = out.splitlines()
|
|
||||||
|
|
||||||
def search_paths(lines):
|
|
||||||
seeking = True
|
|
||||||
for line in lines:
|
|
||||||
if seeking:
|
|
||||||
if line.startswith('Search path'):
|
|
||||||
seeking = False
|
|
||||||
continue
|
|
||||||
else:
|
|
||||||
if line.strip().startswith('/'):
|
|
||||||
yield from (
|
|
||||||
path.strip() for path in line.split(':')
|
|
||||||
if path.strip() != '')
|
|
||||||
for root in search_paths(lines):
|
|
||||||
path = os.path.realpath(os.path.join(root, '../iccprofiles'))
|
|
||||||
if os.path.exists(path):
|
|
||||||
return path
|
|
||||||
|
|
||||||
|
|
||||||
def generate_pdfa_def(target_filename, pdfmark, icc='sRGB'):
|
|
||||||
if icc == 'sRGB':
|
|
||||||
icc_profile = os.path.join(_get_postscript_icc_path(), 'srgb.icc')
|
|
||||||
else:
|
|
||||||
raise NotImplementedError("Only supporting sRGB")
|
|
||||||
|
|
||||||
ps = _get_pdfa_def(icc_profile, icc, pdfmark)
|
|
||||||
|
|
||||||
# Since PostScript might not handle UTF-8 (it's hard to get a clear
|
|
||||||
# answer), insist on ascii
|
|
||||||
with open(target_filename, 'w', encoding='ascii') as f:
|
|
||||||
f.write(ps)
|
|
||||||
@@ -1,61 +0,0 @@
|
|||||||
#!/usr/bin/env python3
|
|
||||||
# © 2015 James R. Barlow: github.com/jbarlow83
|
|
||||||
|
|
||||||
from subprocess import STDOUT, CalledProcessError, check_output
|
|
||||||
import sys
|
|
||||||
import os
|
|
||||||
import re
|
|
||||||
from functools import lru_cache
|
|
||||||
|
|
||||||
|
|
||||||
@lru_cache(maxsize=1)
|
|
||||||
def version():
|
|
||||||
args_tess = [
|
|
||||||
'tesseract',
|
|
||||||
'--version'
|
|
||||||
]
|
|
||||||
try:
|
|
||||||
versions = check_output(
|
|
||||||
args_tess, close_fds=True, universal_newlines=True,
|
|
||||||
stderr=STDOUT)
|
|
||||||
except CalledProcessError:
|
|
||||||
print("Could not find Tesseract executable on system PATH.")
|
|
||||||
sys.exit(1)
|
|
||||||
|
|
||||||
tesseract_version = re.match(r'tesseract\s(.+)', versions).group(1)
|
|
||||||
return tesseract_version
|
|
||||||
|
|
||||||
|
|
||||||
@lru_cache(maxsize=1)
|
|
||||||
def languages():
|
|
||||||
args_tess = [
|
|
||||||
'tesseract',
|
|
||||||
'--list-langs'
|
|
||||||
]
|
|
||||||
langs = check_output(
|
|
||||||
args_tess, close_fds=True, universal_newlines=True,
|
|
||||||
stderr=STDOUT)
|
|
||||||
return set(lang.strip() for lang in langs.splitlines()[1:])
|
|
||||||
|
|
||||||
|
|
||||||
HOCR_TEMPLATE = '''<?xml version="1.0" encoding="UTF-8"?>
|
|
||||||
<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN"
|
|
||||||
"http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd">
|
|
||||||
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en">
|
|
||||||
<head>
|
|
||||||
<title></title>
|
|
||||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
|
||||||
<meta name='ocr-system' content='tesseract 3.02.02' />
|
|
||||||
<meta name='ocr-capabilities' content='ocr_page ocr_carea ocr_par ocr_line ocrx_word'/>
|
|
||||||
</head>
|
|
||||||
<body>
|
|
||||||
<div class='ocr_page' id='page_1' title='image "x.tif"; bbox 0 0 {0} {1}; ppageno 0'>
|
|
||||||
<div class='ocr_carea' id='block_1_1' title="bbox 0 1 {0} {1}">
|
|
||||||
<p class='ocr_par' dir='ltr' id='par_1' title="bbox 0 1 {0} {1}">
|
|
||||||
<span class='ocr_line' id='line_1' title="bbox 0 1 {0} {1}"><span class='ocrx_word' id='word_1' title="bbox 0 1 {0} {1}"> </span>
|
|
||||||
</span>
|
|
||||||
</p>
|
|
||||||
</div>
|
|
||||||
</div>
|
|
||||||
</body>
|
|
||||||
</html>'''
|
|
||||||
@@ -1,107 +0,0 @@
|
|||||||
#!/usr/bin/env python3
|
|
||||||
# © 2015 James R. Barlow: github.com/jbarlow83
|
|
||||||
|
|
||||||
from ocrmypdf import pageinfo
|
|
||||||
from reportlab.pdfgen.canvas import Canvas
|
|
||||||
from PIL import Image
|
|
||||||
from tempfile import NamedTemporaryFile
|
|
||||||
from contextlib import suppress
|
|
||||||
import os
|
|
||||||
import sys
|
|
||||||
import shutil
|
|
||||||
import pytest
|
|
||||||
from pkg_resources import Requirement, resource_filename
|
|
||||||
|
|
||||||
req = Requirement.parse('ocrmypdf')
|
|
||||||
|
|
||||||
TEST_OUTPUT = os.path.join(os.path.dirname(__file__), 'output')
|
|
||||||
|
|
||||||
|
|
||||||
def setup_module():
|
|
||||||
with suppress(FileNotFoundError):
|
|
||||||
shutil.rmtree(TEST_OUTPUT)
|
|
||||||
with suppress(FileExistsError):
|
|
||||||
os.mkdir(TEST_OUTPUT)
|
|
||||||
|
|
||||||
|
|
||||||
def test_single_page_text():
|
|
||||||
filename = os.path.join(TEST_OUTPUT, 'text.pdf')
|
|
||||||
pdf = Canvas(filename, pagesize=(8*72, 6*72))
|
|
||||||
text = pdf.beginText()
|
|
||||||
text.setFont('Helvetica', 12)
|
|
||||||
text.setTextOrigin(1*72, 3*72)
|
|
||||||
text.textLine("Methink'st thou art a general offence and every"
|
|
||||||
" man should beat thee.")
|
|
||||||
pdf.drawText(text)
|
|
||||||
pdf.showPage()
|
|
||||||
pdf.save()
|
|
||||||
|
|
||||||
pdfinfo = pageinfo.pdf_get_all_pageinfo(filename)
|
|
||||||
|
|
||||||
assert len(pdfinfo) == 1
|
|
||||||
page = pdfinfo[0]
|
|
||||||
|
|
||||||
assert page['has_text']
|
|
||||||
assert len(page['images']) == 0
|
|
||||||
|
|
||||||
|
|
||||||
def test_single_page_image():
|
|
||||||
filename = os.path.join(TEST_OUTPUT, 'image-mono.pdf')
|
|
||||||
pdf = Canvas(filename, pagesize=(72, 72))
|
|
||||||
with NamedTemporaryFile() as im_tmp:
|
|
||||||
im = Image.new('1', (8, 8), 0)
|
|
||||||
for n in range(8):
|
|
||||||
im.putpixel((n, n), 1)
|
|
||||||
im.save(im_tmp.name, format='PNG')
|
|
||||||
# Draw image in a 72x72 pt or 1"x1" area
|
|
||||||
pdf.drawImage(im_tmp.name, 0, 0, width=72, height=72)
|
|
||||||
pdf.showPage()
|
|
||||||
pdf.save()
|
|
||||||
|
|
||||||
pdfinfo = pageinfo.pdf_get_all_pageinfo(filename)
|
|
||||||
|
|
||||||
assert len(pdfinfo) == 1
|
|
||||||
page = pdfinfo[0]
|
|
||||||
|
|
||||||
assert not page['has_text']
|
|
||||||
assert len(page['images']) == 1
|
|
||||||
|
|
||||||
pdfimage = page['images'][0]
|
|
||||||
assert pdfimage['width'] == 8
|
|
||||||
# assert pdfimage['color'] == 'gray'
|
|
||||||
|
|
||||||
# While unexpected, this is correct
|
|
||||||
# PDF spec says /FlateDecode image must have /BitsPerComponent 8
|
|
||||||
# So mono images get upgraded to 8-bit
|
|
||||||
assert pdfimage['bpc'] == 8
|
|
||||||
|
|
||||||
# DPI in a 1"x1" is the image width
|
|
||||||
assert pdfimage['dpi_w'] == 8
|
|
||||||
assert pdfimage['dpi_h'] == 8
|
|
||||||
|
|
||||||
|
|
||||||
def test_single_page_inline_image():
|
|
||||||
filename = os.path.join(TEST_OUTPUT, 'image-mono-inline.pdf')
|
|
||||||
pdf = Canvas(filename, pagesize=(8*72, 6*72))
|
|
||||||
with NamedTemporaryFile() as im_tmp:
|
|
||||||
im = Image.new('1', (8, 8), 0)
|
|
||||||
for n in range(8):
|
|
||||||
im.putpixel((n, n), 1)
|
|
||||||
im.save(im_tmp.name, format='PNG')
|
|
||||||
# Draw image in a 72x72 pt or 1"x1" area
|
|
||||||
pdf.drawInlineImage(im_tmp.name, 0, 0, width=72, height=72)
|
|
||||||
pdf.showPage()
|
|
||||||
pdf.save()
|
|
||||||
|
|
||||||
with pytest.raises(NotImplementedError):
|
|
||||||
pageinfo.pdf_get_all_pageinfo(filename)
|
|
||||||
|
|
||||||
|
|
||||||
def test_jpeg():
|
|
||||||
filename = resource_filename(req, 'tests/resources/c02-22.pdf')
|
|
||||||
|
|
||||||
pdfinfo = pageinfo.pdf_get_all_pageinfo(filename)
|
|
||||||
|
|
||||||
pdfimage = pdfinfo[0]['images'][0]
|
|
||||||
assert pdfimage['enc'] == 'jpeg'
|
|
||||||
|
|
||||||
@@ -1,84 +0,0 @@
|
|||||||
#!/usr/bin/env python3
|
|
||||||
# © 2015 James R. Barlow: github.com/jbarlow83
|
|
||||||
# unpaper documentation:
|
|
||||||
# https://github.com/Flameeyes/unpaper/blob/master/doc/basic-concepts.md
|
|
||||||
|
|
||||||
from subprocess import Popen, PIPE
|
|
||||||
from tempfile import NamedTemporaryFile
|
|
||||||
import sys
|
|
||||||
import os
|
|
||||||
from functools import lru_cache
|
|
||||||
|
|
||||||
|
|
||||||
@lru_cache(maxsize=1)
|
|
||||||
def version():
|
|
||||||
args_unpaper = [
|
|
||||||
'unpaper',
|
|
||||||
'--version'
|
|
||||||
]
|
|
||||||
p_unpaper = Popen(args_unpaper, close_fds=True, universal_newlines=True,
|
|
||||||
stdout=PIPE, stderr=PIPE)
|
|
||||||
version, _ = p_unpaper.communicate(timeout=5)
|
|
||||||
|
|
||||||
return version.strip()
|
|
||||||
|
|
||||||
|
|
||||||
try:
|
|
||||||
from PIL import Image
|
|
||||||
except ImportError:
|
|
||||||
print("Could not find Python3 imaging library", file=sys.stderr)
|
|
||||||
raise
|
|
||||||
|
|
||||||
|
|
||||||
def run(input_file, output_file, dpi, log, mode_args):
|
|
||||||
args_unpaper = [
|
|
||||||
'unpaper',
|
|
||||||
'-v',
|
|
||||||
'--dpi', str(dpi)
|
|
||||||
] + mode_args
|
|
||||||
|
|
||||||
SUFFIXES = {'1': '.pbm', 'L': '.pgm', 'RGB': '.ppm'}
|
|
||||||
suffix = ''
|
|
||||||
|
|
||||||
im = Image.open(input_file)
|
|
||||||
suffix = SUFFIXES[im.mode]
|
|
||||||
with NamedTemporaryFile(suffix=suffix) as input_pnm, \
|
|
||||||
NamedTemporaryFile(suffix=suffix, mode="r+b") as output_pnm:
|
|
||||||
im.save(input_pnm, format='PPM')
|
|
||||||
im.close()
|
|
||||||
|
|
||||||
os.unlink(output_pnm.name)
|
|
||||||
|
|
||||||
args_unpaper.extend([input_pnm.name, output_pnm.name])
|
|
||||||
p_unpaper = Popen(
|
|
||||||
args_unpaper, close_fds=True,
|
|
||||||
universal_newlines=True, stdout=PIPE, stderr=PIPE
|
|
||||||
)
|
|
||||||
out, err = p_unpaper.communicate()
|
|
||||||
log.debug(out)
|
|
||||||
log.debug(err)
|
|
||||||
|
|
||||||
Image.open(output_pnm.name).save(output_file)
|
|
||||||
|
|
||||||
|
|
||||||
def deskew(input_file, output_file, dpi, log):
|
|
||||||
run(input_file, output_file, dpi, log, [
|
|
||||||
'--mask-scan-size', '100', # don't blank out narrow columns
|
|
||||||
'--no-border-align', # don't align visible content to borders
|
|
||||||
'--no-mask-center', # don't center visible content within page
|
|
||||||
'--no-grayfilter', # don't remove light gray areas
|
|
||||||
'--no-blackfilter', # don't remove solid black areas
|
|
||||||
'--no-noisefilter', # don't remove salt and pepper noise
|
|
||||||
'--no-blurfilter' # don't remove blurry objects/debris
|
|
||||||
])
|
|
||||||
|
|
||||||
|
|
||||||
def clean(input_file, output_file, dpi, log):
|
|
||||||
run(input_file, output_file, dpi, log, [
|
|
||||||
'--mask-scan-size', '100', # don't blank out narrow columns
|
|
||||||
'--no-border-align', # don't align visible content to borders
|
|
||||||
'--no-mask-center', # don't center visible content within page
|
|
||||||
'--no-grayfilter', # don't remove light gray areas
|
|
||||||
'--no-blackfilter', # don't remove solid black areas
|
|
||||||
'--no-deskew', # don't deskew
|
|
||||||
])
|
|
||||||
-235
@@ -1,235 +0,0 @@
|
|||||||
<?xml version="1.0" encoding="UTF-8" standalone="no"?>
|
|
||||||
<!DOCTYPE svg PUBLIC "-//W3C//DTD SVG 1.1//EN"
|
|
||||||
"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd">
|
|
||||||
<!-- Generated by graphviz version 2.38.0 (20140413.2041)
|
|
||||||
-->
|
|
||||||
<!-- Title: Pipeline: Pages: 1 -->
|
|
||||||
<svg width="728pt" height="651pt"
|
|
||||||
viewBox="0.00 0.00 728.00 650.53" xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink">
|
|
||||||
<g id="graph0" class="graph" transform="scale(1 1) rotate(0) translate(4 646.53)">
|
|
||||||
<title>Pipeline:</title>
|
|
||||||
<polygon fill="white" stroke="none" points="-4,4 -4,-646.53 724,-646.53 724,4 -4,4"/>
|
|
||||||
<g id="clust1" class="cluster"><title>clustertasks</title>
|
|
||||||
<polygon fill="none" stroke="black" points="8,-8 8,-634.53 712,-634.53 712,-8 8,-8"/>
|
|
||||||
<text text-anchor="middle" x="360" y="-606.53" font-family="Times,serif" font-size="30.00" fill="#ff3232">Pipeline:</text>
|
|
||||||
</g>
|
|
||||||
<!-- t0 -->
|
|
||||||
<g id="node1" class="node"><title>t0</title>
|
|
||||||
<polygon fill="#efa03b" stroke="#006000" points="481.791,-588.53 386.209,-588.53 382.209,-584.53 382.209,-552.53 477.791,-552.53 481.791,-556.53 481.791,-588.53"/>
|
|
||||||
<polyline fill="none" stroke="#006000" points="477.791,-584.53 382.209,-584.53 "/>
|
|
||||||
<polyline fill="none" stroke="#006000" points="477.791,-584.53 477.791,-552.53 "/>
|
|
||||||
<polyline fill="none" stroke="#006000" points="477.791,-584.53 481.791,-588.53 "/>
|
|
||||||
<text text-anchor="middle" x="432" y="-564.53" font-family="Times,serif" font-size="20.00" fill="#006000">repair_pdf</text>
|
|
||||||
</g>
|
|
||||||
<!-- t1 -->
|
|
||||||
<g id="node2" class="node"><title>t1</title>
|
|
||||||
<polygon fill="#efa03b" stroke="black" points="466.782,-509.564 374,-526.497 281.218,-509.564 281.304,-482.165 466.696,-482.165 466.782,-509.564"/>
|
|
||||||
<polygon fill="none" stroke="black" points="470.799,-512.902 374,-530.569 277.201,-512.902 277.311,-478.159 470.689,-478.159 470.799,-512.902"/>
|
|
||||||
<text text-anchor="middle" x="374" y="-495.991" font-family="Times,serif" font-size="20.00">split_pages</text>
|
|
||||||
</g>
|
|
||||||
<!-- t0->t1 -->
|
|
||||||
<g id="edge1" class="edge"><title>t0->t1</title>
|
|
||||||
<path fill="none" stroke="gray" d="M417.064,-552.394C412.296,-546.925 406.86,-540.689 401.493,-534.532"/>
|
|
||||||
<polygon fill="gray" stroke="gray" points="403.996,-532.077 394.787,-526.838 398.719,-536.676 403.996,-532.077"/>
|
|
||||||
</g>
|
|
||||||
<!-- t10 -->
|
|
||||||
<g id="node12" class="node"><title>t10</title>
|
|
||||||
<polygon fill="#efa03b" stroke="#006000" points="704.338,-451.452 493.662,-451.452 489.662,-447.452 489.662,-415.452 700.338,-415.452 704.338,-419.452 704.338,-451.452"/>
|
|
||||||
<polyline fill="none" stroke="#006000" points="700.338,-447.452 489.662,-447.452 "/>
|
|
||||||
<polyline fill="none" stroke="#006000" points="700.338,-447.452 700.338,-415.452 "/>
|
|
||||||
<polyline fill="none" stroke="#006000" points="700.338,-447.452 704.338,-451.452 "/>
|
|
||||||
<text text-anchor="middle" x="597" y="-427.452" font-family="Times,serif" font-size="20.00" fill="#006000">generate_postscript_stub</text>
|
|
||||||
</g>
|
|
||||||
<!-- t0->t10 -->
|
|
||||||
<g id="edge16" class="edge"><title>t0->t10</title>
|
|
||||||
<path fill="none" stroke="gray" d="M453.048,-552.468C461.441,-545.654 471.185,-537.73 480,-530.53 510.227,-505.84 544.753,-477.467 568.421,-457.99"/>
|
|
||||||
<polygon fill="gray" stroke="gray" points="570.818,-460.55 576.315,-451.493 566.369,-455.146 570.818,-460.55"/>
|
|
||||||
</g>
|
|
||||||
<!-- t2 -->
|
|
||||||
<g id="node3" class="node"><title>t2</title>
|
|
||||||
<polygon fill="#efa03b" stroke="black" points="451.555,-451.452 228.445,-451.452 224.445,-447.452 224.445,-415.452 447.555,-415.452 451.555,-419.452 451.555,-451.452"/>
|
|
||||||
<polyline fill="none" stroke="black" points="447.555,-447.452 224.445,-447.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="447.555,-447.452 447.555,-415.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="447.555,-447.452 451.555,-451.452 "/>
|
|
||||||
<text text-anchor="middle" x="338" y="-427.452" font-family="Times,serif" font-size="20.00">rasterize_with_ghostscript</text>
|
|
||||||
</g>
|
|
||||||
<!-- t1->t2 -->
|
|
||||||
<g id="edge2" class="edge"><title>t1->t2</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M361.611,-478.092C358.514,-472.369 355.171,-466.19 352.003,-460.333"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="355.064,-458.636 347.227,-451.506 348.907,-461.967 355.064,-458.636"/>
|
|
||||||
</g>
|
|
||||||
<!-- t11 -->
|
|
||||||
<g id="node10" class="node"><title>t11</title>
|
|
||||||
<polygon fill="#efa03b" stroke="black" points="644.594,-393.452 551.406,-393.452 547.406,-389.452 547.406,-357.452 640.594,-357.452 644.594,-361.452 644.594,-393.452"/>
|
|
||||||
<polyline fill="none" stroke="black" points="640.594,-389.452 547.406,-389.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="640.594,-389.452 640.594,-357.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="640.594,-389.452 644.594,-393.452 "/>
|
|
||||||
<text text-anchor="middle" x="596" y="-369.452" font-family="Times,serif" font-size="20.00">skip_page</text>
|
|
||||||
</g>
|
|
||||||
<!-- t1->t11 -->
|
|
||||||
<g id="edge13" class="edge"><title>t1->t11</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M425.139,-478.007C437.834,-470.731 450.731,-461.83 461,-451.452 473.874,-438.442 466.946,-427.178 481,-415.452 490.212,-407.766 513.965,-399.205 537.508,-392.067"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="538.592,-395.396 547.189,-389.203 536.607,-388.683 538.592,-395.396"/>
|
|
||||||
</g>
|
|
||||||
<!-- t9 -->
|
|
||||||
<g id="node11" class="node"><title>t9</title>
|
|
||||||
<polygon fill="#efa03b" stroke="black" points="272.496,-277.452 19.5039,-277.452 15.5039,-273.452 15.5039,-241.452 268.496,-241.452 272.496,-245.452 272.496,-277.452"/>
|
|
||||||
<polyline fill="none" stroke="black" points="268.496,-273.452 15.5039,-273.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="268.496,-273.452 268.496,-241.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="268.496,-273.452 272.496,-277.452 "/>
|
|
||||||
<text text-anchor="middle" x="144" y="-253.452" font-family="Times,serif" font-size="20.00">tesseract_ocr_and_render_pdf</text>
|
|
||||||
</g>
|
|
||||||
<!-- t1->t9 -->
|
|
||||||
<g id="edge15" class="edge"><title>t1->t9</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M277.341,-486.044C254.708,-478.702 232.21,-467.759 215,-451.452 168.207,-407.114 152.056,-329.127 146.633,-287.855"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="150.084,-287.235 145.422,-277.721 143.133,-288.065 150.084,-287.235"/>
|
|
||||||
</g>
|
|
||||||
<!-- t3 -->
|
|
||||||
<g id="node4" class="node"><title>t3</title>
|
|
||||||
<polygon fill="#efa03b" stroke="black" points="458.999,-393.452 291.001,-393.452 287.001,-389.452 287.001,-357.452 454.999,-357.452 458.999,-361.452 458.999,-393.452"/>
|
|
||||||
<polyline fill="none" stroke="black" points="454.999,-389.452 287.001,-389.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="454.999,-389.452 454.999,-357.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="454.999,-389.452 458.999,-393.452 "/>
|
|
||||||
<text text-anchor="middle" x="373" y="-369.452" font-family="Times,serif" font-size="20.00">preprocess_deskew</text>
|
|
||||||
</g>
|
|
||||||
<!-- t2->t3 -->
|
|
||||||
<g id="edge3" class="edge"><title>t2->t3</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M348.691,-415.346C351.337,-411.112 354.23,-406.485 357.066,-401.946"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="360.042,-403.788 362.374,-393.453 354.106,-400.078 360.042,-403.788"/>
|
|
||||||
</g>
|
|
||||||
<!-- t6 -->
|
|
||||||
<g id="node6" class="node"><title>t6</title>
|
|
||||||
<polygon fill="#efa03b" stroke="black" points="664.375,-277.452 477.625,-277.452 473.625,-273.452 473.625,-241.452 660.375,-241.452 664.375,-245.452 664.375,-277.452"/>
|
|
||||||
<polyline fill="none" stroke="black" points="660.375,-273.452 473.625,-273.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="660.375,-273.452 660.375,-241.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="660.375,-273.452 664.375,-277.452 "/>
|
|
||||||
<text text-anchor="middle" x="569" y="-253.452" font-family="Times,serif" font-size="20.00">select_image_for_pdf</text>
|
|
||||||
</g>
|
|
||||||
<!-- t2->t6 -->
|
|
||||||
<g id="edge7" class="edge"><title>t2->t6</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M423.032,-415.374C438.812,-409.925 454.543,-402.785 468,-393.452 508.201,-365.571 539.287,-316.662 555.808,-286.558"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="559.024,-287.965 560.653,-277.496 552.851,-284.664 559.024,-287.965"/>
|
|
||||||
</g>
|
|
||||||
<!-- t4 -->
|
|
||||||
<g id="node5" class="node"><title>t4</title>
|
|
||||||
<polygon fill="#efa03b" stroke="black" points="449.705,-335.452 300.295,-335.452 296.295,-331.452 296.295,-299.452 445.705,-299.452 449.705,-303.452 449.705,-335.452"/>
|
|
||||||
<polyline fill="none" stroke="black" points="445.705,-331.452 296.295,-331.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="445.705,-331.452 445.705,-299.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="445.705,-331.452 449.705,-335.452 "/>
|
|
||||||
<text text-anchor="middle" x="373" y="-311.452" font-family="Times,serif" font-size="20.00">preprocess_clean</text>
|
|
||||||
</g>
|
|
||||||
<!-- t3->t4 -->
|
|
||||||
<g id="edge4" class="edge"><title>t3->t4</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M373,-357.346C373,-353.655 373,-349.665 373,-345.695"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="376.5,-345.453 373,-335.453 369.5,-345.453 376.5,-345.453"/>
|
|
||||||
</g>
|
|
||||||
<!-- t3->t6 -->
|
|
||||||
<g id="edge6" class="edge"><title>t3->t6</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M414.399,-357.412C428.785,-351.039 444.852,-343.405 459,-335.452 466.909,-331.007 506.49,-303.805 535.921,-283.431"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="538.073,-286.198 544.299,-277.625 534.086,-280.444 538.073,-286.198"/>
|
|
||||||
</g>
|
|
||||||
<!-- t4->t6 -->
|
|
||||||
<g id="edge5" class="edge"><title>t4->t6</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M432.605,-299.422C453.701,-293.395 477.617,-286.561 499.465,-280.319"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="500.687,-283.61 509.341,-277.497 498.764,-276.879 500.687,-283.61"/>
|
|
||||||
</g>
|
|
||||||
<!-- t5 -->
|
|
||||||
<g id="node7" class="node"><title>t5</title>
|
|
||||||
<polygon fill="#efa03b" stroke="black" points="455.922,-277.452 294.078,-277.452 290.078,-273.452 290.078,-241.452 451.922,-241.452 455.922,-245.452 455.922,-277.452"/>
|
|
||||||
<polyline fill="none" stroke="black" points="451.922,-273.452 290.078,-273.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="451.922,-273.452 451.922,-241.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="451.922,-273.452 455.922,-277.452 "/>
|
|
||||||
<text text-anchor="middle" x="373" y="-253.452" font-family="Times,serif" font-size="20.00">ocr_tesseract_hocr</text>
|
|
||||||
</g>
|
|
||||||
<!-- t4->t5 -->
|
|
||||||
<g id="edge8" class="edge"><title>t4->t5</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M373,-299.346C373,-295.655 373,-291.665 373,-287.695"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="376.5,-287.453 373,-277.453 369.5,-287.453 376.5,-287.453"/>
|
|
||||||
</g>
|
|
||||||
<!-- t4->t9 -->
|
|
||||||
<g id="edge14" class="edge"><title>t4->t9</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M303.36,-299.422C278.157,-293.259 249.509,-286.253 223.521,-279.898"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="224.249,-276.473 213.704,-277.497 222.586,-283.273 224.249,-276.473"/>
|
|
||||||
</g>
|
|
||||||
<!-- t7 -->
|
|
||||||
<g id="node8" class="node"><title>t7</title>
|
|
||||||
<polygon fill="#efa03b" stroke="black" points="426.366,-219.452 269.634,-219.452 265.634,-215.452 265.634,-183.452 422.366,-183.452 426.366,-187.452 426.366,-219.452"/>
|
|
||||||
<polyline fill="none" stroke="black" points="422.366,-215.452 265.634,-215.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="422.366,-215.452 422.366,-183.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="422.366,-215.452 426.366,-219.452 "/>
|
|
||||||
<text text-anchor="middle" x="346" y="-195.452" font-family="Times,serif" font-size="20.00">render_hocr_page</text>
|
|
||||||
</g>
|
|
||||||
<!-- t6->t7 -->
|
|
||||||
<g id="edge9" class="edge"><title>t6->t7</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M501.184,-241.422C476.75,-235.286 448.99,-228.315 423.772,-221.982"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="424.429,-218.538 413.877,-219.497 422.724,-225.328 424.429,-218.538"/>
|
|
||||||
</g>
|
|
||||||
<!-- t8 -->
|
|
||||||
<g id="node9" class="node"><title>t8</title>
|
|
||||||
<polygon fill="#efa03b" stroke="black" points="663.742,-219.452 448.258,-219.452 444.258,-215.452 444.258,-183.452 659.742,-183.452 663.742,-187.452 663.742,-219.452"/>
|
|
||||||
<polyline fill="none" stroke="black" points="659.742,-215.452 444.258,-215.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="659.742,-215.452 659.742,-183.452 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="659.742,-215.452 663.742,-219.452 "/>
|
|
||||||
<text text-anchor="middle" x="554" y="-195.452" font-family="Times,serif" font-size="20.00">render_hocr_debug_page</text>
|
|
||||||
</g>
|
|
||||||
<!-- t6->t8 -->
|
|
||||||
<g id="edge11" class="edge"><title>t6->t8</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M564.418,-241.346C563.4,-237.546 562.298,-233.43 561.203,-229.345"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="564.522,-228.207 558.554,-219.453 557.761,-230.018 564.522,-228.207"/>
|
|
||||||
</g>
|
|
||||||
<!-- t5->t7 -->
|
|
||||||
<g id="edge10" class="edge"><title>t5->t7</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M364.752,-241.346C362.816,-237.329 360.708,-232.958 358.629,-228.645"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="361.693,-226.941 354.197,-219.453 355.387,-229.981 361.693,-226.941"/>
|
|
||||||
</g>
|
|
||||||
<!-- t5->t8 -->
|
|
||||||
<g id="edge12" class="edge"><title>t5->t8</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M428.289,-241.346C447.587,-235.375 469.416,-228.622 489.404,-222.437"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="490.53,-225.753 499.049,-219.453 488.461,-219.065 490.53,-225.753"/>
|
|
||||||
</g>
|
|
||||||
<!-- t12 -->
|
|
||||||
<g id="node13" class="node"><title>t12</title>
|
|
||||||
<polygon fill="#efa03b" stroke="black" points="473.845,-105.456 554,-78.0208 634.155,-105.456 634.08,-149.848 473.92,-149.848 473.845,-105.456"/>
|
|
||||||
<polygon fill="none" stroke="black" points="469.836,-102.581 554,-73.7729 638.164,-102.581 638.078,-153.869 469.922,-153.869 469.836,-102.581"/>
|
|
||||||
<text text-anchor="middle" x="554" y="-111.726" font-family="Times,serif" font-size="20.00">merge_pages</text>
|
|
||||||
</g>
|
|
||||||
<!-- t7->t12 -->
|
|
||||||
<g id="edge20" class="edge"><title>t7->t12</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M389.35,-183.419C409.937,-175.33 435.398,-165.326 460.092,-155.624"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="461.479,-158.839 469.507,-151.924 458.919,-152.324 461.479,-158.839"/>
|
|
||||||
</g>
|
|
||||||
<!-- t8->t12 -->
|
|
||||||
<g id="edge19" class="edge"><title>t8->t12</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M554,-183.12C554,-177.585 554,-171.177 554,-164.592"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="557.5,-164.201 554,-154.201 550.5,-164.201 557.5,-164.201"/>
|
|
||||||
</g>
|
|
||||||
<!-- t11->t12 -->
|
|
||||||
<g id="edge17" class="edge"><title>t11->t12</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M615.939,-357.13C634.767,-339.333 661.745,-309.733 673,-277.452 686.754,-238.003 694.347,-219.364 673,-183.452 666.345,-172.256 656.999,-162.875 646.44,-155.047"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="648.378,-152.132 638.152,-149.36 644.417,-157.904 648.378,-152.132"/>
|
|
||||||
</g>
|
|
||||||
<!-- t9->t12 -->
|
|
||||||
<g id="edge18" class="edge"><title>t9->t12</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M164.711,-241.235C186.394,-224.08 222.088,-198.224 257,-183.452 321.895,-155.994 399.893,-139.571 459.67,-130.141"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="460.504,-133.554 469.855,-128.574 459.439,-126.635 460.504,-133.554"/>
|
|
||||||
</g>
|
|
||||||
<!-- t10->t12 -->
|
|
||||||
<g id="edge22" class="edge"><title>t10->t12</title>
|
|
||||||
<path fill="none" stroke="gray" d="M629.379,-415.337C660.552,-396.255 703,-362.233 703,-318.452 703,-318.452 703,-318.452 703,-258.452 703,-224.066 705.581,-210.223 684,-183.452 673.976,-171.017 660.952,-160.773 647.026,-152.396"/>
|
|
||||||
<polygon fill="gray" stroke="gray" points="648.682,-149.312 638.258,-147.419 645.227,-155.4 648.682,-149.312"/>
|
|
||||||
</g>
|
|
||||||
<!-- t13 -->
|
|
||||||
<g id="node14" class="node"><title>t13</title>
|
|
||||||
<polygon fill="#efa03b" stroke="black" points="616.338,-52 495.662,-52 491.662,-48 491.662,-16 612.338,-16 616.338,-20 616.338,-52"/>
|
|
||||||
<polyline fill="none" stroke="black" points="612.338,-48 491.662,-48 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="612.338,-48 612.338,-16 "/>
|
|
||||||
<polyline fill="none" stroke="black" points="612.338,-48 616.338,-52 "/>
|
|
||||||
<text text-anchor="middle" x="554" y="-28" font-family="Times,serif" font-size="20.00">validate_pdfa</text>
|
|
||||||
</g>
|
|
||||||
<!-- t12->t13 -->
|
|
||||||
<g id="edge21" class="edge"><title>t12->t13</title>
|
|
||||||
<path fill="none" stroke="#0044a0" d="M554,-73.9482C554,-69.9654 554,-66.007 554,-62.2247"/>
|
|
||||||
<polygon fill="#0044a0" stroke="#0044a0" points="557.5,-62.1573 554,-52.1573 550.5,-62.1574 557.5,-62.1573"/>
|
|
||||||
</g>
|
|
||||||
</g>
|
|
||||||
</svg>
|
|
||||||
|
Before Width: | Height: | Size: 16 KiB |
@@ -1,214 +0,0 @@
|
|||||||
#!/usr/bin/env python3
|
|
||||||
# © 2015 James R. Barlow: github.com/jbarlow83
|
|
||||||
|
|
||||||
from setuptools import setup
|
|
||||||
from subprocess import Popen, STDOUT, check_output, CalledProcessError
|
|
||||||
from string import Template
|
|
||||||
import re
|
|
||||||
import sys
|
|
||||||
|
|
||||||
|
|
||||||
missing_program = '''
|
|
||||||
The program '{program}' could not be executed or was not found on your
|
|
||||||
system PATH.
|
|
||||||
'''
|
|
||||||
|
|
||||||
unknown_version = '''
|
|
||||||
OCRmyPDF requires '{program}' {need_version} or higher. Your system has
|
|
||||||
'{program}' but we cannot tell what version is installed. Contact the
|
|
||||||
package maintainer.
|
|
||||||
'''
|
|
||||||
|
|
||||||
old_version = '''
|
|
||||||
OCRmyPDF requires '{program}' {need_version} or higher. Your system appears
|
|
||||||
to have {found_version}. Please update this program.
|
|
||||||
'''
|
|
||||||
|
|
||||||
okay_its_optional = '''
|
|
||||||
This program is OPTIONAL, so installation of OCRmyPDF can proceed, but
|
|
||||||
some functionality may be missing.
|
|
||||||
'''
|
|
||||||
|
|
||||||
not_okay_its_required = '''
|
|
||||||
This program is REQUIRED for OCRmyPDF to work. Installation will abort.
|
|
||||||
'''
|
|
||||||
|
|
||||||
osx_install_advice = '''
|
|
||||||
If you have homebrew installed, try these command to install the missing
|
|
||||||
packages:
|
|
||||||
brew update
|
|
||||||
brew upgrade
|
|
||||||
brew install {package}
|
|
||||||
'''
|
|
||||||
|
|
||||||
linux_install_advice = '''
|
|
||||||
On systems with the aptitude package manager (Debian, Ubuntu), try these
|
|
||||||
commands:
|
|
||||||
sudo apt-get update
|
|
||||||
sudo apt-get install {package}
|
|
||||||
|
|
||||||
On RPM-based systems (Red Hat, Fedora), search for instructions on
|
|
||||||
installing the RPM for {package}.
|
|
||||||
'''
|
|
||||||
|
|
||||||
|
|
||||||
def _error_trailer(program, package, optional):
|
|
||||||
if program == 'java':
|
|
||||||
return # You're fucked
|
|
||||||
|
|
||||||
if optional:
|
|
||||||
print(okay_its_optional.format(**locals()), file=sys.stderr)
|
|
||||||
else:
|
|
||||||
print(not_okay_its_required.format(**locals()), file=sys.stderr)
|
|
||||||
if sys.platform.startswith('darwin'):
|
|
||||||
print(osx_install_advice.format(**locals()), file=sys.stderr)
|
|
||||||
elif sys.platform.startswith('linux'):
|
|
||||||
print(linux_install_advice.format(**locals()), file=sys.stderr)
|
|
||||||
|
|
||||||
|
|
||||||
def error_missing_program(
|
|
||||||
program,
|
|
||||||
package,
|
|
||||||
optional
|
|
||||||
):
|
|
||||||
print(missing_program.format(**locals()), file=sys.stderr)
|
|
||||||
_error_trailer(**locals())
|
|
||||||
|
|
||||||
|
|
||||||
def error_unknown_version(
|
|
||||||
program,
|
|
||||||
package,
|
|
||||||
optional
|
|
||||||
):
|
|
||||||
print(unknown_version.format(**locals()), file=sys.stderr)
|
|
||||||
_error_trailer(**locals())
|
|
||||||
|
|
||||||
|
|
||||||
def error_old_version(
|
|
||||||
program,
|
|
||||||
package,
|
|
||||||
optional,
|
|
||||||
need_version
|
|
||||||
):
|
|
||||||
print(old_version.format(**locals()), file=sys.stderr)
|
|
||||||
_error_trailer(**locals())
|
|
||||||
|
|
||||||
|
|
||||||
def check_external_program(
|
|
||||||
program,
|
|
||||||
minimum_version,
|
|
||||||
package,
|
|
||||||
version_check_args=['--version'],
|
|
||||||
version_scrape_regex=re.compile(r'(\d+\.\d+(?:\.\d+)?)'),
|
|
||||||
optional=False):
|
|
||||||
|
|
||||||
print('Checking for {program} >= {minimum_version}...'.format(
|
|
||||||
program=program, minimum_version=minimum_version))
|
|
||||||
try:
|
|
||||||
result = check_output(
|
|
||||||
[program] + version_check_args,
|
|
||||||
universal_newlines=True, stderr=STDOUT)
|
|
||||||
except CalledProcessError:
|
|
||||||
error_missing_program(program, package, optional)
|
|
||||||
if not optional:
|
|
||||||
sys.exit(1)
|
|
||||||
|
|
||||||
try:
|
|
||||||
version = version_scrape_regex.search(result).group(1)
|
|
||||||
except AttributeError:
|
|
||||||
error_unknown_version(program, package, optional, minimum_version)
|
|
||||||
if not optional:
|
|
||||||
sys.exit(1)
|
|
||||||
|
|
||||||
if version < minimum_version:
|
|
||||||
error_old_version(program, package, optional, minimum_version)
|
|
||||||
|
|
||||||
print('Found {program} {version}'.format(
|
|
||||||
program=program, version=version))
|
|
||||||
|
|
||||||
command = next((arg for arg in sys.argv[1:] if not arg.startswith('-')), '')
|
|
||||||
|
|
||||||
if command.startswith('install') or \
|
|
||||||
command in ['check', 'test', 'nosetests', 'easy_install', 'egg_info']:
|
|
||||||
check_external_program(
|
|
||||||
program='tesseract',
|
|
||||||
minimum_version='3.02.02',
|
|
||||||
package='tesseract'
|
|
||||||
)
|
|
||||||
check_external_program(
|
|
||||||
program='gs',
|
|
||||||
minimum_version='9.14',
|
|
||||||
package='ghostscript'
|
|
||||||
)
|
|
||||||
check_external_program(
|
|
||||||
program='unpaper',
|
|
||||||
minimum_version='6.1',
|
|
||||||
package='unpaper',
|
|
||||||
optional=True
|
|
||||||
)
|
|
||||||
# Deprecated
|
|
||||||
check_external_program(
|
|
||||||
program='pdfseparate',
|
|
||||||
minimum_version='0.29.0',
|
|
||||||
package='poppler',
|
|
||||||
version_check_args=['-v']
|
|
||||||
)
|
|
||||||
check_external_program(
|
|
||||||
program='java',
|
|
||||||
minimum_version='1.5.0',
|
|
||||||
package='Java Runtime Environment',
|
|
||||||
version_check_args=['-version']
|
|
||||||
)
|
|
||||||
check_external_program(
|
|
||||||
program='mutool',
|
|
||||||
minimum_version='1.7a',
|
|
||||||
version_check_args=['-v'],
|
|
||||||
version_scrape_regex=re.compile(r'(\d+\.\d+[a-z]+)'),
|
|
||||||
package='mupdf-tools'
|
|
||||||
)
|
|
||||||
|
|
||||||
setup(
|
|
||||||
name='ocrmypdf',
|
|
||||||
version='3.0rc2',
|
|
||||||
description='OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched',
|
|
||||||
url='https://github.com/fritz-hh/OCRmyPDF',
|
|
||||||
author='J. R. Barlow',
|
|
||||||
author_email='jim@purplerock.ca',
|
|
||||||
license='Public Domain',
|
|
||||||
packages=['ocrmypdf'],
|
|
||||||
keywords=['PDF', 'OCR', 'optical character recognition', 'PDF/A', 'scanning'],
|
|
||||||
classifiers=[
|
|
||||||
"Programming Language :: Python :: 3",
|
|
||||||
"Development Status :: 4 - Beta",
|
|
||||||
"Environment :: Console",
|
|
||||||
"Intended Audience :: End Users/Desktop",
|
|
||||||
"Intended Audience :: Science/Research",
|
|
||||||
"Intended Audience :: System Administrators",
|
|
||||||
"License :: Public Domain",
|
|
||||||
"Operating System :: MacOS :: MacOS X",
|
|
||||||
"Operating System :: POSIX",
|
|
||||||
"Operating System :: POSIX :: BSD",
|
|
||||||
"Operating System :: POSIX :: Linux",
|
|
||||||
"Topic :: Scientific/Engineering :: Image Recognition",
|
|
||||||
"Topic :: Text Processing :: Indexing",
|
|
||||||
"Topic :: Text Processing :: Linguistic",
|
|
||||||
],
|
|
||||||
install_requires=[
|
|
||||||
'ruffus>=2.6.3',
|
|
||||||
'Pillow>=2.7.0',
|
|
||||||
'lxml>=3.4.2',
|
|
||||||
'reportlab>=3.1.44',
|
|
||||||
'PyPDF2>=1.25.1'
|
|
||||||
],
|
|
||||||
entry_points={
|
|
||||||
'console_scripts': [
|
|
||||||
'ocrmypdf = ocrmypdf.main:run_pipeline'
|
|
||||||
],
|
|
||||||
},
|
|
||||||
eager_resources=[
|
|
||||||
'ocrmypdf/jhove/bin/*.jar',
|
|
||||||
'ocrmypdf/jhove/conf/*.conf',
|
|
||||||
'ocrmypdf/jhove/lib/*.jar'
|
|
||||||
],
|
|
||||||
include_package_data=True,
|
|
||||||
zip_safe=False)
|
|
||||||
@@ -0,0 +1,53 @@
|
|||||||
|
#! /bin/bash
|
||||||
|
|
||||||
|
set -x
|
||||||
|
set -e
|
||||||
|
|
||||||
|
chmod +x OCRmyPDF*.AppImage
|
||||||
|
|
||||||
|
# run OCRmyPDF to test if the AppImage can ocr a test file
|
||||||
|
run_appimage()
|
||||||
|
{
|
||||||
|
echo ""
|
||||||
|
./OCRmyPDF*.AppImage --help
|
||||||
|
echo ""
|
||||||
|
./OCRmyPDF*.AppImage --list-programs
|
||||||
|
echo ""
|
||||||
|
./OCRmyPDF*.AppImage --list-licenses
|
||||||
|
echo ""
|
||||||
|
./OCRmyPDF*.AppImage ocrmypdf -l deu -s -d --jbig2-lossy --optimize 1 "$TRAVIS_BUILD_DIR"/test/test.pdf output.pdf
|
||||||
|
echo ""
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# check AppImage for common issues
|
||||||
|
run_appimagelint()
|
||||||
|
{
|
||||||
|
wget https://github.com/TheAssassin/appimagelint/releases/download/continuous/appimagelint-x86_64.AppImage
|
||||||
|
chmod +x appimagelint-x86_64.AppImage
|
||||||
|
./appimagelint-x86_64.AppImage OCRmyPDF*.AppImage
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# extract the OCRmyPDF AppImage, install pytest & test requirements and run pytest
|
||||||
|
run_pytest()
|
||||||
|
{
|
||||||
|
git clone --depth=1 --branch "v$OCRMYPDF_VERSION" https://github.com/jbarlow83/OCRmyPDF.git
|
||||||
|
./OCRmyPDF*.AppImage --appimage-extract
|
||||||
|
|
||||||
|
pushd squashfs-root
|
||||||
|
./AppRun python3 -m pip install pytest
|
||||||
|
./AppRun python3 -m pip install -r ../OCRmyPDF/requirements/test.txt
|
||||||
|
./AppRun python3 -m pytest ../OCRmyPDF -n auto
|
||||||
|
popd
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
run_appimage
|
||||||
|
|
||||||
|
run_appimagelint
|
||||||
|
|
||||||
|
# run_pytest
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
Binary file not shown.
Binary file not shown.
|
Before Width: | Height: | Size: 1.4 MiB |
@@ -1,17 +0,0 @@
|
|||||||
All test resources must come from free public domain sources for
|
|
||||||
copyright reasons.
|
|
||||||
|
|
||||||
Test files do not necessarily produce perfect (or even good) OCR
|
|
||||||
results.
|
|
||||||
|
|
||||||
+-------------------+--------------------------------------------------------------------------------+
|
|
||||||
| File | Source |
|
|
||||||
+===================+================================================================================+
|
|
||||||
| graph.pdf | Wikimedia |
|
|
||||||
+-------------------+--------------------------------------------------------------------------------+
|
|
||||||
| c02-22.pdf | Project Gutenberg: https://www.gutenberg.org/files/76/76-h/images/c02-22.jpg |
|
|
||||||
+-------------------+--------------------------------------------------------------------------------+
|
|
||||||
| LinnSequencer.jpg | Wikimedia: https://upload.wikimedia.org/wikipedia/en/b/b7/LinnSequencer_hardware_MIDI_sequencer_brochure_page_2_300dpi.jpg |
|
|
||||||
+-------------------+--------------------------------------------------------------------------------+
|
|
||||||
| congress.jpg | http://www.baxleystamps.com/litho/meiji/courts_1871.jpg |
|
|
||||||
+-------------------+--------------------------------------------------------------------------------+
|
|
||||||
File diff suppressed because one or more lines are too long
Binary file not shown.
Binary file not shown.
|
Before Width: | Height: | Size: 188 KiB |
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -1,170 +0,0 @@
|
|||||||
#!/usr/bin/env python3
|
|
||||||
# © 2015 James R. Barlow: github.com/jbarlow83
|
|
||||||
|
|
||||||
from __future__ import print_function
|
|
||||||
from subprocess import Popen, PIPE, check_output
|
|
||||||
import os
|
|
||||||
import shutil
|
|
||||||
from contextlib import suppress
|
|
||||||
import sys
|
|
||||||
from unittest.mock import patch, create_autospec
|
|
||||||
import pytest
|
|
||||||
from ocrmypdf.pageinfo import pdf_get_all_pageinfo
|
|
||||||
|
|
||||||
|
|
||||||
if sys.version_info.major < 3:
|
|
||||||
print("Requires Python 3.4+")
|
|
||||||
sys.exit(1)
|
|
||||||
|
|
||||||
TESTS_ROOT = os.path.abspath(os.path.dirname(__file__))
|
|
||||||
PROJECT_ROOT = os.path.dirname(TESTS_ROOT)
|
|
||||||
OCRMYPDF = os.path.join(PROJECT_ROOT, 'OCRmyPDF.sh')
|
|
||||||
TEST_RESOURCES = os.path.join(PROJECT_ROOT, 'tests', 'resources')
|
|
||||||
TEST_OUTPUT = os.path.join(PROJECT_ROOT, 'tests', 'output')
|
|
||||||
|
|
||||||
|
|
||||||
def setup_module():
|
|
||||||
with suppress(FileNotFoundError):
|
|
||||||
shutil.rmtree(TEST_OUTPUT)
|
|
||||||
with suppress(FileExistsError):
|
|
||||||
os.mkdir(TEST_OUTPUT)
|
|
||||||
|
|
||||||
|
|
||||||
def run_ocrmypdf_sh(input_file, output_file, *args):
|
|
||||||
sh_args = ['sh', OCRMYPDF] + list(args) + [input_file, output_file]
|
|
||||||
sh = Popen(
|
|
||||||
sh_args, close_fds=True, stdout=PIPE, stderr=PIPE,
|
|
||||||
universal_newlines=True)
|
|
||||||
out, err = sh.communicate()
|
|
||||||
return sh, out, err
|
|
||||||
|
|
||||||
|
|
||||||
def check_ocrmypdf(input_basename, output_basename, *args):
|
|
||||||
input_file = os.path.join(TEST_RESOURCES, input_basename)
|
|
||||||
output_file = os.path.join(TEST_OUTPUT, output_basename)
|
|
||||||
|
|
||||||
sh, _, err = run_ocrmypdf_sh(input_file, output_file, *args)
|
|
||||||
assert sh.returncode == 0, err
|
|
||||||
assert os.path.exists(output_file), "Output file not created"
|
|
||||||
assert os.stat(output_file).st_size > 100, "PDF too small or empty"
|
|
||||||
return output_file
|
|
||||||
|
|
||||||
|
|
||||||
def test_quick():
|
|
||||||
check_ocrmypdf('c02-22.pdf', 'test_quick.pdf')
|
|
||||||
|
|
||||||
|
|
||||||
def test_deskew():
|
|
||||||
# Run with deskew
|
|
||||||
deskewed_pdf = check_ocrmypdf('skew.pdf', 'test_deskew.pdf', '-d')
|
|
||||||
|
|
||||||
# Now render as an image again and use Leptonica to find the skew angle
|
|
||||||
# to confirm that it was deskewed
|
|
||||||
from ocrmypdf.ghostscript import rasterize_pdf
|
|
||||||
import logging
|
|
||||||
log = logging.getLogger()
|
|
||||||
|
|
||||||
deskewed_png = os.path.join(TEST_OUTPUT, 'deskewed.png')
|
|
||||||
|
|
||||||
rasterize_pdf(
|
|
||||||
deskewed_pdf,
|
|
||||||
deskewed_png,
|
|
||||||
xres=150,
|
|
||||||
yres=150,
|
|
||||||
raster_device='pngmono',
|
|
||||||
log=log)
|
|
||||||
|
|
||||||
from ocrmypdf.leptonica import pixRead, pixDestroy, pixFindSkew
|
|
||||||
pix = pixRead(deskewed_png)
|
|
||||||
skew_angle, skew_confidence = pixFindSkew(pix)
|
|
||||||
pix = pixDestroy(pix)
|
|
||||||
|
|
||||||
print(skew_angle)
|
|
||||||
assert -0.5 < skew_angle < 0.5, "Deskewing failed"
|
|
||||||
|
|
||||||
|
|
||||||
def test_clean():
|
|
||||||
check_ocrmypdf('skew.pdf', 'test_clean.pdf', '-c')
|
|
||||||
|
|
||||||
|
|
||||||
def test_metadata():
|
|
||||||
pdf = check_ocrmypdf(
|
|
||||||
'c02-22.pdf', 'test_metadata.pdf',
|
|
||||||
'--title', 'Du siehst den Wald vor lauter Bäumen nicht.',
|
|
||||||
'--author', '孔子',
|
|
||||||
'--subject', 'U+1030C is: 𐌌')
|
|
||||||
|
|
||||||
out_pdfinfo = check_output(['pdfinfo', pdf], universal_newlines=True)
|
|
||||||
lines_pdfinfo = out_pdfinfo.splitlines()
|
|
||||||
pdfinfo = {}
|
|
||||||
for line in lines_pdfinfo:
|
|
||||||
k, v = line.strip().split(':', maxsplit=1)
|
|
||||||
pdfinfo[k.strip()] = v.strip()
|
|
||||||
|
|
||||||
assert pdfinfo['Title'] == 'Du siehst den Wald vor lauter Bäumen nicht.'
|
|
||||||
assert pdfinfo['Author'] == '孔子'
|
|
||||||
assert pdfinfo['Subject'] == 'U+1030C is: 𐌌'
|
|
||||||
assert pdfinfo.get('Keywords', '') == ''
|
|
||||||
|
|
||||||
|
|
||||||
def check_oversample(renderer):
|
|
||||||
oversampled_pdf = check_ocrmypdf(
|
|
||||||
'skew.pdf', 'test_oversample_%s.pdf' % renderer, '--oversample', '300',
|
|
||||||
'--pdf-renderer', renderer)
|
|
||||||
|
|
||||||
pdfinfo = pdf_get_all_pageinfo(oversampled_pdf)
|
|
||||||
|
|
||||||
print(pdfinfo[0]['xres'])
|
|
||||||
assert abs(pdfinfo[0]['xres'] - 300) < 1
|
|
||||||
|
|
||||||
|
|
||||||
def test_oversample():
|
|
||||||
yield check_oversample, 'hocr'
|
|
||||||
yield check_oversample, 'tesseract'
|
|
||||||
|
|
||||||
|
|
||||||
def test_repeat_ocr():
|
|
||||||
sh, _, _ = run_ocrmypdf_sh('graph_ocred.pdf', 'wontwork.pdf')
|
|
||||||
assert sh.returncode != 0
|
|
||||||
|
|
||||||
|
|
||||||
def test_force_ocr():
|
|
||||||
out = check_ocrmypdf('graph_ocred.pdf', 'test_force.pdf', '-f')
|
|
||||||
pdfinfo = pdf_get_all_pageinfo(out)
|
|
||||||
assert pdfinfo[0]['has_text']
|
|
||||||
|
|
||||||
|
|
||||||
def test_skip_ocr():
|
|
||||||
check_ocrmypdf('graph_ocred.pdf', 'test_skip.pdf', '-s')
|
|
||||||
|
|
||||||
|
|
||||||
def check_ocr_timeout(renderer):
|
|
||||||
out = check_ocrmypdf('skew.pdf', 'test_timeout_%s.pdf' % renderer,
|
|
||||||
'--tesseract-timeout', '1.0')
|
|
||||||
pdfinfo = pdf_get_all_pageinfo(out)
|
|
||||||
assert pdfinfo[0]['has_text'] == False
|
|
||||||
|
|
||||||
|
|
||||||
def test_ocr_timeout():
|
|
||||||
yield check_ocr_timeout, 'hocr'
|
|
||||||
yield check_ocr_timeout, 'tesseract'
|
|
||||||
|
|
||||||
|
|
||||||
def test_skip_big():
|
|
||||||
out = check_ocrmypdf('enormous.pdf', 'test_enormous.pdf',
|
|
||||||
'--skip-big', '10')
|
|
||||||
pdfinfo = pdf_get_all_pageinfo(out)
|
|
||||||
assert pdfinfo[0]['has_text'] == False
|
|
||||||
|
|
||||||
|
|
||||||
def check_maximum_options(renderer):
|
|
||||||
check_ocrmypdf(
|
|
||||||
'multipage.pdf', 'test_multipage%s.pdf' % renderer,
|
|
||||||
'-d', '-c', '-i', '-g', '-f', '-k', '--oversample', '300',
|
|
||||||
'--skip-big', '10', '--title', 'Too Many Weird Files',
|
|
||||||
'--author', 'py.test', '--pdf-renderer', renderer)
|
|
||||||
|
|
||||||
|
|
||||||
def test_maximum_options():
|
|
||||||
yield check_maximum_options, 'hocr'
|
|
||||||
yield check_maximum_options, 'tesseract'
|
|
||||||
Reference in New Issue
Block a user