Tidy docs
This commit is contained in:
+26
-7
@@ -5,8 +5,8 @@ Please always read this file before installing the package
|
||||
|
||||
Download software here: https://github.com/fritz-hh/OCRmyPDF/tags
|
||||
|
||||
v3.0-rc2:
|
||||
=========
|
||||
v3.0:
|
||||
=====
|
||||
|
||||
New features
|
||||
------------
|
||||
@@ -34,7 +34,8 @@ Changes
|
||||
- Now uses Ghostscript 9.14's improved color conversion model
|
||||
- All "tasks" in the pipeline can be executed in parallel on any
|
||||
available CPUs, increasing performance
|
||||
- The ``-o DPI`` argument has been phased out, in favor of ``--oversample DPI``
|
||||
- The ``-o DPI`` argument has been phased out, in favor of ``--oversample DPI``, in
|
||||
case we need ``-o OUTPUTFILE`` in the future
|
||||
- Removed several dependencies, so it's easier to install. We no
|
||||
longer use:
|
||||
|
||||
@@ -45,19 +46,36 @@ Changes
|
||||
- MuPDF_ tools
|
||||
- shell scripts
|
||||
|
||||
- Some new external dependencies are required:
|
||||
- Some new external dependencies are required or optional, compared to v2.x:
|
||||
|
||||
- Ghostscript 9.14+
|
||||
- qpdf 5.0.0+
|
||||
- qpdf_ 5.0.0+
|
||||
- Unpaper_ 6.1 (optional)
|
||||
- some automatically managed Python dependencies
|
||||
- some automatically managed Python packages
|
||||
|
||||
.. _ruffus: http://www.ruffus.org.uk/index.html
|
||||
.. _parallel: https://www.gnu.org/software/parallel/
|
||||
.. _ImageMagick: http://www.imagemagick.org/script/index.php
|
||||
.. _MuPDF: http://mupdf.com/docs/
|
||||
.. _qpdf: http://qpdf.sourceforge.net/
|
||||
.. _Unpaper: https://github.com/Flameeyes/unpaper
|
||||
|
||||
|
||||
Release candidates
|
||||
------------------
|
||||
|
||||
- rc4:
|
||||
|
||||
- dropped MuPDF in favour of qpdf
|
||||
- fixed some installer issues and errors in installation instructions
|
||||
- improve performance: run Ghostscript with multithreaded rendering
|
||||
- improve performance: use multiple cores by default
|
||||
- bug fix: checking for wrong exception on process timeout
|
||||
|
||||
- rc3: skipping version number intentionally to avoid confusion with Tesseract
|
||||
- rc2: first release for public testing to test-PyPI, Github
|
||||
- rc1: testing release process
|
||||
|
||||
Compatibility notes
|
||||
-------------------
|
||||
|
||||
@@ -91,7 +109,7 @@ Notes
|
||||
|
||||
|
||||
v2.2-stable (2014-09-29):
|
||||
=======
|
||||
=========================
|
||||
|
||||
New features
|
||||
------------
|
||||
@@ -115,6 +133,7 @@ Tested with
|
||||
|
||||
- Operating system: FreeBSD 9.2
|
||||
- Dependencies:
|
||||
|
||||
- parallel 20140822
|
||||
- poppler-utils 0.24.5
|
||||
- ImageMagick 6.8.9-4 2014-09-17
|
||||
|
||||
-96
@@ -1,96 +0,0 @@
|
||||
Recoding in 5 python modules
|
||||
==================
|
||||
|
||||
- Less platform dependent implementation
|
||||
- Higher versality (wrt addition of new intput / output file types)
|
||||
|
||||
The functionality of each module is described below:
|
||||
|
||||
Normalize inputs (inputs can be a pdf file, an image, a folder containing images)
|
||||
----------------
|
||||
|
||||
- For pdf:
|
||||
- Identify if page needs to be ocred (see -s and -f parameters)
|
||||
- If the page needs to be OCRed:
|
||||
- Extract the image corresponding to the page and save it in a tmp folder. 3 approaches to extract images:
|
||||
- extract raw image from pdf and rotate it according to pdf page rotation
|
||||
- if not possible: identify resolution and rasterize
|
||||
- if not possible: use default resolution and rasterize
|
||||
- If not:
|
||||
- Save the page AS-IS in the tmp folder that should contained the final page
|
||||
- For image(s):
|
||||
- Just copy the images with standardized name into the tmp folder containing pages to be OCRed
|
||||
|
||||
Preprocess normalized inputs (perform jobs in parallel)
|
||||
----------------------------
|
||||
|
||||
- Orientation (if requested by user)
|
||||
- Correct orientation
|
||||
- Skew angle (if requested by user)
|
||||
- Correct skew angle
|
||||
- Cleaning (if requested by user)
|
||||
- Clean image
|
||||
|
||||
Perform OCR (perform jobs in parallel)
|
||||
-----------
|
||||
|
||||
- Perform OCR and save resulting hocr file for each respective page (perfom jobs in parallel)
|
||||
|
||||
Generate output for each page
|
||||
-----------------------------
|
||||
|
||||
- For pdf (if output file has a "pdf" extension):
|
||||
- Generate pdf pages from hocr files (note: pdf pages can already exist if OCR has been skipped for them)
|
||||
- For txt
|
||||
- generate txt file for each page (containing txt located into hocr file)
|
||||
|
||||
Build final output
|
||||
------------------
|
||||
|
||||
- For pdf:
|
||||
- Concatenate pdf pages
|
||||
- Convert to pdf/1-a
|
||||
- Verify conformity to pdf/1-a
|
||||
- For txt
|
||||
- Concatenate all txt files into the final output txt file
|
||||
|
||||
|
||||
|
||||
Tmp folder structure
|
||||
=========================
|
||||
|
||||
- tmp_xxxxx/
|
||||
- a_raw_images (either from images or extracted from pdf file)
|
||||
- b_preprocessed_images (after deswing anf cleaning)
|
||||
- c_ocr_out
|
||||
- d_output_pages (one file per page (at first 1 pdf file per page. Later on other formats might be supported)
|
||||
- e_output_final (concatenate output pages and conversion into PDF/1-a standard, Later on support other formats)
|
||||
|
||||
ocrmypdf arguments
|
||||
==================
|
||||
|
||||
ocrmypdf [-h] [-v] [-k] [-g] [-o dpi] [-f|-s] [-r] [-d] [-c] [-i] [-l lan1[+lan2...]] [-C] inputpath outputfile1 [outputfile2...]
|
||||
|
||||
- Overall parameters
|
||||
- [-h] : Display this help message
|
||||
- [-v] : Increase the verbosity (this option can be used more than once) (e.g. -vvv)
|
||||
- [-k] : Do not delete the temporary files
|
||||
- [-g] : Activate debug mode (max verbosity, keep tmp files, generate debug pages)
|
||||
- Normalization parameters:
|
||||
- [-o dpi] : If page resolution is lower x dpi, provide OCR engine with an oversampled image. (Can improve OCR results)
|
||||
- [-f] : Force to OCR the whole document, even if some page already contain font data (only for pdf inputs)
|
||||
- [-s] : If pages contain font data, do not OCR that page, but include the page (as is) in the final output (only for pdf inputs)
|
||||
- Prepocessing parameters:
|
||||
- [-r] : Correct orientation
|
||||
- [-d] : Deskew each page
|
||||
- [-c] : Clean each page
|
||||
- [-i] : Incorporate cleaned image in final output
|
||||
- OCR parameters:
|
||||
- [-l lan1[+lan2...]] : Document language(s). Multiple languages may be specified, separated by '+' characters.
|
||||
- [-C cfg] : Pass an additional cofg file to the tesseract OCR engine. (this option can be used more than once)
|
||||
- output generation parameters:
|
||||
- None by now
|
||||
- input files:
|
||||
- inputpath : path to image, pdf file or folder to be processed
|
||||
- output files:
|
||||
- outputfile1 [outputfile2 ...] : *.pdf file or *.txt file to be generated (argumenst can be repeated if both pdf and txt file should be generated
|
||||
Reference in New Issue
Block a user