docs: document Tesseract configs/ requirement (closes #1567)
When TESSDATA_PREFIX points at a hand-assembled tessdata folder lacking the configs/ subdirectory (e.g. files pulled from tessdata_best), Tesseract prints "read_params_file: Can't open hocr/txt" and produces no output. The runtime now surfaces a clear error (v17.5.0); add the matching documentation: a new errors.md entry and a note on the TESSDATA_PREFIX docs.
This commit is contained in:
@@ -201,6 +201,13 @@ include:
|
|||||||
Overrides the path to Tesseract's data files. This can allow
|
Overrides the path to Tesseract's data files. This can allow
|
||||||
simultaneous installation of the "best" and "fast" training data
|
simultaneous installation of the "best" and "fast" training data
|
||||||
sets. OCRmyPDF does not manage this environment variable.
|
sets. OCRmyPDF does not manage this environment variable.
|
||||||
|
|
||||||
|
If you point ``TESSDATA_PREFIX`` at a hand-assembled ``tessdata``
|
||||||
|
folder (for example, individual ``.traineddata`` files downloaded
|
||||||
|
from tessdata_best), make sure it also contains the ``configs/``
|
||||||
|
subdirectory with the ``hocr`` and ``txt`` files. OCRmyPDF requires
|
||||||
|
these; without them Tesseract produces no output. See
|
||||||
|
:ref:`Tesseract cannot open its config file <tesseract-config-missing>`.
|
||||||
```
|
```
|
||||||
|
|
||||||
```{eval-rst}
|
```{eval-rst}
|
||||||
|
|||||||
@@ -49,3 +49,31 @@ pdftk input.pdf cat output output.pdf
|
|||||||
|
|
||||||
Sometimes Acrobat can repair PDFs with its [Preflight
|
Sometimes Acrobat can repair PDFs with its [Preflight
|
||||||
tool](https://helpx.adobe.com/acrobat/using/correcting-problem-areas-preflight-tool.html).
|
tool](https://helpx.adobe.com/acrobat/using/correcting-problem-areas-preflight-tool.html).
|
||||||
|
|
||||||
|
(tesseract-config-missing)=
|
||||||
|
|
||||||
|
## Tesseract cannot open its config file \'hocr\' or \'txt\'
|
||||||
|
|
||||||
|
:::{code}
|
||||||
|
ERROR - Tesseract cannot open its config file 'hocr'.
|
||||||
|
:::
|
||||||
|
|
||||||
|
OCRmyPDF asks Tesseract to produce `hocr` and `txt` output. Tesseract
|
||||||
|
reads the instructions for these output formats from configuration files
|
||||||
|
named `hocr` and `txt` that live in the `configs/` subdirectory of its
|
||||||
|
`tessdata` folder. If those files are missing, Tesseract prints
|
||||||
|
`read_params_file: Can't open hocr`, exits without error, and produces no
|
||||||
|
output.
|
||||||
|
|
||||||
|
This usually happens when a `tessdata` directory was assembled by hand --
|
||||||
|
for example, by downloading individual `.traineddata` files from
|
||||||
|
[tessdata_best](https://github.com/tesseract-ocr/tessdata_best) and
|
||||||
|
pointing `TESSDATA_PREFIX` at them -- because those repositories do not
|
||||||
|
include the `configs/` directory. A complete Tesseract installation from
|
||||||
|
your operating system\'s package manager includes it.
|
||||||
|
|
||||||
|
To fix this, ensure the `configs/hocr` and `configs/txt` files exist in
|
||||||
|
the `tessdata` directory that Tesseract is using. Copying the `configs/`
|
||||||
|
directory from a full Tesseract installation is sufficient. See
|
||||||
|
{envvar}`TESSDATA_PREFIX` for more on selecting an alternate `tessdata`
|
||||||
|
folder.
|
||||||
|
|||||||
Reference in New Issue
Block a user