James R. Barlow
75dcb90621
Niceties: when environment variable overload is used clarify we're not checking the PATH
2018-01-10 16:33:03 -08:00
James R. Barlow
882fc2257c
Add --max-image-mpixels argument to support Pillow 5.0
2018-01-10 15:43:59 -08:00
James R. Barlow
6bf1f970a0
Fix some parameter validation for --output-type pdfa-1 and pdfa-2
2018-01-10 11:50:08 -08:00
James R. Barlow
7d451f101f
Detect old versions of Ghostscript and warn about them ( #208 )
2018-01-10 11:47:39 -08:00
James R. Barlow
91b42cbfa8
Fix issue in sandwich renderer when skipping OCR on a rotated and deskewed page
...
If OCR is skipped due to --tesseract-timeout or similar, and the skip page is rotated with /Rotate, and the skip page was deskewed or had other image processing, then the skip page was created with the wrong dimensions causing the output page to be cropped.
2018-01-09 00:17:53 -08:00
James R. Barlow
a40689a0ff
tesseract: handle return of bytes properly in error cases
2017-11-29 14:35:26 -08:00
James R. Barlow
44a45fc3fb
Add "bad UTF8 output from Tesseract" test
2017-11-29 14:08:07 -08:00
James R. Barlow
ec4bb5359a
Read tesseract's output as binary to avoid UnicodeDecodeErrors if it messes up
2017-11-29 13:44:40 -08:00
James R. Barlow
d2217632df
Rename _verify_python3_env
2017-11-29 13:43:18 -08:00
James R. Barlow
2cc044feed
Move qpdf complaint to after options checking so that it won't break ocrmypdf --version
2017-11-29 13:42:55 -08:00
James R. Barlow
c5a1d22e81
That fixed it. Complain about old versions of qpdf now
2017-11-29 12:53:34 -08:00
James R. Barlow
d472860e3b
Try to diagnose travis-only failure of qpdf test
2017-11-29 11:41:06 -08:00
James R. Barlow
0b6af8d965
Clarify hocrtransform license/copyright
2017-11-27 13:41:46 -08:00
James R. Barlow
56614fcaa4
Add support and tests for handling page count > ulimit - fixes issue #181
2017-11-27 00:32:35 -08:00
James R. Barlow
2040ae4856
Fix issue #200 , uncommon but valid decimal syntax treated as error
...
Also replace check_output() calls with run() in qpdf.py
2017-11-26 22:52:43 -08:00
James R. Barlow
31d0eaac8e
Ensure intermediate metadata holder PDF has same version as its input file
...
While not known to cause problems, the absence of this could confuse parsers
2017-11-24 00:09:13 -08:00
James R. Barlow
40aa82ab41
Check that the locale is sane before allowing OCR to proceed
2017-11-16 17:18:02 -08:00
James R. Barlow
3ef766bb93
Remove bare 'except:'
2017-11-01 01:52:13 -07:00
James R. Barlow
44b5a18462
Declare our __version__ properly
2017-11-01 01:49:51 -07:00
James R. Barlow
fcbf34a4d3
Refactor obtaining version from subprocesses
...
Issue #196 raised the need to deal with linker warnings on --version.
2017-11-01 01:44:36 -07:00
James R. Barlow
c7b8b6e18b
Fix issue #194 - --sidecar creates blank txt file
2017-10-26 18:15:31 -07:00
James R. Barlow
cc5578488a
Remove workaround for OSD crash and use explicit -l osd
...
Suggested by amitdo in https://github.com/tesseract-ocr/tesseract/issues/1167
2017-10-12 12:45:18 -07:00
James R. Barlow
4b7135f0e5
Add option to produce PDF/A-1B
2017-10-11 14:32:58 -07:00
James R. Barlow
51defa6d66
Import cleanup and some pylint fixes
2017-10-11 12:58:35 -07:00
James R. Barlow
7d73098d6e
ghostscript.py: fix missing imports
2017-10-11 12:39:16 -07:00
James R. Barlow
aa8f534b45
pipeline: fix variable not defined in __init__
2017-10-11 11:17:33 -07:00
James R. Barlow
37dc03eec6
Add workaround for tess4 form feed behavior change
2017-10-10 23:33:54 -07:00
James R. Barlow
5372656893
Don't say tess4 support is experimental - it's pretty good now
2017-10-09 16:17:42 -07:00
James R. Barlow
87c2ed8b27
Improve clarity of --pdf-renderer=tesseract deprecation warning
2017-09-12 14:34:53 -07:00
James R. Barlow
1467d118ab
Add more leptonica functions
2017-09-06 00:27:02 -07:00
James R. Barlow
82ebd8ef1a
Fix missing error message about trying to use sandwich on old tesseract
2017-09-01 12:50:36 -07:00
James R. Barlow
9bb42c0229
Wrong error type used for missing language
2017-08-24 01:07:23 -07:00
James R. Barlow
93a954ef9f
Fix missing import for Py3.5
2017-07-26 23:40:01 -07:00
James R. Barlow
0b012697e5
Whitelist the Latin-1 languages that work with HOCR
...
Omitted French because the rare 'oe' and 'ÿ' glyphs are not in Latin-1.
Basically steer people away from HOCR renderer but avoid a potential
disruptive behavior change.
2017-07-26 21:03:18 -07:00
James R. Barlow
58e357c992
Report location of attempted output_file that fails to write
2017-07-22 17:49:56 -07:00
James R. Barlow
71fbad83ad
Fix py3.5 test
2017-07-21 17:01:06 -07:00
James R. Barlow
1aa34f5d2e
Make some interfaces accepting of both str-paths and Path objects
2017-07-21 13:28:30 -07:00
James R. Barlow
dfa1d88ce9
Fix missing user_words/user_patterns from textonly_pdf case
2017-07-20 17:14:04 -07:00
James R. Barlow
dd38519f07
Merge branch 'feature/user-words' into develop
...
# Conflicts:
# ocrmypdf/exec/tesseract.py
2017-07-20 16:25:20 -07:00
James R. Barlow
cd1a99a0de
Refactor int(os.path.basename(s)[0:6]) -> page_number(s)
2017-06-26 13:29:40 -07:00
James R. Barlow
48e3b267fc
Accept PDFs with whitespace ahead of %PDF marker
...
Noticed in @aagahi 's fork
2017-06-26 13:17:47 -07:00
James R. Barlow
2c24f67deb
Rename “tess4” renderer to “sandwich” and make it default in Tess 3.05.01
...
Tesseract 3.05.01 backported the textonly_pdf=1 which allows the use
of this superior PDF renderer prior to 4.00 alpha. This means that
the tess4 name is no longer accurate, so call it a sandwich because of
its merge-preserve characteristic. Preserve the tess4 name. Fix the
documentation and tests to reflect this.
Make it the default, because it’s better. It does not have the issues
the “tesseract” renderer does prior to Tess 3.05.00 with rendering
PDFs that Ghostscript corrupts, and it produces better output without
re-rastering.
Deprecate some old stuff to avoid the test suite growing obscenely
large.
2017-06-13 13:09:12 -07:00
James R. Barlow
3232643809
Support “textonly PDF” renderer in Tesseract 3.05.01
2017-06-13 10:18:08 -07:00
James R. Barlow
d3c54fbbde
For —rotate-pages, rasterize preview at half DPI instead of 200 DPI
...
Ensures that time is not wasted on previews at higher resolution than
the input as was sometimes the case
2017-05-29 13:01:18 -07:00
James R. Barlow
1d57bcc99e
Fix Ghostscript rasterizing of UserUnit pages and related sizing issues
2017-05-29 12:14:10 -07:00
James R. Barlow
facdd13879
Ghostscript: refactor image output resizing
2017-05-29 11:42:27 -07:00
James R. Barlow
6e891f91d3
ghostscript, qpdf: Restore API backward compatibility
2017-05-29 11:13:06 -07:00
James R. Barlow
9b50ede977
Partially solve ghostscript rasterize_pdf producing wrong file size
...
Kludge. Assumes JPEG for now. Messy.
2017-05-25 01:17:43 -07:00
James R. Barlow
82cf010333
Error out if trying to produce PDF/A >200” due to Ghostscript limitation
2017-05-25 00:07:29 -07:00
James R. Barlow
6ff6c8614f
—output-type=pdf now outputs /UserUnit PDFs at the correct size
...
This currently distorts the output size because Tesseract assumes it
knows the DPI better than we do.
Does not work for Ghostscript, because it emerges that Ghostscript
honors /UserUnit for rasterizing but not in pdfwrite (resolve/wontfix).
https://bugs.ghostscript.com/show_bug.cgi?id=690781
Ghostscript’s output would need to be patched in a PDF/A safe way for
this to work. Temporary route may be to block Ghostscript if
/UserUnit.
2017-05-24 23:26:07 -07:00