James R. Barlow
|
5cc23dbf24
|
pdfinfo: more robustness
|
2018-07-04 17:12:30 -07:00 |
|
James R. Barlow
|
216d60ea2c
|
pdfinfo: improve the regex
|
2018-07-04 00:59:32 -07:00 |
|
James R. Barlow
|
8b0496d35e
|
Fix invalid XML characters choking parser
|
2018-07-03 22:51:59 -07:00 |
|
James R. Barlow
|
b0dbaeafc5
|
Cleanup unused imports
|
2018-06-23 01:47:53 -07:00 |
|
James R. Barlow
|
2530d1791b
|
Fix several pylint errors and warnings
|
2018-06-23 00:54:22 -07:00 |
|
James R. Barlow
|
76e7e8dbbb
|
Replace several uses of str(path) with fspath(path)
Helps make it more explicit. Did not do this to tests because use of paths
is more involved there.
|
2018-06-22 21:00:47 -07:00 |
|
James R. Barlow
|
324598e992
|
Remove helpers.universal_open()
This helper function only had a single usage, this was always an awkward
way to support Python 3.5 that I'd forget to use.
|
2018-06-22 17:56:20 -07:00 |
|
James R. Barlow
|
8c84c515b6
|
Use Ghostscript for text region detection
Ghostscript txtwrite seems to be quite effective at the task.
Eliminates dependency on fitz
|
2018-06-13 00:58:09 -07:00 |
|
James R. Barlow
|
1dfbbdebf4
|
Adjust for pikepdf API change
|
2018-06-08 22:47:56 -07:00 |
|
James R. Barlow
|
9608b22d34
|
Remove all uses of PyPDF2 except PDF/A check
Leave PDF/A check alone for now, since pikepdf has no equivalent.
|
2018-05-26 02:07:18 -07:00 |
|
James R. Barlow
|
8ba4968c48
|
pdfinfo: more robustness
|
2018-05-26 01:54:25 -07:00 |
|
James R. Barlow
|
ffdd78f1a5
|
pdfinfo: Fix text_operators type not changed in related commit
|
2018-05-25 02:10:39 -07:00 |
|
James R. Barlow
|
ad9f8ca78e
|
pdfinfo: reinstate stack normalization for q/Q
|
2018-05-25 01:28:26 -07:00 |
|
James R. Barlow
|
59e786eb3c
|
Remove old code to deal with single page only things
|
2018-05-25 00:32:55 -07:00 |
|
James R. Barlow
|
6d0461435f
|
Use OperandGrouper whitelist
|
2018-05-24 22:52:33 -07:00 |
|
James R. Barlow
|
16f70ff054
|
Main changeset for pikepdf-based refactor pdfinfo
|
2018-05-24 22:22:01 -07:00 |
|
James R. Barlow
|
83f35e00f3
|
Start removing PyPDF2
|
2018-05-21 01:28:21 -07:00 |
|
James R. Barlow
|
c4ab01d63d
|
Fix "AttributeError: 'ImageInfo' object has no attribute '_type'"
Also deal with 'fixme' imagemask comment.
Also fix bpc incorrectly set to 8 by default on stencil masks.
|
2018-05-12 12:14:57 -07:00 |
|
James R. Barlow
|
601863f9e9
|
Return to PyMuPDF 1.12.5
|
2018-05-10 18:47:10 -07:00 |
|
James R. Barlow
|
63032d304d
|
Revert "Since PyMuPDF 1.13.3 corrupts text, pin 1.12.5 and work around it"
This reverts commit b0ce7c63dd.
|
2018-05-10 16:27:17 -07:00 |
|
James R. Barlow
|
a57ecede78
|
Refactor textareas to remove duplicate code
|
2018-05-10 16:26:52 -07:00 |
|
James R. Barlow
|
b0ce7c63dd
|
Since PyMuPDF 1.13.3 corrupts text, pin 1.12.5 and work around it
|
2018-05-10 16:10:24 -07:00 |
|
James R. Barlow
|
45336c7c28
|
textareas: filter out images
|
2018-05-10 01:17:28 -07:00 |
|
James R. Barlow
|
20aabb2e83
|
When deciding if there is a text on a page, ignore the margins
Margins may include watermarks or digital stamps on otherwise
text-free pages.
|
2018-05-10 01:16:11 -07:00 |
|
James R. Barlow
|
da80d3f354
|
Add unconditional (for now) whiteout of text areas
|
2018-05-07 17:37:46 -07:00 |
|
James R. Barlow
|
e0f3f07907
|
Fix text reported as found on all pages when PyMuPDF is not available
|
2018-03-30 00:10:53 -07:00 |
|
James R. Barlow
|
d0271d5049
|
More debug messages on repair; update notes
|
2018-03-28 00:39:38 -07:00 |
|
James R. Barlow
|
5becfcf8ea
|
Refactor fitz ImportError trap
|
2018-03-27 21:38:02 -07:00 |
|
James R. Barlow
|
6a4df78bc0
|
Add _naive_find_text to search for text when fitz is not available
|
2018-03-27 13:36:17 -07:00 |
|
James R. Barlow
|
3e444f6a90
|
Make fitz optional
|
2018-03-26 13:22:09 -07:00 |
|
James R. Barlow
|
45dbff6401
|
Fix table of contents not preserved in PDF/A
|
2018-03-26 02:23:19 -07:00 |
|
James R. Barlow
|
af085b79dd
|
Move ocrmypdf to src/ocrmypdf
|
2018-03-24 23:59:08 -07:00 |
|