From 2642c1b3d31c8c17f4dda0d741c5c06fd47926c7 Mon Sep 17 00:00:00 2001 From: fritz-hh Date: Fri, 26 Apr 2013 18:00:58 +0300 Subject: [PATCH 1/3] Update README.md --- README.md | 1 + 1 file changed, 1 insertion(+) diff --git a/README.md b/README.md index abd75938..0a5e59a2 100644 --- a/README.md +++ b/README.md @@ -9,6 +9,7 @@ Features -------- - Generate a searchable PDF/A file from a PDF file containing only images +- Place OCRed text accurately below the image to easy copy / paste - Keep the exact resolution of the original embedded images - If requested deskew and / or clean the image before performing OCR - Validate the generated file against the PDF/A specification using jhove From 4d80709cfd63e145dd1ae9b325800675d18866ae Mon Sep 17 00:00:00 2001 From: fritz-hh Date: Fri, 26 Apr 2013 18:43:15 +0300 Subject: [PATCH 2/3] Update README.md --- README.md | 1 + 1 file changed, 1 insertion(+) diff --git a/README.md b/README.md index 0a5e59a2..62b8231d 100644 --- a/README.md +++ b/README.md @@ -21,6 +21,7 @@ Motivation I searched the web for a free command line tool to OCR PDF files on linux/unix: I found many, but none of them were really satisfying. - Either they produced PDF files with misplaced text under the image (making copy/paste impossible) +- Or they did not display correctly some escaped html characters located in the hocr file produced by the OCR engine - Or they changed the resolution of the embedded images - Or they generated PDF file having a ridiculous big size - Or they crashed when trying to OCR some of my PDF files From 5ec875325e0ea33644ddd3af0a669b3966c0919e Mon Sep 17 00:00:00 2001 From: fritz-hh Date: Fri, 26 Apr 2013 18:46:18 +0300 Subject: [PATCH 3/3] Update README.md --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 62b8231d..5b29fabb 100644 --- a/README.md +++ b/README.md @@ -37,6 +37,6 @@ Download OCRmyPDF here: https://github.com/fritz-hh/OCRmyPDF/tags Copy the file in onto your linux/unix machine and extract it. -Run: "sh ./OCRmyPDF.sh" to get the script usage +Run: "sh ./OCRmyPDF.sh -h" to get the script usage If not yet installed, the script will notify you about dependencies that need to be installed