Fix flaky tagged-PDF skip-text test under Ghostscript 10.x

The test asserted that --mode skip preserves the structure tree, but with
the default --output-type auto the output runs through Ghostscript PDF/A
conversion, which discards /StructTreeRoot on Ghostscript 10.x (9.x kept
it). This failed on macOS CI and locally while passing on the Ubuntu
runners' Ghostscript 9.55. Pin the test to --output-type pdf so it exercises
OCRmyPDF's own structure-tree handling without the version-dependent GS step,
and document the caveat in advanced.md.
This commit is contained in:
James R. Barlow
2026-07-01 00:10:57 -07:00
parent 72ce05768e
commit 8de7b05fb9
2 changed files with 15 additions and 1 deletions
+10
View File
@@ -135,6 +135,16 @@ OCRmyPDF cannot rebuild a structure tree to match newly recognized text. When
the structure tree no longer corresponds to the page content, so it is discarded.
`--mode skip` leaves text pages untouched, so their structural markup is preserved.
:::{note}
Preservation under `--mode skip` only holds when the output is not converted to
PDF/A. PDF/A conversion is performed by Ghostscript, and Ghostscript 10.x discards
the structure tree during conversion (Ghostscript 9.x preserved it). Because the
default `--output-type auto` may fall back to Ghostscript, use
`--output-type pdf` if you need to guarantee that a Tagged PDF's structural markup
survives. For best results, install veraPDF so that speculative PDF/A
conversion can sidestep this issue entirely in most real cases.
:::
### Time and image size limits
By default, OCRmyPDF permits tesseract to run for three minutes (180