Fix flaky tagged-PDF skip-text test under Ghostscript 10.x

The test asserted that --mode skip preserves the structure tree, but with
the default --output-type auto the output runs through Ghostscript PDF/A
conversion, which discards /StructTreeRoot on Ghostscript 10.x (9.x kept
it). This failed on macOS CI and locally while passing on the Ubuntu
runners' Ghostscript 9.55. Pin the test to --output-type pdf so it exercises
OCRmyPDF's own structure-tree handling without the version-dependent GS step,
and document the caveat in advanced.md.
This commit is contained in:
James R. Barlow
2026-07-01 00:10:57 -07:00
parent 72ce05768e
commit 8de7b05fb9
2 changed files with 15 additions and 1 deletions
+5 -1
View File
@@ -54,10 +54,14 @@ def test_tagged_pdf_mode_ignore_with_skip_text(resources, outpdf, caplog):
outpdf,
tagged_pdf_mode='ignore',
skip_text=True, # Tagged PDF has text, so skip pages with text
# output_type=pdf avoids the Ghostscript PDF/A step, whose treatment of
# the structure tree is version-dependent (Ghostscript >= 10 discards it,
# 9.x preserves it). We only want to assert OCRmyPDF's own behavior here.
output_type='pdf',
plugins=['tests/plugins/tesseract_noop.py'],
)
assert 'structural markup' in caplog.text
# skip-text leaves the text pages untouched, so the structure tree remains valid
# skip-text leaves the text pages untouched, so OCRmyPDF keeps the structure tree
with pikepdf.open(outpdf) as pdf:
assert Name.StructTreeRoot in pdf.Root