ToolPipo Lab: PDF text layer health / synthetic fixtures
21 existing self-created PDF fixtures. No real documents or personal information.
Inspect locally with /tools/pdf-text-layer-health-checker/ (PDF.js 6.3.289).
Reproduce: node scripts/measure-pdf-text-layer-article.mjs
Then: python scripts/package-pdf-text-layer-article.py
Then: node scripts/render-pdf-text-layer-article.js
All per-page counts, input bytes and SHA256 values are in results.json.
Tool options and measured source commit/hashes are recorded there.
Native/image-only/OCR-like examples display synthetic HELLO lettering.
OCR-like means an intentionally authored image plus invisible text, NOT actual OCR.
Images are rasterized with Bitstream Vera; password PDF reuses the licensed
ToolPipoFixtureSans embedded font fixture. See FONT-LICENSE.txt.
No standalone font binary is distributed.
o-password.pdf is intentionally protected; password: fixture-password.
p-malformed.pdf is deliberately broken. These two are expected rejection tests.
Counts do not prove correct search/copy, OCR accuracy, reading order or PDF safety.
PDF.js can omit off-page text and Unicode before the Tool receives it.
Raster image paint operations are not counts of unique physical images.
ZIP metadata is fixed. The published results.json adds the final ZIP hash;
the bundled results.json omits download metadata to avoid circular hashing.
