Мусор в оглавлении сканов: фильтр поворота, проверка заголовков, вход для DjVu

Причиной мусорных заголовков у Брикмана оказался не порог кегля, а повёрнутый
на 90° текст: боковые врезки и подписи к таблицам распознаются в кашу
(«aoHegoduenueArdng») и попадают в h1, потому что кегль у них крупный. Замер:
105 строк из 17 300, все до одной брак. Блоки с неgоризонтальным направлением
строки выбрасываются целиком (is_rotated).

Остаток — подписи внутри иллюстраций — понижается до <p>, а не удаляется:
looks_garbled() в bookhtml.py, шесть признаков структуры, любых двух хватает.
Проверка повторяется после склейки соседних заголовков: по отдельности «LF»,
«FF» и «TIT» проходят как аббревиатуры, склеенные — обрывок таблицы.
Вместе: 99 h1 у Брикмана против 48.

djvu2html.py — третий вход конвейера. Родной текстовый слой DjVu чище нашего
OCR (у Праты 10,9 против 2,8 слов со смесью алфавитов на 10 000), а через PDF
он не проходит: ddjvu -format=pdf кладёт страницы картинками.

Самопроверки: djvu2html.py --selftest, pdf2html.py --selftest --selftest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XCiPWRi3Jqy4saD2LwU1Dp
This commit is contained in:
chesirecatt
2026-08-25 19:03:08 +03:00
parent 356b9ddb22
commit 0320d2d331
5 changed files with 356 additions and 6 deletions
+27 -6
View File
@@ -15,12 +15,27 @@ tag, and 2817 body paragraphs became `<h2>`.
properties, and emits semantic XHTML. calibre is then used only as the
XHTML→EPUB packer.
## Two entry points
## Three entry points
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB. Both emit the
same normalized XHTML — one block per line, `<pre>` for code listings — and
everything downstream (translation, packing, audiobook) is identical. Shared
document skeleton and CSS live in `scripts/bookhtml.py`.
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB,
`scripts/djvu2html.py` for DjVu. All three emit the same normalized XHTML — one
block per line, `<pre>` for code listings — and everything downstream
(translation, packing, audiobook) is identical. The shared document skeleton,
CSS and the junk-heading check `looks_garbled()` live in `scripts/bookhtml.py`.
**A DjVu with its own text layer must not be re-OCR'd.** The layer that is
already in the file beats a fresh `ocrmypdf` run — measured on Prata's C++ 6th
Russian edition, 1244 pages: words mixing Latin and Cyrillic inside one word
dropped from 433 to 109 per 391k words (10.9 → 2.8 per 10 000). That layer does
not survive a trip through PDF: `ddjvu -format=pdf` writes the pages as images
and `pdftotext` then returns nothing at all. `djvu2html.py` reads `djvutxt
--detail=line` directly.
Styling is the price: a DjVu text layer carries no italic or bold at all (the
scan never had them). Headings come from line height, listings from punctuation.
Line height there is the glyph bounding box, not the type size — a line holding
a descender is 10% taller than its neighbour — so the thresholds are set from
measured gaps (`H1`/`H2` in the script), not from "slightly above body text".
**Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws
away the semantic markup that is already there and re-derives it from font
@@ -123,7 +138,13 @@ EOF
is being promoted;
- many double spaces means line joining is off;
- list the extracted headings (`grep -o '<h1[^>]*>.\{0,60\}'`) and read them —
they become the TOC, and a wrong one is obvious at a glance;
they become the TOC, and a wrong one is obvious at a glance. On a scan most of
the junk there comes from **rotated** text, not from bad thresholds: sideways
captions and table stubs OCR into mush (`aoHegoduenueArdng`) and land in
headings because their type is large. `pdf2html.py` drops any block whose line
direction is not horizontal (`is_rotated()`, measured on Brikman: 105 lines out
of 17 300, every one of them garbage), and demotes what is left of the mush to
`<p>` rather than deleting it. That pair took Brikman from 99 `<h1>` to 48;
- read one full page of body text and confirm paragraphs merge across page
breaks and hyphenated words are rejoined.