Мусор в оглавлении сканов: фильтр поворота, проверка заголовков, вход для DjVu
Причиной мусорных заголовков у Брикмана оказался не порог кегля, а повёрнутый на 90° текст: боковые врезки и подписи к таблицам распознаются в кашу («aoHegoduenueArdng») и попадают в h1, потому что кегль у них крупный. Замер: 105 строк из 17 300, все до одной брак. Блоки с неgоризонтальным направлением строки выбрасываются целиком (is_rotated). Остаток — подписи внутри иллюстраций — понижается до <p>, а не удаляется: looks_garbled() в bookhtml.py, шесть признаков структуры, любых двух хватает. Проверка повторяется после склейки соседних заголовков: по отдельности «LF», «FF» и «TIT» проходят как аббревиатуры, склеенные — обрывок таблицы. Вместе: 99 h1 у Брикмана против 48. djvu2html.py — третий вход конвейера. Родной текстовый слой DjVu чище нашего OCR (у Праты 10,9 против 2,8 слов со смесью алфавитов на 10 000), а через PDF он не проходит: ddjvu -format=pdf кладёт страницы картинками. Самопроверки: djvu2html.py --selftest, pdf2html.py --selftest --selftest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XCiPWRi3Jqy4saD2LwU1Dp
This commit is contained in:
@@ -15,12 +15,27 @@ tag, and 2817 body paragraphs became `<h2>`.
|
||||
properties, and emits semantic XHTML. calibre is then used only as the
|
||||
XHTML→EPUB packer.
|
||||
|
||||
## Two entry points
|
||||
## Three entry points
|
||||
|
||||
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB. Both emit the
|
||||
same normalized XHTML — one block per line, `<pre>` for code listings — and
|
||||
everything downstream (translation, packing, audiobook) is identical. Shared
|
||||
document skeleton and CSS live in `scripts/bookhtml.py`.
|
||||
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB,
|
||||
`scripts/djvu2html.py` for DjVu. All three emit the same normalized XHTML — one
|
||||
block per line, `<pre>` for code listings — and everything downstream
|
||||
(translation, packing, audiobook) is identical. The shared document skeleton,
|
||||
CSS and the junk-heading check `looks_garbled()` live in `scripts/bookhtml.py`.
|
||||
|
||||
**A DjVu with its own text layer must not be re-OCR'd.** The layer that is
|
||||
already in the file beats a fresh `ocrmypdf` run — measured on Prata's C++ 6th
|
||||
Russian edition, 1244 pages: words mixing Latin and Cyrillic inside one word
|
||||
dropped from 433 to 109 per 391k words (10.9 → 2.8 per 10 000). That layer does
|
||||
not survive a trip through PDF: `ddjvu -format=pdf` writes the pages as images
|
||||
and `pdftotext` then returns nothing at all. `djvu2html.py` reads `djvutxt
|
||||
--detail=line` directly.
|
||||
|
||||
Styling is the price: a DjVu text layer carries no italic or bold at all (the
|
||||
scan never had them). Headings come from line height, listings from punctuation.
|
||||
Line height there is the glyph bounding box, not the type size — a line holding
|
||||
a descender is 10% taller than its neighbour — so the thresholds are set from
|
||||
measured gaps (`H1`/`H2` in the script), not from "slightly above body text".
|
||||
|
||||
**Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws
|
||||
away the semantic markup that is already there and re-derives it from font
|
||||
@@ -123,7 +138,13 @@ EOF
|
||||
is being promoted;
|
||||
- many double spaces means line joining is off;
|
||||
- list the extracted headings (`grep -o '<h1[^>]*>.\{0,60\}'`) and read them —
|
||||
they become the TOC, and a wrong one is obvious at a glance;
|
||||
they become the TOC, and a wrong one is obvious at a glance. On a scan most of
|
||||
the junk there comes from **rotated** text, not from bad thresholds: sideways
|
||||
captions and table stubs OCR into mush (`aoHegoduenueArdng`) and land in
|
||||
headings because their type is large. `pdf2html.py` drops any block whose line
|
||||
direction is not horizontal (`is_rotated()`, measured on Brikman: 105 lines out
|
||||
of 17 300, every one of them garbage), and demotes what is left of the mush to
|
||||
`<p>` rather than deleting it. That pair took Brikman from 99 `<h1>` to 48;
|
||||
- read one full page of body text and confirm paragraphs merge across page
|
||||
breaks and hyphenated words are rejoined.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user