Восстановление структуры EPUB по CSS для дампов из PDF

Дамп, собранный конвертером из PDF, теряет всю семантику: ни заголовков, ни
листингов, только абзацы с обфусцированными классами. Такие книги уезжали в
перевод одним куском, а код — вместе с текстом.

Теперь при нулевом числе заголовков или листингов разметка восстанавливается из
таблицы стилей книги: три самых крупных кегля становятся уровнями заголовков,
моноширинная гарнитура — листингом, соседние строки листинга склеиваются в один
блок. Книга со своей разметкой в этот путь не попадает.

Проверено на двух настоящих книгах: дамп APoSD — 1833 плоских абзаца стали 29
главами и 58 листингами; Страуструп с собственной разметкой не изменился.
This commit is contained in:
chesirecatt
2026-08-20 19:11:51 +03:00
parent db9fe5e0fa
commit 1bacd38292
3 changed files with 181 additions and 18 deletions
+9
View File
@@ -32,6 +32,15 @@ needed, not PyMuPDF) and convert with:
`python scripts/epub2html.py book.epub out/book.html`
**Dumps converted from PDF carry no semantics at all** — no headings, no `<pre>`,
just `<p class="class_s1e2">` with obfuscated names. The script detects this
(zero headings or zero listings) and rebuilds the structure from the book's own
stylesheet: the three largest font sizes become heading levels, a typewriter or
monospace family becomes `<pre>`, and adjacent listing paragraphs merge back into
one block. A book that carries its own markup never enters this path. Measured on
the dokumen.pub dump of Ousterhout's APoSD: 1833 flat paragraphs became 29
chapters and 58 listings.
It reports which heading level turned out to be the chapter level. In most
EPUBs `<h1>` is the book title and the parts, while chapters are `<h2>` — the
script picks the deepest level that still gives a sane chapter count, because