Восстановление структуры EPUB по CSS для дампов из PDF
Дамп, собранный конвертером из PDF, теряет всю семантику: ни заголовков, ни листингов, только абзацы с обфусцированными классами. Такие книги уезжали в перевод одним куском, а код — вместе с текстом. Теперь при нулевом числе заголовков или листингов разметка восстанавливается из таблицы стилей книги: три самых крупных кегля становятся уровнями заголовков, моноширинная гарнитура — листингом, соседние строки листинга склеиваются в один блок. Книга со своей разметкой в этот путь не попадает. Проверено на двух настоящих книгах: дамп APoSD — 1833 плоских абзаца стали 29 главами и 58 листингами; Страуструп с собственной разметкой не изменился.
This commit is contained in:
@@ -32,6 +32,15 @@ needed, not PyMuPDF) and convert with:
|
||||
|
||||
`python scripts/epub2html.py book.epub out/book.html`
|
||||
|
||||
**Dumps converted from PDF carry no semantics at all** — no headings, no `<pre>`,
|
||||
just `<p class="class_s1e2">` with obfuscated names. The script detects this
|
||||
(zero headings or zero listings) and rebuilds the structure from the book's own
|
||||
stylesheet: the three largest font sizes become heading levels, a typewriter or
|
||||
monospace family becomes `<pre>`, and adjacent listing paragraphs merge back into
|
||||
one block. A book that carries its own markup never enters this path. Measured on
|
||||
the dokumen.pub dump of Ousterhout's APoSD: 1833 flat paragraphs became 29
|
||||
chapters and 58 listings.
|
||||
|
||||
It reports which heading level turned out to be the chapter level. In most
|
||||
EPUBs `<h1>` is the book title and the parts, while chapters are `<h2>` — the
|
||||
script picks the deepest level that still gives a sane chapter count, because
|
||||
|
||||
Reference in New Issue
Block a user