EPUB как второй вход в конвейер; листинги кода не переводятся и не ломают приёмку

- epub2html.py: EPUB -> тот же XHTML, что pdf2html.py, только стандартная
  библиотека. Уровень глав определяется по книге, а не берётся из <h1>:
  в EPUB там обычно название и части.
- pdf2html.py: листинг узнаётся по флагу моноширинного шрифта PyMuPDF,
  переносы и отступы внутри <pre> сохраняются, де-дефисация к коду не
  применяется.
- translate.py: <pre> исключён из проверок языка (английский код утягивал
  долю кириллицы ниже порога приёмки), рабочий каталог привязан к книге,
  устаревшие главы прошлого прогона чистятся, инлайновый <code> переживает
  переводчика.
- bookhtml.py: общий каркас документа и CSS для обоих входов.
- Тесты: test_epub2html.py на собранном в памяти EPUB, четыре новых случая
  в test_translate.py.
This commit is contained in:
chesirecatt
2026-08-20 17:58:09 +03:00
parent 6c81673b4f
commit db9fe5e0fa
7 changed files with 614 additions and 43 deletions
+48 -2
View File
@@ -15,6 +15,30 @@ tag, and 2817 body paragraphs became `<h2>`.
properties, and emits semantic XHTML. calibre is then used only as the
XHTML→EPUB packer.
## Two entry points
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB. Both emit the
same normalized XHTML — one block per line, `<pre>` for code listings — and
everything downstream (translation, packing, audiobook) is identical. Shared
document skeleton and CSS live in `scripts/bookhtml.py`.
**Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws
away the semantic markup that is already there and re-derives it from font
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
the spine, drops the nav document, and maps the book's own headings.
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is
needed, not PyMuPDF) and convert with:
`python scripts/epub2html.py book.epub out/book.html`
It reports which heading level turned out to be the chapter level. In most
EPUBs `<h1>` is the book title and the parts, while chapters are `<h2>` — the
script picks the deepest level that still gives a sane chapter count, because
the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
3rd edition: 7 `<h1>` against 49 `<h2>`, chapters correctly detected as `h2`,
56 chapters out of 656 pages.
## Workflow
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
@@ -35,7 +59,9 @@ XHTML→EPUB packer.
Read off: the body font (largest character count), the italic variant, the
heading sizes, and any secondary family used for sidebars, journal entries,
chat logs or slides.
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. The
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. Code
listings need no tuning — a monospaced span is recognized by the PyMuPDF font
flag (`flags & 8`), not by font name, and becomes a `<pre>` block. The
defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a
title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary
family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that
@@ -45,7 +71,10 @@ XHTML→EPUB packer.
`python scripts/pdf2html.py book.pdf out/book.html`
6. **Verify before packing** (see Verification). Fix thresholds and re-run until
6. **Verify before packing** (see Verification). Check the `pre:` counter against
the real number of listings in the book — a technical book reporting `pre: 0`
means the mono flag never fired and every listing is about to be reflowed as
prose. Fix thresholds and re-run until
the counts are sane. Cheap to iterate; do not skip to packing.
7. **Pack with calibre**, always passing explicit TOC XPaths — without them
calibre applies its own heuristics and re-breaks the chapters:
@@ -97,6 +126,23 @@ navPoint count for the TOC size.
Only when the user asks for a translated book. It slots between step 6 and
step 7 — translate the XHTML, then pack the translated file with calibre.
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
the result untouched. Nothing extra is needed to protect code — but this also
means a listing that was misclassified as a paragraph upstream *will* be
translated, which is the real reason step 6 checks the `pre:` counter.
For the same reason `<pre>` is cut out of both language checks. A book that is
40% listings translates correctly and would otherwise fail acceptance, because
the English code drags the Cyrillic share below the threshold.
**One workdir per book.** Chapter files are named `chapter_NNN.json` for every
book, and the external repo skips chapters it has marked done — a reused workdir
silently stitches one book's translation onto another's text. The bridge writes
`bridge_source.json` into the workdir on first run and refuses to start if the
directory belongs to a different book. Re-running the same book is unaffected;
that is the resume path.
Translation is delegated to an external project,
[`vetermanve/book_translator`](https://github.com/vetermanve/book_translator)
(DeepSeek or a local Ollama model). **Never modify that repository** — it is