EPUB как второй вход в конвейер; листинги кода не переводятся и не ломают приёмку
- epub2html.py: EPUB -> тот же XHTML, что pdf2html.py, только стандартная библиотека. Уровень глав определяется по книге, а не берётся из <h1>: в EPUB там обычно название и части. - pdf2html.py: листинг узнаётся по флагу моноширинного шрифта PyMuPDF, переносы и отступы внутри <pre> сохраняются, де-дефисация к коду не применяется. - translate.py: <pre> исключён из проверок языка (английский код утягивал долю кириллицы ниже порога приёмки), рабочий каталог привязан к книге, устаревшие главы прошлого прогона чистятся, инлайновый <code> переживает переводчика. - bookhtml.py: общий каркас документа и CSS для обоих входов. - Тесты: test_epub2html.py на собранном в памяти EPUB, четыре новых случая в test_translate.py.
This commit is contained in:
@@ -15,6 +15,30 @@ tag, and 2817 body paragraphs became `<h2>`.
|
||||
properties, and emits semantic XHTML. calibre is then used only as the
|
||||
XHTML→EPUB packer.
|
||||
|
||||
## Two entry points
|
||||
|
||||
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB. Both emit the
|
||||
same normalized XHTML — one block per line, `<pre>` for code listings — and
|
||||
everything downstream (translation, packing, audiobook) is identical. Shared
|
||||
document skeleton and CSS live in `scripts/bookhtml.py`.
|
||||
|
||||
**Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws
|
||||
away the semantic markup that is already there and re-derives it from font
|
||||
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
|
||||
the spine, drops the nav document, and maps the book's own headings.
|
||||
|
||||
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is
|
||||
needed, not PyMuPDF) and convert with:
|
||||
|
||||
`python scripts/epub2html.py book.epub out/book.html`
|
||||
|
||||
It reports which heading level turned out to be the chapter level. In most
|
||||
EPUBs `<h1>` is the book title and the parts, while chapters are `<h2>` — the
|
||||
script picks the deepest level that still gives a sane chapter count, because
|
||||
the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
|
||||
3rd edition: 7 `<h1>` against 49 `<h2>`, chapters correctly detected as `h2`,
|
||||
56 chapters out of 656 pages.
|
||||
|
||||
## Workflow
|
||||
|
||||
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
|
||||
@@ -35,7 +59,9 @@ XHTML→EPUB packer.
|
||||
Read off: the body font (largest character count), the italic variant, the
|
||||
heading sizes, and any secondary family used for sidebars, journal entries,
|
||||
chat logs or slides.
|
||||
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. The
|
||||
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. Code
|
||||
listings need no tuning — a monospaced span is recognized by the PyMuPDF font
|
||||
flag (`flags & 8`), not by font name, and becomes a `<pre>` block. The
|
||||
defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a
|
||||
title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary
|
||||
family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that
|
||||
@@ -45,7 +71,10 @@ XHTML→EPUB packer.
|
||||
|
||||
`python scripts/pdf2html.py book.pdf out/book.html`
|
||||
|
||||
6. **Verify before packing** (see Verification). Fix thresholds and re-run until
|
||||
6. **Verify before packing** (see Verification). Check the `pre:` counter against
|
||||
the real number of listings in the book — a technical book reporting `pre: 0`
|
||||
means the mono flag never fired and every listing is about to be reflowed as
|
||||
prose. Fix thresholds and re-run until
|
||||
the counts are sane. Cheap to iterate; do not skip to packing.
|
||||
7. **Pack with calibre**, always passing explicit TOC XPaths — without them
|
||||
calibre applies its own heuristics and re-breaks the chapters:
|
||||
@@ -97,6 +126,23 @@ navPoint count for the TOC size.
|
||||
Only when the user asks for a translated book. It slots between step 6 and
|
||||
step 7 — translate the XHTML, then pack the translated file with calibre.
|
||||
|
||||
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
|
||||
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
|
||||
the result untouched. Nothing extra is needed to protect code — but this also
|
||||
means a listing that was misclassified as a paragraph upstream *will* be
|
||||
translated, which is the real reason step 6 checks the `pre:` counter.
|
||||
|
||||
For the same reason `<pre>` is cut out of both language checks. A book that is
|
||||
40% listings translates correctly and would otherwise fail acceptance, because
|
||||
the English code drags the Cyrillic share below the threshold.
|
||||
|
||||
**One workdir per book.** Chapter files are named `chapter_NNN.json` for every
|
||||
book, and the external repo skips chapters it has marked done — a reused workdir
|
||||
silently stitches one book's translation onto another's text. The bridge writes
|
||||
`bridge_source.json` into the workdir on first run and refuses to start if the
|
||||
directory belongs to a different book. Re-running the same book is unaffected;
|
||||
that is the resume path.
|
||||
|
||||
Translation is delegated to an external project,
|
||||
[`vetermanve/book_translator`](https://github.com/vetermanve/book_translator)
|
||||
(DeepSeek or a local Ollama model). **Never modify that repository** — it is
|
||||
|
||||
Reference in New Issue
Block a user