Пять правок по итогам четырёх книг из to_read

Извлечение (pdf2html.py), каждая правка вызвана конкретной книгой:

* колонтитул опознаётся по повтору текста в верхнем поле на многих полосах, а
  не по одной позиции. Позиционный признак рубит и настоящие заголовки: у
  Вернона («DDD Distilled») глава начинается ровно с верха полосы и короче 80
  знаков, так пропали 14 заголовков из 15. Это ровно тот второй проход, о
  котором говорил ponytail-комментарий на прежнем фильтре;
* блок режется, если шапка набрана крупнее следующего за ней текста. У Вернона
  85 подзаголовков лежали в одном блоке с первым абзацем раздела: оглавление
  теряло второй уровень, а абзац начинался с заголовка без точки;
* порог раздела снижен с 1.15 до 1.10 кегля текста. У Вернона подзаголовок
  17.2 при тексте 15.0 — ровно на пять сотых ниже прежнего порога, и все 85
  разделов уезжали в прозу. Тот же множитель, что у разреза блока, иначе
  разрез и классификация расходятся;
* кандидатом в главы не может быть кегль, у которого в блоках нет букв. У
  Нейгарда («Release it!») номер главы набран кеглем 100 отдельным блоком
  («1», «2», …), и главой объявлялся он, а настоящие заголовки в 30 пунктов
  уезжали в разделы — оглавление выходило пустым. Отбор по наличию букв, а не
  по длине: «Preface» — семь знаков, порог по длине отсекал бы и его.

Упаковка (pack_epub.py): XML 1.0 запрещает управляющие символы, а из PDF они
приезжают регулярно. У Бейера («Site Reliability Engineering») 33 главы из 45
не разбирались как XML из-за одного такого знака в начале абзаца, и читалка
спотыкается на таком файле молча. Чистим на упаковке, а не надеемся на
извлечение. Тест на это добавлен и проверен: с обезвреженной правкой он падает.

SKILL.md приведён в соответствие с репозиторием. Шаг упаковки с 25.08 делает
pack_epub.py, а текст всё ещё требовал calibre на PATH — из-за этого проверка
инструментов давала ложный отказ. Замерено 05.09.2026: ebook-convert
отсутствует на всех трёх здешних машинах, и все четыре книги собрались без
него. За calibre остался только .azw3 для Kindle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XCiPWRi3Jqy4saD2LwU1Dp
This commit is contained in:
chesirecatt
2026-09-05 18:23:05 +03:00
parent af5c243687
commit 33845bd47a
4 changed files with 133 additions and 17 deletions
+25 -13
View File
@@ -42,8 +42,8 @@ away the semantic markup that is already there and re-derives it from font
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
the spine, drops the nav document, and maps the book's own headings.
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is
needed, not PyMuPDF) and convert with:
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (no PyMuPDF
needed) and convert with:
`python scripts/epub2html.py book.epub out/book.html`
@@ -68,7 +68,10 @@ the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
embedded fonts / no extractable text means a scan — stop and say OCR is
needed; this skill does not apply.
2. **Check tooling.** `ebook-convert` must be on PATH. For PyMuPDF, prefer an
2. **Check tooling.** Packing needs nothing but the standard library
(`scripts/pack_epub.py`); calibre is optional and only for `.azw3`. Measured
2026-09-05: `ebook-convert` was absent on all three machines at hand, and the
conversion went through end to end without it. For PyMuPDF, prefer an
existing interpreter that has it; otherwise build a throwaway venv in the
scratchpad — do not install into the system Python:
@@ -100,21 +103,29 @@ the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
means the mono flag never fired and every listing is about to be reflowed as
prose. Fix thresholds and re-run until
the counts are sane. Cheap to iterate; do not skip to packing.
7. **Pack with calibre**, always passing explicit TOC XPaths — without them
calibre applies its own heuristics and re-breaks the chapters:
7. **Pack with `scripts/pack_epub.py`** — no calibre, no root, splits on `<h1>`
and writes the conservative EPUB 2 + NCX that PocketBook and KOReader both
read:
```bash
ebook-convert out/book.html "Title.epub" \
--title="Title" --authors="Author" --language=en \
--cover=out/images/cover.jpg \
--level1-toc='//h:h1' --level2-toc='//h:h2' \
--page-breaks-before='//h:h1' \
--no-default-epub-cover
ebook-convert "Title.epub" "Title.azw3"
python scripts/pack_epub.py out/book.html "Title.epub" \
--title="Title" --author="Author" --lang=en --cover=cover.jpg
```
Take title/author from the user or the book's own title page — PDF metadata
is often an ASIN or a filename.
calibre is needed only for `.azw3` (Kindle over USB), and only if it happens
to be installed:
```bash
ebook-convert "Title.epub" "Title.azw3"
```
⚠️ Do **not** pack the XHTML with `ebook-convert` directly: without explicit
`--level1-toc='//h:h1' --page-breaks-before='//h:h1'` it applies its own
heuristics and re-breaks the chapters. `pack_epub.py` has no such mode to
get wrong.
8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle
(Amazon converts server-side), `.azw3` for USB copy into `documents/`.
Delete intermediate artifacts left in the user's directories.
@@ -181,7 +192,8 @@ for them. Extract those by hand and verify by grepping the existing translation.
## Optional stage: translation
Only when the user asks for a translated book. It slots between step 6 and
step 7 — translate the XHTML, then pack the translated file with calibre.
step 7 — translate the XHTML, then pack the translated file with
`pack_epub.py`.
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into