Пять правок по итогам четырёх книг из to_read
Извлечение (pdf2html.py), каждая правка вызвана конкретной книгой: * колонтитул опознаётся по повтору текста в верхнем поле на многих полосах, а не по одной позиции. Позиционный признак рубит и настоящие заголовки: у Вернона («DDD Distilled») глава начинается ровно с верха полосы и короче 80 знаков, так пропали 14 заголовков из 15. Это ровно тот второй проход, о котором говорил ponytail-комментарий на прежнем фильтре; * блок режется, если шапка набрана крупнее следующего за ней текста. У Вернона 85 подзаголовков лежали в одном блоке с первым абзацем раздела: оглавление теряло второй уровень, а абзац начинался с заголовка без точки; * порог раздела снижен с 1.15 до 1.10 кегля текста. У Вернона подзаголовок 17.2 при тексте 15.0 — ровно на пять сотых ниже прежнего порога, и все 85 разделов уезжали в прозу. Тот же множитель, что у разреза блока, иначе разрез и классификация расходятся; * кандидатом в главы не может быть кегль, у которого в блоках нет букв. У Нейгарда («Release it!») номер главы набран кеглем 100 отдельным блоком («1», «2», …), и главой объявлялся он, а настоящие заголовки в 30 пунктов уезжали в разделы — оглавление выходило пустым. Отбор по наличию букв, а не по длине: «Preface» — семь знаков, порог по длине отсекал бы и его. Упаковка (pack_epub.py): XML 1.0 запрещает управляющие символы, а из PDF они приезжают регулярно. У Бейера («Site Reliability Engineering») 33 главы из 45 не разбирались как XML из-за одного такого знака в начале абзаца, и читалка спотыкается на таком файле молча. Чистим на упаковке, а не надеемся на извлечение. Тест на это добавлен и проверен: с обезвреженной правкой он падает. SKILL.md приведён в соответствие с репозиторием. Шаг упаковки с 25.08 делает pack_epub.py, а текст всё ещё требовал calibre на PATH — из-за этого проверка инструментов давала ложный отказ. Замерено 05.09.2026: ebook-convert отсутствует на всех трёх здешних машинах, и все четыре книги собрались без него. За calibre остался только .azw3 для Kindle. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XCiPWRi3Jqy4saD2LwU1Dp
This commit is contained in:
@@ -42,8 +42,8 @@ away the semantic markup that is already there and re-derives it from font
|
||||
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
|
||||
the spine, drops the nav document, and maps the book's own headings.
|
||||
|
||||
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is
|
||||
needed, not PyMuPDF) and convert with:
|
||||
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (no PyMuPDF
|
||||
needed) and convert with:
|
||||
|
||||
`python scripts/epub2html.py book.epub out/book.html`
|
||||
|
||||
@@ -68,7 +68,10 @@ the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
|
||||
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
|
||||
embedded fonts / no extractable text means a scan — stop and say OCR is
|
||||
needed; this skill does not apply.
|
||||
2. **Check tooling.** `ebook-convert` must be on PATH. For PyMuPDF, prefer an
|
||||
2. **Check tooling.** Packing needs nothing but the standard library
|
||||
(`scripts/pack_epub.py`); calibre is optional and only for `.azw3`. Measured
|
||||
2026-09-05: `ebook-convert` was absent on all three machines at hand, and the
|
||||
conversion went through end to end without it. For PyMuPDF, prefer an
|
||||
existing interpreter that has it; otherwise build a throwaway venv in the
|
||||
scratchpad — do not install into the system Python:
|
||||
|
||||
@@ -100,21 +103,29 @@ the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
|
||||
means the mono flag never fired and every listing is about to be reflowed as
|
||||
prose. Fix thresholds and re-run until
|
||||
the counts are sane. Cheap to iterate; do not skip to packing.
|
||||
7. **Pack with calibre**, always passing explicit TOC XPaths — without them
|
||||
calibre applies its own heuristics and re-breaks the chapters:
|
||||
7. **Pack with `scripts/pack_epub.py`** — no calibre, no root, splits on `<h1>`
|
||||
and writes the conservative EPUB 2 + NCX that PocketBook and KOReader both
|
||||
read:
|
||||
|
||||
```bash
|
||||
ebook-convert out/book.html "Title.epub" \
|
||||
--title="Title" --authors="Author" --language=en \
|
||||
--cover=out/images/cover.jpg \
|
||||
--level1-toc='//h:h1' --level2-toc='//h:h2' \
|
||||
--page-breaks-before='//h:h1' \
|
||||
--no-default-epub-cover
|
||||
ebook-convert "Title.epub" "Title.azw3"
|
||||
python scripts/pack_epub.py out/book.html "Title.epub" \
|
||||
--title="Title" --author="Author" --lang=en --cover=cover.jpg
|
||||
```
|
||||
|
||||
Take title/author from the user or the book's own title page — PDF metadata
|
||||
is often an ASIN or a filename.
|
||||
|
||||
calibre is needed only for `.azw3` (Kindle over USB), and only if it happens
|
||||
to be installed:
|
||||
|
||||
```bash
|
||||
ebook-convert "Title.epub" "Title.azw3"
|
||||
```
|
||||
|
||||
⚠️ Do **not** pack the XHTML with `ebook-convert` directly: without explicit
|
||||
`--level1-toc='//h:h1' --page-breaks-before='//h:h1'` it applies its own
|
||||
heuristics and re-breaks the chapters. `pack_epub.py` has no such mode to
|
||||
get wrong.
|
||||
8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle
|
||||
(Amazon converts server-side), `.azw3` for USB copy into `documents/`.
|
||||
Delete intermediate artifacts left in the user's directories.
|
||||
@@ -181,7 +192,8 @@ for them. Extract those by hand and verify by grepping the existing translation.
|
||||
## Optional stage: translation
|
||||
|
||||
Only when the user asks for a translated book. It slots between step 6 and
|
||||
step 7 — translate the XHTML, then pack the translated file with calibre.
|
||||
step 7 — translate the XHTML, then pack the translated file with
|
||||
`pack_epub.py`.
|
||||
|
||||
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
|
||||
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
|
||||
|
||||
Reference in New Issue
Block a user