Files
pdf2epub/SKILL.md
T
chesirecatt 33845bd47a Пять правок по итогам четырёх книг из to_read
Извлечение (pdf2html.py), каждая правка вызвана конкретной книгой:

* колонтитул опознаётся по повтору текста в верхнем поле на многих полосах, а
  не по одной позиции. Позиционный признак рубит и настоящие заголовки: у
  Вернона («DDD Distilled») глава начинается ровно с верха полосы и короче 80
  знаков, так пропали 14 заголовков из 15. Это ровно тот второй проход, о
  котором говорил ponytail-комментарий на прежнем фильтре;
* блок режется, если шапка набрана крупнее следующего за ней текста. У Вернона
  85 подзаголовков лежали в одном блоке с первым абзацем раздела: оглавление
  теряло второй уровень, а абзац начинался с заголовка без точки;
* порог раздела снижен с 1.15 до 1.10 кегля текста. У Вернона подзаголовок
  17.2 при тексте 15.0 — ровно на пять сотых ниже прежнего порога, и все 85
  разделов уезжали в прозу. Тот же множитель, что у разреза блока, иначе
  разрез и классификация расходятся;
* кандидатом в главы не может быть кегль, у которого в блоках нет букв. У
  Нейгарда («Release it!») номер главы набран кеглем 100 отдельным блоком
  («1», «2», …), и главой объявлялся он, а настоящие заголовки в 30 пунктов
  уезжали в разделы — оглавление выходило пустым. Отбор по наличию букв, а не
  по длине: «Preface» — семь знаков, порог по длине отсекал бы и его.

Упаковка (pack_epub.py): XML 1.0 запрещает управляющие символы, а из PDF они
приезжают регулярно. У Бейера («Site Reliability Engineering») 33 главы из 45
не разбирались как XML из-за одного такого знака в начале абзаца, и читалка
спотыкается на таком файле молча. Чистим на упаковке, а не надеемся на
извлечение. Тест на это добавлен и проверен: с обезвреженной правкой он падает.

SKILL.md приведён в соответствие с репозиторием. Шаг упаковки с 25.08 делает
pack_epub.py, а текст всё ещё требовал calibre на PATH — из-за этого проверка
инструментов давала ложный отказ. Замерено 05.09.2026: ebook-convert
отсутствует на всех трёх здешних машинах, и все четыре книги собрались без
него. За calibre остался только .azw3 для Kindle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XCiPWRi3Jqy4saD2LwU1Dp
2026-09-05 18:23:05 +03:00

397 lines
21 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: pdf-to-kindle
description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. Also covers translating such a book into another language before packing it, delegated to the external book_translator project, and generating a TTS audiobook from it with an integrity check for dropped fragments and a split into per-chapter tagged tracks.
---
# PDF to Kindle
`ebook-convert book.pdf book.epub` silently destroys formatting: poppler infers
style from font names and misses abbreviated ones like `MinionPro-It`, and
calibre labels body text as `<h2>` when the most common font is not the body
font. Measured on a 399-page novel: 435 italic runs in the PDF became 1 `<i>`
tag, and 2817 body paragraphs became `<h2>`.
`scripts/pdf2html.py` extracts spans with PyMuPDF, where style comes from span
properties, and emits semantic XHTML. calibre is then used only as the
XHTML→EPUB packer.
## Three entry points
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB,
`scripts/djvu2html.py` for DjVu. All three emit the same normalized XHTML — one
block per line, `<pre>` for code listings — and everything downstream
(translation, packing, audiobook) is identical. The shared document skeleton,
CSS and the junk-heading check `looks_garbled()` live in `scripts/bookhtml.py`.
**A DjVu with its own text layer must not be re-OCR'd.** The layer that is
already in the file beats a fresh `ocrmypdf` run — measured on Prata's C++ 6th
Russian edition, 1244 pages: words mixing Latin and Cyrillic inside one word
dropped from 433 to 109 per 391k words (10.9 → 2.8 per 10 000). That layer does
not survive a trip through PDF: `ddjvu -format=pdf` writes the pages as images
and `pdftotext` then returns nothing at all. `djvu2html.py` reads `djvutxt
--detail=line` directly.
Styling is the price: a DjVu text layer carries no italic or bold at all (the
scan never had them). Headings come from line height, listings from punctuation.
Line height there is the glyph bounding box, not the type size — a line holding
a descender is 10% taller than its neighbour — so the thresholds are set from
measured gaps (`H1`/`H2` in the script), not from "slightly above body text".
**Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws
away the semantic markup that is already there and re-derives it from font
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
the spine, drops the nav document, and maps the book's own headings.
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (no PyMuPDF
needed) and convert with:
`python scripts/epub2html.py book.epub out/book.html`
**Dumps converted from PDF carry no semantics at all** — no headings, no `<pre>`,
just `<p class="class_s1e2">` with obfuscated names. The script detects this
(zero headings or zero listings) and rebuilds the structure from the book's own
stylesheet: the three largest font sizes become heading levels, a typewriter or
monospace family becomes `<pre>`, and adjacent listing paragraphs merge back into
one block. A book that carries its own markup never enters this path. Measured on
the dokumen.pub dump of Ousterhout's APoSD: 1833 flat paragraphs became 29
chapters and 58 listings.
It reports which heading level turned out to be the chapter level. In most
EPUBs `<h1>` is the book title and the parts, while chapters are `<h2>` — the
script picks the deepest level that still gives a sane chapter count, because
the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
3rd edition: 7 `<h1>` against 49 `<h2>`, chapters correctly detected as `h2`,
56 chapters out of 656 pages.
## Workflow
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
embedded fonts / no extractable text means a scan — stop and say OCR is
needed; this skill does not apply.
2. **Check tooling.** Packing needs nothing but the standard library
(`scripts/pack_epub.py`); calibre is optional and only for `.azw3`. Measured
2026-09-05: `ebook-convert` was absent on all three machines at hand, and the
conversion went through end to end without it. For PyMuPDF, prefer an
existing interpreter that has it; otherwise build a throwaway venv in the
scratchpad — do not install into the system Python:
```bash
python3 -m venv "$SCRATCH/venv" && "$SCRATCH/venv/bin/pip" -q install pymupdf
```
3. **Profile the fonts** before converting anything:
`python scripts/pdf2html.py --fonts book.pdf`
Read off: the body font (largest character count), the italic variant, the
heading sizes, and any secondary family used for sidebars, journal entries,
chat logs or slides.
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. Code
listings need no tuning — a monospaced span is recognized by the PyMuPDF font
flag (`flags & 8`), not by font name, and becomes a `<pre>` block. The
defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a
title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary
family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that
separates an indented first line from a continuation line — verify it
against the real `x0` values, not by assumption.
5. **Convert to XHTML:**
`python scripts/pdf2html.py book.pdf out/book.html`
6. **Verify before packing** (see Verification). Check the `pre:` counter against
the real number of listings in the book — a technical book reporting `pre: 0`
means the mono flag never fired and every listing is about to be reflowed as
prose. Fix thresholds and re-run until
the counts are sane. Cheap to iterate; do not skip to packing.
7. **Pack with `scripts/pack_epub.py`** — no calibre, no root, splits on `<h1>`
and writes the conservative EPUB 2 + NCX that PocketBook and KOReader both
read:
```bash
python scripts/pack_epub.py out/book.html "Title.epub" \
--title="Title" --author="Author" --lang=en --cover=cover.jpg
```
Take title/author from the user or the book's own title page — PDF metadata
is often an ASIN or a filename.
calibre is needed only for `.azw3` (Kindle over USB), and only if it happens
to be installed:
```bash
ebook-convert "Title.epub" "Title.azw3"
```
⚠️ Do **not** pack the XHTML with `ebook-convert` directly: without explicit
`--level1-toc='//h:h1' --page-breaks-before='//h:h1'` it applies its own
heuristics and re-breaks the chapters. `pack_epub.py` has no such mode to
get wrong.
8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle
(Amazon converts server-side), `.azw3` for USB copy into `documents/`.
Delete intermediate artifacts left in the user's directories.
## Verification
Never report success on the converter's own summary line alone. Check:
```bash
python3 - <<'EOF'
import re
t = open('out/book.html').read()
print('italic:', t.count('<i>'), 'bold:', t.count('<b>'))
print('h1:', len(re.findall(r'<h1', t)))
print('double spaces:', re.sub('<[^>]+>', '', t).count(' '))
EOF
```
- italic count near zero on a novel means `style()` missed the font-name pattern;
- an `h1` count in the hundreds means a size threshold is too low and body text
is being promoted;
- many double spaces means line joining is off;
- list the extracted headings (`grep -o '<h1[^>]*>.\{0,60\}'`) and read them —
they become the TOC, and a wrong one is obvious at a glance. On a scan most of
the junk there comes from **rotated** text, not from bad thresholds: sideways
captions and table stubs OCR into mush (`aoHegoduenueArdng`) and land in
headings because their type is large. `pdf2html.py` drops any block whose line
direction is not horizontal (`is_rotated()`, measured on Brikman: 105 lines out
of 17 300, every one of them garbage), and demotes what is left of the mush to
`<p>` rather than deleting it. That pair took Brikman from 99 `<h1>` to 48;
- read one full page of body text and confirm paragraphs merge across page
breaks and hyphenated words are rejoined.
After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx`
navPoint count for the TOC size.
## Glossary from existing translations
For a book in a series that already has published translations, `scripts/glossary.py`
mines a bilingual glossary so the machine translation does not invent new spellings
for names the reader already knows. Feed it pairs of editions of the *same* volume:
```
python scripts/glossary.py --en vol12.fb2 --ru vol12.ru.fb2 \
--en vol15.epub --ru vol15.ru.fb2 --score 0.6
```
Two signals, and both are needed. Position: paragraph indices do not line up
(Russian editions split dialogue, giving 2–3× more paragraphs), so offsets are
measured as a **share of characters**, where the texts track each other closely.
Transliteration: a proper name in Russian is nearly always a transliteration, so
the Cyrillic candidate is romanized and compared to the English term — this is
what turns the output from noise into a usable list.
Two mirrored filters remove the rest of the junk: a candidate whose head word
also appears lowercase in the same text is a sentence-initial common word, not a
name — applied on both sides. Measured on four Dresden Files volumes: 71 pairs,
of which two were wrong.
⚠️ Concept terms (`White Council` → `Белый Совет`, `Spire` → `Копьё`) do **not**
come out of the transliteration path and the positional one alone is too noisy
for them. Extract those by hand and verify by grepping the existing translation.
## Optional stage: translation
Only when the user asks for a translated book. It slots between step 6 and
step 7 — translate the XHTML, then pack the translated file with
`pack_epub.py`.
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
the result untouched. Nothing extra is needed to protect code — but this also
means a listing that was misclassified as a paragraph upstream *will* be
translated, which is the real reason step 6 checks the `pre:` counter.
For the same reason `<pre>` is cut out of both language checks. A book that is
40% listings translates correctly and would otherwise fail acceptance, because
the English code drags the Cyrillic share below the threshold.
**One workdir per book.** Chapter files are named `chapter_NNN.json` for every
book, and the external repo skips chapters it has marked done — a reused workdir
silently stitches one book's translation onto another's text. The bridge writes
`bridge_source.json` into the workdir on first run and refuses to start if the
directory belongs to a different book. Re-running the same book is unaffected;
that is the resume path.
Translation is delegated to an external project,
[`vetermanve/book_translator`](https://github.com/vetermanve/book_translator)
(DeepSeek or a local Ollama model). **Never modify that repository** — it is
driven through its documented CLI and file formats only. `scripts/translate.py`
is the bridge: it writes the repo's input format, shells out to
`03_translate_parallel.py --all`, and reads the repo's output format back.
1. **Check the book actually needs translating.** The script does this first
and refuses to burn API credits on a book already in the target language:
`python scripts/translate.py book.html out.html --repo <clone> --workdir <dir> --check-only`
It reports block/chapter/character counts and the detected source language.
`--force` overrides the refusal.
2. **Set up the external repo once** (a plain clone; never edit it) and its
credentials — a `.env` in the working directory with `USE_LOCAL_MODEL=false`
and `DEEPSEEK_API_KEY=…`, or `USE_LOCAL_MODEL=true` plus `OLLAMA_MODEL` for a
local model. **Its `requirements.txt` is incomplete** — `pip install openai
pyyaml python-dotenv` as well. Both `openai` and `pyyaml` are imported at
runtime and missing from that file; without `pyyaml` every single request
dies inside `_create_system_prompt` *before* reaching the API, and the tool
reports "API запросов: 0, ошибок: N" while writing `[UNTRANSLATED]` stubs.
Write the `.env` under `umask 077`, and never echo the key into logs or
command output.
3. **Translate:**
`python scripts/translate.py book.html book.ru.html --repo <clone> --workdir <dir> --workers 12`
⚠️ **Not resumable, despite what the external repo claims.** Its
`progress/translation_progress.json` stays `{"chapters": {}}` and `--all`
re-translates every chapter, including ones already sitting in
`translations/`. Measured 2026-08-21 on Stroustrup: a re-run to repair 4
failed chapters re-did all 29. Budget a full book on every restart, and
prefer getting one clean run over patching a partial one.
The one thing that *does* skip work: deleting a chapter from `extracted/`
before the run. It is then never sent, and `rebuild` keeps the original
English for it — the right treatment for an index.
**Numbered exercise items come back untranslated.** Measured on Stroustrup
2026-08-21: 182 paragraphs of 10151 (1.8%) shaped `[2] Expanding on what you
have learned…` were echoed back in English. The document-wide Cyrillic check
cannot see this — 1.8% drowns in it — so `rebuild` counts them separately as
"осталось на языке оригинала". If the count is high, add an explicit line to the
prompt that numbered items are prose and must be translated too.
4. **Read the reported counts.** "без перевода" above zero means a chapter came
back with a different paragraph count and kept its original text; "разметка
потеряна" counts paragraphs where the model mangled the inline-tag markers
and the italics were dropped rather than corrupted. Both are expected to be
near zero — a large number means the translator misbehaved, not that the
bridge is broken. The script then checks the output's actual language and
exits non-zero if the text is still the source language or contains
`[UNTRANSLATED]` stubs — **matching paragraph counts do not prove anything
was translated**, since the external tool substitutes the original on API
failure. Do not pack a file that failed this check.
**Re-running after a failure needs the state cleared:** the external repo's
`progress/` directory marks those chapters complete and will skip them.
Delete `<workdir>/progress`, `<workdir>/context`, and
`<workdir>/translations` before the retry.
5. Pack `book.ru.html` with `--language=ru` and translated `--title`/`--authors`.
How formatting survives a translator that only speaks plain text: inline `<i>`
and `<b>` become `⟦i⟧…⟦/i⟧` markers before the text leaves, and are restored
after. Every paragraph's markers are balance-checked on the way back. Images,
`<h1>`/`<h2>` structure, and block order never leave this side — only the text
of each block round-trips, and blocks are reassembled by index.
`scripts/test_translate.py` covers the marker round-trip, the broken-marker
fallback, language detection, and a full split/rebuild against a faked
translator response — no network, no API key. Run it after touching the bridge.
## Optional stage: audiobook
Also delegated to `book_translator` (`05_create_audiobook.py`, Microsoft
edge-tts — free, needs internet). Two wrappers live in `scripts/audiobook.py`;
the external repo is still never edited.
1. **Strip markup markers first.** The translated JSON still holds the
`⟦i⟧` markers from the translation stage — TTS would read them aloud:
`python scripts/audiobook.py prep <workdir>/translations <workdir>/translations_tts`
Point the external script at the *stripped* copy; the original keeps its
italics for the EPUB.
2. **Synthesize:** `05_create_audiobook.py --translations-dir <…>/translations_tts
--voice dmitry --rate '+0%'`. Fragments are **one per paragraph** (the
`--paragraphs-per-group` flag is not used by the loop), named
`chapter_NNN_intro.mp3` / `chapter_NNN_para_NNNN.mp3` under
`audiobook/temp_audio/`.
3. **Verify before the temp files are deleted** — `cleanup_temp_files()` wipes
`temp_audio/`, and after that only the merged file can be checked:
`python scripts/audiobook.py verify <workdir>/translations_tts <workdir>/audiobook`
It flags chapters missing fragments, zero-byte/undecodable mp3s, and
chapters whose duration falls short of what their character count predicts.
The seconds-per-character baseline is the median across chapters, so it
self-calibrates to whatever voice and `--rate` were used. Exit code is
non-zero when anything is wrong.
4. **Re-running fills gaps cheaply** — the external script skips any fragment
file that already exists *and is non-empty*, so delete the bad ones `verify`
named and run it again. It never re-checks that an existing file is sane,
which is exactly why step 3 exists.
5. **Split into chapter tracks**, also before the temp files go:
```bash
python scripts/audiobook.py split <workdir>/translations_tts <workdir>/audiobook \
<workdir>/tracks --album "Название" --author "Автор" --gap 0.3
```
The external stage only ever produces one merged `audiobook_complete.mp3`
with no chapter marks, which is bad for players. `split` rebuilds per-chapter
`NNN - Title.mp3` from the same fragments, with ID3 album/artist/track/title
tags, and refuses a chapter whose fragments are incomplete (`--force`
overrides). Concatenation is stream-copy, falling back to a re-encode only if
the mp3 streams don't line up; `--gap` inserts silence between paragraphs,
generated to match the fragments' own codec parameters so the copy path
stays viable. Hand the result to the `prepare-audiobooks` skill for covers
and library layout.
### Prefer local Silero over edge-tts
edge-tts drops fragments silently under load — measured 1 of 49 on one run and
7 of 49 on the next, at *fewer* workers, so it is volume- not concurrency-bound.
`scripts/silero_render.py` replaces the external synthesis stage entirely with a
local model: no network, no dropouts, no `repair` cycle, free.
```bash
pip install torch numpy --index-url https://download.pytorch.org/whl/cpu
curl -O https://models.silero.ai/models/tts/ru/v5_5_ru.pt # 145 МБ
python scripts/silero_render.py v5_5_ru.pt <workdir>/translations_tts <out> \
--album "Название" --author "Автор" --speaker eugene --tempo 0.87 \
--skip 0 1 2 3 37
```
**Check `lscpu | grep avx2` before choosing the host.** PyTorch needs AVX2;
without it inference is ~30x slower — measured 3.4x realtime on a Celeron N5095
(SSE4 only) versus 97x on a Ryzen 7 5800H. Use 8 threads, not 16: hyperthreading
loses (97x vs 83x).
`scripts/tts_normalize.py` prepares the text — Latin script is what makes a
Russian voice sound worst, and a full book carries far more of it than a sample
chapter suggests (37 unique tokens in one chapter, 728 across the book). It maps
named entities and acronyms by hand with stress marks, transliterates the rest
by rule, spells out numbers, and drops URLs. It also generates Russian chapter
titles, since translated headings often stay English and the intro fragment
would otherwise read them aloud in Latin.
Voice tempo: SSML `<prosody rate>` quantizes to named levels, so percentages
cluster instead of stepping evenly. For fine control use `--tempo`, which
time-stretches with ffmpeg and preserves pitch.
The phonetics stage (`07_extract_terms.py` + `08_generate_phonetics.py`) is
worth running first for a technical book **when using edge-tts**; with
`tts_normalize.py` it is redundant.
**Ordering constraint for the whole stage:** `verify` and `split` both read
`audiobook/temp_audio/`, and the external script's `cleanup_temp_files()`
deletes it right after merging. Run both before that, or the fragments are gone
and only the merged file's total duration can be checked.
`scripts/test_audiobook.py` builds real silent mp3s with ffmpeg and checks that
`verify` catches a missing fragment, an empty file, and passes a clean book.
Needs ffmpeg; no network.
## What the script does
- style from span font names (`-It`, `Italic`, `Bold`, `Semibold`) → `<i>`/`<b>`;
- headings by font size, sidebars by font family;
- one PDF block = one paragraph; an unindented block that opens a page is
appended to the previous page's paragraph;
- trailing hyphens dropped, adjacent `</i><i>` runs merged, doubled spaces collapsed;
- images written to `images/`, cover rendered from page 1 at 150 dpi;
- absolute positioning is deliberately discarded so text reflows at any font size.
## Limits
- Headers/footers are not stripped — calibre-made PDFs have none. A typeset PDF
with running heads and page numbers needs blocks filtered by `y` coordinate
near the margins; add that before trusting the output.
- Multi-column layouts are not handled; block order would need column sorting.
- Thresholds are per-layout constants. Always re-run `--fonts` for a new book.