33845bd47a
Извлечение (pdf2html.py), каждая правка вызвана конкретной книгой: * колонтитул опознаётся по повтору текста в верхнем поле на многих полосах, а не по одной позиции. Позиционный признак рубит и настоящие заголовки: у Вернона («DDD Distilled») глава начинается ровно с верха полосы и короче 80 знаков, так пропали 14 заголовков из 15. Это ровно тот второй проход, о котором говорил ponytail-комментарий на прежнем фильтре; * блок режется, если шапка набрана крупнее следующего за ней текста. У Вернона 85 подзаголовков лежали в одном блоке с первым абзацем раздела: оглавление теряло второй уровень, а абзац начинался с заголовка без точки; * порог раздела снижен с 1.15 до 1.10 кегля текста. У Вернона подзаголовок 17.2 при тексте 15.0 — ровно на пять сотых ниже прежнего порога, и все 85 разделов уезжали в прозу. Тот же множитель, что у разреза блока, иначе разрез и классификация расходятся; * кандидатом в главы не может быть кегль, у которого в блоках нет букв. У Нейгарда («Release it!») номер главы набран кеглем 100 отдельным блоком («1», «2», …), и главой объявлялся он, а настоящие заголовки в 30 пунктов уезжали в разделы — оглавление выходило пустым. Отбор по наличию букв, а не по длине: «Preface» — семь знаков, порог по длине отсекал бы и его. Упаковка (pack_epub.py): XML 1.0 запрещает управляющие символы, а из PDF они приезжают регулярно. У Бейера («Site Reliability Engineering») 33 главы из 45 не разбирались как XML из-за одного такого знака в начале абзаца, и читалка спотыкается на таком файле молча. Чистим на упаковке, а не надеемся на извлечение. Тест на это добавлен и проверен: с обезвреженной правкой он падает. SKILL.md приведён в соответствие с репозиторием. Шаг упаковки с 25.08 делает pack_epub.py, а текст всё ещё требовал calibre на PATH — из-за этого проверка инструментов давала ложный отказ. Замерено 05.09.2026: ebook-convert отсутствует на всех трёх здешних машинах, и все четыре книги собрались без него. За calibre остался только .azw3 для Kindle. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XCiPWRi3Jqy4saD2LwU1Dp
397 lines
21 KiB
Markdown
397 lines
21 KiB
Markdown
---
|
||
name: pdf-to-kindle
|
||
description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. Also covers translating such a book into another language before packing it, delegated to the external book_translator project, and generating a TTS audiobook from it with an integrity check for dropped fragments and a split into per-chapter tagged tracks.
|
||
---
|
||
|
||
# PDF to Kindle
|
||
|
||
`ebook-convert book.pdf book.epub` silently destroys formatting: poppler infers
|
||
style from font names and misses abbreviated ones like `MinionPro-It`, and
|
||
calibre labels body text as `<h2>` when the most common font is not the body
|
||
font. Measured on a 399-page novel: 435 italic runs in the PDF became 1 `<i>`
|
||
tag, and 2817 body paragraphs became `<h2>`.
|
||
|
||
`scripts/pdf2html.py` extracts spans with PyMuPDF, where style comes from span
|
||
properties, and emits semantic XHTML. calibre is then used only as the
|
||
XHTML→EPUB packer.
|
||
|
||
## Three entry points
|
||
|
||
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB,
|
||
`scripts/djvu2html.py` for DjVu. All three emit the same normalized XHTML — one
|
||
block per line, `<pre>` for code listings — and everything downstream
|
||
(translation, packing, audiobook) is identical. The shared document skeleton,
|
||
CSS and the junk-heading check `looks_garbled()` live in `scripts/bookhtml.py`.
|
||
|
||
**A DjVu with its own text layer must not be re-OCR'd.** The layer that is
|
||
already in the file beats a fresh `ocrmypdf` run — measured on Prata's C++ 6th
|
||
Russian edition, 1244 pages: words mixing Latin and Cyrillic inside one word
|
||
dropped from 433 to 109 per 391k words (10.9 → 2.8 per 10 000). That layer does
|
||
not survive a trip through PDF: `ddjvu -format=pdf` writes the pages as images
|
||
and `pdftotext` then returns nothing at all. `djvu2html.py` reads `djvutxt
|
||
--detail=line` directly.
|
||
|
||
Styling is the price: a DjVu text layer carries no italic or bold at all (the
|
||
scan never had them). Headings come from line height, listings from punctuation.
|
||
Line height there is the glyph bounding box, not the type size — a line holding
|
||
a descender is 10% taller than its neighbour — so the thresholds are set from
|
||
measured gaps (`H1`/`H2` in the script), not from "slightly above body text".
|
||
|
||
**Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws
|
||
away the semantic markup that is already there and re-derives it from font
|
||
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
|
||
the spine, drops the nav document, and maps the book's own headings.
|
||
|
||
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (no PyMuPDF
|
||
needed) and convert with:
|
||
|
||
`python scripts/epub2html.py book.epub out/book.html`
|
||
|
||
**Dumps converted from PDF carry no semantics at all** — no headings, no `<pre>`,
|
||
just `<p class="class_s1e2">` with obfuscated names. The script detects this
|
||
(zero headings or zero listings) and rebuilds the structure from the book's own
|
||
stylesheet: the three largest font sizes become heading levels, a typewriter or
|
||
monospace family becomes `<pre>`, and adjacent listing paragraphs merge back into
|
||
one block. A book that carries its own markup never enters this path. Measured on
|
||
the dokumen.pub dump of Ousterhout's APoSD: 1833 flat paragraphs became 29
|
||
chapters and 58 listings.
|
||
|
||
It reports which heading level turned out to be the chapter level. In most
|
||
EPUBs `<h1>` is the book title and the parts, while chapters are `<h2>` — the
|
||
script picks the deepest level that still gives a sane chapter count, because
|
||
the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
|
||
3rd edition: 7 `<h1>` against 49 `<h2>`, chapters correctly detected as `h2`,
|
||
56 chapters out of 656 pages.
|
||
|
||
## Workflow
|
||
|
||
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
|
||
embedded fonts / no extractable text means a scan — stop and say OCR is
|
||
needed; this skill does not apply.
|
||
2. **Check tooling.** Packing needs nothing but the standard library
|
||
(`scripts/pack_epub.py`); calibre is optional and only for `.azw3`. Measured
|
||
2026-09-05: `ebook-convert` was absent on all three machines at hand, and the
|
||
conversion went through end to end without it. For PyMuPDF, prefer an
|
||
existing interpreter that has it; otherwise build a throwaway venv in the
|
||
scratchpad — do not install into the system Python:
|
||
|
||
```bash
|
||
python3 -m venv "$SCRATCH/venv" && "$SCRATCH/venv/bin/pip" -q install pymupdf
|
||
```
|
||
|
||
3. **Profile the fonts** before converting anything:
|
||
|
||
`python scripts/pdf2html.py --fonts book.pdf`
|
||
|
||
Read off: the body font (largest character count), the italic variant, the
|
||
heading sizes, and any secondary family used for sidebars, journal entries,
|
||
chat logs or slides.
|
||
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. Code
|
||
listings need no tuning — a monospaced span is recognized by the PyMuPDF font
|
||
flag (`flags & 8`), not by font name, and becomes a `<pre>` block. The
|
||
defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a
|
||
title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary
|
||
family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that
|
||
separates an indented first line from a continuation line — verify it
|
||
against the real `x0` values, not by assumption.
|
||
5. **Convert to XHTML:**
|
||
|
||
`python scripts/pdf2html.py book.pdf out/book.html`
|
||
|
||
6. **Verify before packing** (see Verification). Check the `pre:` counter against
|
||
the real number of listings in the book — a technical book reporting `pre: 0`
|
||
means the mono flag never fired and every listing is about to be reflowed as
|
||
prose. Fix thresholds and re-run until
|
||
the counts are sane. Cheap to iterate; do not skip to packing.
|
||
7. **Pack with `scripts/pack_epub.py`** — no calibre, no root, splits on `<h1>`
|
||
and writes the conservative EPUB 2 + NCX that PocketBook and KOReader both
|
||
read:
|
||
|
||
```bash
|
||
python scripts/pack_epub.py out/book.html "Title.epub" \
|
||
--title="Title" --author="Author" --lang=en --cover=cover.jpg
|
||
```
|
||
|
||
Take title/author from the user or the book's own title page — PDF metadata
|
||
is often an ASIN or a filename.
|
||
|
||
calibre is needed only for `.azw3` (Kindle over USB), and only if it happens
|
||
to be installed:
|
||
|
||
```bash
|
||
ebook-convert "Title.epub" "Title.azw3"
|
||
```
|
||
|
||
⚠️ Do **not** pack the XHTML with `ebook-convert` directly: without explicit
|
||
`--level1-toc='//h:h1' --page-breaks-before='//h:h1'` it applies its own
|
||
heuristics and re-breaks the chapters. `pack_epub.py` has no such mode to
|
||
get wrong.
|
||
8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle
|
||
(Amazon converts server-side), `.azw3` for USB copy into `documents/`.
|
||
Delete intermediate artifacts left in the user's directories.
|
||
|
||
## Verification
|
||
|
||
Never report success on the converter's own summary line alone. Check:
|
||
|
||
```bash
|
||
python3 - <<'EOF'
|
||
import re
|
||
t = open('out/book.html').read()
|
||
print('italic:', t.count('<i>'), 'bold:', t.count('<b>'))
|
||
print('h1:', len(re.findall(r'<h1', t)))
|
||
print('double spaces:', re.sub('<[^>]+>', '', t).count(' '))
|
||
EOF
|
||
```
|
||
|
||
- italic count near zero on a novel means `style()` missed the font-name pattern;
|
||
- an `h1` count in the hundreds means a size threshold is too low and body text
|
||
is being promoted;
|
||
- many double spaces means line joining is off;
|
||
- list the extracted headings (`grep -o '<h1[^>]*>.\{0,60\}'`) and read them —
|
||
they become the TOC, and a wrong one is obvious at a glance. On a scan most of
|
||
the junk there comes from **rotated** text, not from bad thresholds: sideways
|
||
captions and table stubs OCR into mush (`aoHegoduenueArdng`) and land in
|
||
headings because their type is large. `pdf2html.py` drops any block whose line
|
||
direction is not horizontal (`is_rotated()`, measured on Brikman: 105 lines out
|
||
of 17 300, every one of them garbage), and demotes what is left of the mush to
|
||
`<p>` rather than deleting it. That pair took Brikman from 99 `<h1>` to 48;
|
||
- read one full page of body text and confirm paragraphs merge across page
|
||
breaks and hyphenated words are rejoined.
|
||
|
||
After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx`
|
||
navPoint count for the TOC size.
|
||
|
||
## Glossary from existing translations
|
||
|
||
For a book in a series that already has published translations, `scripts/glossary.py`
|
||
mines a bilingual glossary so the machine translation does not invent new spellings
|
||
for names the reader already knows. Feed it pairs of editions of the *same* volume:
|
||
|
||
```
|
||
python scripts/glossary.py --en vol12.fb2 --ru vol12.ru.fb2 \
|
||
--en vol15.epub --ru vol15.ru.fb2 --score 0.6
|
||
```
|
||
|
||
Two signals, and both are needed. Position: paragraph indices do not line up
|
||
(Russian editions split dialogue, giving 2–3× more paragraphs), so offsets are
|
||
measured as a **share of characters**, where the texts track each other closely.
|
||
Transliteration: a proper name in Russian is nearly always a transliteration, so
|
||
the Cyrillic candidate is romanized and compared to the English term — this is
|
||
what turns the output from noise into a usable list.
|
||
|
||
Two mirrored filters remove the rest of the junk: a candidate whose head word
|
||
also appears lowercase in the same text is a sentence-initial common word, not a
|
||
name — applied on both sides. Measured on four Dresden Files volumes: 71 pairs,
|
||
of which two were wrong.
|
||
|
||
⚠️ Concept terms (`White Council` → `Белый Совет`, `Spire` → `Копьё`) do **not**
|
||
come out of the transliteration path and the positional one alone is too noisy
|
||
for them. Extract those by hand and verify by grepping the existing translation.
|
||
|
||
## Optional stage: translation
|
||
|
||
Only when the user asks for a translated book. It slots between step 6 and
|
||
step 7 — translate the XHTML, then pack the translated file with
|
||
`pack_epub.py`.
|
||
|
||
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
|
||
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
|
||
the result untouched. Nothing extra is needed to protect code — but this also
|
||
means a listing that was misclassified as a paragraph upstream *will* be
|
||
translated, which is the real reason step 6 checks the `pre:` counter.
|
||
|
||
For the same reason `<pre>` is cut out of both language checks. A book that is
|
||
40% listings translates correctly and would otherwise fail acceptance, because
|
||
the English code drags the Cyrillic share below the threshold.
|
||
|
||
**One workdir per book.** Chapter files are named `chapter_NNN.json` for every
|
||
book, and the external repo skips chapters it has marked done — a reused workdir
|
||
silently stitches one book's translation onto another's text. The bridge writes
|
||
`bridge_source.json` into the workdir on first run and refuses to start if the
|
||
directory belongs to a different book. Re-running the same book is unaffected;
|
||
that is the resume path.
|
||
|
||
Translation is delegated to an external project,
|
||
[`vetermanve/book_translator`](https://github.com/vetermanve/book_translator)
|
||
(DeepSeek or a local Ollama model). **Never modify that repository** — it is
|
||
driven through its documented CLI and file formats only. `scripts/translate.py`
|
||
is the bridge: it writes the repo's input format, shells out to
|
||
`03_translate_parallel.py --all`, and reads the repo's output format back.
|
||
|
||
1. **Check the book actually needs translating.** The script does this first
|
||
and refuses to burn API credits on a book already in the target language:
|
||
|
||
`python scripts/translate.py book.html out.html --repo <clone> --workdir <dir> --check-only`
|
||
|
||
It reports block/chapter/character counts and the detected source language.
|
||
`--force` overrides the refusal.
|
||
2. **Set up the external repo once** (a plain clone; never edit it) and its
|
||
credentials — a `.env` in the working directory with `USE_LOCAL_MODEL=false`
|
||
and `DEEPSEEK_API_KEY=…`, or `USE_LOCAL_MODEL=true` plus `OLLAMA_MODEL` for a
|
||
local model. **Its `requirements.txt` is incomplete** — `pip install openai
|
||
pyyaml python-dotenv` as well. Both `openai` and `pyyaml` are imported at
|
||
runtime and missing from that file; without `pyyaml` every single request
|
||
dies inside `_create_system_prompt` *before* reaching the API, and the tool
|
||
reports "API запросов: 0, ошибок: N" while writing `[UNTRANSLATED]` stubs.
|
||
Write the `.env` under `umask 077`, and never echo the key into logs or
|
||
command output.
|
||
3. **Translate:**
|
||
|
||
`python scripts/translate.py book.html book.ru.html --repo <clone> --workdir <dir> --workers 12`
|
||
|
||
⚠️ **Not resumable, despite what the external repo claims.** Its
|
||
`progress/translation_progress.json` stays `{"chapters": {}}` and `--all`
|
||
re-translates every chapter, including ones already sitting in
|
||
`translations/`. Measured 2026-08-21 on Stroustrup: a re-run to repair 4
|
||
failed chapters re-did all 29. Budget a full book on every restart, and
|
||
prefer getting one clean run over patching a partial one.
|
||
|
||
The one thing that *does* skip work: deleting a chapter from `extracted/`
|
||
before the run. It is then never sent, and `rebuild` keeps the original
|
||
English for it — the right treatment for an index.
|
||
**Numbered exercise items come back untranslated.** Measured on Stroustrup
|
||
2026-08-21: 182 paragraphs of 10151 (1.8%) shaped `[2] Expanding on what you
|
||
have learned…` were echoed back in English. The document-wide Cyrillic check
|
||
cannot see this — 1.8% drowns in it — so `rebuild` counts them separately as
|
||
"осталось на языке оригинала". If the count is high, add an explicit line to the
|
||
prompt that numbered items are prose and must be translated too.
|
||
|
||
4. **Read the reported counts.** "без перевода" above zero means a chapter came
|
||
back with a different paragraph count and kept its original text; "разметка
|
||
потеряна" counts paragraphs where the model mangled the inline-tag markers
|
||
and the italics were dropped rather than corrupted. Both are expected to be
|
||
near zero — a large number means the translator misbehaved, not that the
|
||
bridge is broken. The script then checks the output's actual language and
|
||
exits non-zero if the text is still the source language or contains
|
||
`[UNTRANSLATED]` stubs — **matching paragraph counts do not prove anything
|
||
was translated**, since the external tool substitutes the original on API
|
||
failure. Do not pack a file that failed this check.
|
||
**Re-running after a failure needs the state cleared:** the external repo's
|
||
`progress/` directory marks those chapters complete and will skip them.
|
||
Delete `<workdir>/progress`, `<workdir>/context`, and
|
||
`<workdir>/translations` before the retry.
|
||
5. Pack `book.ru.html` with `--language=ru` and translated `--title`/`--authors`.
|
||
|
||
How formatting survives a translator that only speaks plain text: inline `<i>`
|
||
and `<b>` become `⟦i⟧…⟦/i⟧` markers before the text leaves, and are restored
|
||
after. Every paragraph's markers are balance-checked on the way back. Images,
|
||
`<h1>`/`<h2>` structure, and block order never leave this side — only the text
|
||
of each block round-trips, and blocks are reassembled by index.
|
||
|
||
`scripts/test_translate.py` covers the marker round-trip, the broken-marker
|
||
fallback, language detection, and a full split/rebuild against a faked
|
||
translator response — no network, no API key. Run it after touching the bridge.
|
||
|
||
## Optional stage: audiobook
|
||
|
||
Also delegated to `book_translator` (`05_create_audiobook.py`, Microsoft
|
||
edge-tts — free, needs internet). Two wrappers live in `scripts/audiobook.py`;
|
||
the external repo is still never edited.
|
||
|
||
1. **Strip markup markers first.** The translated JSON still holds the
|
||
`⟦i⟧` markers from the translation stage — TTS would read them aloud:
|
||
|
||
`python scripts/audiobook.py prep <workdir>/translations <workdir>/translations_tts`
|
||
|
||
Point the external script at the *stripped* copy; the original keeps its
|
||
italics for the EPUB.
|
||
2. **Synthesize:** `05_create_audiobook.py --translations-dir <…>/translations_tts
|
||
--voice dmitry --rate '+0%'`. Fragments are **one per paragraph** (the
|
||
`--paragraphs-per-group` flag is not used by the loop), named
|
||
`chapter_NNN_intro.mp3` / `chapter_NNN_para_NNNN.mp3` under
|
||
`audiobook/temp_audio/`.
|
||
3. **Verify before the temp files are deleted** — `cleanup_temp_files()` wipes
|
||
`temp_audio/`, and after that only the merged file can be checked:
|
||
|
||
`python scripts/audiobook.py verify <workdir>/translations_tts <workdir>/audiobook`
|
||
|
||
It flags chapters missing fragments, zero-byte/undecodable mp3s, and
|
||
chapters whose duration falls short of what their character count predicts.
|
||
The seconds-per-character baseline is the median across chapters, so it
|
||
self-calibrates to whatever voice and `--rate` were used. Exit code is
|
||
non-zero when anything is wrong.
|
||
4. **Re-running fills gaps cheaply** — the external script skips any fragment
|
||
file that already exists *and is non-empty*, so delete the bad ones `verify`
|
||
named and run it again. It never re-checks that an existing file is sane,
|
||
which is exactly why step 3 exists.
|
||
5. **Split into chapter tracks**, also before the temp files go:
|
||
|
||
```bash
|
||
python scripts/audiobook.py split <workdir>/translations_tts <workdir>/audiobook \
|
||
<workdir>/tracks --album "Название" --author "Автор" --gap 0.3
|
||
```
|
||
|
||
The external stage only ever produces one merged `audiobook_complete.mp3`
|
||
with no chapter marks, which is bad for players. `split` rebuilds per-chapter
|
||
`NNN - Title.mp3` from the same fragments, with ID3 album/artist/track/title
|
||
tags, and refuses a chapter whose fragments are incomplete (`--force`
|
||
overrides). Concatenation is stream-copy, falling back to a re-encode only if
|
||
the mp3 streams don't line up; `--gap` inserts silence between paragraphs,
|
||
generated to match the fragments' own codec parameters so the copy path
|
||
stays viable. Hand the result to the `prepare-audiobooks` skill for covers
|
||
and library layout.
|
||
|
||
### Prefer local Silero over edge-tts
|
||
|
||
edge-tts drops fragments silently under load — measured 1 of 49 on one run and
|
||
7 of 49 on the next, at *fewer* workers, so it is volume- not concurrency-bound.
|
||
`scripts/silero_render.py` replaces the external synthesis stage entirely with a
|
||
local model: no network, no dropouts, no `repair` cycle, free.
|
||
|
||
```bash
|
||
pip install torch numpy --index-url https://download.pytorch.org/whl/cpu
|
||
curl -O https://models.silero.ai/models/tts/ru/v5_5_ru.pt # 145 МБ
|
||
python scripts/silero_render.py v5_5_ru.pt <workdir>/translations_tts <out> \
|
||
--album "Название" --author "Автор" --speaker eugene --tempo 0.87 \
|
||
--skip 0 1 2 3 37
|
||
```
|
||
|
||
**Check `lscpu | grep avx2` before choosing the host.** PyTorch needs AVX2;
|
||
without it inference is ~30x slower — measured 3.4x realtime on a Celeron N5095
|
||
(SSE4 only) versus 97x on a Ryzen 7 5800H. Use 8 threads, not 16: hyperthreading
|
||
loses (97x vs 83x).
|
||
|
||
`scripts/tts_normalize.py` prepares the text — Latin script is what makes a
|
||
Russian voice sound worst, and a full book carries far more of it than a sample
|
||
chapter suggests (37 unique tokens in one chapter, 728 across the book). It maps
|
||
named entities and acronyms by hand with stress marks, transliterates the rest
|
||
by rule, spells out numbers, and drops URLs. It also generates Russian chapter
|
||
titles, since translated headings often stay English and the intro fragment
|
||
would otherwise read them aloud in Latin.
|
||
|
||
Voice tempo: SSML `<prosody rate>` quantizes to named levels, so percentages
|
||
cluster instead of stepping evenly. For fine control use `--tempo`, which
|
||
time-stretches with ffmpeg and preserves pitch.
|
||
|
||
The phonetics stage (`07_extract_terms.py` + `08_generate_phonetics.py`) is
|
||
worth running first for a technical book **when using edge-tts**; with
|
||
`tts_normalize.py` it is redundant.
|
||
|
||
**Ordering constraint for the whole stage:** `verify` and `split` both read
|
||
`audiobook/temp_audio/`, and the external script's `cleanup_temp_files()`
|
||
deletes it right after merging. Run both before that, or the fragments are gone
|
||
and only the merged file's total duration can be checked.
|
||
|
||
`scripts/test_audiobook.py` builds real silent mp3s with ffmpeg and checks that
|
||
`verify` catches a missing fragment, an empty file, and passes a clean book.
|
||
Needs ffmpeg; no network.
|
||
|
||
## What the script does
|
||
|
||
- style from span font names (`-It`, `Italic`, `Bold`, `Semibold`) → `<i>`/`<b>`;
|
||
- headings by font size, sidebars by font family;
|
||
- one PDF block = one paragraph; an unindented block that opens a page is
|
||
appended to the previous page's paragraph;
|
||
- trailing hyphens dropped, adjacent `</i><i>` runs merged, doubled spaces collapsed;
|
||
- images written to `images/`, cover rendered from page 1 at 150 dpi;
|
||
- absolute positioning is deliberately discarded so text reflows at any font size.
|
||
|
||
## Limits
|
||
|
||
- Headers/footers are not stripped — calibre-made PDFs have none. A typeset PDF
|
||
with running heads and page numbers needs blocks filtered by `y` coordinate
|
||
near the margins; add that before trusting the output.
|
||
- Multi-column layouts are not handled; block order would need column sorting.
|
||
- Thresholds are per-layout constants. Always re-run `--fonts` for a new book.
|