--- name: pdf-to-kindle description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. Also covers translating such a book into another language before packing it, delegated to the external book_translator project, and generating a TTS audiobook from it with an integrity check for dropped fragments and a split into per-chapter tagged tracks. --- # PDF to Kindle `ebook-convert book.pdf book.epub` silently destroys formatting: poppler infers style from font names and misses abbreviated ones like `MinionPro-It`, and calibre labels body text as `

` when the most common font is not the body font. Measured on a 399-page novel: 435 italic runs in the PDF became 1 `` tag, and 2817 body paragraphs became `

`. `scripts/pdf2html.py` extracts spans with PyMuPDF, where style comes from span properties, and emits semantic XHTML. calibre is then used only as the XHTML→EPUB packer. ## Workflow 1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No embedded fonts / no extractable text means a scan — stop and say OCR is needed; this skill does not apply. 2. **Check tooling.** `ebook-convert` must be on PATH. For PyMuPDF, prefer an existing interpreter that has it; otherwise build a throwaway venv in the scratchpad — do not install into the system Python: ```bash python3 -m venv "$SCRATCH/venv" && "$SCRATCH/venv/bin/pip" -q install pymupdf ``` 3. **Profile the fonts** before converting anything: `python scripts/pdf2html.py --fonts book.pdf` Read off: the body font (largest character count), the italic variant, the heading sizes, and any secondary family used for sidebars, journal entries, chat logs or slides. 4. **Tune the thresholds** in `block_kind()` and `style()` to that output. The defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that separates an indented first line from a continuation line — verify it against the real `x0` values, not by assumption. 5. **Convert to XHTML:** `python scripts/pdf2html.py book.pdf out/book.html` 6. **Verify before packing** (see Verification). Fix thresholds and re-run until the counts are sane. Cheap to iterate; do not skip to packing. 7. **Pack with calibre**, always passing explicit TOC XPaths — without them calibre applies its own heuristics and re-breaks the chapters: ```bash ebook-convert out/book.html "Title.epub" \ --title="Title" --authors="Author" --language=en \ --cover=out/images/cover.jpg \ --level1-toc='//h:h1' --level2-toc='//h:h2' \ --page-breaks-before='//h:h1' \ --no-default-epub-cover ebook-convert "Title.epub" "Title.azw3" ``` Take title/author from the user or the book's own title page — PDF metadata is often an ASIN or a filename. 8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle (Amazon converts server-side), `.azw3` for USB copy into `documents/`. Delete intermediate artifacts left in the user's directories. ## Verification Never report success on the converter's own summary line alone. Check: ```bash python3 - <<'EOF' import re t = open('out/book.html').read() print('italic:', t.count(''), 'bold:', t.count('')) print('h1:', len(re.findall(r']+>', '', t).count(' ')) EOF ``` - italic count near zero on a novel means `style()` missed the font-name pattern; - an `h1` count in the hundreds means a size threshold is too low and body text is being promoted; - many double spaces means line joining is off; - list the extracted headings (`grep -o ']*>.\{0,60\}'`) and read them — they become the TOC, and a wrong one is obvious at a glance; - read one full page of body text and confirm paragraphs merge across page breaks and hyphenated words are rejoined. After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx` navPoint count for the TOC size. ## Optional stage: translation Only when the user asks for a translated book. It slots between step 6 and step 7 — translate the XHTML, then pack the translated file with calibre. Translation is delegated to an external project, [`vetermanve/book_translator`](https://github.com/vetermanve/book_translator) (DeepSeek or a local Ollama model). **Never modify that repository** — it is driven through its documented CLI and file formats only. `scripts/translate.py` is the bridge: it writes the repo's input format, shells out to `03_translate_parallel.py --all`, and reads the repo's output format back. 1. **Check the book actually needs translating.** The script does this first and refuses to burn API credits on a book already in the target language: `python scripts/translate.py book.html out.html --repo --workdir --check-only` It reports block/chapter/character counts and the detected source language. `--force` overrides the refusal. 2. **Set up the external repo once** (a plain clone; never edit it) and its credentials — a `.env` in the working directory with `USE_LOCAL_MODEL=false` and `DEEPSEEK_API_KEY=…`, or `USE_LOCAL_MODEL=true` plus `OLLAMA_MODEL` for a local model. **Its `requirements.txt` is incomplete** — `pip install openai pyyaml python-dotenv` as well. Both `openai` and `pyyaml` are imported at runtime and missing from that file; without `pyyaml` every single request dies inside `_create_system_prompt` *before* reaching the API, and the tool reports "API запросов: 0, ошибок: N" while writing `[UNTRANSLATED]` stubs. Write the `.env` under `umask 077`, and never echo the key into logs or command output. 3. **Translate:** `python scripts/translate.py book.html book.ru.html --repo --workdir --workers 12` Resumable: the external repo tracks completed chapters and skips them on a re-run, so an interrupted run costs nothing to restart. 4. **Read the reported counts.** "без перевода" above zero means a chapter came back with a different paragraph count and kept its original text; "разметка потеряна" counts paragraphs where the model mangled the inline-tag markers and the italics were dropped rather than corrupted. Both are expected to be near zero — a large number means the translator misbehaved, not that the bridge is broken. The script then checks the output's actual language and exits non-zero if the text is still the source language or contains `[UNTRANSLATED]` stubs — **matching paragraph counts do not prove anything was translated**, since the external tool substitutes the original on API failure. Do not pack a file that failed this check. **Re-running after a failure needs the state cleared:** the external repo's `progress/` directory marks those chapters complete and will skip them. Delete `/progress`, `/context`, and `/translations` before the retry. 5. Pack `book.ru.html` with `--language=ru` and translated `--title`/`--authors`. How formatting survives a translator that only speaks plain text: inline `` and `` become `⟦i⟧…⟦/i⟧` markers before the text leaves, and are restored after. Every paragraph's markers are balance-checked on the way back. Images, `

`/`

` structure, and block order never leave this side — only the text of each block round-trips, and blocks are reassembled by index. `scripts/test_translate.py` covers the marker round-trip, the broken-marker fallback, language detection, and a full split/rebuild against a faked translator response — no network, no API key. Run it after touching the bridge. ## Optional stage: audiobook Also delegated to `book_translator` (`05_create_audiobook.py`, Microsoft edge-tts — free, needs internet). Two wrappers live in `scripts/audiobook.py`; the external repo is still never edited. 1. **Strip markup markers first.** The translated JSON still holds the `⟦i⟧` markers from the translation stage — TTS would read them aloud: `python scripts/audiobook.py prep /translations /translations_tts` Point the external script at the *stripped* copy; the original keeps its italics for the EPUB. 2. **Synthesize:** `05_create_audiobook.py --translations-dir <…>/translations_tts --voice dmitry --rate '+0%'`. Fragments are **one per paragraph** (the `--paragraphs-per-group` flag is not used by the loop), named `chapter_NNN_intro.mp3` / `chapter_NNN_para_NNNN.mp3` under `audiobook/temp_audio/`. 3. **Verify before the temp files are deleted** — `cleanup_temp_files()` wipes `temp_audio/`, and after that only the merged file can be checked: `python scripts/audiobook.py verify /translations_tts /audiobook` It flags chapters missing fragments, zero-byte/undecodable mp3s, and chapters whose duration falls short of what their character count predicts. The seconds-per-character baseline is the median across chapters, so it self-calibrates to whatever voice and `--rate` were used. Exit code is non-zero when anything is wrong. 4. **Re-running fills gaps cheaply** — the external script skips any fragment file that already exists *and is non-empty*, so delete the bad ones `verify` named and run it again. It never re-checks that an existing file is sane, which is exactly why step 3 exists. 5. **Split into chapter tracks**, also before the temp files go: ```bash python scripts/audiobook.py split /translations_tts /audiobook \ /tracks --album "Название" --author "Автор" --gap 0.3 ``` The external stage only ever produces one merged `audiobook_complete.mp3` with no chapter marks, which is bad for players. `split` rebuilds per-chapter `NNN - Title.mp3` from the same fragments, with ID3 album/artist/track/title tags, and refuses a chapter whose fragments are incomplete (`--force` overrides). Concatenation is stream-copy, falling back to a re-encode only if the mp3 streams don't line up; `--gap` inserts silence between paragraphs, generated to match the fragments' own codec parameters so the copy path stays viable. Hand the result to the `prepare-audiobooks` skill for covers and library layout. The phonetics stage (`07_extract_terms.py` + `08_generate_phonetics.py`) is worth running first for a technical book, or the Russian voice will mangle every English term. **Ordering constraint for the whole stage:** `verify` and `split` both read `audiobook/temp_audio/`, and the external script's `cleanup_temp_files()` deletes it right after merging. Run both before that, or the fragments are gone and only the merged file's total duration can be checked. `scripts/test_audiobook.py` builds real silent mp3s with ffmpeg and checks that `verify` catches a missing fragment, an empty file, and passes a clean book. Needs ffmpeg; no network. ## What the script does - style from span font names (`-It`, `Italic`, `Bold`, `Semibold`) → ``/``; - headings by font size, sidebars by font family; - one PDF block = one paragraph; an unindented block that opens a page is appended to the previous page's paragraph; - trailing hyphens dropped, adjacent `` runs merged, doubled spaces collapsed; - images written to `images/`, cover rendered from page 1 at 150 dpi; - absolute positioning is deliberately discarded so text reflows at any font size. ## Limits - Headers/footers are not stripped — calibre-made PDFs have none. A typeset PDF with running heads and page numbers needs blocks filtered by `y` coordinate near the margins; add that before trusting the output. - Multi-column layouts are not handled; block order would need column sorting. - Thresholds are per-layout constants. Always re-run `--fonts` for a new book.