---
name: pdf-to-kindle
description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. Also covers translating such a book into another language before packing it, delegated to the external book_translator project.
---
# PDF to Kindle
`ebook-convert book.pdf book.epub` silently destroys formatting: poppler infers
style from font names and misses abbreviated ones like `MinionPro-It`, and
calibre labels body text as `
` when the most common font is not the body
font. Measured on a 399-page novel: 435 italic runs in the PDF became 1 ``
tag, and 2817 body paragraphs became ``.
`scripts/pdf2html.py` extracts spans with PyMuPDF, where style comes from span
properties, and emits semantic XHTML. calibre is then used only as the
XHTML→EPUB packer.
## Workflow
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
embedded fonts / no extractable text means a scan — stop and say OCR is
needed; this skill does not apply.
2. **Check tooling.** `ebook-convert` must be on PATH. For PyMuPDF, prefer an
existing interpreter that has it; otherwise build a throwaway venv in the
scratchpad — do not install into the system Python:
```bash
python3 -m venv "$SCRATCH/venv" && "$SCRATCH/venv/bin/pip" -q install pymupdf
```
3. **Profile the fonts** before converting anything:
`python scripts/pdf2html.py --fonts book.pdf`
Read off: the body font (largest character count), the italic variant, the
heading sizes, and any secondary family used for sidebars, journal entries,
chat logs or slides.
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. The
defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a
title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary
family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that
separates an indented first line from a continuation line — verify it
against the real `x0` values, not by assumption.
5. **Convert to XHTML:**
`python scripts/pdf2html.py book.pdf out/book.html`
6. **Verify before packing** (see Verification). Fix thresholds and re-run until
the counts are sane. Cheap to iterate; do not skip to packing.
7. **Pack with calibre**, always passing explicit TOC XPaths — without them
calibre applies its own heuristics and re-breaks the chapters:
```bash
ebook-convert out/book.html "Title.epub" \
--title="Title" --authors="Author" --language=en \
--cover=out/images/cover.jpg \
--level1-toc='//h:h1' --level2-toc='//h:h2' \
--page-breaks-before='//h:h1' \
--no-default-epub-cover
ebook-convert "Title.epub" "Title.azw3"
```
Take title/author from the user or the book's own title page — PDF metadata
is often an ASIN or a filename.
8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle
(Amazon converts server-side), `.azw3` for USB copy into `documents/`.
Delete intermediate artifacts left in the user's directories.
## Verification
Never report success on the converter's own summary line alone. Check:
```bash
python3 - <<'EOF'
import re
t = open('out/book.html').read()
print('italic:', t.count(''), 'bold:', t.count(''))
print('h1:', len(re.findall(r']+>', '', t).count(' '))
EOF
```
- italic count near zero on a novel means `style()` missed the font-name pattern;
- an `h1` count in the hundreds means a size threshold is too low and body text
is being promoted;
- many double spaces means line joining is off;
- list the extracted headings (`grep -o ']*>.\{0,60\}'`) and read them —
they become the TOC, and a wrong one is obvious at a glance;
- read one full page of body text and confirm paragraphs merge across page
breaks and hyphenated words are rejoined.
After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx`
navPoint count for the TOC size.
## Optional stage: translation
Only when the user asks for a translated book. It slots between step 6 and
step 7 — translate the XHTML, then pack the translated file with calibre.
Translation is delegated to an external project,
[`vetermanve/book_translator`](https://github.com/vetermanve/book_translator)
(DeepSeek or a local Ollama model). **Never modify that repository** — it is
driven through its documented CLI and file formats only. `scripts/translate.py`
is the bridge: it writes the repo's input format, shells out to
`03_translate_parallel.py --all`, and reads the repo's output format back.
1. **Check the book actually needs translating.** The script does this first
and refuses to burn API credits on a book already in the target language:
`python scripts/translate.py book.html out.html --repo --workdir --check-only`
It reports block/chapter/character counts and the detected source language.
`--force` overrides the refusal.
2. **Set up the external repo once** (a plain clone; never edit it) and its
credentials — a `.env` in the working directory with `USE_LOCAL_MODEL=false`
and `DEEPSEEK_API_KEY=…`, or `USE_LOCAL_MODEL=true` plus `OLLAMA_MODEL` for a
local model. **Its `requirements.txt` is incomplete** — `pip install openai
pyyaml python-dotenv` as well. Both `openai` and `pyyaml` are imported at
runtime and missing from that file; without `pyyaml` every single request
dies inside `_create_system_prompt` *before* reaching the API, and the tool
reports "API запросов: 0, ошибок: N" while writing `[UNTRANSLATED]` stubs.
Write the `.env` under `umask 077`, and never echo the key into logs or
command output.
3. **Translate:**
`python scripts/translate.py book.html book.ru.html --repo --workdir --workers 12`
Resumable: the external repo tracks completed chapters and skips them on a
re-run, so an interrupted run costs nothing to restart.
4. **Read the reported counts.** "без перевода" above zero means a chapter came
back with a different paragraph count and kept its original text; "разметка
потеряна" counts paragraphs where the model mangled the inline-tag markers
and the italics were dropped rather than corrupted. Both are expected to be
near zero — a large number means the translator misbehaved, not that the
bridge is broken. The script then checks the output's actual language and
exits non-zero if the text is still the source language or contains
`[UNTRANSLATED]` stubs — **matching paragraph counts do not prove anything
was translated**, since the external tool substitutes the original on API
failure. Do not pack a file that failed this check.
**Re-running after a failure needs the state cleared:** the external repo's
`progress/` directory marks those chapters complete and will skip them.
Delete `/progress`, `/context`, and
`/translations` before the retry.
5. Pack `book.ru.html` with `--language=ru` and translated `--title`/`--authors`.
How formatting survives a translator that only speaks plain text: inline ``
and `` become `⟦i⟧…⟦/i⟧` markers before the text leaves, and are restored
after. Every paragraph's markers are balance-checked on the way back. Images,
``/`` structure, and block order never leave this side — only the text
of each block round-trips, and blocks are reassembled by index.
`scripts/test_translate.py` covers the marker round-trip, the broken-marker
fallback, language detection, and a full split/rebuild against a faked
translator response — no network, no API key. Run it after touching the bridge.
## What the script does
- style from span font names (`-It`, `Italic`, `Bold`, `Semibold`) → ``/``;
- headings by font size, sidebars by font family;
- one PDF block = one paragraph; an unindented block that opens a page is
appended to the previous page's paragraph;
- trailing hyphens dropped, adjacent `` runs merged, doubled spaces collapsed;
- images written to `images/`, cover rendered from page 1 at 150 dpi;
- absolute positioning is deliberately discarded so text reflows at any font size.
## Limits
- Headers/footers are not stripped — calibre-made PDFs have none. A typeset PDF
with running heads and page numbers needs blocks filtered by `y` coordinate
near the margins; add that before trusting the output.
- Multi-column layouts are not handled; block order would need column sorting.
- Thresholds are per-layout constants. Always re-run `--fonts` for a new book.