--- name: pdf-to-kindle description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. --- # PDF to Kindle `ebook-convert book.pdf book.epub` silently destroys formatting: poppler infers style from font names and misses abbreviated ones like `MinionPro-It`, and calibre labels body text as `

` when the most common font is not the body font. Measured on a 399-page novel: 435 italic runs in the PDF became 1 `` tag, and 2817 body paragraphs became `

`. `scripts/pdf2html.py` extracts spans with PyMuPDF, where style comes from span properties, and emits semantic XHTML. calibre is then used only as the XHTML→EPUB packer. ## Workflow 1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No embedded fonts / no extractable text means a scan — stop and say OCR is needed; this skill does not apply. 2. **Check tooling.** `ebook-convert` must be on PATH. For PyMuPDF, prefer an existing interpreter that has it; otherwise build a throwaway venv in the scratchpad — do not install into the system Python: ```bash python3 -m venv "$SCRATCH/venv" && "$SCRATCH/venv/bin/pip" -q install pymupdf ``` 3. **Profile the fonts** before converting anything: `python scripts/pdf2html.py --fonts book.pdf` Read off: the body font (largest character count), the italic variant, the heading sizes, and any secondary family used for sidebars, journal entries, chat logs or slides. 4. **Tune the thresholds** in `block_kind()` and `style()` to that output. The defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that separates an indented first line from a continuation line — verify it against the real `x0` values, not by assumption. 5. **Convert to XHTML:** `python scripts/pdf2html.py book.pdf out/book.html` 6. **Verify before packing** (see Verification). Fix thresholds and re-run until the counts are sane. Cheap to iterate; do not skip to packing. 7. **Pack with calibre**, always passing explicit TOC XPaths — without them calibre applies its own heuristics and re-breaks the chapters: ```bash ebook-convert out/book.html "Title.epub" \ --title="Title" --authors="Author" --language=en \ --cover=out/images/cover.jpg \ --level1-toc='//h:h1' --level2-toc='//h:h2' \ --page-breaks-before='//h:h1' \ --no-default-epub-cover ebook-convert "Title.epub" "Title.azw3" ``` Take title/author from the user or the book's own title page — PDF metadata is often an ASIN or a filename. 8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle (Amazon converts server-side), `.azw3` for USB copy into `documents/`. Delete intermediate artifacts left in the user's directories. ## Verification Never report success on the converter's own summary line alone. Check: ```bash python3 - <<'EOF' import re t = open('out/book.html').read() print('italic:', t.count(''), 'bold:', t.count('')) print('h1:', len(re.findall(r']+>', '', t).count(' ')) EOF ``` - italic count near zero on a novel means `style()` missed the font-name pattern; - an `h1` count in the hundreds means a size threshold is too low and body text is being promoted; - many double spaces means line joining is off; - list the extracted headings (`grep -o ']*>.\{0,60\}'`) and read them — they become the TOC, and a wrong one is obvious at a glance; - read one full page of body text and confirm paragraphs merge across page breaks and hyphenated words are rejoined. After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx` navPoint count for the TOC size. ## What the script does - style from span font names (`-It`, `Italic`, `Bold`, `Semibold`) → ``/``; - headings by font size, sidebars by font family; - one PDF block = one paragraph; an unindented block that opens a page is appended to the previous page's paragraph; - trailing hyphens dropped, adjacent `` runs merged, doubled spaces collapsed; - images written to `images/`, cover rendered from page 1 at 150 dpi; - absolute positioning is deliberately discarded so text reflows at any font size. ## Limits - Headers/footers are not stripped — calibre-made PDFs have none. A typeset PDF with running heads and page numbers needs blocks filtered by `y` coordinate near the margins; add that before trusting the output. - Multi-column layouts are not handled; block order would need column sorting. - Thresholds are per-layout constants. Always re-run `--fonts` for a new book.