5.1 KiB
name, description
| name | description |
|---|---|
| pdf-to-kindle | Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. |
PDF to Kindle
ebook-convert book.pdf book.epub silently destroys formatting: poppler infers
style from font names and misses abbreviated ones like MinionPro-It, and
calibre labels body text as <h2> when the most common font is not the body
font. Measured on a 399-page novel: 435 italic runs in the PDF became 1 <i>
tag, and 2817 body paragraphs became <h2>.
scripts/pdf2html.py extracts spans with PyMuPDF, where style comes from span
properties, and emits semantic XHTML. calibre is then used only as the
XHTML→EPUB packer.
Workflow
-
Verify a text layer exists.
pdfinfoandpdffontson the file. No embedded fonts / no extractable text means a scan — stop and say OCR is needed; this skill does not apply. -
Check tooling.
ebook-convertmust be on PATH. For PyMuPDF, prefer an existing interpreter that has it; otherwise build a throwaway venv in the scratchpad — do not install into the system Python:python3 -m venv "$SCRATCH/venv" && "$SCRATCH/venv/bin/pip" -q install pymupdf -
Profile the fonts before converting anything:
python scripts/pdf2html.py --fonts book.pdfRead off: the body font (largest character count), the italic variant, the heading sizes, and any secondary family used for sidebars, journal entries, chat logs or slides.
-
Tune the thresholds in
block_kind()andstyle()to that output. The defaults target a 15pt/letter calibre layout:>= 28or a display font is a title,>= 24a chapter/parth1,>= 17anh2subtitle, a secondary family isp.note.INDENT_X(default 88) is the x-coordinate that separates an indented first line from a continuation line — verify it against the realx0values, not by assumption. -
Convert to XHTML:
python scripts/pdf2html.py book.pdf out/book.html -
Verify before packing (see Verification). Fix thresholds and re-run until the counts are sane. Cheap to iterate; do not skip to packing.
-
Pack with calibre, always passing explicit TOC XPaths — without them calibre applies its own heuristics and re-breaks the chapters:
ebook-convert out/book.html "Title.epub" \ --title="Title" --authors="Author" --language=en \ --cover=out/images/cover.jpg \ --level1-toc='//h:h1' --level2-toc='//h:h2' \ --page-breaks-before='//h:h1' \ --no-default-epub-cover ebook-convert "Title.epub" "Title.azw3"Take title/author from the user or the book's own title page — PDF metadata is often an ASIN or a filename.
-
Deliver both files and say what each is for:
.epubfor Send to Kindle (Amazon converts server-side),.azw3for USB copy intodocuments/. Delete intermediate artifacts left in the user's directories.
Verification
Never report success on the converter's own summary line alone. Check:
python3 - <<'EOF'
import re
t = open('out/book.html').read()
print('italic:', t.count('<i>'), 'bold:', t.count('<b>'))
print('h1:', len(re.findall(r'<h1', t)))
print('double spaces:', re.sub('<[^>]+>', '', t).count(' '))
EOF
- italic count near zero on a novel means
style()missed the font-name pattern; - an
h1count in the hundreds means a size threshold is too low and body text is being promoted; - many double spaces means line joining is off;
- list the extracted headings (
grep -o '<h1[^>]*>.\{0,60\}') and read them — they become the TOC, and a wrong one is obvious at a glance; - read one full page of body text and confirm paragraphs merge across page breaks and hyphenated words are rejoined.
After packing, confirm the EPUB: ebook-meta for metadata, and the .ncx
navPoint count for the TOC size.
What the script does
- style from span font names (
-It,Italic,Bold,Semibold) →<i>/<b>; - headings by font size, sidebars by font family;
- one PDF block = one paragraph; an unindented block that opens a page is appended to the previous page's paragraph;
- trailing hyphens dropped, adjacent
</i><i>runs merged, doubled spaces collapsed; - images written to
images/, cover rendered from page 1 at 150 dpi; - absolute positioning is deliberately discarded so text reflows at any font size.
Limits
- Headers/footers are not stripped — calibre-made PDFs have none. A typeset PDF
with running heads and page numbers needs blocks filtered by
ycoordinate near the margins; add that before trusting the output. - Multi-column layouts are not handled; block order would need column sorting.
- Thresholds are per-layout constants. Always re-run
--fontsfor a new book.