Files
pdf2epub/SKILL.md
T

7.8 KiB

name, description
name description
pdf-to-kindle Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. Also covers translating such a book into another language before packing it, delegated to the external book_translator project.

PDF to Kindle

ebook-convert book.pdf book.epub silently destroys formatting: poppler infers style from font names and misses abbreviated ones like MinionPro-It, and calibre labels body text as <h2> when the most common font is not the body font. Measured on a 399-page novel: 435 italic runs in the PDF became 1 <i> tag, and 2817 body paragraphs became <h2>.

scripts/pdf2html.py extracts spans with PyMuPDF, where style comes from span properties, and emits semantic XHTML. calibre is then used only as the XHTML→EPUB packer.

Workflow

  1. Verify a text layer exists. pdfinfo and pdffonts on the file. No embedded fonts / no extractable text means a scan — stop and say OCR is needed; this skill does not apply.

  2. Check tooling. ebook-convert must be on PATH. For PyMuPDF, prefer an existing interpreter that has it; otherwise build a throwaway venv in the scratchpad — do not install into the system Python:

    python3 -m venv "$SCRATCH/venv" && "$SCRATCH/venv/bin/pip" -q install pymupdf
    
  3. Profile the fonts before converting anything:

    python scripts/pdf2html.py --fonts book.pdf

    Read off: the body font (largest character count), the italic variant, the heading sizes, and any secondary family used for sidebars, journal entries, chat logs or slides.

  4. Tune the thresholds in block_kind() and style() to that output. The defaults target a 15pt/letter calibre layout: >= 28 or a display font is a title, >= 24 a chapter/part h1, >= 17 an h2 subtitle, a secondary family is p.note. INDENT_X (default 88) is the x-coordinate that separates an indented first line from a continuation line — verify it against the real x0 values, not by assumption.

  5. Convert to XHTML:

    python scripts/pdf2html.py book.pdf out/book.html

  6. Verify before packing (see Verification). Fix thresholds and re-run until the counts are sane. Cheap to iterate; do not skip to packing.

  7. Pack with calibre, always passing explicit TOC XPaths — without them calibre applies its own heuristics and re-breaks the chapters:

    ebook-convert out/book.html "Title.epub" \
      --title="Title" --authors="Author" --language=en \
      --cover=out/images/cover.jpg \
      --level1-toc='//h:h1' --level2-toc='//h:h2' \
      --page-breaks-before='//h:h1' \
      --no-default-epub-cover
    ebook-convert "Title.epub" "Title.azw3"
    

    Take title/author from the user or the book's own title page — PDF metadata is often an ASIN or a filename.

  8. Deliver both files and say what each is for: .epub for Send to Kindle (Amazon converts server-side), .azw3 for USB copy into documents/. Delete intermediate artifacts left in the user's directories.

Verification

Never report success on the converter's own summary line alone. Check:

python3 - <<'EOF'
import re
t = open('out/book.html').read()
print('italic:', t.count('<i>'), 'bold:', t.count('<b>'))
print('h1:', len(re.findall(r'<h1', t)))
print('double spaces:', re.sub('<[^>]+>', '', t).count('  '))
EOF
  • italic count near zero on a novel means style() missed the font-name pattern;
  • an h1 count in the hundreds means a size threshold is too low and body text is being promoted;
  • many double spaces means line joining is off;
  • list the extracted headings (grep -o '<h1[^>]*>.\{0,60\}') and read them — they become the TOC, and a wrong one is obvious at a glance;
  • read one full page of body text and confirm paragraphs merge across page breaks and hyphenated words are rejoined.

After packing, confirm the EPUB: ebook-meta for metadata, and the .ncx navPoint count for the TOC size.

Optional stage: translation

Only when the user asks for a translated book. It slots between step 6 and step 7 — translate the XHTML, then pack the translated file with calibre.

Translation is delegated to an external project, vetermanve/book_translator (DeepSeek or a local Ollama model). Never modify that repository — it is driven through its documented CLI and file formats only. scripts/translate.py is the bridge: it writes the repo's input format, shells out to 03_translate_parallel.py --all, and reads the repo's output format back.

  1. Check the book actually needs translating. The script does this first and refuses to burn API credits on a book already in the target language:

    python scripts/translate.py book.html out.html --repo <clone> --workdir <dir> --check-only

    It reports block/chapter/character counts and the detected source language. --force overrides the refusal.

  2. Set up the external repo once (a plain clone; never edit it), and its credentials — DEEPSEEK_API_KEY in the environment or a .env in the working directory, or USE_OLLAMA for a local model. Its requirements.txt omits openai, which deepseek_translator.py imports — install it too.

  3. Translate:

    python scripts/translate.py book.html book.ru.html --repo <clone> --workdir <dir> --workers 12

    Resumable: the external repo tracks completed chapters and skips them on a re-run, so an interrupted run costs nothing to restart.

  4. Read the reported counts. "без перевода" above zero means a chapter came back with a different paragraph count and kept its original text; "разметка потеряна" counts paragraphs where the model mangled the inline-tag markers and the italics were dropped rather than corrupted. Both are expected to be near zero — a large number means the translator misbehaved, not that the bridge is broken.

  5. Pack book.ru.html with --language=ru and translated --title/--authors.

How formatting survives a translator that only speaks plain text: inline <i> and <b> become ⟦i⟧…⟦/i⟧ markers before the text leaves, and are restored after. Every paragraph's markers are balance-checked on the way back. Images, <h1>/<h2> structure, and block order never leave this side — only the text of each block round-trips, and blocks are reassembled by index.

scripts/test_translate.py covers the marker round-trip, the broken-marker fallback, language detection, and a full split/rebuild against a faked translator response — no network, no API key. Run it after touching the bridge.

What the script does

  • style from span font names (-It, Italic, Bold, Semibold) → <i>/<b>;
  • headings by font size, sidebars by font family;
  • one PDF block = one paragraph; an unindented block that opens a page is appended to the previous page's paragraph;
  • trailing hyphens dropped, adjacent </i><i> runs merged, doubled spaces collapsed;
  • images written to images/, cover rendered from page 1 at 150 dpi;
  • absolute positioning is deliberately discarded so text reflows at any font size.

Limits

  • Headers/footers are not stripped — calibre-made PDFs have none. A typeset PDF with running heads and page numbers needs blocks filtered by y coordinate near the margins; add that before trusting the output.
  • Multi-column layouts are not handled; block order would need column sorting.
  • Thresholds are per-layout constants. Always re-run --fonts for a new book.