Files
pdf2epub/SKILL.md
T

17 KiB

name, description
name description
pdf-to-kindle Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. Also covers translating such a book into another language before packing it, delegated to the external book_translator project, and generating a TTS audiobook from it with an integrity check for dropped fragments and a split into per-chapter tagged tracks.

PDF to Kindle

ebook-convert book.pdf book.epub silently destroys formatting: poppler infers style from font names and misses abbreviated ones like MinionPro-It, and calibre labels body text as <h2> when the most common font is not the body font. Measured on a 399-page novel: 435 italic runs in the PDF became 1 <i> tag, and 2817 body paragraphs became <h2>.

scripts/pdf2html.py extracts spans with PyMuPDF, where style comes from span properties, and emits semantic XHTML. calibre is then used only as the XHTML→EPUB packer.

Two entry points

scripts/pdf2html.py for PDF, scripts/epub2html.py for EPUB. Both emit the same normalized XHTML — one block per line, <pre> for code listings — and everything downstream (translation, packing, audiobook) is identical. Shared document skeleton and CSS live in scripts/bookhtml.py.

Do not route an EPUB through the PDF path. Converting EPUB→PDF→XHTML throws away the semantic markup that is already there and re-derives it from font sizes. epub2html.py needs no font profiling and no threshold tuning: it reads the spine, drops the nav document, and maps the book's own headings.

Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is needed, not PyMuPDF) and convert with:

python scripts/epub2html.py book.epub out/book.html

Dumps converted from PDF carry no semantics at all — no headings, no <pre>, just <p class="class_s1e2"> with obfuscated names. The script detects this (zero headings or zero listings) and rebuilds the structure from the book's own stylesheet: the three largest font sizes become heading levels, a typewriter or monospace family becomes <pre>, and adjacent listing paragraphs merge back into one block. A book that carries its own markup never enters this path. Measured on the dokumen.pub dump of Ousterhout's APoSD: 1833 flat paragraphs became 29 chapters and 58 listings.

It reports which heading level turned out to be the chapter level. In most EPUBs <h1> is the book title and the parts, while chapters are <h2> — the script picks the deepest level that still gives a sane chapter count, because the translation bridge splits the book on <h1>. Measured on Stroustrup's PPP 3rd edition: 7 <h1> against 49 <h2>, chapters correctly detected as h2, 56 chapters out of 656 pages.

Workflow

  1. Verify a text layer exists. pdfinfo and pdffonts on the file. No embedded fonts / no extractable text means a scan — stop and say OCR is needed; this skill does not apply.

  2. Check tooling. ebook-convert must be on PATH. For PyMuPDF, prefer an existing interpreter that has it; otherwise build a throwaway venv in the scratchpad — do not install into the system Python:

    python3 -m venv "$SCRATCH/venv" && "$SCRATCH/venv/bin/pip" -q install pymupdf
    
  3. Profile the fonts before converting anything:

    python scripts/pdf2html.py --fonts book.pdf

    Read off: the body font (largest character count), the italic variant, the heading sizes, and any secondary family used for sidebars, journal entries, chat logs or slides.

  4. Tune the thresholds in block_kind() and style() to that output. Code listings need no tuning — a monospaced span is recognized by the PyMuPDF font flag (flags & 8), not by font name, and becomes a <pre> block. The defaults target a 15pt/letter calibre layout: >= 28 or a display font is a title, >= 24 a chapter/part h1, >= 17 an h2 subtitle, a secondary family is p.note. INDENT_X (default 88) is the x-coordinate that separates an indented first line from a continuation line — verify it against the real x0 values, not by assumption.

  5. Convert to XHTML:

    python scripts/pdf2html.py book.pdf out/book.html

  6. Verify before packing (see Verification). Check the pre: counter against the real number of listings in the book — a technical book reporting pre: 0 means the mono flag never fired and every listing is about to be reflowed as prose. Fix thresholds and re-run until the counts are sane. Cheap to iterate; do not skip to packing.

  7. Pack with calibre, always passing explicit TOC XPaths — without them calibre applies its own heuristics and re-breaks the chapters:

    ebook-convert out/book.html "Title.epub" \
      --title="Title" --authors="Author" --language=en \
      --cover=out/images/cover.jpg \
      --level1-toc='//h:h1' --level2-toc='//h:h2' \
      --page-breaks-before='//h:h1' \
      --no-default-epub-cover
    ebook-convert "Title.epub" "Title.azw3"
    

    Take title/author from the user or the book's own title page — PDF metadata is often an ASIN or a filename.

  8. Deliver both files and say what each is for: .epub for Send to Kindle (Amazon converts server-side), .azw3 for USB copy into documents/. Delete intermediate artifacts left in the user's directories.

Verification

Never report success on the converter's own summary line alone. Check:

python3 - <<'EOF'
import re
t = open('out/book.html').read()
print('italic:', t.count('<i>'), 'bold:', t.count('<b>'))
print('h1:', len(re.findall(r'<h1', t)))
print('double spaces:', re.sub('<[^>]+>', '', t).count('  '))
EOF
  • italic count near zero on a novel means style() missed the font-name pattern;
  • an h1 count in the hundreds means a size threshold is too low and body text is being promoted;
  • many double spaces means line joining is off;
  • list the extracted headings (grep -o '<h1[^>]*>.\{0,60\}') and read them — they become the TOC, and a wrong one is obvious at a glance;
  • read one full page of body text and confirm paragraphs merge across page breaks and hyphenated words are rejoined.

After packing, confirm the EPUB: ebook-meta for metadata, and the .ncx navPoint count for the TOC size.

Optional stage: translation

Only when the user asks for a translated book. It slots between step 6 and step 7 — translate the XHTML, then pack the translated file with calibre.

Code listings are never translated. The bridge only recognizes <p>, <h1> and <h2>; a <pre> block is not parsed, and anything unparsed is carried into the result untouched. Nothing extra is needed to protect code — but this also means a listing that was misclassified as a paragraph upstream will be translated, which is the real reason step 6 checks the pre: counter.

For the same reason <pre> is cut out of both language checks. A book that is 40% listings translates correctly and would otherwise fail acceptance, because the English code drags the Cyrillic share below the threshold.

One workdir per book. Chapter files are named chapter_NNN.json for every book, and the external repo skips chapters it has marked done — a reused workdir silently stitches one book's translation onto another's text. The bridge writes bridge_source.json into the workdir on first run and refuses to start if the directory belongs to a different book. Re-running the same book is unaffected; that is the resume path.

Translation is delegated to an external project, vetermanve/book_translator (DeepSeek or a local Ollama model). Never modify that repository — it is driven through its documented CLI and file formats only. scripts/translate.py is the bridge: it writes the repo's input format, shells out to 03_translate_parallel.py --all, and reads the repo's output format back.

  1. Check the book actually needs translating. The script does this first and refuses to burn API credits on a book already in the target language:

    python scripts/translate.py book.html out.html --repo <clone> --workdir <dir> --check-only

    It reports block/chapter/character counts and the detected source language. --force overrides the refusal.

  2. Set up the external repo once (a plain clone; never edit it) and its credentials — a .env in the working directory with USE_LOCAL_MODEL=false and DEEPSEEK_API_KEY=…, or USE_LOCAL_MODEL=true plus OLLAMA_MODEL for a local model. Its requirements.txt is incomplete — pip install openai pyyaml python-dotenv as well. Both openai and pyyaml are imported at runtime and missing from that file; without pyyaml every single request dies inside _create_system_prompt before reaching the API, and the tool reports "API запросов: 0, ошибок: N" while writing [UNTRANSLATED] stubs. Write the .env under umask 077, and never echo the key into logs or command output.

  3. Translate:

    python scripts/translate.py book.html book.ru.html --repo <clone> --workdir <dir> --workers 12

    ⚠️ Not resumable, despite what the external repo claims. Its progress/translation_progress.json stays {"chapters": {}} and --all re-translates every chapter, including ones already sitting in translations/. Measured 2026-08-21 on Stroustrup: a re-run to repair 4 failed chapters re-did all 29. Budget a full book on every restart, and prefer getting one clean run over patching a partial one.

    The one thing that does skip work: deleting a chapter from extracted/ before the run. It is then never sent, and rebuild keeps the original English for it — the right treatment for an index.

  4. Read the reported counts. "без перевода" above zero means a chapter came back with a different paragraph count and kept its original text; "разметка потеряна" counts paragraphs where the model mangled the inline-tag markers and the italics were dropped rather than corrupted. Both are expected to be near zero — a large number means the translator misbehaved, not that the bridge is broken. The script then checks the output's actual language and exits non-zero if the text is still the source language or contains [UNTRANSLATED] stubs — matching paragraph counts do not prove anything was translated, since the external tool substitutes the original on API failure. Do not pack a file that failed this check. Re-running after a failure needs the state cleared: the external repo's progress/ directory marks those chapters complete and will skip them. Delete <workdir>/progress, <workdir>/context, and <workdir>/translations before the retry.

  5. Pack book.ru.html with --language=ru and translated --title/--authors.

How formatting survives a translator that only speaks plain text: inline <i> and <b> become ⟦i⟧…⟦/i⟧ markers before the text leaves, and are restored after. Every paragraph's markers are balance-checked on the way back. Images, <h1>/<h2> structure, and block order never leave this side — only the text of each block round-trips, and blocks are reassembled by index.

scripts/test_translate.py covers the marker round-trip, the broken-marker fallback, language detection, and a full split/rebuild against a faked translator response — no network, no API key. Run it after touching the bridge.

Optional stage: audiobook

Also delegated to book_translator (05_create_audiobook.py, Microsoft edge-tts — free, needs internet). Two wrappers live in scripts/audiobook.py; the external repo is still never edited.

  1. Strip markup markers first. The translated JSON still holds the ⟦i⟧ markers from the translation stage — TTS would read them aloud:

    python scripts/audiobook.py prep <workdir>/translations <workdir>/translations_tts

    Point the external script at the stripped copy; the original keeps its italics for the EPUB.

  2. Synthesize: 05_create_audiobook.py --translations-dir <…>/translations_tts --voice dmitry --rate '+0%'. Fragments are one per paragraph (the --paragraphs-per-group flag is not used by the loop), named chapter_NNN_intro.mp3 / chapter_NNN_para_NNNN.mp3 under audiobook/temp_audio/.

  3. Verify before the temp files are deleted — cleanup_temp_files() wipes temp_audio/, and after that only the merged file can be checked:

    python scripts/audiobook.py verify <workdir>/translations_tts <workdir>/audiobook

    It flags chapters missing fragments, zero-byte/undecodable mp3s, and chapters whose duration falls short of what their character count predicts. The seconds-per-character baseline is the median across chapters, so it self-calibrates to whatever voice and --rate were used. Exit code is non-zero when anything is wrong.

  4. Re-running fills gaps cheaply — the external script skips any fragment file that already exists and is non-empty, so delete the bad ones verify named and run it again. It never re-checks that an existing file is sane, which is exactly why step 3 exists.

  5. Split into chapter tracks, also before the temp files go:

    python scripts/audiobook.py split <workdir>/translations_tts <workdir>/audiobook \
      <workdir>/tracks --album "Название" --author "Автор" --gap 0.3
    

    The external stage only ever produces one merged audiobook_complete.mp3 with no chapter marks, which is bad for players. split rebuilds per-chapter NNN - Title.mp3 from the same fragments, with ID3 album/artist/track/title tags, and refuses a chapter whose fragments are incomplete (--force overrides). Concatenation is stream-copy, falling back to a re-encode only if the mp3 streams don't line up; --gap inserts silence between paragraphs, generated to match the fragments' own codec parameters so the copy path stays viable. Hand the result to the prepare-audiobooks skill for covers and library layout.

Prefer local Silero over edge-tts

edge-tts drops fragments silently under load — measured 1 of 49 on one run and 7 of 49 on the next, at fewer workers, so it is volume- not concurrency-bound. scripts/silero_render.py replaces the external synthesis stage entirely with a local model: no network, no dropouts, no repair cycle, free.

pip install torch numpy --index-url https://download.pytorch.org/whl/cpu
curl -O https://models.silero.ai/models/tts/ru/v5_5_ru.pt      # 145 МБ
python scripts/silero_render.py v5_5_ru.pt <workdir>/translations_tts <out> \
  --album "Название" --author "Автор" --speaker eugene --tempo 0.87 \
  --skip 0 1 2 3 37

Check lscpu | grep avx2 before choosing the host. PyTorch needs AVX2; without it inference is ~30x slower — measured 3.4x realtime on a Celeron N5095 (SSE4 only) versus 97x on a Ryzen 7 5800H. Use 8 threads, not 16: hyperthreading loses (97x vs 83x).

scripts/tts_normalize.py prepares the text — Latin script is what makes a Russian voice sound worst, and a full book carries far more of it than a sample chapter suggests (37 unique tokens in one chapter, 728 across the book). It maps named entities and acronyms by hand with stress marks, transliterates the rest by rule, spells out numbers, and drops URLs. It also generates Russian chapter titles, since translated headings often stay English and the intro fragment would otherwise read them aloud in Latin.

Voice tempo: SSML <prosody rate> quantizes to named levels, so percentages cluster instead of stepping evenly. For fine control use --tempo, which time-stretches with ffmpeg and preserves pitch.

The phonetics stage (07_extract_terms.py + 08_generate_phonetics.py) is worth running first for a technical book when using edge-tts; with tts_normalize.py it is redundant.

Ordering constraint for the whole stage: verify and split both read audiobook/temp_audio/, and the external script's cleanup_temp_files() deletes it right after merging. Run both before that, or the fragments are gone and only the merged file's total duration can be checked.

scripts/test_audiobook.py builds real silent mp3s with ffmpeg and checks that verify catches a missing fragment, an empty file, and passes a clean book. Needs ffmpeg; no network.

What the script does

  • style from span font names (-It, Italic, Bold, Semibold) → <i>/<b>;
  • headings by font size, sidebars by font family;
  • one PDF block = one paragraph; an unindented block that opens a page is appended to the previous page's paragraph;
  • trailing hyphens dropped, adjacent </i><i> runs merged, doubled spaces collapsed;
  • images written to images/, cover rendered from page 1 at 150 dpi;
  • absolute positioning is deliberately discarded so text reflows at any font size.

Limits

  • Headers/footers are not stripped — calibre-made PDFs have none. A typeset PDF with running heads and page numbers needs blocks filtered by y coordinate near the margins; add that before trusting the output.
  • Multi-column layouts are not handled; block order would need column sorting.
  • Thresholds are per-layout constants. Always re-run --fonts for a new book.