diff --git a/README.md b/README.md index 537c77e..e39874a 100644 --- a/README.md +++ b/README.md @@ -13,7 +13,7 @@ размечает основной текст как `

` (2817 штук), потому что определяет заголовки по кеглю относительно самого частого. -`pdf2html.py` извлекает текст через PyMuPDF, где начертание берётся из +`scripts/pdf2html.py` извлекает текст через PyMuPDF, где начертание берётся из свойств span'а, и сам раскладывает блоки по семантике. ## Что делает скрипт @@ -29,6 +29,19 @@ - позиционирование не сохраняется намеренно — текст должен течь под любой размер шрифта на читалке. +## Подключение как скил Claude Code + +Репозиторий одновременно является скилом (`SKILL.md` в корне). Подключается +симлинком, чтобы правки и `git pull` сразу применялись к скилу: + +```bash +git clone https://git.katze-lab.ru/chesirecatt/pdf2epub.git ~/Projects/pdf2epub +ln -s ~/Projects/pdf2epub ~/.claude/skills/pdf-to-kindle +``` + +Проверка: скил `pdf-to-kindle` появляется в списке доступных при следующем +запуске сессии. + ## Требования PyMuPDF и calibre: @@ -45,13 +58,13 @@ sudo dnf install calibre # если ebook-convert ещё нет 15pt/letter, для другой книги их правят по этому выводу): ```bash -~/.venvs/pdf2epub/bin/python pdf2html.py --fonts book.pdf +~/.venvs/pdf2epub/bin/python scripts/pdf2html.py --fonts book.pdf ``` Конвертация: ```bash -~/.venvs/pdf2epub/bin/python pdf2html.py book.pdf out/book.html +~/.venvs/pdf2epub/bin/python scripts/pdf2html.py book.pdf out/book.html ebook-convert out/book.html "Название.epub" \ --title="Название" \ diff --git a/SKILL.md b/SKILL.md new file mode 100644 index 0000000..09aec4e --- /dev/null +++ b/SKILL.md @@ -0,0 +1,111 @@ +--- +name: pdf-to-kindle +description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. +--- + +# PDF to Kindle + +`ebook-convert book.pdf book.epub` silently destroys formatting: poppler infers +style from font names and misses abbreviated ones like `MinionPro-It`, and +calibre labels body text as `

` when the most common font is not the body +font. Measured on a 399-page novel: 435 italic runs in the PDF became 1 `` +tag, and 2817 body paragraphs became `

`. + +`scripts/pdf2html.py` extracts spans with PyMuPDF, where style comes from span +properties, and emits semantic XHTML. calibre is then used only as the +XHTML→EPUB packer. + +## Workflow + +1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No + embedded fonts / no extractable text means a scan — stop and say OCR is + needed; this skill does not apply. +2. **Check tooling.** `ebook-convert` must be on PATH. For PyMuPDF, prefer an + existing interpreter that has it; otherwise build a throwaway venv in the + scratchpad — do not install into the system Python: + + ```bash + python3 -m venv "$SCRATCH/venv" && "$SCRATCH/venv/bin/pip" -q install pymupdf + ``` + +3. **Profile the fonts** before converting anything: + + `python scripts/pdf2html.py --fonts book.pdf` + + Read off: the body font (largest character count), the italic variant, the + heading sizes, and any secondary family used for sidebars, journal entries, + chat logs or slides. +4. **Tune the thresholds** in `block_kind()` and `style()` to that output. The + defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a + title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary + family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that + separates an indented first line from a continuation line — verify it + against the real `x0` values, not by assumption. +5. **Convert to XHTML:** + + `python scripts/pdf2html.py book.pdf out/book.html` + +6. **Verify before packing** (see Verification). Fix thresholds and re-run until + the counts are sane. Cheap to iterate; do not skip to packing. +7. **Pack with calibre**, always passing explicit TOC XPaths — without them + calibre applies its own heuristics and re-breaks the chapters: + + ```bash + ebook-convert out/book.html "Title.epub" \ + --title="Title" --authors="Author" --language=en \ + --cover=out/images/cover.jpg \ + --level1-toc='//h:h1' --level2-toc='//h:h2' \ + --page-breaks-before='//h:h1' \ + --no-default-epub-cover + ebook-convert "Title.epub" "Title.azw3" + ``` + + Take title/author from the user or the book's own title page — PDF metadata + is often an ASIN or a filename. +8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle + (Amazon converts server-side), `.azw3` for USB copy into `documents/`. + Delete intermediate artifacts left in the user's directories. + +## Verification + +Never report success on the converter's own summary line alone. Check: + +```bash +python3 - <<'EOF' +import re +t = open('out/book.html').read() +print('italic:', t.count(''), 'bold:', t.count('')) +print('h1:', len(re.findall(r']+>', '', t).count(' ')) +EOF +``` + +- italic count near zero on a novel means `style()` missed the font-name pattern; +- an `h1` count in the hundreds means a size threshold is too low and body text + is being promoted; +- many double spaces means line joining is off; +- list the extracted headings (`grep -o ']*>.\{0,60\}'`) and read them — + they become the TOC, and a wrong one is obvious at a glance; +- read one full page of body text and confirm paragraphs merge across page + breaks and hyphenated words are rejoined. + +After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx` +navPoint count for the TOC size. + +## What the script does + +- style from span font names (`-It`, `Italic`, `Bold`, `Semibold`) → ``/``; +- headings by font size, sidebars by font family; +- one PDF block = one paragraph; an unindented block that opens a page is + appended to the previous page's paragraph; +- trailing hyphens dropped, adjacent `` runs merged, doubled spaces collapsed; +- images written to `images/`, cover rendered from page 1 at 150 dpi; +- absolute positioning is deliberately discarded so text reflows at any font size. + +## Limits + +- Headers/footers are not stripped — calibre-made PDFs have none. A typeset PDF + with running heads and page numbers needs blocks filtered by `y` coordinate + near the margins; add that before trusting the output. +- Multi-column layouts are not handled; block order would need column sorting. +- Thresholds are per-layout constants. Always re-run `--fonts` for a new book. diff --git a/pdf2html.py b/scripts/pdf2html.py similarity index 100% rename from pdf2html.py rename to scripts/pdf2html.py