Оформить репозиторий как скил pdf-to-kindle: SKILL.md, скрипт в scripts/, установка симлинком
This commit is contained in:
@@ -13,7 +13,7 @@
|
|||||||
размечает основной текст как `<h2>` (2817 штук), потому что определяет заголовки
|
размечает основной текст как `<h2>` (2817 штук), потому что определяет заголовки
|
||||||
по кеглю относительно самого частого.
|
по кеглю относительно самого частого.
|
||||||
|
|
||||||
`pdf2html.py` извлекает текст через PyMuPDF, где начертание берётся из
|
`scripts/pdf2html.py` извлекает текст через PyMuPDF, где начертание берётся из
|
||||||
свойств span'а, и сам раскладывает блоки по семантике.
|
свойств span'а, и сам раскладывает блоки по семантике.
|
||||||
|
|
||||||
## Что делает скрипт
|
## Что делает скрипт
|
||||||
@@ -29,6 +29,19 @@
|
|||||||
- позиционирование не сохраняется намеренно — текст должен течь под любой
|
- позиционирование не сохраняется намеренно — текст должен течь под любой
|
||||||
размер шрифта на читалке.
|
размер шрифта на читалке.
|
||||||
|
|
||||||
|
## Подключение как скил Claude Code
|
||||||
|
|
||||||
|
Репозиторий одновременно является скилом (`SKILL.md` в корне). Подключается
|
||||||
|
симлинком, чтобы правки и `git pull` сразу применялись к скилу:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git clone https://git.katze-lab.ru/chesirecatt/pdf2epub.git ~/Projects/pdf2epub
|
||||||
|
ln -s ~/Projects/pdf2epub ~/.claude/skills/pdf-to-kindle
|
||||||
|
```
|
||||||
|
|
||||||
|
Проверка: скил `pdf-to-kindle` появляется в списке доступных при следующем
|
||||||
|
запуске сессии.
|
||||||
|
|
||||||
## Требования
|
## Требования
|
||||||
|
|
||||||
PyMuPDF и calibre:
|
PyMuPDF и calibre:
|
||||||
@@ -45,13 +58,13 @@ sudo dnf install calibre # если ebook-convert ещё нет
|
|||||||
15pt/letter, для другой книги их правят по этому выводу):
|
15pt/letter, для другой книги их правят по этому выводу):
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
~/.venvs/pdf2epub/bin/python pdf2html.py --fonts book.pdf
|
~/.venvs/pdf2epub/bin/python scripts/pdf2html.py --fonts book.pdf
|
||||||
```
|
```
|
||||||
|
|
||||||
Конвертация:
|
Конвертация:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
~/.venvs/pdf2epub/bin/python pdf2html.py book.pdf out/book.html
|
~/.venvs/pdf2epub/bin/python scripts/pdf2html.py book.pdf out/book.html
|
||||||
|
|
||||||
ebook-convert out/book.html "Название.epub" \
|
ebook-convert out/book.html "Название.epub" \
|
||||||
--title="Название" \
|
--title="Название" \
|
||||||
|
|||||||
@@ -0,0 +1,111 @@
|
|||||||
|
---
|
||||||
|
name: pdf-to-kindle
|
||||||
|
description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first.
|
||||||
|
---
|
||||||
|
|
||||||
|
# PDF to Kindle
|
||||||
|
|
||||||
|
`ebook-convert book.pdf book.epub` silently destroys formatting: poppler infers
|
||||||
|
style from font names and misses abbreviated ones like `MinionPro-It`, and
|
||||||
|
calibre labels body text as `<h2>` when the most common font is not the body
|
||||||
|
font. Measured on a 399-page novel: 435 italic runs in the PDF became 1 `<i>`
|
||||||
|
tag, and 2817 body paragraphs became `<h2>`.
|
||||||
|
|
||||||
|
`scripts/pdf2html.py` extracts spans with PyMuPDF, where style comes from span
|
||||||
|
properties, and emits semantic XHTML. calibre is then used only as the
|
||||||
|
XHTML→EPUB packer.
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
|
||||||
|
embedded fonts / no extractable text means a scan — stop and say OCR is
|
||||||
|
needed; this skill does not apply.
|
||||||
|
2. **Check tooling.** `ebook-convert` must be on PATH. For PyMuPDF, prefer an
|
||||||
|
existing interpreter that has it; otherwise build a throwaway venv in the
|
||||||
|
scratchpad — do not install into the system Python:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 -m venv "$SCRATCH/venv" && "$SCRATCH/venv/bin/pip" -q install pymupdf
|
||||||
|
```
|
||||||
|
|
||||||
|
3. **Profile the fonts** before converting anything:
|
||||||
|
|
||||||
|
`python scripts/pdf2html.py --fonts book.pdf`
|
||||||
|
|
||||||
|
Read off: the body font (largest character count), the italic variant, the
|
||||||
|
heading sizes, and any secondary family used for sidebars, journal entries,
|
||||||
|
chat logs or slides.
|
||||||
|
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. The
|
||||||
|
defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a
|
||||||
|
title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary
|
||||||
|
family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that
|
||||||
|
separates an indented first line from a continuation line — verify it
|
||||||
|
against the real `x0` values, not by assumption.
|
||||||
|
5. **Convert to XHTML:**
|
||||||
|
|
||||||
|
`python scripts/pdf2html.py book.pdf out/book.html`
|
||||||
|
|
||||||
|
6. **Verify before packing** (see Verification). Fix thresholds and re-run until
|
||||||
|
the counts are sane. Cheap to iterate; do not skip to packing.
|
||||||
|
7. **Pack with calibre**, always passing explicit TOC XPaths — without them
|
||||||
|
calibre applies its own heuristics and re-breaks the chapters:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ebook-convert out/book.html "Title.epub" \
|
||||||
|
--title="Title" --authors="Author" --language=en \
|
||||||
|
--cover=out/images/cover.jpg \
|
||||||
|
--level1-toc='//h:h1' --level2-toc='//h:h2' \
|
||||||
|
--page-breaks-before='//h:h1' \
|
||||||
|
--no-default-epub-cover
|
||||||
|
ebook-convert "Title.epub" "Title.azw3"
|
||||||
|
```
|
||||||
|
|
||||||
|
Take title/author from the user or the book's own title page — PDF metadata
|
||||||
|
is often an ASIN or a filename.
|
||||||
|
8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle
|
||||||
|
(Amazon converts server-side), `.azw3` for USB copy into `documents/`.
|
||||||
|
Delete intermediate artifacts left in the user's directories.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Never report success on the converter's own summary line alone. Check:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 - <<'EOF'
|
||||||
|
import re
|
||||||
|
t = open('out/book.html').read()
|
||||||
|
print('italic:', t.count('<i>'), 'bold:', t.count('<b>'))
|
||||||
|
print('h1:', len(re.findall(r'<h1', t)))
|
||||||
|
print('double spaces:', re.sub('<[^>]+>', '', t).count(' '))
|
||||||
|
EOF
|
||||||
|
```
|
||||||
|
|
||||||
|
- italic count near zero on a novel means `style()` missed the font-name pattern;
|
||||||
|
- an `h1` count in the hundreds means a size threshold is too low and body text
|
||||||
|
is being promoted;
|
||||||
|
- many double spaces means line joining is off;
|
||||||
|
- list the extracted headings (`grep -o '<h1[^>]*>.\{0,60\}'`) and read them —
|
||||||
|
they become the TOC, and a wrong one is obvious at a glance;
|
||||||
|
- read one full page of body text and confirm paragraphs merge across page
|
||||||
|
breaks and hyphenated words are rejoined.
|
||||||
|
|
||||||
|
After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx`
|
||||||
|
navPoint count for the TOC size.
|
||||||
|
|
||||||
|
## What the script does
|
||||||
|
|
||||||
|
- style from span font names (`-It`, `Italic`, `Bold`, `Semibold`) → `<i>`/`<b>`;
|
||||||
|
- headings by font size, sidebars by font family;
|
||||||
|
- one PDF block = one paragraph; an unindented block that opens a page is
|
||||||
|
appended to the previous page's paragraph;
|
||||||
|
- trailing hyphens dropped, adjacent `</i><i>` runs merged, doubled spaces collapsed;
|
||||||
|
- images written to `images/`, cover rendered from page 1 at 150 dpi;
|
||||||
|
- absolute positioning is deliberately discarded so text reflows at any font size.
|
||||||
|
|
||||||
|
## Limits
|
||||||
|
|
||||||
|
- Headers/footers are not stripped — calibre-made PDFs have none. A typeset PDF
|
||||||
|
with running heads and page numbers needs blocks filtered by `y` coordinate
|
||||||
|
near the margins; add that before trusting the output.
|
||||||
|
- Multi-column layouts are not handled; block order would need column sorting.
|
||||||
|
- Thresholds are per-layout constants. Always re-run `--fonts` for a new book.
|
||||||
Reference in New Issue
Block a user