From 33845bd47a5590ed493ab46f95c242755ec82a2f Mon Sep 17 00:00:00 2001 From: chesirecatt Date: Sat, 5 Sep 2026 18:23:05 +0300 Subject: [PATCH] =?UTF-8?q?=D0=9F=D1=8F=D1=82=D1=8C=20=D0=BF=D1=80=D0=B0?= =?UTF-8?q?=D0=B2=D0=BE=D0=BA=20=D0=BF=D0=BE=20=D0=B8=D1=82=D0=BE=D0=B3?= =?UTF-8?q?=D0=B0=D0=BC=20=D1=87=D0=B5=D1=82=D1=8B=D1=80=D1=91=D1=85=20?= =?UTF-8?q?=D0=BA=D0=BD=D0=B8=D0=B3=20=D0=B8=D0=B7=20to=5Fread?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Извлечение (pdf2html.py), каждая правка вызвана конкретной книгой: * колонтитул опознаётся по повтору текста в верхнем поле на многих полосах, а не по одной позиции. Позиционный признак рубит и настоящие заголовки: у Вернона («DDD Distilled») глава начинается ровно с верха полосы и короче 80 знаков, так пропали 14 заголовков из 15. Это ровно тот второй проход, о котором говорил ponytail-комментарий на прежнем фильтре; * блок режется, если шапка набрана крупнее следующего за ней текста. У Вернона 85 подзаголовков лежали в одном блоке с первым абзацем раздела: оглавление теряло второй уровень, а абзац начинался с заголовка без точки; * порог раздела снижен с 1.15 до 1.10 кегля текста. У Вернона подзаголовок 17.2 при тексте 15.0 — ровно на пять сотых ниже прежнего порога, и все 85 разделов уезжали в прозу. Тот же множитель, что у разреза блока, иначе разрез и классификация расходятся; * кандидатом в главы не может быть кегль, у которого в блоках нет букв. У Нейгарда («Release it!») номер главы набран кеглем 100 отдельным блоком («1», «2», …), и главой объявлялся он, а настоящие заголовки в 30 пунктов уезжали в разделы — оглавление выходило пустым. Отбор по наличию букв, а не по длине: «Preface» — семь знаков, порог по длине отсекал бы и его. Упаковка (pack_epub.py): XML 1.0 запрещает управляющие символы, а из PDF они приезжают регулярно. У Бейера («Site Reliability Engineering») 33 главы из 45 не разбирались как XML из-за одного такого знака в начале абзаца, и читалка спотыкается на таком файле молча. Чистим на упаковке, а не надеемся на извлечение. Тест на это добавлен и проверен: с обезвреженной правкой он падает. SKILL.md приведён в соответствие с репозиторием. Шаг упаковки с 25.08 делает pack_epub.py, а текст всё ещё требовал calibre на PATH — из-за этого проверка инструментов давала ложный отказ. Замерено 05.09.2026: ebook-convert отсутствует на всех трёх здешних машинах, и все четыре книги собрались без него. За calibre остался только .azw3 для Kindle. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01XCiPWRi3Jqy4saD2LwU1Dp --- SKILL.md | 38 ++++++++++++------- scripts/pack_epub.py | 8 +++- scripts/pdf2html.py | 80 +++++++++++++++++++++++++++++++++++++-- scripts/test_pack_epub.py | 24 ++++++++++++ 4 files changed, 133 insertions(+), 17 deletions(-) diff --git a/SKILL.md b/SKILL.md index 73ac0c5..3a64023 100644 --- a/SKILL.md +++ b/SKILL.md @@ -42,8 +42,8 @@ away the semantic markup that is already there and re-derives it from font sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads the spine, drops the nav document, and maps the book's own headings. -Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is -needed, not PyMuPDF) and convert with: +Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (no PyMuPDF +needed) and convert with: `python scripts/epub2html.py book.epub out/book.html` @@ -68,7 +68,10 @@ the translation bridge splits the book on `

`. Measured on Stroustrup's PPP 1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No embedded fonts / no extractable text means a scan — stop and say OCR is needed; this skill does not apply. -2. **Check tooling.** `ebook-convert` must be on PATH. For PyMuPDF, prefer an +2. **Check tooling.** Packing needs nothing but the standard library + (`scripts/pack_epub.py`); calibre is optional and only for `.azw3`. Measured + 2026-09-05: `ebook-convert` was absent on all three machines at hand, and the + conversion went through end to end without it. For PyMuPDF, prefer an existing interpreter that has it; otherwise build a throwaway venv in the scratchpad — do not install into the system Python: @@ -100,21 +103,29 @@ the translation bridge splits the book on `

`. Measured on Stroustrup's PPP means the mono flag never fired and every listing is about to be reflowed as prose. Fix thresholds and re-run until the counts are sane. Cheap to iterate; do not skip to packing. -7. **Pack with calibre**, always passing explicit TOC XPaths — without them - calibre applies its own heuristics and re-breaks the chapters: +7. **Pack with `scripts/pack_epub.py`** — no calibre, no root, splits on `

` + and writes the conservative EPUB 2 + NCX that PocketBook and KOReader both + read: ```bash - ebook-convert out/book.html "Title.epub" \ - --title="Title" --authors="Author" --language=en \ - --cover=out/images/cover.jpg \ - --level1-toc='//h:h1' --level2-toc='//h:h2' \ - --page-breaks-before='//h:h1' \ - --no-default-epub-cover - ebook-convert "Title.epub" "Title.azw3" + python scripts/pack_epub.py out/book.html "Title.epub" \ + --title="Title" --author="Author" --lang=en --cover=cover.jpg ``` Take title/author from the user or the book's own title page — PDF metadata is often an ASIN or a filename. + + calibre is needed only for `.azw3` (Kindle over USB), and only if it happens + to be installed: + + ```bash + ebook-convert "Title.epub" "Title.azw3" + ``` + + ⚠️ Do **not** pack the XHTML with `ebook-convert` directly: without explicit + `--level1-toc='//h:h1' --page-breaks-before='//h:h1'` it applies its own + heuristics and re-breaks the chapters. `pack_epub.py` has no such mode to + get wrong. 8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle (Amazon converts server-side), `.azw3` for USB copy into `documents/`. Delete intermediate artifacts left in the user's directories. @@ -181,7 +192,8 @@ for them. Extract those by hand and verify by grepping the existing translation. ## Optional stage: translation Only when the user asks for a translated book. It slots between step 6 and -step 7 — translate the XHTML, then pack the translated file with calibre. +step 7 — translate the XHTML, then pack the translated file with +`pack_epub.py`. **Code listings are never translated.** The bridge only recognizes `

`, `

` and `

`; a `
` block is not parsed, and anything unparsed is carried into
diff --git a/scripts/pack_epub.py b/scripts/pack_epub.py
index 9304704..1cd75de 100644
--- a/scripts/pack_epub.py
+++ b/scripts/pack_epub.py
@@ -17,6 +17,12 @@ import sys
 import zipfile
 from pathlib import Path
 
+# XML 1.0 запрещает управляющие символы, а из PDF они приезжают регулярно:
+# у Бейера («Site Reliability Engineering») 33 главы из 45 не разбирались как
+# XML из-за одного такого знака в начале абзаца. Читалка спотыкается молча,
+# поэтому чистим на упаковке, а не надеемся на извлечение.
+BAD_XML = re.compile(r"[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]")
+
 H1 = re.compile(r"^]*>(.*?)

$") IMG = re.compile(r'(.*?)", text, re.S) css = css.group(1) if css else "" body = re.search(r"(.*)", text, re.S) diff --git a/scripts/pdf2html.py b/scripts/pdf2html.py index 9555e2c..d6871dc 100644 --- a/scripts/pdf2html.py +++ b/scripts/pdf2html.py @@ -97,6 +97,7 @@ def profile(doc, step=7): """ size_chars = collections.Counter() head_blocks = collections.Counter() + head_len = {} x0 = collections.Counter() for pno in range(1, doc.page_count, step): for b in doc[pno].get_text("dict")["blocks"]: @@ -110,13 +111,25 @@ def profile(doc, step=7): size_chars[size_of(sp)] += len(sp["text"]) top = max(spans, key=lambda sp: len(sp["text"])) head_blocks[size_of(top)] += 1 + head_len.setdefault(size_of(top), []).append( + "".join(sp["text"] for sp in spans).strip()) if not size_chars: sys.exit("в PDF нет извлекаемого текста — нужен OCR, этот скрипт не поможет") body = size_chars.most_common(1)[0][0] # кандидаты в главы: крупнее текста и встречаются не единожды (единичный # размер — это титул, а не уровень заголовка) + # Ещё условие: у кандидата должны быть слова, а не цифры. У Нейгарда номер + # главы набран кеглем 100 отдельным блоком («1», «2», …), и без проверки + # главой объявлялся он, а настоящие заголовки в 30 пунктов уезжали в + # разделы — оглавление выходило пустым. Отбор идёт по наличию букв, а не по + # длине: «Preface» — семь знаков, и порог по длине отсекал бы и его. + def _wordy(sz): + txts = head_len.get(sz) or [] + letters = sum(1 for t in txts if re.search(r"[^\W\d_]", t)) + return bool(txts) and letters * 2 >= len(txts) + cand = sorted((sz for sz, n in head_blocks.items() - if sz >= body * 1.15 and n >= 3), reverse=True) + if sz >= body * 1.15 and n >= 3 and _wordy(sz)), reverse=True) h1 = cand[0] if cand else body * 1.55 left = min(x for x, n in x0.most_common(4)) return body, left + max(4, body * 0.6), h1 @@ -165,7 +178,11 @@ def block_kind(b, body, h1_size): return ("h1", "title") if sz >= h1_size * 0.97: return ("h1", None) - if sz >= body * 1.15: + # 1.10, а не 1.15: у Вернона подзаголовок набран 17.2 при тексте 15.0, то + # есть ровно на пять сотых ниже прежнего порога — и все 85 разделов книги + # уезжали в прозу. Тот же множитель, что у split_by_size, иначе разрез и + # классификация расходятся. + if sz >= body * 1.10: return ("h2", None) return ("p", None) @@ -237,11 +254,65 @@ open_para = None # (tag, cls, text) still collecting pix = doc[0].get_pixmap(dpi=150) pix.save(IMGDIR / "cover.jpg") +def running_heads(doc, band, limit=80, min_pages=5): + """Тексты, повторяющиеся в верхнем поле на многих полосах. + + Позиционный признак в одиночку рубит и настоящие заголовки: у книг, + свёрстанных calibre, глава начинается ровно с верха полосы и короче 80 + знаков. У Вернона так пропали 14 заголовков из 15. Повтор по десяткам + страниц — то, чем колонтитул отличается от заголовка. + """ + seen = collections.Counter() + for page in doc: + for b in page.get_text("dict")["blocks"]: + if b.get("type") == 1 or "lines" not in b: + continue + if b["bbox"][1] >= page.rect.height * band: + continue + txt = "".join(sp["text"] for l in b["lines"] for sp in l["spans"]) + txt = re.sub(r"\d+", "", txt).strip().lower() + if txt and len(txt) < limit: + seen[txt] += 1 + return {t for t, n in seen.items() if n >= min_pages} + + +def split_by_size(b, body): + """Разрезать блок, где шапка набрана крупнее следующего за ней текста. + + У Вернона 85 подзаголовков лежат в одном блоке с первым абзацем раздела: + строка 17.2 и сразу за ней 15.0. Без разреза они становятся частью абзаца, + оглавление теряет второй уровень, а текст начинается с заголовка без точки. + """ + lines = [l for l in b.get("lines", []) + if any(sp["text"].strip() for sp in l["spans"])] + if len(lines) < 2: + return [b] + big = [max(size_of(sp) for sp in l["spans"] if sp["text"].strip()) + > body * 1.12 for l in lines] + if not big[0] or all(big): + return [b] + cut = big.index(False) + if not any(big[:cut]) or any(big[cut:]): + return [b] # разрез только когда крупное строго сверху + head = dict(b, lines=lines[:cut], + bbox=(b["bbox"][0], b["bbox"][1], b["bbox"][2], + lines[cut - 1]["bbox"][3])) + rest = dict(b, lines=lines[cut:], + bbox=(b["bbox"][0], lines[cut]["bbox"][1], b["bbox"][2], + b["bbox"][3])) + return [head, rest] + + +HEADS = running_heads(doc, HEAD_BAND) +print("колонтитулов опознано по повтору:", len(HEADS)) + img_n = 0 for pno, page in enumerate(doc): if pno == 0: continue blocks = sorted(page.get_text("dict")["blocks"], key=lambda b: b["bbox"][1]) + blocks = [part for b in blocks + for part in (split_by_size(b, BODY) if b.get("type") != 1 else [b])] for bi, b in enumerate(blocks): # Повёрнутый на 90° текст: боковые врезки и подписи к таблицам. В книге # их единицы, а распознаются они всегда в кашу («aoHegoduenueArdng») и @@ -280,7 +351,10 @@ for pno, page in enumerate(doc): # отсеивается 21 блок крупного кегля из 511, все до одного — шапки. # ponytail: признак позиционный. Надёжнее — текст, повторяющийся на # десятках полос, но за это платить вторым проходом по документу. - if b["bbox"][1] < page.rect.height * HEAD_BAND and len(txt) < 80: + if (b["bbox"][1] < page.rect.height * HEAD_BAND and len(txt) < 80 + and re.sub(r"\d+", "", bookhtml.strip_tags(txt) + if hasattr(bookhtml, "strip_tags") + else re.sub(r"<[^>]+>", "", txt)).strip().lower() in HEADS): continue if tag in ("h1", "h2") and bookhtml.looks_garbled(txt): tag, cls = "p", None diff --git a/scripts/test_pack_epub.py b/scripts/test_pack_epub.py index 3702847..2ed95a4 100644 --- a/scripts/test_pack_epub.py +++ b/scripts/test_pack_epub.py @@ -74,6 +74,30 @@ def test_cover_used_in_text_is_not_duplicated(tmp: Path): assert opf.count('href="images/cover.jpg"') == 1, opf +def test_control_chars_stripped(tmp: Path): + """Управляющий знак из PDF делает главу неразбираемой как XML. + + У Бейера («Site Reliability Engineering») так падали 33 главы из 45: + один \x02 в начале абзаца, и читалка молча спотыкается на файле. + """ + import xml.etree.ElementTree as ET + + src = tmp / "book.html" + src.write_text(BOOK.replace("

Текст первой главы.

", + "

\x02Текст первой главы.

"), + encoding="utf-8") + (tmp / "images").mkdir() + (tmp / "images" / "fig.png").write_bytes(b"\x89PNG\r\n\x1a\n") + (tmp / "images" / "cover.jpg").write_bytes(b"\xff\xd8\xff") + out = tmp / "book.epub" + pack_epub.build(src, out, {"title": "T", "author": "A", "lang": "ru", + "cover": "cover.jpg", "id": "urn:uuid:test"}) + z = zipfile.ZipFile(out) + for name in z.namelist(): + if name.endswith((".xhtml", ".opf", ".ncx")): + ET.fromstring(z.read(name)) # падает, если знак остался + + def test_no_content(tmp: Path): src = tmp / "empty.html" src.write_text("\n", encoding="utf-8")