Пять правок по итогам четырёх книг из to_read

Извлечение (pdf2html.py), каждая правка вызвана конкретной книгой:

* колонтитул опознаётся по повтору текста в верхнем поле на многих полосах, а
  не по одной позиции. Позиционный признак рубит и настоящие заголовки: у
  Вернона («DDD Distilled») глава начинается ровно с верха полосы и короче 80
  знаков, так пропали 14 заголовков из 15. Это ровно тот второй проход, о
  котором говорил ponytail-комментарий на прежнем фильтре;
* блок режется, если шапка набрана крупнее следующего за ней текста. У Вернона
  85 подзаголовков лежали в одном блоке с первым абзацем раздела: оглавление
  теряло второй уровень, а абзац начинался с заголовка без точки;
* порог раздела снижен с 1.15 до 1.10 кегля текста. У Вернона подзаголовок
  17.2 при тексте 15.0 — ровно на пять сотых ниже прежнего порога, и все 85
  разделов уезжали в прозу. Тот же множитель, что у разреза блока, иначе
  разрез и классификация расходятся;
* кандидатом в главы не может быть кегль, у которого в блоках нет букв. У
  Нейгарда («Release it!») номер главы набран кеглем 100 отдельным блоком
  («1», «2», …), и главой объявлялся он, а настоящие заголовки в 30 пунктов
  уезжали в разделы — оглавление выходило пустым. Отбор по наличию букв, а не
  по длине: «Preface» — семь знаков, порог по длине отсекал бы и его.

Упаковка (pack_epub.py): XML 1.0 запрещает управляющие символы, а из PDF они
приезжают регулярно. У Бейера («Site Reliability Engineering») 33 главы из 45
не разбирались как XML из-за одного такого знака в начале абзаца, и читалка
спотыкается на таком файле молча. Чистим на упаковке, а не надеемся на
извлечение. Тест на это добавлен и проверен: с обезвреженной правкой он падает.

SKILL.md приведён в соответствие с репозиторием. Шаг упаковки с 25.08 делает
pack_epub.py, а текст всё ещё требовал calibre на PATH — из-за этого проверка
инструментов давала ложный отказ. Замерено 05.09.2026: ebook-convert
отсутствует на всех трёх здешних машинах, и все четыре книги собрались без
него. За calibre остался только .azw3 для Kindle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XCiPWRi3Jqy4saD2LwU1Dp
This commit is contained in:
chesirecatt
2026-09-05 18:23:05 +03:00
parent af5c243687
commit 33845bd47a
4 changed files with 133 additions and 17 deletions
+25 -13
View File
@@ -42,8 +42,8 @@ away the semantic markup that is already there and re-derives it from font
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
the spine, drops the nav document, and maps the book's own headings. the spine, drops the nav document, and maps the book's own headings.
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (no PyMuPDF
needed, not PyMuPDF) and convert with: needed) and convert with:
`python scripts/epub2html.py book.epub out/book.html` `python scripts/epub2html.py book.epub out/book.html`
@@ -68,7 +68,10 @@ the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No 1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
embedded fonts / no extractable text means a scan — stop and say OCR is embedded fonts / no extractable text means a scan — stop and say OCR is
needed; this skill does not apply. needed; this skill does not apply.
2. **Check tooling.** `ebook-convert` must be on PATH. For PyMuPDF, prefer an 2. **Check tooling.** Packing needs nothing but the standard library
(`scripts/pack_epub.py`); calibre is optional and only for `.azw3`. Measured
2026-09-05: `ebook-convert` was absent on all three machines at hand, and the
conversion went through end to end without it. For PyMuPDF, prefer an
existing interpreter that has it; otherwise build a throwaway venv in the existing interpreter that has it; otherwise build a throwaway venv in the
scratchpad — do not install into the system Python: scratchpad — do not install into the system Python:
@@ -100,21 +103,29 @@ the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
means the mono flag never fired and every listing is about to be reflowed as means the mono flag never fired and every listing is about to be reflowed as
prose. Fix thresholds and re-run until prose. Fix thresholds and re-run until
the counts are sane. Cheap to iterate; do not skip to packing. the counts are sane. Cheap to iterate; do not skip to packing.
7. **Pack with calibre**, always passing explicit TOC XPaths — without them 7. **Pack with `scripts/pack_epub.py`** — no calibre, no root, splits on `<h1>`
calibre applies its own heuristics and re-breaks the chapters: and writes the conservative EPUB 2 + NCX that PocketBook and KOReader both
read:
```bash ```bash
ebook-convert out/book.html "Title.epub" \ python scripts/pack_epub.py out/book.html "Title.epub" \
--title="Title" --authors="Author" --language=en \ --title="Title" --author="Author" --lang=en --cover=cover.jpg
--cover=out/images/cover.jpg \
--level1-toc='//h:h1' --level2-toc='//h:h2' \
--page-breaks-before='//h:h1' \
--no-default-epub-cover
ebook-convert "Title.epub" "Title.azw3"
``` ```
Take title/author from the user or the book's own title page — PDF metadata Take title/author from the user or the book's own title page — PDF metadata
is often an ASIN or a filename. is often an ASIN or a filename.
calibre is needed only for `.azw3` (Kindle over USB), and only if it happens
to be installed:
```bash
ebook-convert "Title.epub" "Title.azw3"
```
⚠️ Do **not** pack the XHTML with `ebook-convert` directly: without explicit
`--level1-toc='//h:h1' --page-breaks-before='//h:h1'` it applies its own
heuristics and re-breaks the chapters. `pack_epub.py` has no such mode to
get wrong.
8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle 8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle
(Amazon converts server-side), `.azw3` for USB copy into `documents/`. (Amazon converts server-side), `.azw3` for USB copy into `documents/`.
Delete intermediate artifacts left in the user's directories. Delete intermediate artifacts left in the user's directories.
@@ -181,7 +192,8 @@ for them. Extract those by hand and verify by grepping the existing translation.
## Optional stage: translation ## Optional stage: translation
Only when the user asks for a translated book. It slots between step 6 and Only when the user asks for a translated book. It slots between step 6 and
step 7 — translate the XHTML, then pack the translated file with calibre. step 7 — translate the XHTML, then pack the translated file with
`pack_epub.py`.
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>` **Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
+7 -1
View File
@@ -17,6 +17,12 @@ import sys
import zipfile import zipfile
from pathlib import Path from pathlib import Path
# XML 1.0 запрещает управляющие символы, а из PDF они приезжают регулярно:
# у Бейера («Site Reliability Engineering») 33 главы из 45 не разбирались как
# XML из-за одного такого знака в начале абзаца. Читалка спотыкается молча,
# поэтому чистим на упаковке, а не надеемся на извлечение.
BAD_XML = re.compile(r"[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]")
H1 = re.compile(r"^<h1[^>]*>(.*?)</h1>$") H1 = re.compile(r"^<h1[^>]*>(.*?)</h1>$")
IMG = re.compile(r'<img src="images/([^"]+)"') IMG = re.compile(r'<img src="images/([^"]+)"')
NS = 'xmlns="http://www.w3.org/1999/xhtml"' NS = 'xmlns="http://www.w3.org/1999/xhtml"'
@@ -66,7 +72,7 @@ def series_meta(meta):
def build(src, out, meta): def build(src, out, meta):
text = src.read_text(encoding="utf-8") text = BAD_XML.sub("", src.read_text(encoding="utf-8"))
css = re.search(r"<style>(.*?)</style>", text, re.S) css = re.search(r"<style>(.*?)</style>", text, re.S)
css = css.group(1) if css else "" css = css.group(1) if css else ""
body = re.search(r"<body>(.*)</body>", text, re.S) body = re.search(r"<body>(.*)</body>", text, re.S)
+77 -3
View File
@@ -97,6 +97,7 @@ def profile(doc, step=7):
""" """
size_chars = collections.Counter() size_chars = collections.Counter()
head_blocks = collections.Counter() head_blocks = collections.Counter()
head_len = {}
x0 = collections.Counter() x0 = collections.Counter()
for pno in range(1, doc.page_count, step): for pno in range(1, doc.page_count, step):
for b in doc[pno].get_text("dict")["blocks"]: for b in doc[pno].get_text("dict")["blocks"]:
@@ -110,13 +111,25 @@ def profile(doc, step=7):
size_chars[size_of(sp)] += len(sp["text"]) size_chars[size_of(sp)] += len(sp["text"])
top = max(spans, key=lambda sp: len(sp["text"])) top = max(spans, key=lambda sp: len(sp["text"]))
head_blocks[size_of(top)] += 1 head_blocks[size_of(top)] += 1
head_len.setdefault(size_of(top), []).append(
"".join(sp["text"] for sp in spans).strip())
if not size_chars: if not size_chars:
sys.exit("в PDF нет извлекаемого текста — нужен OCR, этот скрипт не поможет") sys.exit("в PDF нет извлекаемого текста — нужен OCR, этот скрипт не поможет")
body = size_chars.most_common(1)[0][0] body = size_chars.most_common(1)[0][0]
# кандидаты в главы: крупнее текста и встречаются не единожды (единичный # кандидаты в главы: крупнее текста и встречаются не единожды (единичный
# размер — это титул, а не уровень заголовка) # размер — это титул, а не уровень заголовка)
# Ещё условие: у кандидата должны быть слова, а не цифры. У Нейгарда номер
# главы набран кеглем 100 отдельным блоком («1», «2», …), и без проверки
# главой объявлялся он, а настоящие заголовки в 30 пунктов уезжали в
# разделы — оглавление выходило пустым. Отбор идёт по наличию букв, а не по
# длине: «Preface» — семь знаков, и порог по длине отсекал бы и его.
def _wordy(sz):
txts = head_len.get(sz) or []
letters = sum(1 for t in txts if re.search(r"[^\W\d_]", t))
return bool(txts) and letters * 2 >= len(txts)
cand = sorted((sz for sz, n in head_blocks.items() cand = sorted((sz for sz, n in head_blocks.items()
if sz >= body * 1.15 and n >= 3), reverse=True) if sz >= body * 1.15 and n >= 3 and _wordy(sz)), reverse=True)
h1 = cand[0] if cand else body * 1.55 h1 = cand[0] if cand else body * 1.55
left = min(x for x, n in x0.most_common(4)) left = min(x for x, n in x0.most_common(4))
return body, left + max(4, body * 0.6), h1 return body, left + max(4, body * 0.6), h1
@@ -165,7 +178,11 @@ def block_kind(b, body, h1_size):
return ("h1", "title") return ("h1", "title")
if sz >= h1_size * 0.97: if sz >= h1_size * 0.97:
return ("h1", None) return ("h1", None)
if sz >= body * 1.15: # 1.10, а не 1.15: у Вернона подзаголовок набран 17.2 при тексте 15.0, то
# есть ровно на пять сотых ниже прежнего порога — и все 85 разделов книги
# уезжали в прозу. Тот же множитель, что у split_by_size, иначе разрез и
# классификация расходятся.
if sz >= body * 1.10:
return ("h2", None) return ("h2", None)
return ("p", None) return ("p", None)
@@ -237,11 +254,65 @@ open_para = None # (tag, cls, text) still collecting
pix = doc[0].get_pixmap(dpi=150) pix = doc[0].get_pixmap(dpi=150)
pix.save(IMGDIR / "cover.jpg") pix.save(IMGDIR / "cover.jpg")
def running_heads(doc, band, limit=80, min_pages=5):
"""Тексты, повторяющиеся в верхнем поле на многих полосах.
Позиционный признак в одиночку рубит и настоящие заголовки: у книг,
свёрстанных calibre, глава начинается ровно с верха полосы и короче 80
знаков. У Вернона так пропали 14 заголовков из 15. Повтор по десяткам
страниц — то, чем колонтитул отличается от заголовка.
"""
seen = collections.Counter()
for page in doc:
for b in page.get_text("dict")["blocks"]:
if b.get("type") == 1 or "lines" not in b:
continue
if b["bbox"][1] >= page.rect.height * band:
continue
txt = "".join(sp["text"] for l in b["lines"] for sp in l["spans"])
txt = re.sub(r"\d+", "", txt).strip().lower()
if txt and len(txt) < limit:
seen[txt] += 1
return {t for t, n in seen.items() if n >= min_pages}
def split_by_size(b, body):
"""Разрезать блок, где шапка набрана крупнее следующего за ней текста.
У Вернона 85 подзаголовков лежат в одном блоке с первым абзацем раздела:
строка 17.2 и сразу за ней 15.0. Без разреза они становятся частью абзаца,
оглавление теряет второй уровень, а текст начинается с заголовка без точки.
"""
lines = [l for l in b.get("lines", [])
if any(sp["text"].strip() for sp in l["spans"])]
if len(lines) < 2:
return [b]
big = [max(size_of(sp) for sp in l["spans"] if sp["text"].strip())
> body * 1.12 for l in lines]
if not big[0] or all(big):
return [b]
cut = big.index(False)
if not any(big[:cut]) or any(big[cut:]):
return [b] # разрез только когда крупное строго сверху
head = dict(b, lines=lines[:cut],
bbox=(b["bbox"][0], b["bbox"][1], b["bbox"][2],
lines[cut - 1]["bbox"][3]))
rest = dict(b, lines=lines[cut:],
bbox=(b["bbox"][0], lines[cut]["bbox"][1], b["bbox"][2],
b["bbox"][3]))
return [head, rest]
HEADS = running_heads(doc, HEAD_BAND)
print("колонтитулов опознано по повтору:", len(HEADS))
img_n = 0 img_n = 0
for pno, page in enumerate(doc): for pno, page in enumerate(doc):
if pno == 0: if pno == 0:
continue continue
blocks = sorted(page.get_text("dict")["blocks"], key=lambda b: b["bbox"][1]) blocks = sorted(page.get_text("dict")["blocks"], key=lambda b: b["bbox"][1])
blocks = [part for b in blocks
for part in (split_by_size(b, BODY) if b.get("type") != 1 else [b])]
for bi, b in enumerate(blocks): for bi, b in enumerate(blocks):
# Повёрнутый на 90° текст: боковые врезки и подписи к таблицам. В книге # Повёрнутый на 90° текст: боковые врезки и подписи к таблицам. В книге
# их единицы, а распознаются они всегда в кашу («aoHegoduenueArdng») и # их единицы, а распознаются они всегда в кашу («aoHegoduenueArdng») и
@@ -280,7 +351,10 @@ for pno, page in enumerate(doc):
# отсеивается 21 блок крупного кегля из 511, все до одного — шапки. # отсеивается 21 блок крупного кегля из 511, все до одного — шапки.
# ponytail: признак позиционный. Надёжнее — текст, повторяющийся на # ponytail: признак позиционный. Надёжнее — текст, повторяющийся на
# десятках полос, но за это платить вторым проходом по документу. # десятках полос, но за это платить вторым проходом по документу.
if b["bbox"][1] < page.rect.height * HEAD_BAND and len(txt) < 80: if (b["bbox"][1] < page.rect.height * HEAD_BAND and len(txt) < 80
and re.sub(r"\d+", "", bookhtml.strip_tags(txt)
if hasattr(bookhtml, "strip_tags")
else re.sub(r"<[^>]+>", "", txt)).strip().lower() in HEADS):
continue continue
if tag in ("h1", "h2") and bookhtml.looks_garbled(txt): if tag in ("h1", "h2") and bookhtml.looks_garbled(txt):
tag, cls = "p", None tag, cls = "p", None
+24
View File
@@ -74,6 +74,30 @@ def test_cover_used_in_text_is_not_duplicated(tmp: Path):
assert opf.count('href="images/cover.jpg"') == 1, opf assert opf.count('href="images/cover.jpg"') == 1, opf
def test_control_chars_stripped(tmp: Path):
"""Управляющий знак из PDF делает главу неразбираемой как XML.
У Бейера («Site Reliability Engineering») так падали 33 главы из 45:
один \x02 в начале абзаца, и читалка молча спотыкается на файле.
"""
import xml.etree.ElementTree as ET
src = tmp / "book.html"
src.write_text(BOOK.replace("<p>Текст первой главы.</p>",
"<p>\x02Текст первой главы.</p>"),
encoding="utf-8")
(tmp / "images").mkdir()
(tmp / "images" / "fig.png").write_bytes(b"\x89PNG\r\n\x1a\n")
(tmp / "images" / "cover.jpg").write_bytes(b"\xff\xd8\xff")
out = tmp / "book.epub"
pack_epub.build(src, out, {"title": "T", "author": "A", "lang": "ru",
"cover": "cover.jpg", "id": "urn:uuid:test"})
z = zipfile.ZipFile(out)
for name in z.namelist():
if name.endswith((".xhtml", ".opf", ".ncx")):
ET.fromstring(z.read(name)) # падает, если знак остался
def test_no_content(tmp: Path): def test_no_content(tmp: Path):
src = tmp / "empty.html" src = tmp / "empty.html"
src.write_text("<html><body>\n</body></html>", encoding="utf-8") src.write_text("<html><body>\n</body></html>", encoding="utf-8")