Пять правок по итогам четырёх книг из to_read
Извлечение (pdf2html.py), каждая правка вызвана конкретной книгой: * колонтитул опознаётся по повтору текста в верхнем поле на многих полосах, а не по одной позиции. Позиционный признак рубит и настоящие заголовки: у Вернона («DDD Distilled») глава начинается ровно с верха полосы и короче 80 знаков, так пропали 14 заголовков из 15. Это ровно тот второй проход, о котором говорил ponytail-комментарий на прежнем фильтре; * блок режется, если шапка набрана крупнее следующего за ней текста. У Вернона 85 подзаголовков лежали в одном блоке с первым абзацем раздела: оглавление теряло второй уровень, а абзац начинался с заголовка без точки; * порог раздела снижен с 1.15 до 1.10 кегля текста. У Вернона подзаголовок 17.2 при тексте 15.0 — ровно на пять сотых ниже прежнего порога, и все 85 разделов уезжали в прозу. Тот же множитель, что у разреза блока, иначе разрез и классификация расходятся; * кандидатом в главы не может быть кегль, у которого в блоках нет букв. У Нейгарда («Release it!») номер главы набран кеглем 100 отдельным блоком («1», «2», …), и главой объявлялся он, а настоящие заголовки в 30 пунктов уезжали в разделы — оглавление выходило пустым. Отбор по наличию букв, а не по длине: «Preface» — семь знаков, порог по длине отсекал бы и его. Упаковка (pack_epub.py): XML 1.0 запрещает управляющие символы, а из PDF они приезжают регулярно. У Бейера («Site Reliability Engineering») 33 главы из 45 не разбирались как XML из-за одного такого знака в начале абзаца, и читалка спотыкается на таком файле молча. Чистим на упаковке, а не надеемся на извлечение. Тест на это добавлен и проверен: с обезвреженной правкой он падает. SKILL.md приведён в соответствие с репозиторием. Шаг упаковки с 25.08 делает pack_epub.py, а текст всё ещё требовал calibre на PATH — из-за этого проверка инструментов давала ложный отказ. Замерено 05.09.2026: ebook-convert отсутствует на всех трёх здешних машинах, и все четыре книги собрались без него. За calibre остался только .azw3 для Kindle. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XCiPWRi3Jqy4saD2LwU1Dp
This commit is contained in:
@@ -42,8 +42,8 @@ away the semantic markup that is already there and re-derives it from font
|
||||
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
|
||||
the spine, drops the nav document, and maps the book's own headings.
|
||||
|
||||
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is
|
||||
needed, not PyMuPDF) and convert with:
|
||||
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (no PyMuPDF
|
||||
needed) and convert with:
|
||||
|
||||
`python scripts/epub2html.py book.epub out/book.html`
|
||||
|
||||
@@ -68,7 +68,10 @@ the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
|
||||
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
|
||||
embedded fonts / no extractable text means a scan — stop and say OCR is
|
||||
needed; this skill does not apply.
|
||||
2. **Check tooling.** `ebook-convert` must be on PATH. For PyMuPDF, prefer an
|
||||
2. **Check tooling.** Packing needs nothing but the standard library
|
||||
(`scripts/pack_epub.py`); calibre is optional and only for `.azw3`. Measured
|
||||
2026-09-05: `ebook-convert` was absent on all three machines at hand, and the
|
||||
conversion went through end to end without it. For PyMuPDF, prefer an
|
||||
existing interpreter that has it; otherwise build a throwaway venv in the
|
||||
scratchpad — do not install into the system Python:
|
||||
|
||||
@@ -100,21 +103,29 @@ the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
|
||||
means the mono flag never fired and every listing is about to be reflowed as
|
||||
prose. Fix thresholds and re-run until
|
||||
the counts are sane. Cheap to iterate; do not skip to packing.
|
||||
7. **Pack with calibre**, always passing explicit TOC XPaths — without them
|
||||
calibre applies its own heuristics and re-breaks the chapters:
|
||||
7. **Pack with `scripts/pack_epub.py`** — no calibre, no root, splits on `<h1>`
|
||||
and writes the conservative EPUB 2 + NCX that PocketBook and KOReader both
|
||||
read:
|
||||
|
||||
```bash
|
||||
ebook-convert out/book.html "Title.epub" \
|
||||
--title="Title" --authors="Author" --language=en \
|
||||
--cover=out/images/cover.jpg \
|
||||
--level1-toc='//h:h1' --level2-toc='//h:h2' \
|
||||
--page-breaks-before='//h:h1' \
|
||||
--no-default-epub-cover
|
||||
ebook-convert "Title.epub" "Title.azw3"
|
||||
python scripts/pack_epub.py out/book.html "Title.epub" \
|
||||
--title="Title" --author="Author" --lang=en --cover=cover.jpg
|
||||
```
|
||||
|
||||
Take title/author from the user or the book's own title page — PDF metadata
|
||||
is often an ASIN or a filename.
|
||||
|
||||
calibre is needed only for `.azw3` (Kindle over USB), and only if it happens
|
||||
to be installed:
|
||||
|
||||
```bash
|
||||
ebook-convert "Title.epub" "Title.azw3"
|
||||
```
|
||||
|
||||
⚠️ Do **not** pack the XHTML with `ebook-convert` directly: without explicit
|
||||
`--level1-toc='//h:h1' --page-breaks-before='//h:h1'` it applies its own
|
||||
heuristics and re-breaks the chapters. `pack_epub.py` has no such mode to
|
||||
get wrong.
|
||||
8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle
|
||||
(Amazon converts server-side), `.azw3` for USB copy into `documents/`.
|
||||
Delete intermediate artifacts left in the user's directories.
|
||||
@@ -181,7 +192,8 @@ for them. Extract those by hand and verify by grepping the existing translation.
|
||||
## Optional stage: translation
|
||||
|
||||
Only when the user asks for a translated book. It slots between step 6 and
|
||||
step 7 — translate the XHTML, then pack the translated file with calibre.
|
||||
step 7 — translate the XHTML, then pack the translated file with
|
||||
`pack_epub.py`.
|
||||
|
||||
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
|
||||
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
|
||||
|
||||
@@ -17,6 +17,12 @@ import sys
|
||||
import zipfile
|
||||
from pathlib import Path
|
||||
|
||||
# XML 1.0 запрещает управляющие символы, а из PDF они приезжают регулярно:
|
||||
# у Бейера («Site Reliability Engineering») 33 главы из 45 не разбирались как
|
||||
# XML из-за одного такого знака в начале абзаца. Читалка спотыкается молча,
|
||||
# поэтому чистим на упаковке, а не надеемся на извлечение.
|
||||
BAD_XML = re.compile(r"[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]")
|
||||
|
||||
H1 = re.compile(r"^<h1[^>]*>(.*?)</h1>$")
|
||||
IMG = re.compile(r'<img src="images/([^"]+)"')
|
||||
NS = 'xmlns="http://www.w3.org/1999/xhtml"'
|
||||
@@ -66,7 +72,7 @@ def series_meta(meta):
|
||||
|
||||
|
||||
def build(src, out, meta):
|
||||
text = src.read_text(encoding="utf-8")
|
||||
text = BAD_XML.sub("", src.read_text(encoding="utf-8"))
|
||||
css = re.search(r"<style>(.*?)</style>", text, re.S)
|
||||
css = css.group(1) if css else ""
|
||||
body = re.search(r"<body>(.*)</body>", text, re.S)
|
||||
|
||||
+77
-3
@@ -97,6 +97,7 @@ def profile(doc, step=7):
|
||||
"""
|
||||
size_chars = collections.Counter()
|
||||
head_blocks = collections.Counter()
|
||||
head_len = {}
|
||||
x0 = collections.Counter()
|
||||
for pno in range(1, doc.page_count, step):
|
||||
for b in doc[pno].get_text("dict")["blocks"]:
|
||||
@@ -110,13 +111,25 @@ def profile(doc, step=7):
|
||||
size_chars[size_of(sp)] += len(sp["text"])
|
||||
top = max(spans, key=lambda sp: len(sp["text"]))
|
||||
head_blocks[size_of(top)] += 1
|
||||
head_len.setdefault(size_of(top), []).append(
|
||||
"".join(sp["text"] for sp in spans).strip())
|
||||
if not size_chars:
|
||||
sys.exit("в PDF нет извлекаемого текста — нужен OCR, этот скрипт не поможет")
|
||||
body = size_chars.most_common(1)[0][0]
|
||||
# кандидаты в главы: крупнее текста и встречаются не единожды (единичный
|
||||
# размер — это титул, а не уровень заголовка)
|
||||
# Ещё условие: у кандидата должны быть слова, а не цифры. У Нейгарда номер
|
||||
# главы набран кеглем 100 отдельным блоком («1», «2», …), и без проверки
|
||||
# главой объявлялся он, а настоящие заголовки в 30 пунктов уезжали в
|
||||
# разделы — оглавление выходило пустым. Отбор идёт по наличию букв, а не по
|
||||
# длине: «Preface» — семь знаков, и порог по длине отсекал бы и его.
|
||||
def _wordy(sz):
|
||||
txts = head_len.get(sz) or []
|
||||
letters = sum(1 for t in txts if re.search(r"[^\W\d_]", t))
|
||||
return bool(txts) and letters * 2 >= len(txts)
|
||||
|
||||
cand = sorted((sz for sz, n in head_blocks.items()
|
||||
if sz >= body * 1.15 and n >= 3), reverse=True)
|
||||
if sz >= body * 1.15 and n >= 3 and _wordy(sz)), reverse=True)
|
||||
h1 = cand[0] if cand else body * 1.55
|
||||
left = min(x for x, n in x0.most_common(4))
|
||||
return body, left + max(4, body * 0.6), h1
|
||||
@@ -165,7 +178,11 @@ def block_kind(b, body, h1_size):
|
||||
return ("h1", "title")
|
||||
if sz >= h1_size * 0.97:
|
||||
return ("h1", None)
|
||||
if sz >= body * 1.15:
|
||||
# 1.10, а не 1.15: у Вернона подзаголовок набран 17.2 при тексте 15.0, то
|
||||
# есть ровно на пять сотых ниже прежнего порога — и все 85 разделов книги
|
||||
# уезжали в прозу. Тот же множитель, что у split_by_size, иначе разрез и
|
||||
# классификация расходятся.
|
||||
if sz >= body * 1.10:
|
||||
return ("h2", None)
|
||||
return ("p", None)
|
||||
|
||||
@@ -237,11 +254,65 @@ open_para = None # (tag, cls, text) still collecting
|
||||
pix = doc[0].get_pixmap(dpi=150)
|
||||
pix.save(IMGDIR / "cover.jpg")
|
||||
|
||||
def running_heads(doc, band, limit=80, min_pages=5):
|
||||
"""Тексты, повторяющиеся в верхнем поле на многих полосах.
|
||||
|
||||
Позиционный признак в одиночку рубит и настоящие заголовки: у книг,
|
||||
свёрстанных calibre, глава начинается ровно с верха полосы и короче 80
|
||||
знаков. У Вернона так пропали 14 заголовков из 15. Повтор по десяткам
|
||||
страниц — то, чем колонтитул отличается от заголовка.
|
||||
"""
|
||||
seen = collections.Counter()
|
||||
for page in doc:
|
||||
for b in page.get_text("dict")["blocks"]:
|
||||
if b.get("type") == 1 or "lines" not in b:
|
||||
continue
|
||||
if b["bbox"][1] >= page.rect.height * band:
|
||||
continue
|
||||
txt = "".join(sp["text"] for l in b["lines"] for sp in l["spans"])
|
||||
txt = re.sub(r"\d+", "", txt).strip().lower()
|
||||
if txt and len(txt) < limit:
|
||||
seen[txt] += 1
|
||||
return {t for t, n in seen.items() if n >= min_pages}
|
||||
|
||||
|
||||
def split_by_size(b, body):
|
||||
"""Разрезать блок, где шапка набрана крупнее следующего за ней текста.
|
||||
|
||||
У Вернона 85 подзаголовков лежат в одном блоке с первым абзацем раздела:
|
||||
строка 17.2 и сразу за ней 15.0. Без разреза они становятся частью абзаца,
|
||||
оглавление теряет второй уровень, а текст начинается с заголовка без точки.
|
||||
"""
|
||||
lines = [l for l in b.get("lines", [])
|
||||
if any(sp["text"].strip() for sp in l["spans"])]
|
||||
if len(lines) < 2:
|
||||
return [b]
|
||||
big = [max(size_of(sp) for sp in l["spans"] if sp["text"].strip())
|
||||
> body * 1.12 for l in lines]
|
||||
if not big[0] or all(big):
|
||||
return [b]
|
||||
cut = big.index(False)
|
||||
if not any(big[:cut]) or any(big[cut:]):
|
||||
return [b] # разрез только когда крупное строго сверху
|
||||
head = dict(b, lines=lines[:cut],
|
||||
bbox=(b["bbox"][0], b["bbox"][1], b["bbox"][2],
|
||||
lines[cut - 1]["bbox"][3]))
|
||||
rest = dict(b, lines=lines[cut:],
|
||||
bbox=(b["bbox"][0], lines[cut]["bbox"][1], b["bbox"][2],
|
||||
b["bbox"][3]))
|
||||
return [head, rest]
|
||||
|
||||
|
||||
HEADS = running_heads(doc, HEAD_BAND)
|
||||
print("колонтитулов опознано по повтору:", len(HEADS))
|
||||
|
||||
img_n = 0
|
||||
for pno, page in enumerate(doc):
|
||||
if pno == 0:
|
||||
continue
|
||||
blocks = sorted(page.get_text("dict")["blocks"], key=lambda b: b["bbox"][1])
|
||||
blocks = [part for b in blocks
|
||||
for part in (split_by_size(b, BODY) if b.get("type") != 1 else [b])]
|
||||
for bi, b in enumerate(blocks):
|
||||
# Повёрнутый на 90° текст: боковые врезки и подписи к таблицам. В книге
|
||||
# их единицы, а распознаются они всегда в кашу («aoHegoduenueArdng») и
|
||||
@@ -280,7 +351,10 @@ for pno, page in enumerate(doc):
|
||||
# отсеивается 21 блок крупного кегля из 511, все до одного — шапки.
|
||||
# ponytail: признак позиционный. Надёжнее — текст, повторяющийся на
|
||||
# десятках полос, но за это платить вторым проходом по документу.
|
||||
if b["bbox"][1] < page.rect.height * HEAD_BAND and len(txt) < 80:
|
||||
if (b["bbox"][1] < page.rect.height * HEAD_BAND and len(txt) < 80
|
||||
and re.sub(r"\d+", "", bookhtml.strip_tags(txt)
|
||||
if hasattr(bookhtml, "strip_tags")
|
||||
else re.sub(r"<[^>]+>", "", txt)).strip().lower() in HEADS):
|
||||
continue
|
||||
if tag in ("h1", "h2") and bookhtml.looks_garbled(txt):
|
||||
tag, cls = "p", None
|
||||
|
||||
@@ -74,6 +74,30 @@ def test_cover_used_in_text_is_not_duplicated(tmp: Path):
|
||||
assert opf.count('href="images/cover.jpg"') == 1, opf
|
||||
|
||||
|
||||
def test_control_chars_stripped(tmp: Path):
|
||||
"""Управляющий знак из PDF делает главу неразбираемой как XML.
|
||||
|
||||
У Бейера («Site Reliability Engineering») так падали 33 главы из 45:
|
||||
один \x02 в начале абзаца, и читалка молча спотыкается на файле.
|
||||
"""
|
||||
import xml.etree.ElementTree as ET
|
||||
|
||||
src = tmp / "book.html"
|
||||
src.write_text(BOOK.replace("<p>Текст первой главы.</p>",
|
||||
"<p>\x02Текст первой главы.</p>"),
|
||||
encoding="utf-8")
|
||||
(tmp / "images").mkdir()
|
||||
(tmp / "images" / "fig.png").write_bytes(b"\x89PNG\r\n\x1a\n")
|
||||
(tmp / "images" / "cover.jpg").write_bytes(b"\xff\xd8\xff")
|
||||
out = tmp / "book.epub"
|
||||
pack_epub.build(src, out, {"title": "T", "author": "A", "lang": "ru",
|
||||
"cover": "cover.jpg", "id": "urn:uuid:test"})
|
||||
z = zipfile.ZipFile(out)
|
||||
for name in z.namelist():
|
||||
if name.endswith((".xhtml", ".opf", ".ncx")):
|
||||
ET.fromstring(z.read(name)) # падает, если знак остался
|
||||
|
||||
|
||||
def test_no_content(tmp: Path):
|
||||
src = tmp / "empty.html"
|
||||
src.write_text("<html><body>\n</body></html>", encoding="utf-8")
|
||||
|
||||
Reference in New Issue
Block a user