EPUB как второй вход в конвейер; листинги кода не переводятся и не ломают приёмку

- epub2html.py: EPUB -> тот же XHTML, что pdf2html.py, только стандартная
  библиотека. Уровень глав определяется по книге, а не берётся из <h1>:
  в EPUB там обычно название и части.
- pdf2html.py: листинг узнаётся по флагу моноширинного шрифта PyMuPDF,
  переносы и отступы внутри <pre> сохраняются, де-дефисация к коду не
  применяется.
- translate.py: <pre> исключён из проверок языка (английский код утягивал
  долю кириллицы ниже порога приёмки), рабочий каталог привязан к книге,
  устаревшие главы прошлого прогона чистятся, инлайновый <code> переживает
  переводчика.
- bookhtml.py: общий каркас документа и CSS для обоих входов.
- Тесты: test_epub2html.py на собранном в памяти EPUB, четыре новых случая
  в test_translate.py.
This commit is contained in:
chesirecatt
2026-08-20 17:58:09 +03:00
parent 6c81673b4f
commit db9fe5e0fa
7 changed files with 614 additions and 43 deletions
+48 -2
View File
@@ -15,6 +15,30 @@ tag, and 2817 body paragraphs became `<h2>`.
properties, and emits semantic XHTML. calibre is then used only as the
XHTML→EPUB packer.
## Two entry points
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB. Both emit the
same normalized XHTML — one block per line, `<pre>` for code listings — and
everything downstream (translation, packing, audiobook) is identical. Shared
document skeleton and CSS live in `scripts/bookhtml.py`.
**Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws
away the semantic markup that is already there and re-derives it from font
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
the spine, drops the nav document, and maps the book's own headings.
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is
needed, not PyMuPDF) and convert with:
`python scripts/epub2html.py book.epub out/book.html`
It reports which heading level turned out to be the chapter level. In most
EPUBs `<h1>` is the book title and the parts, while chapters are `<h2>` — the
script picks the deepest level that still gives a sane chapter count, because
the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
3rd edition: 7 `<h1>` against 49 `<h2>`, chapters correctly detected as `h2`,
56 chapters out of 656 pages.
## Workflow
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
@@ -35,7 +59,9 @@ XHTML→EPUB packer.
Read off: the body font (largest character count), the italic variant, the
heading sizes, and any secondary family used for sidebars, journal entries,
chat logs or slides.
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. The
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. Code
listings need no tuning — a monospaced span is recognized by the PyMuPDF font
flag (`flags & 8`), not by font name, and becomes a `<pre>` block. The
defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a
title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary
family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that
@@ -45,7 +71,10 @@ XHTML→EPUB packer.
`python scripts/pdf2html.py book.pdf out/book.html`
6. **Verify before packing** (see Verification). Fix thresholds and re-run until
6. **Verify before packing** (see Verification). Check the `pre:` counter against
the real number of listings in the book — a technical book reporting `pre: 0`
means the mono flag never fired and every listing is about to be reflowed as
prose. Fix thresholds and re-run until
the counts are sane. Cheap to iterate; do not skip to packing.
7. **Pack with calibre**, always passing explicit TOC XPaths — without them
calibre applies its own heuristics and re-breaks the chapters:
@@ -97,6 +126,23 @@ navPoint count for the TOC size.
Only when the user asks for a translated book. It slots between step 6 and
step 7 — translate the XHTML, then pack the translated file with calibre.
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
the result untouched. Nothing extra is needed to protect code — but this also
means a listing that was misclassified as a paragraph upstream *will* be
translated, which is the real reason step 6 checks the `pre:` counter.
For the same reason `<pre>` is cut out of both language checks. A book that is
40% listings translates correctly and would otherwise fail acceptance, because
the English code drags the Cyrillic share below the threshold.
**One workdir per book.** Chapter files are named `chapter_NNN.json` for every
book, and the external repo skips chapters it has marked done — a reused workdir
silently stitches one book's translation onto another's text. The bridge writes
`bridge_source.json` into the workdir on first run and refuses to start if the
directory belongs to a different book. Re-running the same book is unaffected;
that is the resume path.
Translation is delegated to an external project,
[`vetermanve/book_translator`](https://github.com/vetermanve/book_translator)
(DeepSeek or a local Ollama model). **Never modify that repository** — it is
+44
View File
@@ -0,0 +1,44 @@
#!/usr/bin/env python3
"""Общая сборка XHTML для pdf2html.py и epub2html.py.
Формат намеренно жёсткий: один блок — одна строка (кроме <pre>, где переносы
значимы). На этом держится мост translate.py — он правит блоки по номеру
строки, а всё, чего не узнал, доносит до результата нетронутым.
"""
import html
CSS = """
body { margin: 0 1em; }
h1 { text-align: center; margin: 2em 0 0.2em; page-break-before: always; }
h1.title { page-break-before: avoid; }
h2 { text-align: center; font-style: italic; font-weight: normal;
font-size: 1.1em; margin: 0.2em 0 1.5em; }
p { text-indent: 1.2em; margin: 0; text-align: justify; }
p.note { text-indent: 0; margin: 1em 2em; font-size: 0.9em;
font-family: sans-serif; }
p.li { text-indent: 0; margin: 0.3em 0 0.3em 1.5em; text-align: left; }
p.row { text-indent: 0; margin: 0.3em 0; text-align: left;
font-size: 0.9em; }
p.figure { text-indent: 0; text-align: center; margin: 1em 0; }
img { max-width: 100%; }
pre { font-family: monospace; font-size: 0.75em; margin: 1em 0;
white-space: pre-wrap; word-wrap: break-word; text-align: left;
text-indent: 0; }
code { font-family: monospace; font-size: 0.9em; }
"""
def document(title, parts):
"""parts: [(tag, cls, inner_html)]; tag 'figure' — псевдотег для картинки."""
buf = ['<?xml version="1.0" encoding="utf-8"?>',
'<html xmlns="http://www.w3.org/1999/xhtml"><head>',
"<title>%s</title>" % html.escape(title),
"<style>%s</style></head><body>" % CSS]
for tag, cls, txt in parts:
if tag == "figure":
buf.append('<p class="figure">%s</p>' % txt)
else:
c = ' class="%s"' % cls if cls else ""
buf.append("<%s%s>%s</%s>" % (tag, c, txt, tag))
buf.append("</body></html>")
return "\n".join(buf)
+267
View File
@@ -0,0 +1,267 @@
#!/usr/bin/env python3
"""EPUB -> тот же XHTML, что выдаёт pdf2html.py (один блок = одна строка).
EPUB уже размечен семантически, поэтому здесь не разведка шрифтов, а
нормализация: привести чужую вёрстку к формату, который понимает мост
translate.py, и вычистить всё, что переводчику видеть не нужно.
Листинги остаются в <pre> и не переводятся вовсе: BLOCK-регулярка моста их не
узнаёт, а всё неузнанное доходит до результата нетронутым.
Только стандартная библиотека — ни calibre, ни lxml здесь не нужны.
"""
import argparse
import collections
import html
import posixpath
import re
import sys
import zipfile
from html.parser import HTMLParser
from pathlib import Path
from xml.etree import ElementTree
import bookhtml
CONTAINER = "META-INF/container.xml"
OPF_NS = "{http://www.idpf.org/2007/opf}"
CONT_NS = "{urn:oasis:names:tc:opendocument:xmlns:container}"
# tag источника -> (tag результата, css-класс)
BLOCK_MAP = {
"h1": ("h1", None), "h2": ("h2", None), "h3": ("h3", None),
"h4": ("h4", None), "h5": ("h5", None), "h6": ("h6", None),
"p": ("p", None), "pre": ("pre", None), "li": ("p", "li"),
"tr": ("p", "row"), "figcaption": ("p", "note"),
"dt": ("p", "note"), "dd": ("p", "note"),
}
INLINE_MAP = {"i": "i", "em": "i", "cite": "i", "b": "b", "strong": "b",
"code": "code", "kbd": "code", "samp": "code", "tt": "code",
"var": "code"}
DROP = {"script", "style", "head", "title", "nav"}
IMG_EXT = {"image/jpeg": ".jpg", "image/png": ".png", "image/gif": ".gif",
"image/svg+xml": ".svg", "image/webp": ".webp"}
class Reader(HTMLParser):
"""Собирает блоки из одного XHTML-документа книги."""
def __init__(self, on_image):
super().__init__(convert_charrefs=True)
self.on_image = on_image
self.blocks = []
self.cur = None # (tag, cls, [куски])
self.open_inline = [] # незакрытые инлайновые теги текущего блока
self.drop = 0
# -- служебное ---------------------------------------------------------
def flush(self):
if not self.cur:
return
tag, cls, parts = self.cur
self.cur = None
for t in reversed(self.open_inline):
parts.append("</%s>" % t) # чужая вёрстка бывает несбалансированной
self.open_inline = []
txt = "".join(parts)
if tag == "pre":
txt = txt.strip("\n").rstrip()
else:
txt = re.sub(r"\s+", " ", txt).strip()
txt = re.sub(r"<(i|b|code)>(\s*)</\1>", r"\2", txt) # пустая разметка
if txt.strip():
self.blocks.append((tag, cls, txt))
def start(self, tag, cls):
self.flush()
self.cur = (tag, cls, [])
# -- HTMLParser --------------------------------------------------------
def handle_starttag(self, tag, attrs):
if tag in DROP:
self.drop += 1
return
if self.drop:
return
a = dict(attrs)
if tag in ("img", "image"):
src = a.get("src") or a.get("{http://www.w3.org/1999/xlink}href") \
or a.get("xlink:href") or a.get("href")
name = self.on_image(src) if src else None
if name:
self.flush()
self.blocks.append(("figure", None,
'<img src="images/%s"/>' % name))
return
if tag == "br":
if self.cur:
self.cur[2].append("\n" if self.cur[0] == "pre" else " ")
return
if tag in BLOCK_MAP:
self.start(*BLOCK_MAP[tag])
return
if tag in ("td", "th") and self.cur and self.cur[0] == "p" \
and self.cur[1] == "row" and self.cur[2]:
self.cur[2].append(" | ")
return
if tag in INLINE_MAP and self.cur and self.cur[0] != "pre":
# Внутри листинга инлайн не нужен: книги размечают код сплошным
# <strong>, и на e-ink это страница жирного текста.
out = INLINE_MAP[tag]
self.open_inline.append(out)
self.cur[2].append("<%s>" % out)
def handle_endtag(self, tag):
if tag in DROP:
self.drop = max(0, self.drop - 1)
return
if self.drop:
return
if tag in BLOCK_MAP:
self.flush()
return
if tag in INLINE_MAP and self.cur and self.cur[0] != "pre":
out = INLINE_MAP[tag]
if out in self.open_inline:
self.open_inline.remove(out)
self.cur[2].append("</%s>" % out)
def handle_data(self, data):
if self.drop:
return
if not self.cur:
if not data.strip():
return
self.start("p", None) # текст вне блочного тега не теряем
self.cur[2].append(html.escape(data, quote=False))
def close(self):
super().close()
self.flush()
def spine_documents(zf):
"""[(путь в архиве, media-type)] в порядке чтения + путь к OPF."""
root = ElementTree.fromstring(zf.read(CONTAINER))
opf_path = root.find(".//%srootfile" % CONT_NS).get("full-path")
opf = ElementTree.fromstring(zf.read(opf_path))
base = posixpath.dirname(opf_path)
items, cover = {}, None
for it in opf.iter("%sitem" % OPF_NS):
href = posixpath.normpath(posixpath.join(base, it.get("href")))
props = it.get("properties") or ""
items[it.get("id")] = (href, it.get("media-type"), props)
if "cover-image" in props:
cover = href
if cover is None: # EPUB 2: обложка объявляется через <meta name="cover">
meta = opf.find(".//%smeta[@name='cover']" % OPF_NS)
if meta is not None and meta.get("content") in items:
cover = items[meta.get("content")][0]
docs = []
for ref in opf.iter("%sitemref" % OPF_NS):
item = items.get(ref.get("idref"))
if not item:
continue
href, mtype, props = item
if "nav" in props or re.search(r"(toc|nav|content)s?\.x?html?$", href, re.I):
continue
docs.append(href)
return docs, items, cover
# Больше этого числа глав дробить незачем: переводчик держит контекст в пределах
# главы, слишком мелкая нарезка его обесценивает.
MAX_CHAPTERS = 150
def chapter_level(parts):
"""Каким уровнем заголовка в этой книге размечены главы.
Мост режет книгу на главы по <h1>, а в EPUB <h1> обычно занят названием
книги и частями: у Страуструпа их 7 на 656 страниц, тогда как главы — это
<h2> (49 штук). Берём самый глубокий уровень, при котором число глав ещё
остаётся вменяемым.
"""
counts = collections.Counter(t for t, _, _ in parts if re.fullmatch(r"h[1-6]", t))
best, total = 1, 0
for lvl in range(1, 7):
total += counts.get("h%d" % lvl, 0)
if total > MAX_CHAPTERS:
break
if total:
best = lvl
return best
def convert(src, out):
imgdir = out.parent / "images"
imgdir.mkdir(parents=True, exist_ok=True)
parts, seen, counter = [], {}, [0]
with zipfile.ZipFile(src) as zf:
names = set(zf.namelist())
docs, items, cover = spine_documents(zf)
def extract(path, stem=None):
if path in seen:
return seen[path]
if path not in names:
return None
ext = posixpath.splitext(path)[1] or ".img"
if stem:
name = stem + ext
else:
counter[0] += 1
name = "img%03d%s" % (counter[0], ext)
(imgdir / name).write_bytes(zf.read(path))
seen[path] = name
return name
if cover:
extract(cover, stem="cover")
for doc in docs:
if doc not in names:
print("нет в архиве, пропущен: %s" % doc, file=sys.stderr)
continue
here = posixpath.dirname(doc)
def on_image(src_attr, here=here):
target = posixpath.normpath(posixpath.join(here, src_attr.split("#")[0]))
return extract(target)
r = Reader(on_image)
r.feed(zf.read(doc).decode("utf-8", "replace"))
r.close()
parts.extend(r.blocks)
title = out.stem
lvl = chapter_level(parts)
parts = [(("h1" if int(t[1]) <= lvl else "h2", c, x)
if re.fullmatch(r"h[1-6]", t) else (t, c, x))
for t, c, x in parts]
h1 = sum(1 for p in parts if p[0] == "h1")
print("главы размечены h%d, глав получилось: %d" % (lvl, h1))
if h1 < 3:
print("ВНИМАНИЕ: заголовков-глав всего %d — переводчик будет держать "
"контекст крупными кусками, проверь результат внимательнее" % h1)
out.write_text(bookhtml.document(title, parts), encoding="utf-8")
print("blocks: %d, images: %d" % (len(parts), len(seen)))
print("h1: %d p: %d pre: %d"
% (h1, sum(1 for p in parts if p[0] == "p"),
sum(1 for p in parts if p[0] == "pre")))
def main():
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument("source", type=Path, help="исходный .epub")
ap.add_argument("output", type=Path, help="куда писать XHTML")
args = ap.parse_args()
if not args.source.exists():
sys.exit("нет файла %s" % args.source)
convert(args.source, args.output)
if __name__ == "__main__":
main()
+34 -32
View File
@@ -7,6 +7,11 @@ from pathlib import Path
import pymupdf
import bookhtml
if len(sys.argv) < 3 or sys.argv[1] in ("-h", "--help"):
sys.exit("usage: pdf2html.py book.pdf out.html | pdf2html.py --fonts book.pdf")
if sys.argv[1] == "--fonts": # разведка: какие шрифты/кегли в PDF
import collections
@@ -29,6 +34,7 @@ IMGDIR.mkdir(parents=True, exist_ok=True)
# Пороги подобраны под вёрстку 15pt/letter. Для другой книги сначала
# посмотреть реальные шрифты и кегли: pdf2html.py --fonts file.pdf
INDENT_X = 88 # x0 первой строки: больше — абзац с красной строки
MONO = 8 # бит моноширинного шрифта в span["flags"] (PyMuPDF)
def style(font):
@@ -36,10 +42,20 @@ def style(font):
return ("bold" in f or "semibold" in f, "-it" in f or "italic" in f)
def span_html(s):
def is_mono(s):
"""Моноширинный шрифт = листинг кода. Признак берётся из флагов PyMuPDF, а
не из имени шрифта: имена у каждого издательства свои, флаг одинаковый."""
return bool(s["flags"] & MONO)
def span_html(s, plain=False):
t = html.escape(s["text"])
if not t:
return ""
if plain: # внутри <pre> курсив и полужирный только мешают
return t
if is_mono(s):
return "<code>%s</code>" % t
bold, ital = style(s["font"])
if ital:
t = "<i>%s</i>" % t
@@ -59,6 +75,8 @@ def block_kind(b):
return ("h1", "title")
if sz >= 24:
return ("h1", None)
if all(is_mono(sp) for sp in spans):
return ("pre", None)
if sz >= 17:
return ("h2", None)
if f.startswith("MyriadPro"):
@@ -66,7 +84,13 @@ def block_kind(b):
return ("p", None)
def block_text(b):
def block_text(b, pre=False):
if pre:
# Перенос строки в коде значим, висящий дефис — это минус, а не перенос
# слова. Ни склейки строк, ни де-дефисации здесь быть не должно.
rows = ["".join(span_html(s, plain=True) for s in l["spans"])
for l in b["lines"]]
return "\n".join(rows).rstrip()
out = []
for i, l in enumerate(b["lines"]):
line = "".join(span_html(s) for s in l["spans"])
@@ -82,6 +106,7 @@ def block_text(b):
txt = "".join(out)
txt = re.sub(r"</i>(\s*)<i>", r"\1", txt)
txt = re.sub(r"</b>(\s*)<b>", r"\1", txt)
txt = re.sub(r"(?<=\S) {2,}(?=\S)", " ", txt)
return txt.strip()
@@ -112,7 +137,7 @@ for pno, page in enumerate(doc):
if not kind:
continue
tag, cls = kind
txt = block_text(b)
txt = block_text(b, pre=(tag == "pre"))
if not txt:
continue
indented = b["lines"][0]["bbox"][0] >= INDENT_X
@@ -134,36 +159,13 @@ for pno, page in enumerate(doc):
if open_para:
parts.append(open_para)
CSS = """
body { margin: 0 1em; }
h1 { text-align: center; margin: 2em 0 0.2em; page-break-before: always; }
h1.title { page-break-before: avoid; }
h2 { text-align: center; font-style: italic; font-weight: normal;
font-size: 1.1em; margin: 0.2em 0 1.5em; }
p { text-indent: 1.2em; margin: 0; text-align: justify; }
p.note { text-indent: 0; margin: 1em 2em; font-size: 0.9em;
font-family: sans-serif; }
p.figure { text-indent: 0; text-align: center; margin: 1em 0; }
img { max-width: 100%; }
"""
buf = ['<?xml version="1.0" encoding="utf-8"?>',
'<html xmlns="http://www.w3.org/1999/xhtml"><head>',
"<title>%s</title>" % html.escape(doc.metadata.get("title") or SRC.stem),
"<style>%s</style></head><body>" % CSS]
for tag, cls, txt in parts:
if tag == "figure":
buf.append('<p class="figure">%s</p>' % txt)
else:
c = ' class="%s"' % cls if cls else ""
buf.append("<%s%s>%s</%s>" % (tag, c, txt, tag))
buf.append("</body></html>")
doc_html = "\n".join(buf)
# merge tag runs split at page/line boundaries, collapse doubled spaces
doc_html = re.sub(r"</(i|b)>(\s*)<\1>", r"\2", doc_html)
doc_html = re.sub(r"(?<=\S) {2,}(?=\S)", " ", doc_html)
doc_html = bookhtml.document(doc.metadata.get("title") or SRC.stem, parts)
# merge tag runs split at page/line boundaries. Пробелы здесь уже не трогаем:
# внутри <pre> они значимы, а в прозе схлопнуты в block_text.
doc_html = re.sub(r"</(i|b|code)>(\s*)<\1>", r"\2", doc_html)
OUT.write_text(doc_html, encoding="utf-8")
print("blocks:", len(parts), "images:", img_n)
print("h1:", sum(1 for p in parts if p[0] == "h1"),
"p:", sum(1 for p in parts if p[0] == "p"))
"p:", sum(1 for p in parts if p[0] == "p"),
"pre:", sum(1 for p in parts if p[0] == "pre"))
+119
View File
@@ -0,0 +1,119 @@
#!/usr/bin/env python3
"""Самопроверка конвертера EPUB без сети: python3 test_epub2html.py"""
import tempfile
import zipfile
from pathlib import Path
import epub2html
CONTAINER = """<?xml version="1.0"?>
<container version="1.0" xmlns="urn:oasis:names:tc:opendocument:xmlns:container">
<rootfiles><rootfile full-path="OEBPS/content.opf"
media-type="application/oebps-package+xml"/></rootfiles>
</container>"""
OPF = """<?xml version="1.0"?>
<package xmlns="http://www.idpf.org/2007/opf" version="3.0">
<metadata/>
<manifest>
<item id="c1" href="ch1.xhtml" media-type="application/xhtml+xml"/>
<item id="c2" href="ch2.xhtml" media-type="application/xhtml+xml"/>
<item id="c3" href="ch3.xhtml" media-type="application/xhtml+xml"/>
<item id="nav" href="nav.xhtml" media-type="application/xhtml+xml"
properties="nav"/>
<item id="pic" href="img/fig.png" media-type="image/png"/>
<item id="cov" href="img/cover.png" media-type="image/png"
properties="cover-image"/>
</manifest>
<spine>
<itemref idref="nav"/><itemref idref="c1"/>
<itemref idref="c2"/><itemref idref="c3"/>
</spine>
</package>"""
CH1 = """<?xml version="1.0" encoding="utf-8"?>
<html xmlns="http://www.w3.org/1999/xhtml"><head><title>skip me</title></head>
<body>
<h2>Chapter One</h2>
<p>The <em>lazy</em> fox calls <code>fetch_data()</code> twice&nbsp;a day.</p>
<pre>def main():
x = 1 - 2
return x</pre>
<ul><li>first item</li><li>second item</li></ul>
<table><tr><th>Name</th><th>Value</th></tr><tr><td>alpha</td><td>1</td></tr></table>
<p>Broken <b>markup that never closes.</p>
<p><img src="img/fig.png" alt="figure"/></p>
<script>var noise = 1;</script>
</body></html>"""
CH2 = """<?xml version="1.0" encoding="utf-8"?>
<html xmlns="http://www.w3.org/1999/xhtml"><body>
<h2>Chapter Two</h2><p>Second chapter body text.</p>
</body></html>"""
CH3 = """<?xml version="1.0" encoding="utf-8"?>
<html xmlns="http://www.w3.org/1999/xhtml"><body>
<h2>Chapter Three</h2><div>Loose text outside any block tag.</div>
</body></html>"""
NAV = """<?xml version="1.0" encoding="utf-8"?>
<html xmlns="http://www.w3.org/1999/xhtml"><body>
<nav><ol><li>Chapter One</li></ol></nav></body></html>"""
PNG = bytes.fromhex("89504e470d0a1a0a") # достаточно как содержимое файла
def build_epub(path):
with zipfile.ZipFile(path, "w") as z:
z.writestr("mimetype", "application/epub+zip")
z.writestr("META-INF/container.xml", CONTAINER)
z.writestr("OEBPS/content.opf", OPF)
z.writestr("OEBPS/ch1.xhtml", CH1)
z.writestr("OEBPS/ch2.xhtml", CH2)
z.writestr("OEBPS/ch3.xhtml", CH3)
z.writestr("OEBPS/nav.xhtml", NAV)
z.writestr("OEBPS/img/fig.png", PNG)
z.writestr("OEBPS/img/cover.png", PNG)
def test_convert(tmp: Path):
src = tmp / "book.epub"
build_epub(src)
out = tmp / "book.html"
epub2html.convert(src, out)
text = out.read_text(encoding="utf-8")
lines = text.splitlines()
# Листинг: переносы строк, отступы и минусы не тронуты
pre = text[text.index("<pre>"):text.index("</pre>")]
assert "def main():\n x = 1 - 2\n return x" in pre, pre
assert "<code>" not in pre, "внутри листинга инлайновая разметка не нужна"
# Проза: инлайн переведён в наш набор тегов, сущности раскрыты
para = next(l for l in lines if "fox" in l)
assert "<i>lazy</i>" in para and "<code>fetch_data()</code>" in para, para
assert "&nbsp;" not in para and " " not in para, para
assert '<p class="li">first item</p>' in lines
assert '<p class="row">Name | Value</p>' in lines
assert '<p class="row">alpha | 1</p>' in lines
assert '<img src="images/img001.png"/>' in text
assert (out.parent / "images" / "cover.png").exists(), "обложка не извлечена"
# Несбалансированная чужая разметка закрывается на границе блока
broken = next(l for l in lines if "never closes" in l)
assert broken.count("<b>") == broken.count("</b>") == 1, broken
assert "noise" not in text, "<script> не должен попадать в книгу"
assert "skip me" not in text, "<title> документа — не текст книги"
assert "Loose text outside any block tag." in text, "текст вне блока потерян"
# nav-документ выкинут, h2 подняты до h1 — иначе книга уедет одной главой
assert text.count("Chapter One") == 1, "оглавление попало в текст"
assert text.count("<h1>") == 3 and "<h2>" not in text
if __name__ == "__main__":
with tempfile.TemporaryDirectory() as d:
test_convert(Path(d))
print("OK")
+67 -4
View File
@@ -1,6 +1,7 @@
#!/usr/bin/env python3
"""Самопроверка моста без сети: python3 test_translate.py"""
import json
import re
from pathlib import Path
import translate as t
@@ -17,6 +18,14 @@ def test_marks_roundtrip():
assert "&amp;" in body, "спецсимволы должны экранироваться обратно"
def test_inline_code_survives():
"""Имена функций в прозе не должны терять разметку по дороге к модели."""
marked = t.to_marks("call <code>fetch()</code> twice")
assert marked == "call ⟦code⟧fetch()⟦/code⟧ twice"
body, ok = t.from_marks(marked)
assert ok and "<code>fetch()</code>" in body
def test_broken_marks_drop_tags():
body, ok = t.from_marks("текст ⟦i⟧без закрытия")
assert not ok and "⟦" not in body and "<i>" not in body
@@ -71,6 +80,58 @@ def test_split_and_rebuild(tmp: Path):
assert result.count("<h1>") == 2
def test_code_listings_are_left_alone(tmp: Path):
"""Листинги не переводятся и не участвуют в приёмке по языку: иначе книга,
где кода много, честно переведётся и завалит проверку."""
listing = "\n".join([
"<pre>def compute_total(items, discount_rate, shipping_cost):",
" subtotal = sum(item.price * item.quantity for item in items)",
" return subtotal * (1 - discount_rate) + shipping_cost # not translated",
"</pre>",
])
prose = ("<p>Текст главы про обработку данных, списки значений и способы "
"хранения промежуточных результатов между запусками.</p>")
src = tmp / "code.html"
src.write_text("<h1>Глава</h1>\n" + (listing + "\n" + prose + "\n") * 3,
encoding="utf-8")
blocks = t.parse_blocks(src)
assert [b[1] for b in blocks] == ["h1", "p", "p", "p"], \
"листинг не должен попасть в перевод"
# латиницы в файле больше, чем кириллицы, — но вся она в <pre>
raw = re.sub("<[^>]+>", "", src.read_text(encoding="utf-8"))
assert t.detect_language(raw)[0] == "en", "образец должен быть латинским целиком"
assert t.verify_output(src, "ru"), "код не должен заваливать приёмку"
def test_workdir_belongs_to_one_book(tmp: Path):
"""Имена chapter_NNN.json у всех книг одинаковы: чужой workdir склеит
перевод одной книги с текстом другой."""
wd = tmp / "wd"
wd.mkdir()
a = [(0, "p", "", "First book text.")]
b = [(0, "p", "", "Совершенно другая книга.")]
t.claim_workdir(wd, Path("a.html"), a)
t.claim_workdir(wd, Path("a.html"), a) # повторный запуск той же книги — ок
try:
t.claim_workdir(wd, Path("b.html"), b)
except SystemExit as e:
assert "занят другой книгой" in str(e)
else:
raise AssertionError("чужая книга в занятом workdir должна отвергаться")
def test_stale_chapters_removed(tmp: Path):
"""Прошлый прогон был длиннее — его главы иначе уедут в перевод."""
extracted = tmp / "extracted"
extracted.mkdir()
(extracted / "chapter_009.json").write_text("{}", encoding="utf-8")
t.write_input([[(0, "h1", "", "Глава")]], extracted)
assert not (extracted / "chapter_009.json").exists()
assert (extracted / "chapter_000.json").exists()
def test_verify_output_catches_untranslated(tmp: Path):
"""Регресс: первый боевой прогон отрапортовал успех на непереведённом файле."""
bad = tmp / "bad.html"
@@ -90,10 +151,12 @@ if __name__ == "__main__":
import tempfile
test_marks_roundtrip()
test_inline_code_survives()
test_broken_marks_drop_tags()
test_language_detection()
with tempfile.TemporaryDirectory() as d:
test_split_and_rebuild(Path(d))
with tempfile.TemporaryDirectory() as d:
test_verify_output_catches_untranslated(Path(d))
for case in (test_split_and_rebuild, test_code_listings_are_left_alone,
test_workdir_belongs_to_one_book, test_stale_chapters_removed,
test_verify_output_catches_untranslated):
with tempfile.TemporaryDirectory() as d:
case(Path(d))
print("OK")
+35 -5
View File
@@ -11,6 +11,7 @@
поломке абзац сохраняется переведённым, но без внутренней разметки.
"""
import argparse
import hashlib
import html
import json
import re
@@ -19,10 +20,15 @@ import sys
from pathlib import Path
BLOCK = re.compile(r"^<(p|h1|h2)([^>]*)>(.*)</\1>$")
TAG = re.compile(r"</?([ib])>")
MARK = re.compile(r"⟦(/?)([ib])⟧")
INLINE = ("i", "b", "code")
TAG = re.compile(r"</?(%s)>" % "|".join(INLINE))
MARK = re.compile(r"⟦(/?)(%s)⟧" % "|".join(INLINE))
# картинки и пустые блоки не переводим
SKIP = re.compile(r"<img\b")
# Листинги кода не переводятся вовсе: BLOCK их не узнаёт, поэтому строки <pre>
# доходят до результата нетронутыми. Здесь <pre> вырезается только из проверок
# языка — иначе английский код утянул бы долю кириллицы ниже порога приёмки.
PRE = re.compile(r"<pre\b.*?</pre>", re.S)
def detect_language(text):
@@ -44,7 +50,7 @@ def to_marks(s):
def from_marks(s):
"""Маркеры -> теги. Возвращает (html, ok): ok=False если разметка разъехалась."""
escaped = html.escape(s, quote=False)
depth = {"i": 0, "b": 0}
depth = dict.fromkeys(INLINE, 0)
ok = True
for slash, tag in MARK.findall(escaped):
depth[tag] += -1 if slash else 1
@@ -82,6 +88,8 @@ def split_chapters(blocks):
def write_input(chapters, extracted):
extracted.mkdir(parents=True, exist_ok=True)
for stale in extracted.glob("chapter_*.json"):
stale.unlink() # главы прошлого прогона иначе уедут в перевод
index = []
for n, ch in enumerate(chapters):
paragraphs = [to_marks(b[3]) for b in ch]
@@ -109,6 +117,25 @@ def write_input(chapters, extracted):
return index
def claim_workdir(workdir, source, blocks):
"""Рабочий каталог принадлежит одной книге. Имена chapter_NNN.json у всех
книг одинаковы, а внешний репозиторий пропускает главы, помеченные
готовыми, — без этой привязки вторая книга молча соберётся из перевода
первой."""
digest = hashlib.sha256(
"\n".join(b[3] for b in blocks).encode("utf-8")).hexdigest()[:16]
claim = workdir / "bridge_source.json"
now = {"source": source.name, "blocks": len(blocks), "sha256": digest}
if claim.exists():
was = json.loads(claim.read_text(encoding="utf-8"))
if was.get("sha256") != digest:
sys.exit("рабочий каталог занят другой книгой (%s, блоков %s). "
"Возьми чистый --workdir." % (was.get("source"), was.get("blocks")))
else:
claim.write_text(json.dumps(now, ensure_ascii=False, indent=2),
encoding="utf-8")
def run_translator(repo, workdir, extracted, workers):
script = repo / "03_translate_parallel.py"
if not script.exists():
@@ -150,7 +177,7 @@ def verify_output(out, target):
"""Совпадение числа абзацев ещё не значит, что перевод состоялся: внешний
скрипт при ошибке API молча подставляет оригинал. Проверяем язык результата.
"""
text = re.sub("<[^>]+>", "", out.read_text(encoding="utf-8"))
text = re.sub("<[^>]+>", "", PRE.sub(" ", out.read_text(encoding="utf-8")))
lang, share = detect_language(text)
stub = text.count("[UNTRANSLATED]")
print("результат: %s (кириллица %.0f%%), заглушек [UNTRANSLATED]: %d"
@@ -186,7 +213,9 @@ def main():
plain = " ".join(re.sub("<[^>]+>", "", b[3]) for b in blocks)
lang, share = detect_language(plain)
chapters = split_chapters(blocks)
print("блоков: %d, глав: %d, знаков: %d" % (len(blocks), len(chapters), len(plain)))
pre_n = len(PRE.findall(args.source.read_text(encoding="utf-8")))
print("блоков: %d, глав: %d, знаков: %d, листингов <pre> (не переводятся): %d"
% (len(blocks), len(chapters), len(plain), pre_n))
print("язык источника: %s (кириллица %.0f%%)" % (lang, share * 100))
if lang == args.target and not args.force:
@@ -199,6 +228,7 @@ def main():
return
args.workdir.mkdir(parents=True, exist_ok=True)
claim_workdir(args.workdir, args.source, blocks)
extracted = (args.workdir / "extracted").resolve()
write_input(chapters, extracted)
run_translator(args.repo.resolve(), args.workdir.resolve(), extracted, args.workers)