EPUB как второй вход в конвейер; листинги кода не переводятся и не ломают приёмку
- epub2html.py: EPUB -> тот же XHTML, что pdf2html.py, только стандартная библиотека. Уровень глав определяется по книге, а не берётся из <h1>: в EPUB там обычно название и части. - pdf2html.py: листинг узнаётся по флагу моноширинного шрифта PyMuPDF, переносы и отступы внутри <pre> сохраняются, де-дефисация к коду не применяется. - translate.py: <pre> исключён из проверок языка (английский код утягивал долю кириллицы ниже порога приёмки), рабочий каталог привязан к книге, устаревшие главы прошлого прогона чистятся, инлайновый <code> переживает переводчика. - bookhtml.py: общий каркас документа и CSS для обоих входов. - Тесты: test_epub2html.py на собранном в памяти EPUB, четыре новых случая в test_translate.py.
This commit is contained in:
@@ -15,6 +15,30 @@ tag, and 2817 body paragraphs became `<h2>`.
|
||||
properties, and emits semantic XHTML. calibre is then used only as the
|
||||
XHTML→EPUB packer.
|
||||
|
||||
## Two entry points
|
||||
|
||||
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB. Both emit the
|
||||
same normalized XHTML — one block per line, `<pre>` for code listings — and
|
||||
everything downstream (translation, packing, audiobook) is identical. Shared
|
||||
document skeleton and CSS live in `scripts/bookhtml.py`.
|
||||
|
||||
**Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws
|
||||
away the semantic markup that is already there and re-derives it from font
|
||||
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
|
||||
the spine, drops the nav document, and maps the book's own headings.
|
||||
|
||||
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is
|
||||
needed, not PyMuPDF) and convert with:
|
||||
|
||||
`python scripts/epub2html.py book.epub out/book.html`
|
||||
|
||||
It reports which heading level turned out to be the chapter level. In most
|
||||
EPUBs `<h1>` is the book title and the parts, while chapters are `<h2>` — the
|
||||
script picks the deepest level that still gives a sane chapter count, because
|
||||
the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
|
||||
3rd edition: 7 `<h1>` against 49 `<h2>`, chapters correctly detected as `h2`,
|
||||
56 chapters out of 656 pages.
|
||||
|
||||
## Workflow
|
||||
|
||||
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
|
||||
@@ -35,7 +59,9 @@ XHTML→EPUB packer.
|
||||
Read off: the body font (largest character count), the italic variant, the
|
||||
heading sizes, and any secondary family used for sidebars, journal entries,
|
||||
chat logs or slides.
|
||||
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. The
|
||||
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. Code
|
||||
listings need no tuning — a monospaced span is recognized by the PyMuPDF font
|
||||
flag (`flags & 8`), not by font name, and becomes a `<pre>` block. The
|
||||
defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a
|
||||
title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary
|
||||
family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that
|
||||
@@ -45,7 +71,10 @@ XHTML→EPUB packer.
|
||||
|
||||
`python scripts/pdf2html.py book.pdf out/book.html`
|
||||
|
||||
6. **Verify before packing** (see Verification). Fix thresholds and re-run until
|
||||
6. **Verify before packing** (see Verification). Check the `pre:` counter against
|
||||
the real number of listings in the book — a technical book reporting `pre: 0`
|
||||
means the mono flag never fired and every listing is about to be reflowed as
|
||||
prose. Fix thresholds and re-run until
|
||||
the counts are sane. Cheap to iterate; do not skip to packing.
|
||||
7. **Pack with calibre**, always passing explicit TOC XPaths — without them
|
||||
calibre applies its own heuristics and re-breaks the chapters:
|
||||
@@ -97,6 +126,23 @@ navPoint count for the TOC size.
|
||||
Only when the user asks for a translated book. It slots between step 6 and
|
||||
step 7 — translate the XHTML, then pack the translated file with calibre.
|
||||
|
||||
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
|
||||
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
|
||||
the result untouched. Nothing extra is needed to protect code — but this also
|
||||
means a listing that was misclassified as a paragraph upstream *will* be
|
||||
translated, which is the real reason step 6 checks the `pre:` counter.
|
||||
|
||||
For the same reason `<pre>` is cut out of both language checks. A book that is
|
||||
40% listings translates correctly and would otherwise fail acceptance, because
|
||||
the English code drags the Cyrillic share below the threshold.
|
||||
|
||||
**One workdir per book.** Chapter files are named `chapter_NNN.json` for every
|
||||
book, and the external repo skips chapters it has marked done — a reused workdir
|
||||
silently stitches one book's translation onto another's text. The bridge writes
|
||||
`bridge_source.json` into the workdir on first run and refuses to start if the
|
||||
directory belongs to a different book. Re-running the same book is unaffected;
|
||||
that is the resume path.
|
||||
|
||||
Translation is delegated to an external project,
|
||||
[`vetermanve/book_translator`](https://github.com/vetermanve/book_translator)
|
||||
(DeepSeek or a local Ollama model). **Never modify that repository** — it is
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Общая сборка XHTML для pdf2html.py и epub2html.py.
|
||||
|
||||
Формат намеренно жёсткий: один блок — одна строка (кроме <pre>, где переносы
|
||||
значимы). На этом держится мост translate.py — он правит блоки по номеру
|
||||
строки, а всё, чего не узнал, доносит до результата нетронутым.
|
||||
"""
|
||||
import html
|
||||
|
||||
CSS = """
|
||||
body { margin: 0 1em; }
|
||||
h1 { text-align: center; margin: 2em 0 0.2em; page-break-before: always; }
|
||||
h1.title { page-break-before: avoid; }
|
||||
h2 { text-align: center; font-style: italic; font-weight: normal;
|
||||
font-size: 1.1em; margin: 0.2em 0 1.5em; }
|
||||
p { text-indent: 1.2em; margin: 0; text-align: justify; }
|
||||
p.note { text-indent: 0; margin: 1em 2em; font-size: 0.9em;
|
||||
font-family: sans-serif; }
|
||||
p.li { text-indent: 0; margin: 0.3em 0 0.3em 1.5em; text-align: left; }
|
||||
p.row { text-indent: 0; margin: 0.3em 0; text-align: left;
|
||||
font-size: 0.9em; }
|
||||
p.figure { text-indent: 0; text-align: center; margin: 1em 0; }
|
||||
img { max-width: 100%; }
|
||||
pre { font-family: monospace; font-size: 0.75em; margin: 1em 0;
|
||||
white-space: pre-wrap; word-wrap: break-word; text-align: left;
|
||||
text-indent: 0; }
|
||||
code { font-family: monospace; font-size: 0.9em; }
|
||||
"""
|
||||
|
||||
|
||||
def document(title, parts):
|
||||
"""parts: [(tag, cls, inner_html)]; tag 'figure' — псевдотег для картинки."""
|
||||
buf = ['<?xml version="1.0" encoding="utf-8"?>',
|
||||
'<html xmlns="http://www.w3.org/1999/xhtml"><head>',
|
||||
"<title>%s</title>" % html.escape(title),
|
||||
"<style>%s</style></head><body>" % CSS]
|
||||
for tag, cls, txt in parts:
|
||||
if tag == "figure":
|
||||
buf.append('<p class="figure">%s</p>' % txt)
|
||||
else:
|
||||
c = ' class="%s"' % cls if cls else ""
|
||||
buf.append("<%s%s>%s</%s>" % (tag, c, txt, tag))
|
||||
buf.append("</body></html>")
|
||||
return "\n".join(buf)
|
||||
@@ -0,0 +1,267 @@
|
||||
#!/usr/bin/env python3
|
||||
"""EPUB -> тот же XHTML, что выдаёт pdf2html.py (один блок = одна строка).
|
||||
|
||||
EPUB уже размечен семантически, поэтому здесь не разведка шрифтов, а
|
||||
нормализация: привести чужую вёрстку к формату, который понимает мост
|
||||
translate.py, и вычистить всё, что переводчику видеть не нужно.
|
||||
|
||||
Листинги остаются в <pre> и не переводятся вовсе: BLOCK-регулярка моста их не
|
||||
узнаёт, а всё неузнанное доходит до результата нетронутым.
|
||||
|
||||
Только стандартная библиотека — ни calibre, ни lxml здесь не нужны.
|
||||
"""
|
||||
import argparse
|
||||
import collections
|
||||
import html
|
||||
import posixpath
|
||||
import re
|
||||
import sys
|
||||
import zipfile
|
||||
from html.parser import HTMLParser
|
||||
from pathlib import Path
|
||||
from xml.etree import ElementTree
|
||||
|
||||
import bookhtml
|
||||
|
||||
CONTAINER = "META-INF/container.xml"
|
||||
OPF_NS = "{http://www.idpf.org/2007/opf}"
|
||||
CONT_NS = "{urn:oasis:names:tc:opendocument:xmlns:container}"
|
||||
|
||||
# tag источника -> (tag результата, css-класс)
|
||||
BLOCK_MAP = {
|
||||
"h1": ("h1", None), "h2": ("h2", None), "h3": ("h3", None),
|
||||
"h4": ("h4", None), "h5": ("h5", None), "h6": ("h6", None),
|
||||
"p": ("p", None), "pre": ("pre", None), "li": ("p", "li"),
|
||||
"tr": ("p", "row"), "figcaption": ("p", "note"),
|
||||
"dt": ("p", "note"), "dd": ("p", "note"),
|
||||
}
|
||||
INLINE_MAP = {"i": "i", "em": "i", "cite": "i", "b": "b", "strong": "b",
|
||||
"code": "code", "kbd": "code", "samp": "code", "tt": "code",
|
||||
"var": "code"}
|
||||
DROP = {"script", "style", "head", "title", "nav"}
|
||||
IMG_EXT = {"image/jpeg": ".jpg", "image/png": ".png", "image/gif": ".gif",
|
||||
"image/svg+xml": ".svg", "image/webp": ".webp"}
|
||||
|
||||
|
||||
class Reader(HTMLParser):
|
||||
"""Собирает блоки из одного XHTML-документа книги."""
|
||||
|
||||
def __init__(self, on_image):
|
||||
super().__init__(convert_charrefs=True)
|
||||
self.on_image = on_image
|
||||
self.blocks = []
|
||||
self.cur = None # (tag, cls, [куски])
|
||||
self.open_inline = [] # незакрытые инлайновые теги текущего блока
|
||||
self.drop = 0
|
||||
|
||||
# -- служебное ---------------------------------------------------------
|
||||
def flush(self):
|
||||
if not self.cur:
|
||||
return
|
||||
tag, cls, parts = self.cur
|
||||
self.cur = None
|
||||
for t in reversed(self.open_inline):
|
||||
parts.append("</%s>" % t) # чужая вёрстка бывает несбалансированной
|
||||
self.open_inline = []
|
||||
txt = "".join(parts)
|
||||
if tag == "pre":
|
||||
txt = txt.strip("\n").rstrip()
|
||||
else:
|
||||
txt = re.sub(r"\s+", " ", txt).strip()
|
||||
txt = re.sub(r"<(i|b|code)>(\s*)</\1>", r"\2", txt) # пустая разметка
|
||||
if txt.strip():
|
||||
self.blocks.append((tag, cls, txt))
|
||||
|
||||
def start(self, tag, cls):
|
||||
self.flush()
|
||||
self.cur = (tag, cls, [])
|
||||
|
||||
# -- HTMLParser --------------------------------------------------------
|
||||
def handle_starttag(self, tag, attrs):
|
||||
if tag in DROP:
|
||||
self.drop += 1
|
||||
return
|
||||
if self.drop:
|
||||
return
|
||||
a = dict(attrs)
|
||||
if tag in ("img", "image"):
|
||||
src = a.get("src") or a.get("{http://www.w3.org/1999/xlink}href") \
|
||||
or a.get("xlink:href") or a.get("href")
|
||||
name = self.on_image(src) if src else None
|
||||
if name:
|
||||
self.flush()
|
||||
self.blocks.append(("figure", None,
|
||||
'<img src="images/%s"/>' % name))
|
||||
return
|
||||
if tag == "br":
|
||||
if self.cur:
|
||||
self.cur[2].append("\n" if self.cur[0] == "pre" else " ")
|
||||
return
|
||||
if tag in BLOCK_MAP:
|
||||
self.start(*BLOCK_MAP[tag])
|
||||
return
|
||||
if tag in ("td", "th") and self.cur and self.cur[0] == "p" \
|
||||
and self.cur[1] == "row" and self.cur[2]:
|
||||
self.cur[2].append(" | ")
|
||||
return
|
||||
if tag in INLINE_MAP and self.cur and self.cur[0] != "pre":
|
||||
# Внутри листинга инлайн не нужен: книги размечают код сплошным
|
||||
# <strong>, и на e-ink это страница жирного текста.
|
||||
out = INLINE_MAP[tag]
|
||||
self.open_inline.append(out)
|
||||
self.cur[2].append("<%s>" % out)
|
||||
|
||||
def handle_endtag(self, tag):
|
||||
if tag in DROP:
|
||||
self.drop = max(0, self.drop - 1)
|
||||
return
|
||||
if self.drop:
|
||||
return
|
||||
if tag in BLOCK_MAP:
|
||||
self.flush()
|
||||
return
|
||||
if tag in INLINE_MAP and self.cur and self.cur[0] != "pre":
|
||||
out = INLINE_MAP[tag]
|
||||
if out in self.open_inline:
|
||||
self.open_inline.remove(out)
|
||||
self.cur[2].append("</%s>" % out)
|
||||
|
||||
def handle_data(self, data):
|
||||
if self.drop:
|
||||
return
|
||||
if not self.cur:
|
||||
if not data.strip():
|
||||
return
|
||||
self.start("p", None) # текст вне блочного тега не теряем
|
||||
self.cur[2].append(html.escape(data, quote=False))
|
||||
|
||||
def close(self):
|
||||
super().close()
|
||||
self.flush()
|
||||
|
||||
|
||||
def spine_documents(zf):
|
||||
"""[(путь в архиве, media-type)] в порядке чтения + путь к OPF."""
|
||||
root = ElementTree.fromstring(zf.read(CONTAINER))
|
||||
opf_path = root.find(".//%srootfile" % CONT_NS).get("full-path")
|
||||
opf = ElementTree.fromstring(zf.read(opf_path))
|
||||
base = posixpath.dirname(opf_path)
|
||||
items, cover = {}, None
|
||||
for it in opf.iter("%sitem" % OPF_NS):
|
||||
href = posixpath.normpath(posixpath.join(base, it.get("href")))
|
||||
props = it.get("properties") or ""
|
||||
items[it.get("id")] = (href, it.get("media-type"), props)
|
||||
if "cover-image" in props:
|
||||
cover = href
|
||||
if cover is None: # EPUB 2: обложка объявляется через <meta name="cover">
|
||||
meta = opf.find(".//%smeta[@name='cover']" % OPF_NS)
|
||||
if meta is not None and meta.get("content") in items:
|
||||
cover = items[meta.get("content")][0]
|
||||
docs = []
|
||||
for ref in opf.iter("%sitemref" % OPF_NS):
|
||||
item = items.get(ref.get("idref"))
|
||||
if not item:
|
||||
continue
|
||||
href, mtype, props = item
|
||||
if "nav" in props or re.search(r"(toc|nav|content)s?\.x?html?$", href, re.I):
|
||||
continue
|
||||
docs.append(href)
|
||||
return docs, items, cover
|
||||
|
||||
|
||||
# Больше этого числа глав дробить незачем: переводчик держит контекст в пределах
|
||||
# главы, слишком мелкая нарезка его обесценивает.
|
||||
MAX_CHAPTERS = 150
|
||||
|
||||
|
||||
def chapter_level(parts):
|
||||
"""Каким уровнем заголовка в этой книге размечены главы.
|
||||
|
||||
Мост режет книгу на главы по <h1>, а в EPUB <h1> обычно занят названием
|
||||
книги и частями: у Страуструпа их 7 на 656 страниц, тогда как главы — это
|
||||
<h2> (49 штук). Берём самый глубокий уровень, при котором число глав ещё
|
||||
остаётся вменяемым.
|
||||
"""
|
||||
counts = collections.Counter(t for t, _, _ in parts if re.fullmatch(r"h[1-6]", t))
|
||||
best, total = 1, 0
|
||||
for lvl in range(1, 7):
|
||||
total += counts.get("h%d" % lvl, 0)
|
||||
if total > MAX_CHAPTERS:
|
||||
break
|
||||
if total:
|
||||
best = lvl
|
||||
return best
|
||||
|
||||
|
||||
def convert(src, out):
|
||||
imgdir = out.parent / "images"
|
||||
imgdir.mkdir(parents=True, exist_ok=True)
|
||||
parts, seen, counter = [], {}, [0]
|
||||
|
||||
with zipfile.ZipFile(src) as zf:
|
||||
names = set(zf.namelist())
|
||||
docs, items, cover = spine_documents(zf)
|
||||
|
||||
def extract(path, stem=None):
|
||||
if path in seen:
|
||||
return seen[path]
|
||||
if path not in names:
|
||||
return None
|
||||
ext = posixpath.splitext(path)[1] or ".img"
|
||||
if stem:
|
||||
name = stem + ext
|
||||
else:
|
||||
counter[0] += 1
|
||||
name = "img%03d%s" % (counter[0], ext)
|
||||
(imgdir / name).write_bytes(zf.read(path))
|
||||
seen[path] = name
|
||||
return name
|
||||
|
||||
if cover:
|
||||
extract(cover, stem="cover")
|
||||
|
||||
for doc in docs:
|
||||
if doc not in names:
|
||||
print("нет в архиве, пропущен: %s" % doc, file=sys.stderr)
|
||||
continue
|
||||
here = posixpath.dirname(doc)
|
||||
|
||||
def on_image(src_attr, here=here):
|
||||
target = posixpath.normpath(posixpath.join(here, src_attr.split("#")[0]))
|
||||
return extract(target)
|
||||
|
||||
r = Reader(on_image)
|
||||
r.feed(zf.read(doc).decode("utf-8", "replace"))
|
||||
r.close()
|
||||
parts.extend(r.blocks)
|
||||
|
||||
title = out.stem
|
||||
lvl = chapter_level(parts)
|
||||
parts = [(("h1" if int(t[1]) <= lvl else "h2", c, x)
|
||||
if re.fullmatch(r"h[1-6]", t) else (t, c, x))
|
||||
for t, c, x in parts]
|
||||
h1 = sum(1 for p in parts if p[0] == "h1")
|
||||
print("главы размечены h%d, глав получилось: %d" % (lvl, h1))
|
||||
if h1 < 3:
|
||||
print("ВНИМАНИЕ: заголовков-глав всего %d — переводчик будет держать "
|
||||
"контекст крупными кусками, проверь результат внимательнее" % h1)
|
||||
|
||||
out.write_text(bookhtml.document(title, parts), encoding="utf-8")
|
||||
print("blocks: %d, images: %d" % (len(parts), len(seen)))
|
||||
print("h1: %d p: %d pre: %d"
|
||||
% (h1, sum(1 for p in parts if p[0] == "p"),
|
||||
sum(1 for p in parts if p[0] == "pre")))
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser(description=__doc__)
|
||||
ap.add_argument("source", type=Path, help="исходный .epub")
|
||||
ap.add_argument("output", type=Path, help="куда писать XHTML")
|
||||
args = ap.parse_args()
|
||||
if not args.source.exists():
|
||||
sys.exit("нет файла %s" % args.source)
|
||||
convert(args.source, args.output)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
+34
-32
@@ -7,6 +7,11 @@ from pathlib import Path
|
||||
|
||||
import pymupdf
|
||||
|
||||
import bookhtml
|
||||
|
||||
if len(sys.argv) < 3 or sys.argv[1] in ("-h", "--help"):
|
||||
sys.exit("usage: pdf2html.py book.pdf out.html | pdf2html.py --fonts book.pdf")
|
||||
|
||||
if sys.argv[1] == "--fonts": # разведка: какие шрифты/кегли в PDF
|
||||
import collections
|
||||
|
||||
@@ -29,6 +34,7 @@ IMGDIR.mkdir(parents=True, exist_ok=True)
|
||||
# Пороги подобраны под вёрстку 15pt/letter. Для другой книги сначала
|
||||
# посмотреть реальные шрифты и кегли: pdf2html.py --fonts file.pdf
|
||||
INDENT_X = 88 # x0 первой строки: больше — абзац с красной строки
|
||||
MONO = 8 # бит моноширинного шрифта в span["flags"] (PyMuPDF)
|
||||
|
||||
|
||||
def style(font):
|
||||
@@ -36,10 +42,20 @@ def style(font):
|
||||
return ("bold" in f or "semibold" in f, "-it" in f or "italic" in f)
|
||||
|
||||
|
||||
def span_html(s):
|
||||
def is_mono(s):
|
||||
"""Моноширинный шрифт = листинг кода. Признак берётся из флагов PyMuPDF, а
|
||||
не из имени шрифта: имена у каждого издательства свои, флаг одинаковый."""
|
||||
return bool(s["flags"] & MONO)
|
||||
|
||||
|
||||
def span_html(s, plain=False):
|
||||
t = html.escape(s["text"])
|
||||
if not t:
|
||||
return ""
|
||||
if plain: # внутри <pre> курсив и полужирный только мешают
|
||||
return t
|
||||
if is_mono(s):
|
||||
return "<code>%s</code>" % t
|
||||
bold, ital = style(s["font"])
|
||||
if ital:
|
||||
t = "<i>%s</i>" % t
|
||||
@@ -59,6 +75,8 @@ def block_kind(b):
|
||||
return ("h1", "title")
|
||||
if sz >= 24:
|
||||
return ("h1", None)
|
||||
if all(is_mono(sp) for sp in spans):
|
||||
return ("pre", None)
|
||||
if sz >= 17:
|
||||
return ("h2", None)
|
||||
if f.startswith("MyriadPro"):
|
||||
@@ -66,7 +84,13 @@ def block_kind(b):
|
||||
return ("p", None)
|
||||
|
||||
|
||||
def block_text(b):
|
||||
def block_text(b, pre=False):
|
||||
if pre:
|
||||
# Перенос строки в коде значим, висящий дефис — это минус, а не перенос
|
||||
# слова. Ни склейки строк, ни де-дефисации здесь быть не должно.
|
||||
rows = ["".join(span_html(s, plain=True) for s in l["spans"])
|
||||
for l in b["lines"]]
|
||||
return "\n".join(rows).rstrip()
|
||||
out = []
|
||||
for i, l in enumerate(b["lines"]):
|
||||
line = "".join(span_html(s) for s in l["spans"])
|
||||
@@ -82,6 +106,7 @@ def block_text(b):
|
||||
txt = "".join(out)
|
||||
txt = re.sub(r"</i>(\s*)<i>", r"\1", txt)
|
||||
txt = re.sub(r"</b>(\s*)<b>", r"\1", txt)
|
||||
txt = re.sub(r"(?<=\S) {2,}(?=\S)", " ", txt)
|
||||
return txt.strip()
|
||||
|
||||
|
||||
@@ -112,7 +137,7 @@ for pno, page in enumerate(doc):
|
||||
if not kind:
|
||||
continue
|
||||
tag, cls = kind
|
||||
txt = block_text(b)
|
||||
txt = block_text(b, pre=(tag == "pre"))
|
||||
if not txt:
|
||||
continue
|
||||
indented = b["lines"][0]["bbox"][0] >= INDENT_X
|
||||
@@ -134,36 +159,13 @@ for pno, page in enumerate(doc):
|
||||
if open_para:
|
||||
parts.append(open_para)
|
||||
|
||||
CSS = """
|
||||
body { margin: 0 1em; }
|
||||
h1 { text-align: center; margin: 2em 0 0.2em; page-break-before: always; }
|
||||
h1.title { page-break-before: avoid; }
|
||||
h2 { text-align: center; font-style: italic; font-weight: normal;
|
||||
font-size: 1.1em; margin: 0.2em 0 1.5em; }
|
||||
p { text-indent: 1.2em; margin: 0; text-align: justify; }
|
||||
p.note { text-indent: 0; margin: 1em 2em; font-size: 0.9em;
|
||||
font-family: sans-serif; }
|
||||
p.figure { text-indent: 0; text-align: center; margin: 1em 0; }
|
||||
img { max-width: 100%; }
|
||||
"""
|
||||
|
||||
buf = ['<?xml version="1.0" encoding="utf-8"?>',
|
||||
'<html xmlns="http://www.w3.org/1999/xhtml"><head>',
|
||||
"<title>%s</title>" % html.escape(doc.metadata.get("title") or SRC.stem),
|
||||
"<style>%s</style></head><body>" % CSS]
|
||||
for tag, cls, txt in parts:
|
||||
if tag == "figure":
|
||||
buf.append('<p class="figure">%s</p>' % txt)
|
||||
else:
|
||||
c = ' class="%s"' % cls if cls else ""
|
||||
buf.append("<%s%s>%s</%s>" % (tag, c, txt, tag))
|
||||
buf.append("</body></html>")
|
||||
doc_html = "\n".join(buf)
|
||||
# merge tag runs split at page/line boundaries, collapse doubled spaces
|
||||
doc_html = re.sub(r"</(i|b)>(\s*)<\1>", r"\2", doc_html)
|
||||
doc_html = re.sub(r"(?<=\S) {2,}(?=\S)", " ", doc_html)
|
||||
doc_html = bookhtml.document(doc.metadata.get("title") or SRC.stem, parts)
|
||||
# merge tag runs split at page/line boundaries. Пробелы здесь уже не трогаем:
|
||||
# внутри <pre> они значимы, а в прозе схлопнуты в block_text.
|
||||
doc_html = re.sub(r"</(i|b|code)>(\s*)<\1>", r"\2", doc_html)
|
||||
OUT.write_text(doc_html, encoding="utf-8")
|
||||
|
||||
print("blocks:", len(parts), "images:", img_n)
|
||||
print("h1:", sum(1 for p in parts if p[0] == "h1"),
|
||||
"p:", sum(1 for p in parts if p[0] == "p"))
|
||||
"p:", sum(1 for p in parts if p[0] == "p"),
|
||||
"pre:", sum(1 for p in parts if p[0] == "pre"))
|
||||
|
||||
@@ -0,0 +1,119 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Самопроверка конвертера EPUB без сети: python3 test_epub2html.py"""
|
||||
import tempfile
|
||||
import zipfile
|
||||
from pathlib import Path
|
||||
|
||||
import epub2html
|
||||
|
||||
CONTAINER = """<?xml version="1.0"?>
|
||||
<container version="1.0" xmlns="urn:oasis:names:tc:opendocument:xmlns:container">
|
||||
<rootfiles><rootfile full-path="OEBPS/content.opf"
|
||||
media-type="application/oebps-package+xml"/></rootfiles>
|
||||
</container>"""
|
||||
|
||||
OPF = """<?xml version="1.0"?>
|
||||
<package xmlns="http://www.idpf.org/2007/opf" version="3.0">
|
||||
<metadata/>
|
||||
<manifest>
|
||||
<item id="c1" href="ch1.xhtml" media-type="application/xhtml+xml"/>
|
||||
<item id="c2" href="ch2.xhtml" media-type="application/xhtml+xml"/>
|
||||
<item id="c3" href="ch3.xhtml" media-type="application/xhtml+xml"/>
|
||||
<item id="nav" href="nav.xhtml" media-type="application/xhtml+xml"
|
||||
properties="nav"/>
|
||||
<item id="pic" href="img/fig.png" media-type="image/png"/>
|
||||
<item id="cov" href="img/cover.png" media-type="image/png"
|
||||
properties="cover-image"/>
|
||||
</manifest>
|
||||
<spine>
|
||||
<itemref idref="nav"/><itemref idref="c1"/>
|
||||
<itemref idref="c2"/><itemref idref="c3"/>
|
||||
</spine>
|
||||
</package>"""
|
||||
|
||||
CH1 = """<?xml version="1.0" encoding="utf-8"?>
|
||||
<html xmlns="http://www.w3.org/1999/xhtml"><head><title>skip me</title></head>
|
||||
<body>
|
||||
<h2>Chapter One</h2>
|
||||
<p>The <em>lazy</em> fox calls <code>fetch_data()</code> twice a day.</p>
|
||||
<pre>def main():
|
||||
x = 1 - 2
|
||||
return x</pre>
|
||||
<ul><li>first item</li><li>second item</li></ul>
|
||||
<table><tr><th>Name</th><th>Value</th></tr><tr><td>alpha</td><td>1</td></tr></table>
|
||||
<p>Broken <b>markup that never closes.</p>
|
||||
<p><img src="img/fig.png" alt="figure"/></p>
|
||||
<script>var noise = 1;</script>
|
||||
</body></html>"""
|
||||
|
||||
CH2 = """<?xml version="1.0" encoding="utf-8"?>
|
||||
<html xmlns="http://www.w3.org/1999/xhtml"><body>
|
||||
<h2>Chapter Two</h2><p>Second chapter body text.</p>
|
||||
</body></html>"""
|
||||
|
||||
CH3 = """<?xml version="1.0" encoding="utf-8"?>
|
||||
<html xmlns="http://www.w3.org/1999/xhtml"><body>
|
||||
<h2>Chapter Three</h2><div>Loose text outside any block tag.</div>
|
||||
</body></html>"""
|
||||
|
||||
NAV = """<?xml version="1.0" encoding="utf-8"?>
|
||||
<html xmlns="http://www.w3.org/1999/xhtml"><body>
|
||||
<nav><ol><li>Chapter One</li></ol></nav></body></html>"""
|
||||
|
||||
PNG = bytes.fromhex("89504e470d0a1a0a") # достаточно как содержимое файла
|
||||
|
||||
|
||||
def build_epub(path):
|
||||
with zipfile.ZipFile(path, "w") as z:
|
||||
z.writestr("mimetype", "application/epub+zip")
|
||||
z.writestr("META-INF/container.xml", CONTAINER)
|
||||
z.writestr("OEBPS/content.opf", OPF)
|
||||
z.writestr("OEBPS/ch1.xhtml", CH1)
|
||||
z.writestr("OEBPS/ch2.xhtml", CH2)
|
||||
z.writestr("OEBPS/ch3.xhtml", CH3)
|
||||
z.writestr("OEBPS/nav.xhtml", NAV)
|
||||
z.writestr("OEBPS/img/fig.png", PNG)
|
||||
z.writestr("OEBPS/img/cover.png", PNG)
|
||||
|
||||
|
||||
def test_convert(tmp: Path):
|
||||
src = tmp / "book.epub"
|
||||
build_epub(src)
|
||||
out = tmp / "book.html"
|
||||
epub2html.convert(src, out)
|
||||
text = out.read_text(encoding="utf-8")
|
||||
lines = text.splitlines()
|
||||
|
||||
# Листинг: переносы строк, отступы и минусы не тронуты
|
||||
pre = text[text.index("<pre>"):text.index("</pre>")]
|
||||
assert "def main():\n x = 1 - 2\n return x" in pre, pre
|
||||
assert "<code>" not in pre, "внутри листинга инлайновая разметка не нужна"
|
||||
|
||||
# Проза: инлайн переведён в наш набор тегов, сущности раскрыты
|
||||
para = next(l for l in lines if "fox" in l)
|
||||
assert "<i>lazy</i>" in para and "<code>fetch_data()</code>" in para, para
|
||||
assert " " not in para and " " not in para, para
|
||||
|
||||
assert '<p class="li">first item</p>' in lines
|
||||
assert '<p class="row">Name | Value</p>' in lines
|
||||
assert '<p class="row">alpha | 1</p>' in lines
|
||||
assert '<img src="images/img001.png"/>' in text
|
||||
assert (out.parent / "images" / "cover.png").exists(), "обложка не извлечена"
|
||||
|
||||
# Несбалансированная чужая разметка закрывается на границе блока
|
||||
broken = next(l for l in lines if "never closes" in l)
|
||||
assert broken.count("<b>") == broken.count("</b>") == 1, broken
|
||||
|
||||
assert "noise" not in text, "<script> не должен попадать в книгу"
|
||||
assert "skip me" not in text, "<title> документа — не текст книги"
|
||||
assert "Loose text outside any block tag." in text, "текст вне блока потерян"
|
||||
|
||||
# nav-документ выкинут, h2 подняты до h1 — иначе книга уедет одной главой
|
||||
assert text.count("Chapter One") == 1, "оглавление попало в текст"
|
||||
assert text.count("<h1>") == 3 and "<h2>" not in text
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
with tempfile.TemporaryDirectory() as d:
|
||||
test_convert(Path(d))
|
||||
print("OK")
|
||||
@@ -1,6 +1,7 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Самопроверка моста без сети: python3 test_translate.py"""
|
||||
import json
|
||||
import re
|
||||
from pathlib import Path
|
||||
|
||||
import translate as t
|
||||
@@ -17,6 +18,14 @@ def test_marks_roundtrip():
|
||||
assert "&" in body, "спецсимволы должны экранироваться обратно"
|
||||
|
||||
|
||||
def test_inline_code_survives():
|
||||
"""Имена функций в прозе не должны терять разметку по дороге к модели."""
|
||||
marked = t.to_marks("call <code>fetch()</code> twice")
|
||||
assert marked == "call ⟦code⟧fetch()⟦/code⟧ twice"
|
||||
body, ok = t.from_marks(marked)
|
||||
assert ok and "<code>fetch()</code>" in body
|
||||
|
||||
|
||||
def test_broken_marks_drop_tags():
|
||||
body, ok = t.from_marks("текст ⟦i⟧без закрытия")
|
||||
assert not ok and "⟦" not in body and "<i>" not in body
|
||||
@@ -71,6 +80,58 @@ def test_split_and_rebuild(tmp: Path):
|
||||
assert result.count("<h1>") == 2
|
||||
|
||||
|
||||
def test_code_listings_are_left_alone(tmp: Path):
|
||||
"""Листинги не переводятся и не участвуют в приёмке по языку: иначе книга,
|
||||
где кода много, честно переведётся и завалит проверку."""
|
||||
listing = "\n".join([
|
||||
"<pre>def compute_total(items, discount_rate, shipping_cost):",
|
||||
" subtotal = sum(item.price * item.quantity for item in items)",
|
||||
" return subtotal * (1 - discount_rate) + shipping_cost # not translated",
|
||||
"</pre>",
|
||||
])
|
||||
prose = ("<p>Текст главы про обработку данных, списки значений и способы "
|
||||
"хранения промежуточных результатов между запусками.</p>")
|
||||
src = tmp / "code.html"
|
||||
src.write_text("<h1>Глава</h1>\n" + (listing + "\n" + prose + "\n") * 3,
|
||||
encoding="utf-8")
|
||||
|
||||
blocks = t.parse_blocks(src)
|
||||
assert [b[1] for b in blocks] == ["h1", "p", "p", "p"], \
|
||||
"листинг не должен попасть в перевод"
|
||||
|
||||
# латиницы в файле больше, чем кириллицы, — но вся она в <pre>
|
||||
raw = re.sub("<[^>]+>", "", src.read_text(encoding="utf-8"))
|
||||
assert t.detect_language(raw)[0] == "en", "образец должен быть латинским целиком"
|
||||
assert t.verify_output(src, "ru"), "код не должен заваливать приёмку"
|
||||
|
||||
|
||||
def test_workdir_belongs_to_one_book(tmp: Path):
|
||||
"""Имена chapter_NNN.json у всех книг одинаковы: чужой workdir склеит
|
||||
перевод одной книги с текстом другой."""
|
||||
wd = tmp / "wd"
|
||||
wd.mkdir()
|
||||
a = [(0, "p", "", "First book text.")]
|
||||
b = [(0, "p", "", "Совершенно другая книга.")]
|
||||
t.claim_workdir(wd, Path("a.html"), a)
|
||||
t.claim_workdir(wd, Path("a.html"), a) # повторный запуск той же книги — ок
|
||||
try:
|
||||
t.claim_workdir(wd, Path("b.html"), b)
|
||||
except SystemExit as e:
|
||||
assert "занят другой книгой" in str(e)
|
||||
else:
|
||||
raise AssertionError("чужая книга в занятом workdir должна отвергаться")
|
||||
|
||||
|
||||
def test_stale_chapters_removed(tmp: Path):
|
||||
"""Прошлый прогон был длиннее — его главы иначе уедут в перевод."""
|
||||
extracted = tmp / "extracted"
|
||||
extracted.mkdir()
|
||||
(extracted / "chapter_009.json").write_text("{}", encoding="utf-8")
|
||||
t.write_input([[(0, "h1", "", "Глава")]], extracted)
|
||||
assert not (extracted / "chapter_009.json").exists()
|
||||
assert (extracted / "chapter_000.json").exists()
|
||||
|
||||
|
||||
def test_verify_output_catches_untranslated(tmp: Path):
|
||||
"""Регресс: первый боевой прогон отрапортовал успех на непереведённом файле."""
|
||||
bad = tmp / "bad.html"
|
||||
@@ -90,10 +151,12 @@ if __name__ == "__main__":
|
||||
import tempfile
|
||||
|
||||
test_marks_roundtrip()
|
||||
test_inline_code_survives()
|
||||
test_broken_marks_drop_tags()
|
||||
test_language_detection()
|
||||
for case in (test_split_and_rebuild, test_code_listings_are_left_alone,
|
||||
test_workdir_belongs_to_one_book, test_stale_chapters_removed,
|
||||
test_verify_output_catches_untranslated):
|
||||
with tempfile.TemporaryDirectory() as d:
|
||||
test_split_and_rebuild(Path(d))
|
||||
with tempfile.TemporaryDirectory() as d:
|
||||
test_verify_output_catches_untranslated(Path(d))
|
||||
case(Path(d))
|
||||
print("OK")
|
||||
|
||||
+35
-5
@@ -11,6 +11,7 @@
|
||||
поломке абзац сохраняется переведённым, но без внутренней разметки.
|
||||
"""
|
||||
import argparse
|
||||
import hashlib
|
||||
import html
|
||||
import json
|
||||
import re
|
||||
@@ -19,10 +20,15 @@ import sys
|
||||
from pathlib import Path
|
||||
|
||||
BLOCK = re.compile(r"^<(p|h1|h2)([^>]*)>(.*)</\1>$")
|
||||
TAG = re.compile(r"</?([ib])>")
|
||||
MARK = re.compile(r"⟦(/?)([ib])⟧")
|
||||
INLINE = ("i", "b", "code")
|
||||
TAG = re.compile(r"</?(%s)>" % "|".join(INLINE))
|
||||
MARK = re.compile(r"⟦(/?)(%s)⟧" % "|".join(INLINE))
|
||||
# картинки и пустые блоки не переводим
|
||||
SKIP = re.compile(r"<img\b")
|
||||
# Листинги кода не переводятся вовсе: BLOCK их не узнаёт, поэтому строки <pre>
|
||||
# доходят до результата нетронутыми. Здесь <pre> вырезается только из проверок
|
||||
# языка — иначе английский код утянул бы долю кириллицы ниже порога приёмки.
|
||||
PRE = re.compile(r"<pre\b.*?</pre>", re.S)
|
||||
|
||||
|
||||
def detect_language(text):
|
||||
@@ -44,7 +50,7 @@ def to_marks(s):
|
||||
def from_marks(s):
|
||||
"""Маркеры -> теги. Возвращает (html, ok): ok=False если разметка разъехалась."""
|
||||
escaped = html.escape(s, quote=False)
|
||||
depth = {"i": 0, "b": 0}
|
||||
depth = dict.fromkeys(INLINE, 0)
|
||||
ok = True
|
||||
for slash, tag in MARK.findall(escaped):
|
||||
depth[tag] += -1 if slash else 1
|
||||
@@ -82,6 +88,8 @@ def split_chapters(blocks):
|
||||
|
||||
def write_input(chapters, extracted):
|
||||
extracted.mkdir(parents=True, exist_ok=True)
|
||||
for stale in extracted.glob("chapter_*.json"):
|
||||
stale.unlink() # главы прошлого прогона иначе уедут в перевод
|
||||
index = []
|
||||
for n, ch in enumerate(chapters):
|
||||
paragraphs = [to_marks(b[3]) for b in ch]
|
||||
@@ -109,6 +117,25 @@ def write_input(chapters, extracted):
|
||||
return index
|
||||
|
||||
|
||||
def claim_workdir(workdir, source, blocks):
|
||||
"""Рабочий каталог принадлежит одной книге. Имена chapter_NNN.json у всех
|
||||
книг одинаковы, а внешний репозиторий пропускает главы, помеченные
|
||||
готовыми, — без этой привязки вторая книга молча соберётся из перевода
|
||||
первой."""
|
||||
digest = hashlib.sha256(
|
||||
"\n".join(b[3] for b in blocks).encode("utf-8")).hexdigest()[:16]
|
||||
claim = workdir / "bridge_source.json"
|
||||
now = {"source": source.name, "blocks": len(blocks), "sha256": digest}
|
||||
if claim.exists():
|
||||
was = json.loads(claim.read_text(encoding="utf-8"))
|
||||
if was.get("sha256") != digest:
|
||||
sys.exit("рабочий каталог занят другой книгой (%s, блоков %s). "
|
||||
"Возьми чистый --workdir." % (was.get("source"), was.get("blocks")))
|
||||
else:
|
||||
claim.write_text(json.dumps(now, ensure_ascii=False, indent=2),
|
||||
encoding="utf-8")
|
||||
|
||||
|
||||
def run_translator(repo, workdir, extracted, workers):
|
||||
script = repo / "03_translate_parallel.py"
|
||||
if not script.exists():
|
||||
@@ -150,7 +177,7 @@ def verify_output(out, target):
|
||||
"""Совпадение числа абзацев ещё не значит, что перевод состоялся: внешний
|
||||
скрипт при ошибке API молча подставляет оригинал. Проверяем язык результата.
|
||||
"""
|
||||
text = re.sub("<[^>]+>", "", out.read_text(encoding="utf-8"))
|
||||
text = re.sub("<[^>]+>", "", PRE.sub(" ", out.read_text(encoding="utf-8")))
|
||||
lang, share = detect_language(text)
|
||||
stub = text.count("[UNTRANSLATED]")
|
||||
print("результат: %s (кириллица %.0f%%), заглушек [UNTRANSLATED]: %d"
|
||||
@@ -186,7 +213,9 @@ def main():
|
||||
plain = " ".join(re.sub("<[^>]+>", "", b[3]) for b in blocks)
|
||||
lang, share = detect_language(plain)
|
||||
chapters = split_chapters(blocks)
|
||||
print("блоков: %d, глав: %d, знаков: %d" % (len(blocks), len(chapters), len(plain)))
|
||||
pre_n = len(PRE.findall(args.source.read_text(encoding="utf-8")))
|
||||
print("блоков: %d, глав: %d, знаков: %d, листингов <pre> (не переводятся): %d"
|
||||
% (len(blocks), len(chapters), len(plain), pre_n))
|
||||
print("язык источника: %s (кириллица %.0f%%)" % (lang, share * 100))
|
||||
|
||||
if lang == args.target and not args.force:
|
||||
@@ -199,6 +228,7 @@ def main():
|
||||
return
|
||||
|
||||
args.workdir.mkdir(parents=True, exist_ok=True)
|
||||
claim_workdir(args.workdir, args.source, blocks)
|
||||
extracted = (args.workdir / "extracted").resolve()
|
||||
write_input(chapters, extracted)
|
||||
run_translator(args.repo.resolve(), args.workdir.resolve(), extracted, args.workers)
|
||||
|
||||
Reference in New Issue
Block a user