EPUB как второй вход в конвейер; листинги кода не переводятся и не ломают приёмку
- epub2html.py: EPUB -> тот же XHTML, что pdf2html.py, только стандартная библиотека. Уровень глав определяется по книге, а не берётся из <h1>: в EPUB там обычно название и части. - pdf2html.py: листинг узнаётся по флагу моноширинного шрифта PyMuPDF, переносы и отступы внутри <pre> сохраняются, де-дефисация к коду не применяется. - translate.py: <pre> исключён из проверок языка (английский код утягивал долю кириллицы ниже порога приёмки), рабочий каталог привязан к книге, устаревшие главы прошлого прогона чистятся, инлайновый <code> переживает переводчика. - bookhtml.py: общий каркас документа и CSS для обоих входов. - Тесты: test_epub2html.py на собранном в памяти EPUB, четыре новых случая в test_translate.py.
This commit is contained in:
@@ -15,6 +15,30 @@ tag, and 2817 body paragraphs became `<h2>`.
|
|||||||
properties, and emits semantic XHTML. calibre is then used only as the
|
properties, and emits semantic XHTML. calibre is then used only as the
|
||||||
XHTML→EPUB packer.
|
XHTML→EPUB packer.
|
||||||
|
|
||||||
|
## Two entry points
|
||||||
|
|
||||||
|
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB. Both emit the
|
||||||
|
same normalized XHTML — one block per line, `<pre>` for code listings — and
|
||||||
|
everything downstream (translation, packing, audiobook) is identical. Shared
|
||||||
|
document skeleton and CSS live in `scripts/bookhtml.py`.
|
||||||
|
|
||||||
|
**Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws
|
||||||
|
away the semantic markup that is already there and re-derives it from font
|
||||||
|
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
|
||||||
|
the spine, drops the nav document, and maps the book's own headings.
|
||||||
|
|
||||||
|
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is
|
||||||
|
needed, not PyMuPDF) and convert with:
|
||||||
|
|
||||||
|
`python scripts/epub2html.py book.epub out/book.html`
|
||||||
|
|
||||||
|
It reports which heading level turned out to be the chapter level. In most
|
||||||
|
EPUBs `<h1>` is the book title and the parts, while chapters are `<h2>` — the
|
||||||
|
script picks the deepest level that still gives a sane chapter count, because
|
||||||
|
the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
|
||||||
|
3rd edition: 7 `<h1>` against 49 `<h2>`, chapters correctly detected as `h2`,
|
||||||
|
56 chapters out of 656 pages.
|
||||||
|
|
||||||
## Workflow
|
## Workflow
|
||||||
|
|
||||||
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
|
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
|
||||||
@@ -35,7 +59,9 @@ XHTML→EPUB packer.
|
|||||||
Read off: the body font (largest character count), the italic variant, the
|
Read off: the body font (largest character count), the italic variant, the
|
||||||
heading sizes, and any secondary family used for sidebars, journal entries,
|
heading sizes, and any secondary family used for sidebars, journal entries,
|
||||||
chat logs or slides.
|
chat logs or slides.
|
||||||
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. The
|
4. **Tune the thresholds** in `block_kind()` and `style()` to that output. Code
|
||||||
|
listings need no tuning — a monospaced span is recognized by the PyMuPDF font
|
||||||
|
flag (`flags & 8`), not by font name, and becomes a `<pre>` block. The
|
||||||
defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a
|
defaults target a 15pt/letter calibre layout: `>= 28` or a display font is a
|
||||||
title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary
|
title, `>= 24` a chapter/part `h1`, `>= 17` an `h2` subtitle, a secondary
|
||||||
family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that
|
family is `p.note`. `INDENT_X` (default 88) is the x-coordinate that
|
||||||
@@ -45,7 +71,10 @@ XHTML→EPUB packer.
|
|||||||
|
|
||||||
`python scripts/pdf2html.py book.pdf out/book.html`
|
`python scripts/pdf2html.py book.pdf out/book.html`
|
||||||
|
|
||||||
6. **Verify before packing** (see Verification). Fix thresholds and re-run until
|
6. **Verify before packing** (see Verification). Check the `pre:` counter against
|
||||||
|
the real number of listings in the book — a technical book reporting `pre: 0`
|
||||||
|
means the mono flag never fired and every listing is about to be reflowed as
|
||||||
|
prose. Fix thresholds and re-run until
|
||||||
the counts are sane. Cheap to iterate; do not skip to packing.
|
the counts are sane. Cheap to iterate; do not skip to packing.
|
||||||
7. **Pack with calibre**, always passing explicit TOC XPaths — without them
|
7. **Pack with calibre**, always passing explicit TOC XPaths — without them
|
||||||
calibre applies its own heuristics and re-breaks the chapters:
|
calibre applies its own heuristics and re-breaks the chapters:
|
||||||
@@ -97,6 +126,23 @@ navPoint count for the TOC size.
|
|||||||
Only when the user asks for a translated book. It slots between step 6 and
|
Only when the user asks for a translated book. It slots between step 6 and
|
||||||
step 7 — translate the XHTML, then pack the translated file with calibre.
|
step 7 — translate the XHTML, then pack the translated file with calibre.
|
||||||
|
|
||||||
|
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
|
||||||
|
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
|
||||||
|
the result untouched. Nothing extra is needed to protect code — but this also
|
||||||
|
means a listing that was misclassified as a paragraph upstream *will* be
|
||||||
|
translated, which is the real reason step 6 checks the `pre:` counter.
|
||||||
|
|
||||||
|
For the same reason `<pre>` is cut out of both language checks. A book that is
|
||||||
|
40% listings translates correctly and would otherwise fail acceptance, because
|
||||||
|
the English code drags the Cyrillic share below the threshold.
|
||||||
|
|
||||||
|
**One workdir per book.** Chapter files are named `chapter_NNN.json` for every
|
||||||
|
book, and the external repo skips chapters it has marked done — a reused workdir
|
||||||
|
silently stitches one book's translation onto another's text. The bridge writes
|
||||||
|
`bridge_source.json` into the workdir on first run and refuses to start if the
|
||||||
|
directory belongs to a different book. Re-running the same book is unaffected;
|
||||||
|
that is the resume path.
|
||||||
|
|
||||||
Translation is delegated to an external project,
|
Translation is delegated to an external project,
|
||||||
[`vetermanve/book_translator`](https://github.com/vetermanve/book_translator)
|
[`vetermanve/book_translator`](https://github.com/vetermanve/book_translator)
|
||||||
(DeepSeek or a local Ollama model). **Never modify that repository** — it is
|
(DeepSeek or a local Ollama model). **Never modify that repository** — it is
|
||||||
|
|||||||
@@ -0,0 +1,44 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Общая сборка XHTML для pdf2html.py и epub2html.py.
|
||||||
|
|
||||||
|
Формат намеренно жёсткий: один блок — одна строка (кроме <pre>, где переносы
|
||||||
|
значимы). На этом держится мост translate.py — он правит блоки по номеру
|
||||||
|
строки, а всё, чего не узнал, доносит до результата нетронутым.
|
||||||
|
"""
|
||||||
|
import html
|
||||||
|
|
||||||
|
CSS = """
|
||||||
|
body { margin: 0 1em; }
|
||||||
|
h1 { text-align: center; margin: 2em 0 0.2em; page-break-before: always; }
|
||||||
|
h1.title { page-break-before: avoid; }
|
||||||
|
h2 { text-align: center; font-style: italic; font-weight: normal;
|
||||||
|
font-size: 1.1em; margin: 0.2em 0 1.5em; }
|
||||||
|
p { text-indent: 1.2em; margin: 0; text-align: justify; }
|
||||||
|
p.note { text-indent: 0; margin: 1em 2em; font-size: 0.9em;
|
||||||
|
font-family: sans-serif; }
|
||||||
|
p.li { text-indent: 0; margin: 0.3em 0 0.3em 1.5em; text-align: left; }
|
||||||
|
p.row { text-indent: 0; margin: 0.3em 0; text-align: left;
|
||||||
|
font-size: 0.9em; }
|
||||||
|
p.figure { text-indent: 0; text-align: center; margin: 1em 0; }
|
||||||
|
img { max-width: 100%; }
|
||||||
|
pre { font-family: monospace; font-size: 0.75em; margin: 1em 0;
|
||||||
|
white-space: pre-wrap; word-wrap: break-word; text-align: left;
|
||||||
|
text-indent: 0; }
|
||||||
|
code { font-family: monospace; font-size: 0.9em; }
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
def document(title, parts):
|
||||||
|
"""parts: [(tag, cls, inner_html)]; tag 'figure' — псевдотег для картинки."""
|
||||||
|
buf = ['<?xml version="1.0" encoding="utf-8"?>',
|
||||||
|
'<html xmlns="http://www.w3.org/1999/xhtml"><head>',
|
||||||
|
"<title>%s</title>" % html.escape(title),
|
||||||
|
"<style>%s</style></head><body>" % CSS]
|
||||||
|
for tag, cls, txt in parts:
|
||||||
|
if tag == "figure":
|
||||||
|
buf.append('<p class="figure">%s</p>' % txt)
|
||||||
|
else:
|
||||||
|
c = ' class="%s"' % cls if cls else ""
|
||||||
|
buf.append("<%s%s>%s</%s>" % (tag, c, txt, tag))
|
||||||
|
buf.append("</body></html>")
|
||||||
|
return "\n".join(buf)
|
||||||
@@ -0,0 +1,267 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""EPUB -> тот же XHTML, что выдаёт pdf2html.py (один блок = одна строка).
|
||||||
|
|
||||||
|
EPUB уже размечен семантически, поэтому здесь не разведка шрифтов, а
|
||||||
|
нормализация: привести чужую вёрстку к формату, который понимает мост
|
||||||
|
translate.py, и вычистить всё, что переводчику видеть не нужно.
|
||||||
|
|
||||||
|
Листинги остаются в <pre> и не переводятся вовсе: BLOCK-регулярка моста их не
|
||||||
|
узнаёт, а всё неузнанное доходит до результата нетронутым.
|
||||||
|
|
||||||
|
Только стандартная библиотека — ни calibre, ни lxml здесь не нужны.
|
||||||
|
"""
|
||||||
|
import argparse
|
||||||
|
import collections
|
||||||
|
import html
|
||||||
|
import posixpath
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import zipfile
|
||||||
|
from html.parser import HTMLParser
|
||||||
|
from pathlib import Path
|
||||||
|
from xml.etree import ElementTree
|
||||||
|
|
||||||
|
import bookhtml
|
||||||
|
|
||||||
|
CONTAINER = "META-INF/container.xml"
|
||||||
|
OPF_NS = "{http://www.idpf.org/2007/opf}"
|
||||||
|
CONT_NS = "{urn:oasis:names:tc:opendocument:xmlns:container}"
|
||||||
|
|
||||||
|
# tag источника -> (tag результата, css-класс)
|
||||||
|
BLOCK_MAP = {
|
||||||
|
"h1": ("h1", None), "h2": ("h2", None), "h3": ("h3", None),
|
||||||
|
"h4": ("h4", None), "h5": ("h5", None), "h6": ("h6", None),
|
||||||
|
"p": ("p", None), "pre": ("pre", None), "li": ("p", "li"),
|
||||||
|
"tr": ("p", "row"), "figcaption": ("p", "note"),
|
||||||
|
"dt": ("p", "note"), "dd": ("p", "note"),
|
||||||
|
}
|
||||||
|
INLINE_MAP = {"i": "i", "em": "i", "cite": "i", "b": "b", "strong": "b",
|
||||||
|
"code": "code", "kbd": "code", "samp": "code", "tt": "code",
|
||||||
|
"var": "code"}
|
||||||
|
DROP = {"script", "style", "head", "title", "nav"}
|
||||||
|
IMG_EXT = {"image/jpeg": ".jpg", "image/png": ".png", "image/gif": ".gif",
|
||||||
|
"image/svg+xml": ".svg", "image/webp": ".webp"}
|
||||||
|
|
||||||
|
|
||||||
|
class Reader(HTMLParser):
|
||||||
|
"""Собирает блоки из одного XHTML-документа книги."""
|
||||||
|
|
||||||
|
def __init__(self, on_image):
|
||||||
|
super().__init__(convert_charrefs=True)
|
||||||
|
self.on_image = on_image
|
||||||
|
self.blocks = []
|
||||||
|
self.cur = None # (tag, cls, [куски])
|
||||||
|
self.open_inline = [] # незакрытые инлайновые теги текущего блока
|
||||||
|
self.drop = 0
|
||||||
|
|
||||||
|
# -- служебное ---------------------------------------------------------
|
||||||
|
def flush(self):
|
||||||
|
if not self.cur:
|
||||||
|
return
|
||||||
|
tag, cls, parts = self.cur
|
||||||
|
self.cur = None
|
||||||
|
for t in reversed(self.open_inline):
|
||||||
|
parts.append("</%s>" % t) # чужая вёрстка бывает несбалансированной
|
||||||
|
self.open_inline = []
|
||||||
|
txt = "".join(parts)
|
||||||
|
if tag == "pre":
|
||||||
|
txt = txt.strip("\n").rstrip()
|
||||||
|
else:
|
||||||
|
txt = re.sub(r"\s+", " ", txt).strip()
|
||||||
|
txt = re.sub(r"<(i|b|code)>(\s*)</\1>", r"\2", txt) # пустая разметка
|
||||||
|
if txt.strip():
|
||||||
|
self.blocks.append((tag, cls, txt))
|
||||||
|
|
||||||
|
def start(self, tag, cls):
|
||||||
|
self.flush()
|
||||||
|
self.cur = (tag, cls, [])
|
||||||
|
|
||||||
|
# -- HTMLParser --------------------------------------------------------
|
||||||
|
def handle_starttag(self, tag, attrs):
|
||||||
|
if tag in DROP:
|
||||||
|
self.drop += 1
|
||||||
|
return
|
||||||
|
if self.drop:
|
||||||
|
return
|
||||||
|
a = dict(attrs)
|
||||||
|
if tag in ("img", "image"):
|
||||||
|
src = a.get("src") or a.get("{http://www.w3.org/1999/xlink}href") \
|
||||||
|
or a.get("xlink:href") or a.get("href")
|
||||||
|
name = self.on_image(src) if src else None
|
||||||
|
if name:
|
||||||
|
self.flush()
|
||||||
|
self.blocks.append(("figure", None,
|
||||||
|
'<img src="images/%s"/>' % name))
|
||||||
|
return
|
||||||
|
if tag == "br":
|
||||||
|
if self.cur:
|
||||||
|
self.cur[2].append("\n" if self.cur[0] == "pre" else " ")
|
||||||
|
return
|
||||||
|
if tag in BLOCK_MAP:
|
||||||
|
self.start(*BLOCK_MAP[tag])
|
||||||
|
return
|
||||||
|
if tag in ("td", "th") and self.cur and self.cur[0] == "p" \
|
||||||
|
and self.cur[1] == "row" and self.cur[2]:
|
||||||
|
self.cur[2].append(" | ")
|
||||||
|
return
|
||||||
|
if tag in INLINE_MAP and self.cur and self.cur[0] != "pre":
|
||||||
|
# Внутри листинга инлайн не нужен: книги размечают код сплошным
|
||||||
|
# <strong>, и на e-ink это страница жирного текста.
|
||||||
|
out = INLINE_MAP[tag]
|
||||||
|
self.open_inline.append(out)
|
||||||
|
self.cur[2].append("<%s>" % out)
|
||||||
|
|
||||||
|
def handle_endtag(self, tag):
|
||||||
|
if tag in DROP:
|
||||||
|
self.drop = max(0, self.drop - 1)
|
||||||
|
return
|
||||||
|
if self.drop:
|
||||||
|
return
|
||||||
|
if tag in BLOCK_MAP:
|
||||||
|
self.flush()
|
||||||
|
return
|
||||||
|
if tag in INLINE_MAP and self.cur and self.cur[0] != "pre":
|
||||||
|
out = INLINE_MAP[tag]
|
||||||
|
if out in self.open_inline:
|
||||||
|
self.open_inline.remove(out)
|
||||||
|
self.cur[2].append("</%s>" % out)
|
||||||
|
|
||||||
|
def handle_data(self, data):
|
||||||
|
if self.drop:
|
||||||
|
return
|
||||||
|
if not self.cur:
|
||||||
|
if not data.strip():
|
||||||
|
return
|
||||||
|
self.start("p", None) # текст вне блочного тега не теряем
|
||||||
|
self.cur[2].append(html.escape(data, quote=False))
|
||||||
|
|
||||||
|
def close(self):
|
||||||
|
super().close()
|
||||||
|
self.flush()
|
||||||
|
|
||||||
|
|
||||||
|
def spine_documents(zf):
|
||||||
|
"""[(путь в архиве, media-type)] в порядке чтения + путь к OPF."""
|
||||||
|
root = ElementTree.fromstring(zf.read(CONTAINER))
|
||||||
|
opf_path = root.find(".//%srootfile" % CONT_NS).get("full-path")
|
||||||
|
opf = ElementTree.fromstring(zf.read(opf_path))
|
||||||
|
base = posixpath.dirname(opf_path)
|
||||||
|
items, cover = {}, None
|
||||||
|
for it in opf.iter("%sitem" % OPF_NS):
|
||||||
|
href = posixpath.normpath(posixpath.join(base, it.get("href")))
|
||||||
|
props = it.get("properties") or ""
|
||||||
|
items[it.get("id")] = (href, it.get("media-type"), props)
|
||||||
|
if "cover-image" in props:
|
||||||
|
cover = href
|
||||||
|
if cover is None: # EPUB 2: обложка объявляется через <meta name="cover">
|
||||||
|
meta = opf.find(".//%smeta[@name='cover']" % OPF_NS)
|
||||||
|
if meta is not None and meta.get("content") in items:
|
||||||
|
cover = items[meta.get("content")][0]
|
||||||
|
docs = []
|
||||||
|
for ref in opf.iter("%sitemref" % OPF_NS):
|
||||||
|
item = items.get(ref.get("idref"))
|
||||||
|
if not item:
|
||||||
|
continue
|
||||||
|
href, mtype, props = item
|
||||||
|
if "nav" in props or re.search(r"(toc|nav|content)s?\.x?html?$", href, re.I):
|
||||||
|
continue
|
||||||
|
docs.append(href)
|
||||||
|
return docs, items, cover
|
||||||
|
|
||||||
|
|
||||||
|
# Больше этого числа глав дробить незачем: переводчик держит контекст в пределах
|
||||||
|
# главы, слишком мелкая нарезка его обесценивает.
|
||||||
|
MAX_CHAPTERS = 150
|
||||||
|
|
||||||
|
|
||||||
|
def chapter_level(parts):
|
||||||
|
"""Каким уровнем заголовка в этой книге размечены главы.
|
||||||
|
|
||||||
|
Мост режет книгу на главы по <h1>, а в EPUB <h1> обычно занят названием
|
||||||
|
книги и частями: у Страуструпа их 7 на 656 страниц, тогда как главы — это
|
||||||
|
<h2> (49 штук). Берём самый глубокий уровень, при котором число глав ещё
|
||||||
|
остаётся вменяемым.
|
||||||
|
"""
|
||||||
|
counts = collections.Counter(t for t, _, _ in parts if re.fullmatch(r"h[1-6]", t))
|
||||||
|
best, total = 1, 0
|
||||||
|
for lvl in range(1, 7):
|
||||||
|
total += counts.get("h%d" % lvl, 0)
|
||||||
|
if total > MAX_CHAPTERS:
|
||||||
|
break
|
||||||
|
if total:
|
||||||
|
best = lvl
|
||||||
|
return best
|
||||||
|
|
||||||
|
|
||||||
|
def convert(src, out):
|
||||||
|
imgdir = out.parent / "images"
|
||||||
|
imgdir.mkdir(parents=True, exist_ok=True)
|
||||||
|
parts, seen, counter = [], {}, [0]
|
||||||
|
|
||||||
|
with zipfile.ZipFile(src) as zf:
|
||||||
|
names = set(zf.namelist())
|
||||||
|
docs, items, cover = spine_documents(zf)
|
||||||
|
|
||||||
|
def extract(path, stem=None):
|
||||||
|
if path in seen:
|
||||||
|
return seen[path]
|
||||||
|
if path not in names:
|
||||||
|
return None
|
||||||
|
ext = posixpath.splitext(path)[1] or ".img"
|
||||||
|
if stem:
|
||||||
|
name = stem + ext
|
||||||
|
else:
|
||||||
|
counter[0] += 1
|
||||||
|
name = "img%03d%s" % (counter[0], ext)
|
||||||
|
(imgdir / name).write_bytes(zf.read(path))
|
||||||
|
seen[path] = name
|
||||||
|
return name
|
||||||
|
|
||||||
|
if cover:
|
||||||
|
extract(cover, stem="cover")
|
||||||
|
|
||||||
|
for doc in docs:
|
||||||
|
if doc not in names:
|
||||||
|
print("нет в архиве, пропущен: %s" % doc, file=sys.stderr)
|
||||||
|
continue
|
||||||
|
here = posixpath.dirname(doc)
|
||||||
|
|
||||||
|
def on_image(src_attr, here=here):
|
||||||
|
target = posixpath.normpath(posixpath.join(here, src_attr.split("#")[0]))
|
||||||
|
return extract(target)
|
||||||
|
|
||||||
|
r = Reader(on_image)
|
||||||
|
r.feed(zf.read(doc).decode("utf-8", "replace"))
|
||||||
|
r.close()
|
||||||
|
parts.extend(r.blocks)
|
||||||
|
|
||||||
|
title = out.stem
|
||||||
|
lvl = chapter_level(parts)
|
||||||
|
parts = [(("h1" if int(t[1]) <= lvl else "h2", c, x)
|
||||||
|
if re.fullmatch(r"h[1-6]", t) else (t, c, x))
|
||||||
|
for t, c, x in parts]
|
||||||
|
h1 = sum(1 for p in parts if p[0] == "h1")
|
||||||
|
print("главы размечены h%d, глав получилось: %d" % (lvl, h1))
|
||||||
|
if h1 < 3:
|
||||||
|
print("ВНИМАНИЕ: заголовков-глав всего %d — переводчик будет держать "
|
||||||
|
"контекст крупными кусками, проверь результат внимательнее" % h1)
|
||||||
|
|
||||||
|
out.write_text(bookhtml.document(title, parts), encoding="utf-8")
|
||||||
|
print("blocks: %d, images: %d" % (len(parts), len(seen)))
|
||||||
|
print("h1: %d p: %d pre: %d"
|
||||||
|
% (h1, sum(1 for p in parts if p[0] == "p"),
|
||||||
|
sum(1 for p in parts if p[0] == "pre")))
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
ap = argparse.ArgumentParser(description=__doc__)
|
||||||
|
ap.add_argument("source", type=Path, help="исходный .epub")
|
||||||
|
ap.add_argument("output", type=Path, help="куда писать XHTML")
|
||||||
|
args = ap.parse_args()
|
||||||
|
if not args.source.exists():
|
||||||
|
sys.exit("нет файла %s" % args.source)
|
||||||
|
convert(args.source, args.output)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
+34
-32
@@ -7,6 +7,11 @@ from pathlib import Path
|
|||||||
|
|
||||||
import pymupdf
|
import pymupdf
|
||||||
|
|
||||||
|
import bookhtml
|
||||||
|
|
||||||
|
if len(sys.argv) < 3 or sys.argv[1] in ("-h", "--help"):
|
||||||
|
sys.exit("usage: pdf2html.py book.pdf out.html | pdf2html.py --fonts book.pdf")
|
||||||
|
|
||||||
if sys.argv[1] == "--fonts": # разведка: какие шрифты/кегли в PDF
|
if sys.argv[1] == "--fonts": # разведка: какие шрифты/кегли в PDF
|
||||||
import collections
|
import collections
|
||||||
|
|
||||||
@@ -29,6 +34,7 @@ IMGDIR.mkdir(parents=True, exist_ok=True)
|
|||||||
# Пороги подобраны под вёрстку 15pt/letter. Для другой книги сначала
|
# Пороги подобраны под вёрстку 15pt/letter. Для другой книги сначала
|
||||||
# посмотреть реальные шрифты и кегли: pdf2html.py --fonts file.pdf
|
# посмотреть реальные шрифты и кегли: pdf2html.py --fonts file.pdf
|
||||||
INDENT_X = 88 # x0 первой строки: больше — абзац с красной строки
|
INDENT_X = 88 # x0 первой строки: больше — абзац с красной строки
|
||||||
|
MONO = 8 # бит моноширинного шрифта в span["flags"] (PyMuPDF)
|
||||||
|
|
||||||
|
|
||||||
def style(font):
|
def style(font):
|
||||||
@@ -36,10 +42,20 @@ def style(font):
|
|||||||
return ("bold" in f or "semibold" in f, "-it" in f or "italic" in f)
|
return ("bold" in f or "semibold" in f, "-it" in f or "italic" in f)
|
||||||
|
|
||||||
|
|
||||||
def span_html(s):
|
def is_mono(s):
|
||||||
|
"""Моноширинный шрифт = листинг кода. Признак берётся из флагов PyMuPDF, а
|
||||||
|
не из имени шрифта: имена у каждого издательства свои, флаг одинаковый."""
|
||||||
|
return bool(s["flags"] & MONO)
|
||||||
|
|
||||||
|
|
||||||
|
def span_html(s, plain=False):
|
||||||
t = html.escape(s["text"])
|
t = html.escape(s["text"])
|
||||||
if not t:
|
if not t:
|
||||||
return ""
|
return ""
|
||||||
|
if plain: # внутри <pre> курсив и полужирный только мешают
|
||||||
|
return t
|
||||||
|
if is_mono(s):
|
||||||
|
return "<code>%s</code>" % t
|
||||||
bold, ital = style(s["font"])
|
bold, ital = style(s["font"])
|
||||||
if ital:
|
if ital:
|
||||||
t = "<i>%s</i>" % t
|
t = "<i>%s</i>" % t
|
||||||
@@ -59,6 +75,8 @@ def block_kind(b):
|
|||||||
return ("h1", "title")
|
return ("h1", "title")
|
||||||
if sz >= 24:
|
if sz >= 24:
|
||||||
return ("h1", None)
|
return ("h1", None)
|
||||||
|
if all(is_mono(sp) for sp in spans):
|
||||||
|
return ("pre", None)
|
||||||
if sz >= 17:
|
if sz >= 17:
|
||||||
return ("h2", None)
|
return ("h2", None)
|
||||||
if f.startswith("MyriadPro"):
|
if f.startswith("MyriadPro"):
|
||||||
@@ -66,7 +84,13 @@ def block_kind(b):
|
|||||||
return ("p", None)
|
return ("p", None)
|
||||||
|
|
||||||
|
|
||||||
def block_text(b):
|
def block_text(b, pre=False):
|
||||||
|
if pre:
|
||||||
|
# Перенос строки в коде значим, висящий дефис — это минус, а не перенос
|
||||||
|
# слова. Ни склейки строк, ни де-дефисации здесь быть не должно.
|
||||||
|
rows = ["".join(span_html(s, plain=True) for s in l["spans"])
|
||||||
|
for l in b["lines"]]
|
||||||
|
return "\n".join(rows).rstrip()
|
||||||
out = []
|
out = []
|
||||||
for i, l in enumerate(b["lines"]):
|
for i, l in enumerate(b["lines"]):
|
||||||
line = "".join(span_html(s) for s in l["spans"])
|
line = "".join(span_html(s) for s in l["spans"])
|
||||||
@@ -82,6 +106,7 @@ def block_text(b):
|
|||||||
txt = "".join(out)
|
txt = "".join(out)
|
||||||
txt = re.sub(r"</i>(\s*)<i>", r"\1", txt)
|
txt = re.sub(r"</i>(\s*)<i>", r"\1", txt)
|
||||||
txt = re.sub(r"</b>(\s*)<b>", r"\1", txt)
|
txt = re.sub(r"</b>(\s*)<b>", r"\1", txt)
|
||||||
|
txt = re.sub(r"(?<=\S) {2,}(?=\S)", " ", txt)
|
||||||
return txt.strip()
|
return txt.strip()
|
||||||
|
|
||||||
|
|
||||||
@@ -112,7 +137,7 @@ for pno, page in enumerate(doc):
|
|||||||
if not kind:
|
if not kind:
|
||||||
continue
|
continue
|
||||||
tag, cls = kind
|
tag, cls = kind
|
||||||
txt = block_text(b)
|
txt = block_text(b, pre=(tag == "pre"))
|
||||||
if not txt:
|
if not txt:
|
||||||
continue
|
continue
|
||||||
indented = b["lines"][0]["bbox"][0] >= INDENT_X
|
indented = b["lines"][0]["bbox"][0] >= INDENT_X
|
||||||
@@ -134,36 +159,13 @@ for pno, page in enumerate(doc):
|
|||||||
if open_para:
|
if open_para:
|
||||||
parts.append(open_para)
|
parts.append(open_para)
|
||||||
|
|
||||||
CSS = """
|
doc_html = bookhtml.document(doc.metadata.get("title") or SRC.stem, parts)
|
||||||
body { margin: 0 1em; }
|
# merge tag runs split at page/line boundaries. Пробелы здесь уже не трогаем:
|
||||||
h1 { text-align: center; margin: 2em 0 0.2em; page-break-before: always; }
|
# внутри <pre> они значимы, а в прозе схлопнуты в block_text.
|
||||||
h1.title { page-break-before: avoid; }
|
doc_html = re.sub(r"</(i|b|code)>(\s*)<\1>", r"\2", doc_html)
|
||||||
h2 { text-align: center; font-style: italic; font-weight: normal;
|
|
||||||
font-size: 1.1em; margin: 0.2em 0 1.5em; }
|
|
||||||
p { text-indent: 1.2em; margin: 0; text-align: justify; }
|
|
||||||
p.note { text-indent: 0; margin: 1em 2em; font-size: 0.9em;
|
|
||||||
font-family: sans-serif; }
|
|
||||||
p.figure { text-indent: 0; text-align: center; margin: 1em 0; }
|
|
||||||
img { max-width: 100%; }
|
|
||||||
"""
|
|
||||||
|
|
||||||
buf = ['<?xml version="1.0" encoding="utf-8"?>',
|
|
||||||
'<html xmlns="http://www.w3.org/1999/xhtml"><head>',
|
|
||||||
"<title>%s</title>" % html.escape(doc.metadata.get("title") or SRC.stem),
|
|
||||||
"<style>%s</style></head><body>" % CSS]
|
|
||||||
for tag, cls, txt in parts:
|
|
||||||
if tag == "figure":
|
|
||||||
buf.append('<p class="figure">%s</p>' % txt)
|
|
||||||
else:
|
|
||||||
c = ' class="%s"' % cls if cls else ""
|
|
||||||
buf.append("<%s%s>%s</%s>" % (tag, c, txt, tag))
|
|
||||||
buf.append("</body></html>")
|
|
||||||
doc_html = "\n".join(buf)
|
|
||||||
# merge tag runs split at page/line boundaries, collapse doubled spaces
|
|
||||||
doc_html = re.sub(r"</(i|b)>(\s*)<\1>", r"\2", doc_html)
|
|
||||||
doc_html = re.sub(r"(?<=\S) {2,}(?=\S)", " ", doc_html)
|
|
||||||
OUT.write_text(doc_html, encoding="utf-8")
|
OUT.write_text(doc_html, encoding="utf-8")
|
||||||
|
|
||||||
print("blocks:", len(parts), "images:", img_n)
|
print("blocks:", len(parts), "images:", img_n)
|
||||||
print("h1:", sum(1 for p in parts if p[0] == "h1"),
|
print("h1:", sum(1 for p in parts if p[0] == "h1"),
|
||||||
"p:", sum(1 for p in parts if p[0] == "p"))
|
"p:", sum(1 for p in parts if p[0] == "p"),
|
||||||
|
"pre:", sum(1 for p in parts if p[0] == "pre"))
|
||||||
|
|||||||
@@ -0,0 +1,119 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Самопроверка конвертера EPUB без сети: python3 test_epub2html.py"""
|
||||||
|
import tempfile
|
||||||
|
import zipfile
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import epub2html
|
||||||
|
|
||||||
|
CONTAINER = """<?xml version="1.0"?>
|
||||||
|
<container version="1.0" xmlns="urn:oasis:names:tc:opendocument:xmlns:container">
|
||||||
|
<rootfiles><rootfile full-path="OEBPS/content.opf"
|
||||||
|
media-type="application/oebps-package+xml"/></rootfiles>
|
||||||
|
</container>"""
|
||||||
|
|
||||||
|
OPF = """<?xml version="1.0"?>
|
||||||
|
<package xmlns="http://www.idpf.org/2007/opf" version="3.0">
|
||||||
|
<metadata/>
|
||||||
|
<manifest>
|
||||||
|
<item id="c1" href="ch1.xhtml" media-type="application/xhtml+xml"/>
|
||||||
|
<item id="c2" href="ch2.xhtml" media-type="application/xhtml+xml"/>
|
||||||
|
<item id="c3" href="ch3.xhtml" media-type="application/xhtml+xml"/>
|
||||||
|
<item id="nav" href="nav.xhtml" media-type="application/xhtml+xml"
|
||||||
|
properties="nav"/>
|
||||||
|
<item id="pic" href="img/fig.png" media-type="image/png"/>
|
||||||
|
<item id="cov" href="img/cover.png" media-type="image/png"
|
||||||
|
properties="cover-image"/>
|
||||||
|
</manifest>
|
||||||
|
<spine>
|
||||||
|
<itemref idref="nav"/><itemref idref="c1"/>
|
||||||
|
<itemref idref="c2"/><itemref idref="c3"/>
|
||||||
|
</spine>
|
||||||
|
</package>"""
|
||||||
|
|
||||||
|
CH1 = """<?xml version="1.0" encoding="utf-8"?>
|
||||||
|
<html xmlns="http://www.w3.org/1999/xhtml"><head><title>skip me</title></head>
|
||||||
|
<body>
|
||||||
|
<h2>Chapter One</h2>
|
||||||
|
<p>The <em>lazy</em> fox calls <code>fetch_data()</code> twice a day.</p>
|
||||||
|
<pre>def main():
|
||||||
|
x = 1 - 2
|
||||||
|
return x</pre>
|
||||||
|
<ul><li>first item</li><li>second item</li></ul>
|
||||||
|
<table><tr><th>Name</th><th>Value</th></tr><tr><td>alpha</td><td>1</td></tr></table>
|
||||||
|
<p>Broken <b>markup that never closes.</p>
|
||||||
|
<p><img src="img/fig.png" alt="figure"/></p>
|
||||||
|
<script>var noise = 1;</script>
|
||||||
|
</body></html>"""
|
||||||
|
|
||||||
|
CH2 = """<?xml version="1.0" encoding="utf-8"?>
|
||||||
|
<html xmlns="http://www.w3.org/1999/xhtml"><body>
|
||||||
|
<h2>Chapter Two</h2><p>Second chapter body text.</p>
|
||||||
|
</body></html>"""
|
||||||
|
|
||||||
|
CH3 = """<?xml version="1.0" encoding="utf-8"?>
|
||||||
|
<html xmlns="http://www.w3.org/1999/xhtml"><body>
|
||||||
|
<h2>Chapter Three</h2><div>Loose text outside any block tag.</div>
|
||||||
|
</body></html>"""
|
||||||
|
|
||||||
|
NAV = """<?xml version="1.0" encoding="utf-8"?>
|
||||||
|
<html xmlns="http://www.w3.org/1999/xhtml"><body>
|
||||||
|
<nav><ol><li>Chapter One</li></ol></nav></body></html>"""
|
||||||
|
|
||||||
|
PNG = bytes.fromhex("89504e470d0a1a0a") # достаточно как содержимое файла
|
||||||
|
|
||||||
|
|
||||||
|
def build_epub(path):
|
||||||
|
with zipfile.ZipFile(path, "w") as z:
|
||||||
|
z.writestr("mimetype", "application/epub+zip")
|
||||||
|
z.writestr("META-INF/container.xml", CONTAINER)
|
||||||
|
z.writestr("OEBPS/content.opf", OPF)
|
||||||
|
z.writestr("OEBPS/ch1.xhtml", CH1)
|
||||||
|
z.writestr("OEBPS/ch2.xhtml", CH2)
|
||||||
|
z.writestr("OEBPS/ch3.xhtml", CH3)
|
||||||
|
z.writestr("OEBPS/nav.xhtml", NAV)
|
||||||
|
z.writestr("OEBPS/img/fig.png", PNG)
|
||||||
|
z.writestr("OEBPS/img/cover.png", PNG)
|
||||||
|
|
||||||
|
|
||||||
|
def test_convert(tmp: Path):
|
||||||
|
src = tmp / "book.epub"
|
||||||
|
build_epub(src)
|
||||||
|
out = tmp / "book.html"
|
||||||
|
epub2html.convert(src, out)
|
||||||
|
text = out.read_text(encoding="utf-8")
|
||||||
|
lines = text.splitlines()
|
||||||
|
|
||||||
|
# Листинг: переносы строк, отступы и минусы не тронуты
|
||||||
|
pre = text[text.index("<pre>"):text.index("</pre>")]
|
||||||
|
assert "def main():\n x = 1 - 2\n return x" in pre, pre
|
||||||
|
assert "<code>" not in pre, "внутри листинга инлайновая разметка не нужна"
|
||||||
|
|
||||||
|
# Проза: инлайн переведён в наш набор тегов, сущности раскрыты
|
||||||
|
para = next(l for l in lines if "fox" in l)
|
||||||
|
assert "<i>lazy</i>" in para and "<code>fetch_data()</code>" in para, para
|
||||||
|
assert " " not in para and " " not in para, para
|
||||||
|
|
||||||
|
assert '<p class="li">first item</p>' in lines
|
||||||
|
assert '<p class="row">Name | Value</p>' in lines
|
||||||
|
assert '<p class="row">alpha | 1</p>' in lines
|
||||||
|
assert '<img src="images/img001.png"/>' in text
|
||||||
|
assert (out.parent / "images" / "cover.png").exists(), "обложка не извлечена"
|
||||||
|
|
||||||
|
# Несбалансированная чужая разметка закрывается на границе блока
|
||||||
|
broken = next(l for l in lines if "never closes" in l)
|
||||||
|
assert broken.count("<b>") == broken.count("</b>") == 1, broken
|
||||||
|
|
||||||
|
assert "noise" not in text, "<script> не должен попадать в книгу"
|
||||||
|
assert "skip me" not in text, "<title> документа — не текст книги"
|
||||||
|
assert "Loose text outside any block tag." in text, "текст вне блока потерян"
|
||||||
|
|
||||||
|
# nav-документ выкинут, h2 подняты до h1 — иначе книга уедет одной главой
|
||||||
|
assert text.count("Chapter One") == 1, "оглавление попало в текст"
|
||||||
|
assert text.count("<h1>") == 3 and "<h2>" not in text
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
with tempfile.TemporaryDirectory() as d:
|
||||||
|
test_convert(Path(d))
|
||||||
|
print("OK")
|
||||||
@@ -1,6 +1,7 @@
|
|||||||
#!/usr/bin/env python3
|
#!/usr/bin/env python3
|
||||||
"""Самопроверка моста без сети: python3 test_translate.py"""
|
"""Самопроверка моста без сети: python3 test_translate.py"""
|
||||||
import json
|
import json
|
||||||
|
import re
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
import translate as t
|
import translate as t
|
||||||
@@ -17,6 +18,14 @@ def test_marks_roundtrip():
|
|||||||
assert "&" in body, "спецсимволы должны экранироваться обратно"
|
assert "&" in body, "спецсимволы должны экранироваться обратно"
|
||||||
|
|
||||||
|
|
||||||
|
def test_inline_code_survives():
|
||||||
|
"""Имена функций в прозе не должны терять разметку по дороге к модели."""
|
||||||
|
marked = t.to_marks("call <code>fetch()</code> twice")
|
||||||
|
assert marked == "call ⟦code⟧fetch()⟦/code⟧ twice"
|
||||||
|
body, ok = t.from_marks(marked)
|
||||||
|
assert ok and "<code>fetch()</code>" in body
|
||||||
|
|
||||||
|
|
||||||
def test_broken_marks_drop_tags():
|
def test_broken_marks_drop_tags():
|
||||||
body, ok = t.from_marks("текст ⟦i⟧без закрытия")
|
body, ok = t.from_marks("текст ⟦i⟧без закрытия")
|
||||||
assert not ok and "⟦" not in body and "<i>" not in body
|
assert not ok and "⟦" not in body and "<i>" not in body
|
||||||
@@ -71,6 +80,58 @@ def test_split_and_rebuild(tmp: Path):
|
|||||||
assert result.count("<h1>") == 2
|
assert result.count("<h1>") == 2
|
||||||
|
|
||||||
|
|
||||||
|
def test_code_listings_are_left_alone(tmp: Path):
|
||||||
|
"""Листинги не переводятся и не участвуют в приёмке по языку: иначе книга,
|
||||||
|
где кода много, честно переведётся и завалит проверку."""
|
||||||
|
listing = "\n".join([
|
||||||
|
"<pre>def compute_total(items, discount_rate, shipping_cost):",
|
||||||
|
" subtotal = sum(item.price * item.quantity for item in items)",
|
||||||
|
" return subtotal * (1 - discount_rate) + shipping_cost # not translated",
|
||||||
|
"</pre>",
|
||||||
|
])
|
||||||
|
prose = ("<p>Текст главы про обработку данных, списки значений и способы "
|
||||||
|
"хранения промежуточных результатов между запусками.</p>")
|
||||||
|
src = tmp / "code.html"
|
||||||
|
src.write_text("<h1>Глава</h1>\n" + (listing + "\n" + prose + "\n") * 3,
|
||||||
|
encoding="utf-8")
|
||||||
|
|
||||||
|
blocks = t.parse_blocks(src)
|
||||||
|
assert [b[1] for b in blocks] == ["h1", "p", "p", "p"], \
|
||||||
|
"листинг не должен попасть в перевод"
|
||||||
|
|
||||||
|
# латиницы в файле больше, чем кириллицы, — но вся она в <pre>
|
||||||
|
raw = re.sub("<[^>]+>", "", src.read_text(encoding="utf-8"))
|
||||||
|
assert t.detect_language(raw)[0] == "en", "образец должен быть латинским целиком"
|
||||||
|
assert t.verify_output(src, "ru"), "код не должен заваливать приёмку"
|
||||||
|
|
||||||
|
|
||||||
|
def test_workdir_belongs_to_one_book(tmp: Path):
|
||||||
|
"""Имена chapter_NNN.json у всех книг одинаковы: чужой workdir склеит
|
||||||
|
перевод одной книги с текстом другой."""
|
||||||
|
wd = tmp / "wd"
|
||||||
|
wd.mkdir()
|
||||||
|
a = [(0, "p", "", "First book text.")]
|
||||||
|
b = [(0, "p", "", "Совершенно другая книга.")]
|
||||||
|
t.claim_workdir(wd, Path("a.html"), a)
|
||||||
|
t.claim_workdir(wd, Path("a.html"), a) # повторный запуск той же книги — ок
|
||||||
|
try:
|
||||||
|
t.claim_workdir(wd, Path("b.html"), b)
|
||||||
|
except SystemExit as e:
|
||||||
|
assert "занят другой книгой" in str(e)
|
||||||
|
else:
|
||||||
|
raise AssertionError("чужая книга в занятом workdir должна отвергаться")
|
||||||
|
|
||||||
|
|
||||||
|
def test_stale_chapters_removed(tmp: Path):
|
||||||
|
"""Прошлый прогон был длиннее — его главы иначе уедут в перевод."""
|
||||||
|
extracted = tmp / "extracted"
|
||||||
|
extracted.mkdir()
|
||||||
|
(extracted / "chapter_009.json").write_text("{}", encoding="utf-8")
|
||||||
|
t.write_input([[(0, "h1", "", "Глава")]], extracted)
|
||||||
|
assert not (extracted / "chapter_009.json").exists()
|
||||||
|
assert (extracted / "chapter_000.json").exists()
|
||||||
|
|
||||||
|
|
||||||
def test_verify_output_catches_untranslated(tmp: Path):
|
def test_verify_output_catches_untranslated(tmp: Path):
|
||||||
"""Регресс: первый боевой прогон отрапортовал успех на непереведённом файле."""
|
"""Регресс: первый боевой прогон отрапортовал успех на непереведённом файле."""
|
||||||
bad = tmp / "bad.html"
|
bad = tmp / "bad.html"
|
||||||
@@ -90,10 +151,12 @@ if __name__ == "__main__":
|
|||||||
import tempfile
|
import tempfile
|
||||||
|
|
||||||
test_marks_roundtrip()
|
test_marks_roundtrip()
|
||||||
|
test_inline_code_survives()
|
||||||
test_broken_marks_drop_tags()
|
test_broken_marks_drop_tags()
|
||||||
test_language_detection()
|
test_language_detection()
|
||||||
|
for case in (test_split_and_rebuild, test_code_listings_are_left_alone,
|
||||||
|
test_workdir_belongs_to_one_book, test_stale_chapters_removed,
|
||||||
|
test_verify_output_catches_untranslated):
|
||||||
with tempfile.TemporaryDirectory() as d:
|
with tempfile.TemporaryDirectory() as d:
|
||||||
test_split_and_rebuild(Path(d))
|
case(Path(d))
|
||||||
with tempfile.TemporaryDirectory() as d:
|
|
||||||
test_verify_output_catches_untranslated(Path(d))
|
|
||||||
print("OK")
|
print("OK")
|
||||||
|
|||||||
+35
-5
@@ -11,6 +11,7 @@
|
|||||||
поломке абзац сохраняется переведённым, но без внутренней разметки.
|
поломке абзац сохраняется переведённым, но без внутренней разметки.
|
||||||
"""
|
"""
|
||||||
import argparse
|
import argparse
|
||||||
|
import hashlib
|
||||||
import html
|
import html
|
||||||
import json
|
import json
|
||||||
import re
|
import re
|
||||||
@@ -19,10 +20,15 @@ import sys
|
|||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
BLOCK = re.compile(r"^<(p|h1|h2)([^>]*)>(.*)</\1>$")
|
BLOCK = re.compile(r"^<(p|h1|h2)([^>]*)>(.*)</\1>$")
|
||||||
TAG = re.compile(r"</?([ib])>")
|
INLINE = ("i", "b", "code")
|
||||||
MARK = re.compile(r"⟦(/?)([ib])⟧")
|
TAG = re.compile(r"</?(%s)>" % "|".join(INLINE))
|
||||||
|
MARK = re.compile(r"⟦(/?)(%s)⟧" % "|".join(INLINE))
|
||||||
# картинки и пустые блоки не переводим
|
# картинки и пустые блоки не переводим
|
||||||
SKIP = re.compile(r"<img\b")
|
SKIP = re.compile(r"<img\b")
|
||||||
|
# Листинги кода не переводятся вовсе: BLOCK их не узнаёт, поэтому строки <pre>
|
||||||
|
# доходят до результата нетронутыми. Здесь <pre> вырезается только из проверок
|
||||||
|
# языка — иначе английский код утянул бы долю кириллицы ниже порога приёмки.
|
||||||
|
PRE = re.compile(r"<pre\b.*?</pre>", re.S)
|
||||||
|
|
||||||
|
|
||||||
def detect_language(text):
|
def detect_language(text):
|
||||||
@@ -44,7 +50,7 @@ def to_marks(s):
|
|||||||
def from_marks(s):
|
def from_marks(s):
|
||||||
"""Маркеры -> теги. Возвращает (html, ok): ok=False если разметка разъехалась."""
|
"""Маркеры -> теги. Возвращает (html, ok): ok=False если разметка разъехалась."""
|
||||||
escaped = html.escape(s, quote=False)
|
escaped = html.escape(s, quote=False)
|
||||||
depth = {"i": 0, "b": 0}
|
depth = dict.fromkeys(INLINE, 0)
|
||||||
ok = True
|
ok = True
|
||||||
for slash, tag in MARK.findall(escaped):
|
for slash, tag in MARK.findall(escaped):
|
||||||
depth[tag] += -1 if slash else 1
|
depth[tag] += -1 if slash else 1
|
||||||
@@ -82,6 +88,8 @@ def split_chapters(blocks):
|
|||||||
|
|
||||||
def write_input(chapters, extracted):
|
def write_input(chapters, extracted):
|
||||||
extracted.mkdir(parents=True, exist_ok=True)
|
extracted.mkdir(parents=True, exist_ok=True)
|
||||||
|
for stale in extracted.glob("chapter_*.json"):
|
||||||
|
stale.unlink() # главы прошлого прогона иначе уедут в перевод
|
||||||
index = []
|
index = []
|
||||||
for n, ch in enumerate(chapters):
|
for n, ch in enumerate(chapters):
|
||||||
paragraphs = [to_marks(b[3]) for b in ch]
|
paragraphs = [to_marks(b[3]) for b in ch]
|
||||||
@@ -109,6 +117,25 @@ def write_input(chapters, extracted):
|
|||||||
return index
|
return index
|
||||||
|
|
||||||
|
|
||||||
|
def claim_workdir(workdir, source, blocks):
|
||||||
|
"""Рабочий каталог принадлежит одной книге. Имена chapter_NNN.json у всех
|
||||||
|
книг одинаковы, а внешний репозиторий пропускает главы, помеченные
|
||||||
|
готовыми, — без этой привязки вторая книга молча соберётся из перевода
|
||||||
|
первой."""
|
||||||
|
digest = hashlib.sha256(
|
||||||
|
"\n".join(b[3] for b in blocks).encode("utf-8")).hexdigest()[:16]
|
||||||
|
claim = workdir / "bridge_source.json"
|
||||||
|
now = {"source": source.name, "blocks": len(blocks), "sha256": digest}
|
||||||
|
if claim.exists():
|
||||||
|
was = json.loads(claim.read_text(encoding="utf-8"))
|
||||||
|
if was.get("sha256") != digest:
|
||||||
|
sys.exit("рабочий каталог занят другой книгой (%s, блоков %s). "
|
||||||
|
"Возьми чистый --workdir." % (was.get("source"), was.get("blocks")))
|
||||||
|
else:
|
||||||
|
claim.write_text(json.dumps(now, ensure_ascii=False, indent=2),
|
||||||
|
encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
def run_translator(repo, workdir, extracted, workers):
|
def run_translator(repo, workdir, extracted, workers):
|
||||||
script = repo / "03_translate_parallel.py"
|
script = repo / "03_translate_parallel.py"
|
||||||
if not script.exists():
|
if not script.exists():
|
||||||
@@ -150,7 +177,7 @@ def verify_output(out, target):
|
|||||||
"""Совпадение числа абзацев ещё не значит, что перевод состоялся: внешний
|
"""Совпадение числа абзацев ещё не значит, что перевод состоялся: внешний
|
||||||
скрипт при ошибке API молча подставляет оригинал. Проверяем язык результата.
|
скрипт при ошибке API молча подставляет оригинал. Проверяем язык результата.
|
||||||
"""
|
"""
|
||||||
text = re.sub("<[^>]+>", "", out.read_text(encoding="utf-8"))
|
text = re.sub("<[^>]+>", "", PRE.sub(" ", out.read_text(encoding="utf-8")))
|
||||||
lang, share = detect_language(text)
|
lang, share = detect_language(text)
|
||||||
stub = text.count("[UNTRANSLATED]")
|
stub = text.count("[UNTRANSLATED]")
|
||||||
print("результат: %s (кириллица %.0f%%), заглушек [UNTRANSLATED]: %d"
|
print("результат: %s (кириллица %.0f%%), заглушек [UNTRANSLATED]: %d"
|
||||||
@@ -186,7 +213,9 @@ def main():
|
|||||||
plain = " ".join(re.sub("<[^>]+>", "", b[3]) for b in blocks)
|
plain = " ".join(re.sub("<[^>]+>", "", b[3]) for b in blocks)
|
||||||
lang, share = detect_language(plain)
|
lang, share = detect_language(plain)
|
||||||
chapters = split_chapters(blocks)
|
chapters = split_chapters(blocks)
|
||||||
print("блоков: %d, глав: %d, знаков: %d" % (len(blocks), len(chapters), len(plain)))
|
pre_n = len(PRE.findall(args.source.read_text(encoding="utf-8")))
|
||||||
|
print("блоков: %d, глав: %d, знаков: %d, листингов <pre> (не переводятся): %d"
|
||||||
|
% (len(blocks), len(chapters), len(plain), pre_n))
|
||||||
print("язык источника: %s (кириллица %.0f%%)" % (lang, share * 100))
|
print("язык источника: %s (кириллица %.0f%%)" % (lang, share * 100))
|
||||||
|
|
||||||
if lang == args.target and not args.force:
|
if lang == args.target and not args.force:
|
||||||
@@ -199,6 +228,7 @@ def main():
|
|||||||
return
|
return
|
||||||
|
|
||||||
args.workdir.mkdir(parents=True, exist_ok=True)
|
args.workdir.mkdir(parents=True, exist_ok=True)
|
||||||
|
claim_workdir(args.workdir, args.source, blocks)
|
||||||
extracted = (args.workdir / "extracted").resolve()
|
extracted = (args.workdir / "extracted").resolve()
|
||||||
write_input(chapters, extracted)
|
write_input(chapters, extracted)
|
||||||
run_translator(args.repo.resolve(), args.workdir.resolve(), extracted, args.workers)
|
run_translator(args.repo.resolve(), args.workdir.resolve(), extracted, args.workers)
|
||||||
|
|||||||
Reference in New Issue
Block a user