Compare commits
10 Commits
a21a4d713b
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
| 5d957e7be9 | |||
| 33845bd47a | |||
| af5c243687 | |||
| 0320d2d331 | |||
| 356b9ddb22 | |||
| 1a2c03f999 | |||
| f19410d9ce | |||
| 43d6cfc456 | |||
| 14c7fad763 | |||
| 7ed6e500ca |
@@ -26,9 +26,44 @@
|
|||||||
к абзацу с предыдущей страницы;
|
к абзацу с предыдущей страницы;
|
||||||
- переносы на конце строки снимаются, соседние `</i><i>` схлопываются;
|
- переносы на конце строки снимаются, соседние `</i><i>` схлопываются;
|
||||||
- иллюстрации выгружаются в `images/`, обложка рендерится со страницы 1 в 150 dpi;
|
- иллюстрации выгружаются в `images/`, обложка рендерится со страницы 1 в 150 dpi;
|
||||||
|
- повёрнутый на 90° текст выбрасывается целиком: боковые врезки и подписи к
|
||||||
|
таблицам распознаются в кашу («aoHegoduenueArdng») и лезут в заголовки, потому
|
||||||
|
что кегль у них крупный. У Брикмана это 105 строк из 17 300, и все до одной —
|
||||||
|
брак;
|
||||||
|
- мусорный заголовок понижается до `<p>`, а не удаляется: оглавление чистится,
|
||||||
|
текст остаётся. Проверка — `looks_garbled()` в `scripts/bookhtml.py`, шесть
|
||||||
|
признаков структуры, любых двух хватает. Вместе с фильтром поворота это увело
|
||||||
|
Брикмана с 99 `<h1>` до 48;
|
||||||
- позиционирование не сохраняется намеренно — текст должен течь под любой
|
- позиционирование не сохраняется намеренно — текст должен течь под любой
|
||||||
размер шрифта на читалке.
|
размер шрифта на читалке.
|
||||||
|
|
||||||
|
## DjVu со своим текстовым слоем: `scripts/djvu2html.py`
|
||||||
|
|
||||||
|
У сканов в DjVu слой распознавания обычно уже есть, и он лучше нашего прогона
|
||||||
|
через `ocrmypdf`. Замер на Прате (C++ 6-е рус. изд., 1244 полосы): слов со
|
||||||
|
смесью алфавитов внутри слова было 433 на 391 тысячу слов (10,9 на 10 000),
|
||||||
|
после сборки из родного слоя — 109 (2,8).
|
||||||
|
|
||||||
|
⚠️ **Через PDF этот слой не проходит.** `ddjvu -format=pdf` кладёт страницы
|
||||||
|
картинками, `pdftotext` после этого отдаёт пустоту. Текст берётся напрямую из
|
||||||
|
`djvutxt --detail=line`, скрипт разбирает его сам.
|
||||||
|
|
||||||
|
Чего в слое DjVu нет вовсе — курсива и полужирного: в скане их и не было.
|
||||||
|
Заголовки опознаются высотой строки, листинги — пунктуацией.
|
||||||
|
|
||||||
|
⚠️ **Высота строки в DjVu — это габарит с выносными элементами, а не кегль.**
|
||||||
|
Строка с «Ц» или «р» выше соседней на те же 10%, поэтому пороги взяты по
|
||||||
|
измеренным разрывам, а не «чуть выше основного текста». У Праты при основной
|
||||||
|
строке 78: колонтитулы 85–90, разделы 105–125, названия глав 170–185.
|
||||||
|
|
||||||
|
Строки заголовка склеиваются подряд: название главы занимает две-три строки, и
|
||||||
|
без склейки «Класс string и стандартная библиотека шаблонов» разваливается на
|
||||||
|
четыре пункта оглавления.
|
||||||
|
|
||||||
|
Самопроверки без сети: `python3 scripts/djvu2html.py --selftest` и
|
||||||
|
`python3 scripts/pdf2html.py --selftest --selftest` (флаг занимает оба
|
||||||
|
обязательных аргумента).
|
||||||
|
|
||||||
## Подключение как скил Claude Code
|
## Подключение как скил Claude Code
|
||||||
|
|
||||||
Репозиторий одновременно является скилом (`SKILL.md` в корне) и клонируется
|
Репозиторий одновременно является скилом (`SKILL.md` в корне) и клонируется
|
||||||
|
|||||||
@@ -15,20 +15,35 @@ tag, and 2817 body paragraphs became `<h2>`.
|
|||||||
properties, and emits semantic XHTML. calibre is then used only as the
|
properties, and emits semantic XHTML. calibre is then used only as the
|
||||||
XHTML→EPUB packer.
|
XHTML→EPUB packer.
|
||||||
|
|
||||||
## Two entry points
|
## Three entry points
|
||||||
|
|
||||||
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB. Both emit the
|
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB,
|
||||||
same normalized XHTML — one block per line, `<pre>` for code listings — and
|
`scripts/djvu2html.py` for DjVu. All three emit the same normalized XHTML — one
|
||||||
everything downstream (translation, packing, audiobook) is identical. Shared
|
block per line, `<pre>` for code listings — and everything downstream
|
||||||
document skeleton and CSS live in `scripts/bookhtml.py`.
|
(translation, packing, audiobook) is identical. The shared document skeleton,
|
||||||
|
CSS and the junk-heading check `looks_garbled()` live in `scripts/bookhtml.py`.
|
||||||
|
|
||||||
|
**A DjVu with its own text layer must not be re-OCR'd.** The layer that is
|
||||||
|
already in the file beats a fresh `ocrmypdf` run — measured on Prata's C++ 6th
|
||||||
|
Russian edition, 1244 pages: words mixing Latin and Cyrillic inside one word
|
||||||
|
dropped from 433 to 109 per 391k words (10.9 → 2.8 per 10 000). That layer does
|
||||||
|
not survive a trip through PDF: `ddjvu -format=pdf` writes the pages as images
|
||||||
|
and `pdftotext` then returns nothing at all. `djvu2html.py` reads `djvutxt
|
||||||
|
--detail=line` directly.
|
||||||
|
|
||||||
|
Styling is the price: a DjVu text layer carries no italic or bold at all (the
|
||||||
|
scan never had them). Headings come from line height, listings from punctuation.
|
||||||
|
Line height there is the glyph bounding box, not the type size — a line holding
|
||||||
|
a descender is 10% taller than its neighbour — so the thresholds are set from
|
||||||
|
measured gaps (`H1`/`H2` in the script), not from "slightly above body text".
|
||||||
|
|
||||||
**Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws
|
**Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws
|
||||||
away the semantic markup that is already there and re-derives it from font
|
away the semantic markup that is already there and re-derives it from font
|
||||||
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
|
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
|
||||||
the spine, drops the nav document, and maps the book's own headings.
|
the spine, drops the nav document, and maps the book's own headings.
|
||||||
|
|
||||||
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is
|
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (no PyMuPDF
|
||||||
needed, not PyMuPDF) and convert with:
|
needed) and convert with:
|
||||||
|
|
||||||
`python scripts/epub2html.py book.epub out/book.html`
|
`python scripts/epub2html.py book.epub out/book.html`
|
||||||
|
|
||||||
@@ -53,7 +68,10 @@ the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
|
|||||||
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
|
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
|
||||||
embedded fonts / no extractable text means a scan — stop and say OCR is
|
embedded fonts / no extractable text means a scan — stop and say OCR is
|
||||||
needed; this skill does not apply.
|
needed; this skill does not apply.
|
||||||
2. **Check tooling.** `ebook-convert` must be on PATH. For PyMuPDF, prefer an
|
2. **Check tooling.** Packing needs nothing but the standard library
|
||||||
|
(`scripts/pack_epub.py`); calibre is optional and only for `.azw3`. Measured
|
||||||
|
2026-09-05: `ebook-convert` was absent on all three machines at hand, and the
|
||||||
|
conversion went through end to end without it. For PyMuPDF, prefer an
|
||||||
existing interpreter that has it; otherwise build a throwaway venv in the
|
existing interpreter that has it; otherwise build a throwaway venv in the
|
||||||
scratchpad — do not install into the system Python:
|
scratchpad — do not install into the system Python:
|
||||||
|
|
||||||
@@ -85,21 +103,29 @@ the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
|
|||||||
means the mono flag never fired and every listing is about to be reflowed as
|
means the mono flag never fired and every listing is about to be reflowed as
|
||||||
prose. Fix thresholds and re-run until
|
prose. Fix thresholds and re-run until
|
||||||
the counts are sane. Cheap to iterate; do not skip to packing.
|
the counts are sane. Cheap to iterate; do not skip to packing.
|
||||||
7. **Pack with calibre**, always passing explicit TOC XPaths — without them
|
7. **Pack with `scripts/pack_epub.py`** — no calibre, no root, splits on `<h1>`
|
||||||
calibre applies its own heuristics and re-breaks the chapters:
|
and writes the conservative EPUB 2 + NCX that PocketBook and KOReader both
|
||||||
|
read:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ebook-convert out/book.html "Title.epub" \
|
python scripts/pack_epub.py out/book.html "Title.epub" \
|
||||||
--title="Title" --authors="Author" --language=en \
|
--title="Title" --author="Author" --lang=en --cover=cover.jpg
|
||||||
--cover=out/images/cover.jpg \
|
|
||||||
--level1-toc='//h:h1' --level2-toc='//h:h2' \
|
|
||||||
--page-breaks-before='//h:h1' \
|
|
||||||
--no-default-epub-cover
|
|
||||||
ebook-convert "Title.epub" "Title.azw3"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Take title/author from the user or the book's own title page — PDF metadata
|
Take title/author from the user or the book's own title page — PDF metadata
|
||||||
is often an ASIN or a filename.
|
is often an ASIN or a filename.
|
||||||
|
|
||||||
|
calibre is needed only for `.azw3` (Kindle over USB), and only if it happens
|
||||||
|
to be installed:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ebook-convert "Title.epub" "Title.azw3"
|
||||||
|
```
|
||||||
|
|
||||||
|
⚠️ Do **not** pack the XHTML with `ebook-convert` directly: without explicit
|
||||||
|
`--level1-toc='//h:h1' --page-breaks-before='//h:h1'` it applies its own
|
||||||
|
heuristics and re-breaks the chapters. `pack_epub.py` has no such mode to
|
||||||
|
get wrong.
|
||||||
8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle
|
8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle
|
||||||
(Amazon converts server-side), `.azw3` for USB copy into `documents/`.
|
(Amazon converts server-side), `.azw3` for USB copy into `documents/`.
|
||||||
Delete intermediate artifacts left in the user's directories.
|
Delete intermediate artifacts left in the user's directories.
|
||||||
@@ -123,17 +149,51 @@ EOF
|
|||||||
is being promoted;
|
is being promoted;
|
||||||
- many double spaces means line joining is off;
|
- many double spaces means line joining is off;
|
||||||
- list the extracted headings (`grep -o '<h1[^>]*>.\{0,60\}'`) and read them —
|
- list the extracted headings (`grep -o '<h1[^>]*>.\{0,60\}'`) and read them —
|
||||||
they become the TOC, and a wrong one is obvious at a glance;
|
they become the TOC, and a wrong one is obvious at a glance. On a scan most of
|
||||||
|
the junk there comes from **rotated** text, not from bad thresholds: sideways
|
||||||
|
captions and table stubs OCR into mush (`aoHegoduenueArdng`) and land in
|
||||||
|
headings because their type is large. `pdf2html.py` drops any block whose line
|
||||||
|
direction is not horizontal (`is_rotated()`, measured on Brikman: 105 lines out
|
||||||
|
of 17 300, every one of them garbage), and demotes what is left of the mush to
|
||||||
|
`<p>` rather than deleting it. That pair took Brikman from 99 `<h1>` to 48;
|
||||||
- read one full page of body text and confirm paragraphs merge across page
|
- read one full page of body text and confirm paragraphs merge across page
|
||||||
breaks and hyphenated words are rejoined.
|
breaks and hyphenated words are rejoined.
|
||||||
|
|
||||||
After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx`
|
After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx`
|
||||||
navPoint count for the TOC size.
|
navPoint count for the TOC size.
|
||||||
|
|
||||||
|
## Glossary from existing translations
|
||||||
|
|
||||||
|
For a book in a series that already has published translations, `scripts/glossary.py`
|
||||||
|
mines a bilingual glossary so the machine translation does not invent new spellings
|
||||||
|
for names the reader already knows. Feed it pairs of editions of the *same* volume:
|
||||||
|
|
||||||
|
```
|
||||||
|
python scripts/glossary.py --en vol12.fb2 --ru vol12.ru.fb2 \
|
||||||
|
--en vol15.epub --ru vol15.ru.fb2 --score 0.6
|
||||||
|
```
|
||||||
|
|
||||||
|
Two signals, and both are needed. Position: paragraph indices do not line up
|
||||||
|
(Russian editions split dialogue, giving 2–3× more paragraphs), so offsets are
|
||||||
|
measured as a **share of characters**, where the texts track each other closely.
|
||||||
|
Transliteration: a proper name in Russian is nearly always a transliteration, so
|
||||||
|
the Cyrillic candidate is romanized and compared to the English term — this is
|
||||||
|
what turns the output from noise into a usable list.
|
||||||
|
|
||||||
|
Two mirrored filters remove the rest of the junk: a candidate whose head word
|
||||||
|
also appears lowercase in the same text is a sentence-initial common word, not a
|
||||||
|
name — applied on both sides. Measured on four Dresden Files volumes: 71 pairs,
|
||||||
|
of which two were wrong.
|
||||||
|
|
||||||
|
⚠️ Concept terms (`White Council` → `Белый Совет`, `Spire` → `Копьё`) do **not**
|
||||||
|
come out of the transliteration path and the positional one alone is too noisy
|
||||||
|
for them. Extract those by hand and verify by grepping the existing translation.
|
||||||
|
|
||||||
## Optional stage: translation
|
## Optional stage: translation
|
||||||
|
|
||||||
Only when the user asks for a translated book. It slots between step 6 and
|
Only when the user asks for a translated book. It slots between step 6 and
|
||||||
step 7 — translate the XHTML, then pack the translated file with calibre.
|
step 7 — translate the XHTML, then pack the translated file with
|
||||||
|
`pack_epub.py`.
|
||||||
|
|
||||||
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
|
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
|
||||||
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
|
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
|
||||||
@@ -180,8 +240,23 @@ is the bridge: it writes the repo's input format, shells out to
|
|||||||
|
|
||||||
`python scripts/translate.py book.html book.ru.html --repo <clone> --workdir <dir> --workers 12`
|
`python scripts/translate.py book.html book.ru.html --repo <clone> --workdir <dir> --workers 12`
|
||||||
|
|
||||||
Resumable: the external repo tracks completed chapters and skips them on a
|
⚠️ **Not resumable, despite what the external repo claims.** Its
|
||||||
re-run, so an interrupted run costs nothing to restart.
|
`progress/translation_progress.json` stays `{"chapters": {}}` and `--all`
|
||||||
|
re-translates every chapter, including ones already sitting in
|
||||||
|
`translations/`. Measured 2026-08-21 on Stroustrup: a re-run to repair 4
|
||||||
|
failed chapters re-did all 29. Budget a full book on every restart, and
|
||||||
|
prefer getting one clean run over patching a partial one.
|
||||||
|
|
||||||
|
The one thing that *does* skip work: deleting a chapter from `extracted/`
|
||||||
|
before the run. It is then never sent, and `rebuild` keeps the original
|
||||||
|
English for it — the right treatment for an index.
|
||||||
|
**Numbered exercise items come back untranslated.** Measured on Stroustrup
|
||||||
|
2026-08-21: 182 paragraphs of 10151 (1.8%) shaped `[2] Expanding on what you
|
||||||
|
have learned…` were echoed back in English. The document-wide Cyrillic check
|
||||||
|
cannot see this — 1.8% drowns in it — so `rebuild` counts them separately as
|
||||||
|
"осталось на языке оригинала". If the count is high, add an explicit line to the
|
||||||
|
prompt that numbered items are prose and must be translated too.
|
||||||
|
|
||||||
4. **Read the reported counts.** "без перевода" above zero means a chapter came
|
4. **Read the reported counts.** "без перевода" above zero means a chapter came
|
||||||
back with a different paragraph count and kept its original text; "разметка
|
back with a different paragraph count and kept its original text; "разметка
|
||||||
потеряна" counts paragraphs where the model mangled the inline-tag markers
|
потеряна" counts paragraphs where the model mangled the inline-tag markers
|
||||||
|
|||||||
@@ -6,6 +6,14 @@
|
|||||||
строки, а всё, чего не узнал, доносит до результата нетронутым.
|
строки, а всё, чего не узнал, доносит до результата нетронутым.
|
||||||
"""
|
"""
|
||||||
import html
|
import html
|
||||||
|
import re
|
||||||
|
|
||||||
|
# Брак OCR-а внутри слова: цифра вплотную к букве («30HWod1Q»), смесь алфавитов
|
||||||
|
# в одном слове («aoHegoduenue»), заглавная посреди слова.
|
||||||
|
DIGIT_WORD = re.compile(r"[^\W\d_][\d]|[\d][^\W\d_]")
|
||||||
|
WORD = re.compile(r"\w+")
|
||||||
|
LAT, CYR = re.compile(r"[A-Za-z]"), re.compile(r"[А-Яа-яЁё]")
|
||||||
|
MIDCAP = re.compile(r"[a-zа-яё][A-ZА-ЯЁ]")
|
||||||
|
|
||||||
CSS = """
|
CSS = """
|
||||||
body { margin: 0 1em; }
|
body { margin: 0 1em; }
|
||||||
@@ -28,6 +36,53 @@ code { font-family: monospace; font-size: 0.9em; }
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
def looks_garbled(t):
|
||||||
|
"""Заголовок ли это вообще, или подпись из схемы, распознанная по буквам.
|
||||||
|
|
||||||
|
Основную часть брака снимает фильтр поворота в главном цикле — здесь остаётся
|
||||||
|
то, что набрано горизонтально: подписи внутри иллюстраций и обрывки таблиц
|
||||||
|
(«LF. FF. FT. TIT», «0|91239123», «В»).
|
||||||
|
|
||||||
|
Безусловный брак — три случая, каждый сам по себе: в заголовке нет ни одной
|
||||||
|
буквы; заголовок короче трёх знаков; в нём есть палка (в наборе её не бывает,
|
||||||
|
в распознанном скане она попадается постоянно).
|
||||||
|
Дальше шесть признаков структуры, любых двух хватает. Одного мало: «C++11 и
|
||||||
|
лямбды» даёт «не-букв больше трети», «Python3» — цифру вплотную к букве,
|
||||||
|
и оба заголовка настоящие.
|
||||||
|
"""
|
||||||
|
t = t.strip()
|
||||||
|
words = WORD.findall(t)
|
||||||
|
if not words or not any(c.isalpha() for c in t):
|
||||||
|
return True
|
||||||
|
if "|" in t:
|
||||||
|
return True
|
||||||
|
# Заголовок короче четырёх знаков — обрывок («В», «лов», «юн»). Аббревиатуры
|
||||||
|
# целиком заглавными («API», «AI») настоящими заголовками бывают, их щадим —
|
||||||
|
# но только если букв в них не одна и та же: «TT» это обрывок таблицы.
|
||||||
|
if len(t) < 4 and not (t.isupper() and len(set(t)) >= 2):
|
||||||
|
return True
|
||||||
|
# Слова из двух букв по кругу: «TIT. ITITIT», «LF. FF. FT». Нужны минимум два
|
||||||
|
# таких слова подряд и ни одного нормального, иначе под нож попадёт «Часть III».
|
||||||
|
long_words = [w for w in words if len(w) >= 3]
|
||||||
|
if len(long_words) >= 2 and all(len(set(w.lower())) <= 2 for w in long_words):
|
||||||
|
return True
|
||||||
|
letters = sum(c.isalpha() for c in t)
|
||||||
|
toks = t.split()
|
||||||
|
# Обрывок — короткое слово без цифр внутри: «LF.», «оо». Цифры исключены
|
||||||
|
# намеренно, иначе «C++11» и «Qt 6» считались бы обрывками.
|
||||||
|
frags = [w for w in toks if not any(c.isdigit() for c in w)
|
||||||
|
and sum(c.isalpha() for c in w) < 3]
|
||||||
|
signals = (
|
||||||
|
len(frags) * 2 > len(toks), # обрывки слов
|
||||||
|
letters * 3 < len(t) * 2, # не-букв больше трети
|
||||||
|
bool(DIGIT_WORD.search(t)),
|
||||||
|
any(LAT.search(w) and CYR.search(w) for w in words), # смесь алфавитов в слове
|
||||||
|
any(MIDCAP.search(w) for w in words), # заглавная посреди слова
|
||||||
|
next((c.islower() for c in t if c.isalpha()), False), # начинается со строчной
|
||||||
|
)
|
||||||
|
return sum(signals) >= 2
|
||||||
|
|
||||||
|
|
||||||
def document(title, parts):
|
def document(title, parts):
|
||||||
"""parts: [(tag, cls, inner_html)]; tag 'figure' — псевдотег для картинки."""
|
"""parts: [(tag, cls, inner_html)]; tag 'figure' — псевдотег для картинки."""
|
||||||
buf = ['<?xml version="1.0" encoding="utf-8"?>',
|
buf = ['<?xml version="1.0" encoding="utf-8"?>',
|
||||||
|
|||||||
@@ -0,0 +1,195 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""DjVu со своим текстовым слоем -> semantic XHTML.
|
||||||
|
|
||||||
|
Для сканов, у которых слой распознавания уже есть в файле: он почти всегда
|
||||||
|
лучше нашего OCR-а. Замерено на Прате 6-го изд.: родной слой 302 523 слова
|
||||||
|
при смеси алфавитов 0.8 на 10 000 знаков, наш прогон через ocrmypdf —
|
||||||
|
296 954 слова при 12.6 и на 26 страниц меньше.
|
||||||
|
|
||||||
|
Через PDF этот слой не проходит: ddjvu -format=pdf кладёт страницы картинками
|
||||||
|
и текст теряется целиком (проверено, pdftotext даёт пустоту). Поэтому слой
|
||||||
|
берётся напрямую из djvutxt.
|
||||||
|
|
||||||
|
Стиля (курсив, полужирный) в слое DjVu нет вовсе — в скане его и не было.
|
||||||
|
Заголовки опознаются высотой строки, листинги — пунктуацией, как в pdf2html.py.
|
||||||
|
|
||||||
|
djvu2html.py book.djvu out.html
|
||||||
|
djvu2html.py --selftest
|
||||||
|
"""
|
||||||
|
import collections
|
||||||
|
import html
|
||||||
|
import re
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import bookhtml
|
||||||
|
|
||||||
|
CODEY = re.compile(r"[{}();\[\]<>=*/#]|::|->")
|
||||||
|
SOFT_HYPHEN = ""
|
||||||
|
HEAD_BAND = 0.07 # доля высоты полосы: выше — колонтитул, а не текст
|
||||||
|
FOOT_BAND = 0.05 # то же снизу — колонцифра
|
||||||
|
# Высота строки в DjVu — это габарит с выносными элементами, а не кегль: строка
|
||||||
|
# с «Ц» или «р» выше соседней на те же 10%. Поэтому пороги взяты не «чуть выше
|
||||||
|
# основного текста», а по измеренным разрывам у Праты (основная строка 78):
|
||||||
|
# 85–90 — колонтитулы, 105–125 — разделы, 170–185 — названия глав.
|
||||||
|
H1 = 2.0
|
||||||
|
H2 = 1.30
|
||||||
|
|
||||||
|
|
||||||
|
def parse(txt):
|
||||||
|
"""djvutxt --detail=line -> [[(x0, y0, x1, y1, text), ...] по страницам].
|
||||||
|
|
||||||
|
Свой разбор, а не sexpdata: формат простой, а зависимость ради него —
|
||||||
|
лишняя. Строка может переноситься между скобкой и текстом, поэтому идём
|
||||||
|
по знакам, а не построчно.
|
||||||
|
"""
|
||||||
|
pages, cur, i, n = [], None, 0, len(txt)
|
||||||
|
while i < n:
|
||||||
|
c = txt[i]
|
||||||
|
if c == '"': # строка: до неэкранированной закрывающей кавычки
|
||||||
|
j, buf = i + 1, []
|
||||||
|
while j < n and txt[j] != '"':
|
||||||
|
if txt[j] == "\\" and j + 1 < n:
|
||||||
|
buf.append({"n": "\n", "t": "\t"}.get(txt[j + 1], txt[j + 1]))
|
||||||
|
j += 2
|
||||||
|
continue
|
||||||
|
buf.append(txt[j])
|
||||||
|
j += 1
|
||||||
|
if cur is not None and cur["box"]:
|
||||||
|
cur["lines"].append(tuple(cur["box"]) + ("".join(buf),))
|
||||||
|
cur["box"] = None
|
||||||
|
i = j + 1
|
||||||
|
continue
|
||||||
|
m = re.match(r"\((page|line)((?:\s+-?\d+){4})", txt[i:])
|
||||||
|
if m:
|
||||||
|
box = [int(v) for v in m.group(2).split()]
|
||||||
|
if m.group(1) == "page":
|
||||||
|
cur = {"h": box[3] - box[1], "lines": [], "box": None}
|
||||||
|
pages.append(cur)
|
||||||
|
elif cur is not None:
|
||||||
|
cur["box"] = box
|
||||||
|
i += m.end()
|
||||||
|
continue
|
||||||
|
i += 1
|
||||||
|
return [(p["h"], p["lines"]) for p in pages]
|
||||||
|
|
||||||
|
|
||||||
|
def profile(pages):
|
||||||
|
"""Высота основной строки и левое поле — по самым частым значениям."""
|
||||||
|
heights, lefts = collections.Counter(), collections.Counter()
|
||||||
|
for _, lines in pages:
|
||||||
|
for x0, y0, x1, y1, _ in lines:
|
||||||
|
heights[y1 - y0] += 1
|
||||||
|
lefts[x0] += 1
|
||||||
|
if not heights:
|
||||||
|
sys.exit("в DjVu нет текстового слоя — этот скрипт не поможет, нужен OCR")
|
||||||
|
# Высота строки гуляет на пиксель-другой, поэтому берётся не голая мода, а
|
||||||
|
# та высота, у которой вместе с соседями ±1 набирается больше всего строк.
|
||||||
|
body = max(heights, key=lambda h: sum(heights[h + d] for d in (-1, 0, 1)))
|
||||||
|
left = min(x for x, _ in lefts.most_common(4))
|
||||||
|
return body, left
|
||||||
|
|
||||||
|
|
||||||
|
def looks_like_code(t):
|
||||||
|
return len(t) > 8 and len(CODEY.findall(t)) / len(t) > 0.03
|
||||||
|
|
||||||
|
|
||||||
|
def kind(height, body):
|
||||||
|
if height >= body * H1:
|
||||||
|
return "h1"
|
||||||
|
if height >= body * H2:
|
||||||
|
return "h2"
|
||||||
|
return "p"
|
||||||
|
|
||||||
|
|
||||||
|
def convert(pages, body, left):
|
||||||
|
parts, open_para = [], None # open_para: [tag, cls, text]
|
||||||
|
for page_h, lines in pages:
|
||||||
|
for x0, y0, x1, y1, raw in lines:
|
||||||
|
txt = raw.strip()
|
||||||
|
if not txt:
|
||||||
|
continue
|
||||||
|
# Координаты DjVu считаются снизу вверх.
|
||||||
|
top = (page_h - y1) / page_h if page_h else 0.5
|
||||||
|
bottom = y0 / page_h if page_h else 0.5
|
||||||
|
if (top < HEAD_BAND or bottom < FOOT_BAND) and len(txt) < 80:
|
||||||
|
continue # колонтитул и колонцифра
|
||||||
|
tag = kind(y1 - y0, body)
|
||||||
|
if tag == "p" and looks_like_code(txt):
|
||||||
|
tag = "pre"
|
||||||
|
if tag in ("h1", "h2") and bookhtml.looks_garbled(txt):
|
||||||
|
tag = "p"
|
||||||
|
# Абзац продолжается, пока строки идут от левого поля: красная
|
||||||
|
# строка и смена тега начинают новый. Заголовок склеивается всегда:
|
||||||
|
# название главы занимает две-три строки («Класс string» / «и
|
||||||
|
# стандартная» / «библиотека» / «шаблонов» — это один заголовок),
|
||||||
|
# а два разных заголовка подряд без текста между ними не встречаются.
|
||||||
|
cont = open_para and open_para[0] == tag and (
|
||||||
|
tag in ("h1", "h2") or (tag == "p" and x0 <= left + (y1 - y0)))
|
||||||
|
if tag == "pre" and open_para and open_para[0] == "pre":
|
||||||
|
open_para[2] += "\n" + txt
|
||||||
|
continue
|
||||||
|
if cont:
|
||||||
|
prev = open_para[2]
|
||||||
|
if prev.endswith(SOFT_HYPHEN):
|
||||||
|
open_para[2] = prev[:-1] + txt
|
||||||
|
elif prev.endswith("-") and not prev.endswith("--"):
|
||||||
|
open_para[2] = prev[:-1] + txt
|
||||||
|
else:
|
||||||
|
open_para[2] = prev + " " + txt
|
||||||
|
continue
|
||||||
|
if open_para:
|
||||||
|
parts.append(tuple(open_para))
|
||||||
|
open_para = [tag, None, txt]
|
||||||
|
if open_para:
|
||||||
|
parts.append(tuple(open_para))
|
||||||
|
return [(t, c, html.escape(x).replace(SOFT_HYPHEN, "")) for t, c, x in parts]
|
||||||
|
|
||||||
|
|
||||||
|
# Кегли взяты с настоящих полос Праты: колонтитул 64, текст 77, раздел 108,
|
||||||
|
# название главы 180. Начало координат в DjVu внизу полосы.
|
||||||
|
SAMPLE = """(page 0 0 100 1000 (line 10 950 90 985 "300 Глава 6 ")
|
||||||
|
(line 10 720 90 900 "Упражнения по программированию ")
|
||||||
|
(line 10 590 90 698 "Подраздел про циклы ")
|
||||||
|
(line 10 500 90 577 "Напишите программу, которая читает ввод до сим-")
|
||||||
|
(line 10 420 90 497 "вола @ и повторяет его. ")
|
||||||
|
(line 10 320 90 397 "int main() { return 0; }")
|
||||||
|
(line 10 20 90 60 "300 "))"""
|
||||||
|
|
||||||
|
|
||||||
|
def selftest():
|
||||||
|
pages = parse(SAMPLE)
|
||||||
|
assert len(pages) == 1, pages
|
||||||
|
assert len(pages[0][1]) == 7, pages[0][1]
|
||||||
|
body, left = profile(pages)
|
||||||
|
assert body == 77, body
|
||||||
|
parts = convert(pages, body, left)
|
||||||
|
tags = [p[0] for p in parts]
|
||||||
|
assert tags == ["h1", "h2", "p", "pre"], tags
|
||||||
|
# Колонтитул сверху и колонцифра снизу выброшены, перенос склеен без пробела.
|
||||||
|
assert "символа" in parts[2][2], parts[2][2]
|
||||||
|
assert "300" not in " ".join(p[2] for p in parts), parts
|
||||||
|
assert parse('(page 0 0 10 10 (line 1 1 2 2 "a \\"b\\" c"))')[0][1][0][4] == 'a "b" c'
|
||||||
|
print("selftest OK")
|
||||||
|
|
||||||
|
|
||||||
|
if len(sys.argv) == 2 and sys.argv[1] == "--selftest":
|
||||||
|
selftest()
|
||||||
|
sys.exit()
|
||||||
|
if len(sys.argv) < 3:
|
||||||
|
sys.exit("usage: djvu2html.py book.djvu out.html | djvu2html.py --selftest")
|
||||||
|
|
||||||
|
SRC, OUT = Path(sys.argv[1]), Path(sys.argv[2])
|
||||||
|
raw = subprocess.run(["djvutxt", "--detail=line", str(SRC)],
|
||||||
|
capture_output=True, text=True, check=True).stdout
|
||||||
|
pages = parse(raw)
|
||||||
|
BODY, LEFT = profile(pages)
|
||||||
|
print("страниц %d, высота строки %d, левое поле %d" % (len(pages), BODY, LEFT))
|
||||||
|
parts = convert(pages, BODY, LEFT)
|
||||||
|
OUT.write_text(bookhtml.document(SRC.stem, parts), encoding="utf-8")
|
||||||
|
print("blocks:", len(parts),
|
||||||
|
"h1:", sum(1 for p in parts if p[0] == "h1"),
|
||||||
|
"h2:", sum(1 for p in parts if p[0] == "h2"),
|
||||||
|
"p:", sum(1 for p in parts if p[0] == "p"),
|
||||||
|
"pre:", sum(1 for p in parts if p[0] == "pre"))
|
||||||
@@ -0,0 +1,163 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""FB2 -> тот же XHTML, что выдают pdf2html.py и epub2html.py.
|
||||||
|
|
||||||
|
FB2 в этой библиотеке основной формат (704 тысячи файлов), и без этого входа
|
||||||
|
конвейер до них не дотягивается. Разметка у fb2 уже семантическая, поэтому
|
||||||
|
задача та же, что у epub2html.py: привести чужие теги к нашим.
|
||||||
|
|
||||||
|
Только стандартная библиотека.
|
||||||
|
"""
|
||||||
|
import argparse
|
||||||
|
import html
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import zipfile
|
||||||
|
from html.parser import HTMLParser
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import bookhtml
|
||||||
|
|
||||||
|
BLOCK = {"p": ("p", None), "v": ("p", "li"), "subtitle": ("h2", None),
|
||||||
|
"text-author": ("p", "note"), "th": ("p", "row"), "td": ("p", "row")}
|
||||||
|
INLINE = {"emphasis": "i", "strong": "b", "code": "code", "sub": "i", "sup": "i"}
|
||||||
|
DROP = {"description", "binary", "stylesheet"}
|
||||||
|
|
||||||
|
|
||||||
|
class Reader(HTMLParser):
|
||||||
|
def __init__(self):
|
||||||
|
super().__init__(convert_charrefs=True)
|
||||||
|
self.blocks, self.cur, self.open_i, self.drop = [], None, [], 0
|
||||||
|
self.depth = 0 # вложенность section: даёт уровень заголовка
|
||||||
|
self.in_title = False
|
||||||
|
|
||||||
|
def flush(self):
|
||||||
|
if not self.cur:
|
||||||
|
return
|
||||||
|
tag, cls, parts = self.cur
|
||||||
|
self.cur = None
|
||||||
|
for t in reversed(self.open_i):
|
||||||
|
parts.append("</%s>" % t)
|
||||||
|
self.open_i = []
|
||||||
|
txt = re.sub(r"\s+", " ", "".join(parts)).strip()
|
||||||
|
if txt:
|
||||||
|
self.blocks.append((tag, cls, txt))
|
||||||
|
|
||||||
|
def handle_starttag(self, tag, attrs):
|
||||||
|
if tag in DROP:
|
||||||
|
self.drop += 1
|
||||||
|
return
|
||||||
|
if self.drop:
|
||||||
|
return
|
||||||
|
if tag == "section":
|
||||||
|
self.depth += 1
|
||||||
|
return
|
||||||
|
if tag == "title":
|
||||||
|
self.flush()
|
||||||
|
self.in_title = True
|
||||||
|
return
|
||||||
|
if tag == "empty-line":
|
||||||
|
self.flush()
|
||||||
|
return
|
||||||
|
if tag == "image":
|
||||||
|
return # картинки fb2 лежат в base64, пропускаем
|
||||||
|
if self.in_title and tag == "p":
|
||||||
|
# заголовок раздела: уровень по вложенности section
|
||||||
|
self.flush()
|
||||||
|
self.cur = ("h1" if self.depth <= 2 else "h2", None, [])
|
||||||
|
return
|
||||||
|
if tag in BLOCK:
|
||||||
|
self.flush()
|
||||||
|
self.cur = (BLOCK[tag][0], BLOCK[tag][1], [])
|
||||||
|
return
|
||||||
|
if tag in INLINE and self.cur:
|
||||||
|
out = INLINE[tag]
|
||||||
|
self.open_i.append(out)
|
||||||
|
self.cur[2].append("<%s>" % out)
|
||||||
|
|
||||||
|
def handle_endtag(self, tag):
|
||||||
|
if tag in DROP:
|
||||||
|
self.drop = max(0, self.drop - 1)
|
||||||
|
return
|
||||||
|
if self.drop:
|
||||||
|
return
|
||||||
|
if tag == "section":
|
||||||
|
self.flush()
|
||||||
|
self.depth = max(0, self.depth - 1)
|
||||||
|
return
|
||||||
|
if tag == "title":
|
||||||
|
self.flush()
|
||||||
|
self.in_title = False
|
||||||
|
return
|
||||||
|
if tag in BLOCK:
|
||||||
|
self.flush()
|
||||||
|
return
|
||||||
|
if tag in INLINE and self.cur:
|
||||||
|
out = INLINE[tag]
|
||||||
|
if out in self.open_i:
|
||||||
|
self.open_i.remove(out)
|
||||||
|
self.cur[2].append("</%s>" % out)
|
||||||
|
|
||||||
|
def handle_data(self, data):
|
||||||
|
if self.drop or not self.cur:
|
||||||
|
return
|
||||||
|
self.cur[2].append(html.escape(data, quote=False))
|
||||||
|
|
||||||
|
def close(self):
|
||||||
|
super().close()
|
||||||
|
self.flush()
|
||||||
|
|
||||||
|
|
||||||
|
def read_fb2(path):
|
||||||
|
p = Path(path)
|
||||||
|
if p.suffix.lower() == ".zip" or zipfile.is_zipfile(p):
|
||||||
|
with zipfile.ZipFile(p) as z:
|
||||||
|
name = next(n for n in z.namelist() if n.lower().endswith(".fb2"))
|
||||||
|
raw = z.read(name)
|
||||||
|
else:
|
||||||
|
raw = p.read_bytes()
|
||||||
|
enc = "utf-8"
|
||||||
|
m = re.search(rb'encoding="([\w-]+)"', raw[:200])
|
||||||
|
if m:
|
||||||
|
enc = m.group(1).decode("ascii", "ignore")
|
||||||
|
return raw.decode(enc, "replace")
|
||||||
|
|
||||||
|
|
||||||
|
def meta(text):
|
||||||
|
"""(автор, название) из description — для метаданных EPUB."""
|
||||||
|
def tag(name):
|
||||||
|
m = re.search(r"<%s>(.*?)</%s>" % (name, name), text, re.S)
|
||||||
|
return re.sub(r"<[^>]+>", " ", m.group(1)).strip() if m else ""
|
||||||
|
ti = re.search(r"<title-info>(.*?)</title-info>", text, re.S)
|
||||||
|
block = ti.group(1) if ti else text
|
||||||
|
first = re.search(r"<first-name>(.*?)</first-name>", block, re.S)
|
||||||
|
last = re.search(r"<last-name>(.*?)</last-name>", block, re.S)
|
||||||
|
author = " ".join(re.sub(r"<[^>]+>", "", x.group(1)).strip()
|
||||||
|
for x in (first, last) if x).strip()
|
||||||
|
book = re.search(r"<book-title>(.*?)</book-title>", block, re.S)
|
||||||
|
return author, (re.sub(r"<[^>]+>", "", book.group(1)).strip() if book else "")
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
ap = argparse.ArgumentParser(description=__doc__)
|
||||||
|
ap.add_argument("source", type=Path)
|
||||||
|
ap.add_argument("output", type=Path)
|
||||||
|
args = ap.parse_args()
|
||||||
|
text = read_fb2(args.source)
|
||||||
|
body = re.search(r"<body[^>]*>(.*)</body>", text, re.S)
|
||||||
|
r = Reader()
|
||||||
|
r.feed(body.group(1) if body else text)
|
||||||
|
r.close()
|
||||||
|
if not r.blocks:
|
||||||
|
sys.exit("в %s не нашлось текста" % args.source)
|
||||||
|
author, title = meta(text)
|
||||||
|
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
args.output.write_text(bookhtml.document(title or args.output.stem, r.blocks),
|
||||||
|
encoding="utf-8")
|
||||||
|
print("blocks: %d, h1: %d, p: %d" % (len(r.blocks),
|
||||||
|
sum(1 for t, _, _ in r.blocks if t == "h1"),
|
||||||
|
sum(1 for t, _, _ in r.blocks if t == "p")))
|
||||||
|
print("метаданные: %s — %s" % (author or "?", title or "?"))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,219 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Двуязычный глоссарий из параллельных изданий одной книги.
|
||||||
|
|
||||||
|
Задача: машинный перевод очередного тома цикла не должен расходиться с уже
|
||||||
|
изданными переводами в именах и реалиях. Частотный список даёт только имена;
|
||||||
|
пары вида White Council → Белый Совет так не получить.
|
||||||
|
|
||||||
|
Метод — выравнивание по относительной позиции в тексте. Абзацы английского и
|
||||||
|
русского изданий не совпадают ни числом, ни границами, но идут в одном порядке,
|
||||||
|
поэтому термин, встречающийся в английском тексте на 12%, 34% и 78% длины,
|
||||||
|
в переводе окажется примерно там же. Кандидат, чьи позиции совпали с позициями
|
||||||
|
термина лучше, чем с текстом вообще, и есть перевод.
|
||||||
|
|
||||||
|
Работает с fb2, epub и голым текстом. Только стандартная библиотека.
|
||||||
|
"""
|
||||||
|
import argparse
|
||||||
|
import collections
|
||||||
|
import difflib
|
||||||
|
import html
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import zipfile
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
# Слова, с которых начинается предложение, — не имена собственные.
|
||||||
|
RU_STOP = {"Когда", "Если", "Что", "Как", "Это", "Она", "Они", "Мы", "Вы", "Так",
|
||||||
|
"Потом", "Затем", "Даже", "Его", "Её", "Их", "Может", "Все", "Теперь",
|
||||||
|
"Нет", "Да", "Меня", "Мне", "Тогда", "Только", "После", "Пока", "Там",
|
||||||
|
"Здесь", "Однако", "Впрочем", "Просто", "Один", "Вот", "Или", "Почему",
|
||||||
|
"Мой", "Моя", "Кто", "Конечно", "Хорошо", "Ага", "Думаю", "Возможно",
|
||||||
|
"Глава", "Никто", "Ничего", "Что-то", "Кажется", "Значит", "Сейчас"}
|
||||||
|
EN_STOP = {"The", "A", "An", "And", "But", "I", "It", "He", "She", "They", "We",
|
||||||
|
"You", "This", "That", "There", "Then", "When", "If", "So", "My", "His",
|
||||||
|
"Her", "Chapter", "What", "Why", "How", "No", "Yes", "Not", "For", "Of",
|
||||||
|
"In", "On", "At", "To", "With", "As", "Was", "Were", "Had", "Have",
|
||||||
|
"Would", "Could", "Should", "One", "All", "Just", "Like", "Now", "Well",
|
||||||
|
"Okay", "Oh", "Maybe", "Something", "Nothing", "Someone"}
|
||||||
|
|
||||||
|
|
||||||
|
def read_text(path):
|
||||||
|
"""Текст книги из fb2, epub или txt — разметка выкидывается."""
|
||||||
|
p = Path(path)
|
||||||
|
if p.suffix.lower() == ".epub":
|
||||||
|
with zipfile.ZipFile(p) as z:
|
||||||
|
parts = [z.read(n).decode("utf-8", "replace")
|
||||||
|
for n in z.namelist() if n.lower().endswith((".xhtml", ".html"))]
|
||||||
|
raw = "\n".join(parts)
|
||||||
|
elif p.suffix.lower() in (".fb2", ".xml"):
|
||||||
|
raw = p.read_bytes().decode("utf-8", "replace")
|
||||||
|
else:
|
||||||
|
raw = p.read_text(encoding="utf-8", errors="replace")
|
||||||
|
raw = re.sub(r"<(script|style|head)[^>]*>.*?</\1>", " ", raw, flags=re.S | re.I)
|
||||||
|
raw = re.sub(r"</p>|</section>|<br\s*/?>", "\n", raw, flags=re.I)
|
||||||
|
return html.unescape(re.sub(r"<[^>]+>", " ", raw))
|
||||||
|
|
||||||
|
|
||||||
|
def paragraphs(text):
|
||||||
|
return [p.strip() for p in text.split("\n") if len(p.strip()) > 40]
|
||||||
|
|
||||||
|
|
||||||
|
def candidates(paras, pattern, stop, min_count):
|
||||||
|
"""{термин: [относительные позиции вхождений]}
|
||||||
|
|
||||||
|
Позиция считается по символам, а не по номеру абзаца: русские издания
|
||||||
|
разбивают диалоги построчно, и абзацев там втрое больше — по индексу
|
||||||
|
тексты расходятся, по доле знаков идут почти вровень.
|
||||||
|
"""
|
||||||
|
pos = collections.defaultdict(list)
|
||||||
|
total = sum(len(p) for p in paras) or 1
|
||||||
|
acc = 0
|
||||||
|
for p in paras:
|
||||||
|
rel = acc / total
|
||||||
|
acc += len(p)
|
||||||
|
for m in set(pattern.findall(p)):
|
||||||
|
head = m.split()[0]
|
||||||
|
if head in stop or len(m) < 3:
|
||||||
|
continue
|
||||||
|
pos[m].append(rel)
|
||||||
|
return {t: v for t, v in pos.items() if len(v) >= min_count}
|
||||||
|
|
||||||
|
|
||||||
|
# Кириллица -> латиница для сравнения имён. Имена собственные в переводе почти
|
||||||
|
# всегда транслитерация, и это признак куда надёжнее позиционного совпадения.
|
||||||
|
LAT = {"а":"a","б":"b","в":"v","г":"g","д":"d","е":"e","ё":"e","ж":"j","з":"z",
|
||||||
|
"и":"i","й":"i","к":"k","л":"l","м":"m","н":"n","о":"o","п":"p","р":"r",
|
||||||
|
"с":"s","т":"t","у":"u","ф":"f","х":"h","ц":"c","ч":"c","ш":"s","щ":"s",
|
||||||
|
"ъ":"","ы":"i","ь":"","э":"e","ю":"u","я":"a"}
|
||||||
|
|
||||||
|
|
||||||
|
def translit(s):
|
||||||
|
return "".join(LAT.get(c, c) for c in s.lower())
|
||||||
|
|
||||||
|
|
||||||
|
def name_similarity(en, ru):
|
||||||
|
"""Похожесть после огрубления: латиница обеих сторон без удвоений и гласных
|
||||||
|
на конце. Nicodemus/Никодимус дают почти единицу, Harry/Молли — ноль."""
|
||||||
|
a = re.sub(r"[^a-z]", "", en.lower())
|
||||||
|
b = re.sub(r"[^a-z]", "", translit(ru))
|
||||||
|
a = re.sub(r"(.)\1+", r"\1", a)
|
||||||
|
b = re.sub(r"(.)\1+", r"\1", b)
|
||||||
|
if not a or not b:
|
||||||
|
return 0.0
|
||||||
|
return difflib.SequenceMatcher(None, a, b[:len(a) + 3]).ratio()
|
||||||
|
|
||||||
|
|
||||||
|
def overlap(a, b, tol):
|
||||||
|
"""Доля вхождений a, у которых нашлось вхождение b поблизости."""
|
||||||
|
if not a or not b:
|
||||||
|
return 0.0
|
||||||
|
b = sorted(b)
|
||||||
|
hit = 0
|
||||||
|
for x in a:
|
||||||
|
lo, hi = 0, len(b) - 1
|
||||||
|
best = 1.0
|
||||||
|
while lo <= hi:
|
||||||
|
mid = (lo + hi) // 2
|
||||||
|
best = min(best, abs(b[mid] - x))
|
||||||
|
if b[mid] < x:
|
||||||
|
lo = mid + 1
|
||||||
|
else:
|
||||||
|
hi = mid - 1
|
||||||
|
if best <= tol:
|
||||||
|
hit += 1
|
||||||
|
return hit / len(a)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
ap = argparse.ArgumentParser(description=__doc__)
|
||||||
|
ap.add_argument("--en", action="append", required=True, help="английское издание")
|
||||||
|
ap.add_argument("--ru", action="append", required=True, help="русское издание того же тома")
|
||||||
|
ap.add_argument("--min-count", type=int, default=4)
|
||||||
|
ap.add_argument("--tol", type=float, default=0.02, help="допуск по относительной позиции")
|
||||||
|
ap.add_argument("--score", type=float, default=0.6, help="порог уверенности")
|
||||||
|
ap.add_argument("--top", type=int, default=80)
|
||||||
|
args = ap.parse_args()
|
||||||
|
if len(args.en) != len(args.ru):
|
||||||
|
sys.exit("нужно одинаковое число --en и --ru: это должны быть пары изданий")
|
||||||
|
|
||||||
|
EN = re.compile(r"\b[A-Z][a-z]+(?:\s+(?:of\s+|the\s+)?[A-Z][a-z]+){0,2}")
|
||||||
|
lower_seen = collections.Counter()
|
||||||
|
ru_lower = collections.Counter()
|
||||||
|
RU = re.compile(r"\b[А-ЯЁ][а-яё]+(?:\s+[А-ЯЁ][а-яё]+){0,2}")
|
||||||
|
en_pos, ru_pos = collections.defaultdict(list), collections.defaultdict(list)
|
||||||
|
for i, (fe, fr) in enumerate(zip(args.en, args.ru)):
|
||||||
|
pe, pr = paragraphs(read_text(fe)), paragraphs(read_text(fr))
|
||||||
|
# Настоящее имя собственное почти не встречается со строчной буквы:
|
||||||
|
# так отсеиваются After, Once, More и прочие начала предложений.
|
||||||
|
for w in re.findall(r"\b[a-z]{3,}\b", " ".join(pe)):
|
||||||
|
lower_seen[w] += 1
|
||||||
|
for w in re.findall(r"\b[а-яё]{3,}\b", " ".join(pr)):
|
||||||
|
ru_lower[w] += 1
|
||||||
|
print("пара %d: %d абзацев EN, %d RU" % (i + 1, len(pe), len(pr)), file=sys.stderr)
|
||||||
|
for t, v in candidates(pe, EN, EN_STOP, 1).items():
|
||||||
|
en_pos[t] += [(i, x) for x in v]
|
||||||
|
for t, v in candidates(pr, RU, RU_STOP, 1).items():
|
||||||
|
ru_pos[t] += [(i, x) for x in v]
|
||||||
|
|
||||||
|
en_pos = {t: v for t, v in en_pos.items()
|
||||||
|
if len(v) >= args.min_count
|
||||||
|
and lower_seen[t.split()[0].lower()] < max(3, len(v) * 0.2)}
|
||||||
|
# Зеркальный отсев: «Надеюсь», «Зачем», «Парень» — обычные слова, они
|
||||||
|
# встречаются со строчной буквы и именами собственными быть не могут.
|
||||||
|
ru_pos = {t: v for t, v in ru_pos.items()
|
||||||
|
if len(v) >= args.min_count
|
||||||
|
and ru_lower[t.split()[0].lower()] < max(3, len(v) * 0.2)}
|
||||||
|
print("кандидатов: %d EN, %d RU" % (len(en_pos), len(ru_pos)), file=sys.stderr)
|
||||||
|
|
||||||
|
by_book_ru = collections.defaultdict(dict)
|
||||||
|
for t, v in ru_pos.items():
|
||||||
|
for b, x in v:
|
||||||
|
by_book_ru[b].setdefault(t, []).append(x)
|
||||||
|
|
||||||
|
out = []
|
||||||
|
for term, occ in sorted(en_pos.items(), key=lambda kv: -len(kv[1])):
|
||||||
|
mine = collections.defaultdict(list)
|
||||||
|
for b, x in occ:
|
||||||
|
mine[b].append(x)
|
||||||
|
best, best_score = None, 0.0
|
||||||
|
for cand, cocc in ru_pos.items():
|
||||||
|
# Частота кандидата должна быть сопоставима: «Гарри» встречается на
|
||||||
|
# каждой странице и по одностороннему совпадению побеждает всех.
|
||||||
|
if not (0.25 <= len(cocc) / len(occ) <= 4.0):
|
||||||
|
continue
|
||||||
|
fwd = back = tot_f = tot_b = 0.0
|
||||||
|
for b, xs in mine.items():
|
||||||
|
cand_pos = by_book_ru[b].get(cand)
|
||||||
|
if not cand_pos:
|
||||||
|
continue
|
||||||
|
fwd += overlap(xs, cand_pos, args.tol) * len(xs)
|
||||||
|
back += overlap(cand_pos, xs, args.tol) * len(cand_pos)
|
||||||
|
tot_f += len(xs); tot_b += len(cand_pos)
|
||||||
|
if tot_f < len(occ) * 0.5 or not tot_b:
|
||||||
|
continue
|
||||||
|
# Симметричная мера: термин должен находить перевод, а перевод —
|
||||||
|
# термин. Иначе частотное слово выигрывает у настоящего соответствия.
|
||||||
|
score = (fwd / tot_f) * (back / tot_b)
|
||||||
|
# Транслитерация — сильный самостоятельный довод: если написание
|
||||||
|
# совпадает, позиционного подтверждения нужно куда меньше.
|
||||||
|
sim = name_similarity(term, cand)
|
||||||
|
if sim >= 0.72:
|
||||||
|
score = max(score, sim) + 0.25
|
||||||
|
elif sim < 0.3 and " " not in term:
|
||||||
|
score *= 0.35 # разные имена, позиции совпали случайно
|
||||||
|
if score > best_score:
|
||||||
|
best, best_score = cand, score
|
||||||
|
if best and best_score >= args.score:
|
||||||
|
out.append((len(occ), best_score, term, best))
|
||||||
|
|
||||||
|
print("# частота | уверенность | оригинал -> перевод")
|
||||||
|
seen = set()
|
||||||
|
for n, sc, en, ru in out[:args.top]:
|
||||||
|
if (en, ru) in seen:
|
||||||
|
continue
|
||||||
|
seen.add((en, ru))
|
||||||
|
print("%5d %.2f %-30s -> %s" % (n, sc, en, ru))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
+41
-8
@@ -17,6 +17,12 @@ import sys
|
|||||||
import zipfile
|
import zipfile
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
|
# XML 1.0 запрещает управляющие символы, а из PDF они приезжают регулярно:
|
||||||
|
# у Бейера («Site Reliability Engineering») 33 главы из 45 не разбирались как
|
||||||
|
# XML из-за одного такого знака в начале абзаца. Читалка спотыкается молча,
|
||||||
|
# поэтому чистим на упаковке, а не надеемся на извлечение.
|
||||||
|
BAD_XML = re.compile(r"[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]")
|
||||||
|
|
||||||
H1 = re.compile(r"^<h1[^>]*>(.*?)</h1>$")
|
H1 = re.compile(r"^<h1[^>]*>(.*?)</h1>$")
|
||||||
IMG = re.compile(r'<img src="images/([^"]+)"')
|
IMG = re.compile(r'<img src="images/([^"]+)"')
|
||||||
NS = 'xmlns="http://www.w3.org/1999/xhtml"'
|
NS = 'xmlns="http://www.w3.org/1999/xhtml"'
|
||||||
@@ -48,8 +54,25 @@ def split_chapters(lines):
|
|||||||
return [(t, b) for t, b in chapters if any(x.strip() for x in b)]
|
return [(t, b) for t, b in chapters if any(x.strip() for x in b)]
|
||||||
|
|
||||||
|
|
||||||
|
def series_meta(meta):
|
||||||
|
"""Цикл и номер в нём. Пишем в двух видах: calibre-совместимый meta читают
|
||||||
|
почти все читалки, belongs-to-collection — требование EPUB 3."""
|
||||||
|
name = meta.get("series")
|
||||||
|
if not name:
|
||||||
|
return ""
|
||||||
|
idx = meta.get("series_index") or ""
|
||||||
|
out = ['\n <meta name="calibre:series" content="%s"/>' % html.escape(name)]
|
||||||
|
if idx:
|
||||||
|
out.append('<meta name="calibre:series_index" content="%s"/>' % html.escape(str(idx)))
|
||||||
|
out.append('<meta property="belongs-to-collection" id="c01">%s</meta>' % html.escape(name))
|
||||||
|
out.append('<meta refines="#c01" property="collection-type">series</meta>')
|
||||||
|
if idx:
|
||||||
|
out.append('<meta refines="#c01" property="group-position">%s</meta>' % html.escape(str(idx)))
|
||||||
|
return "\n ".join(out)
|
||||||
|
|
||||||
|
|
||||||
def build(src, out, meta):
|
def build(src, out, meta):
|
||||||
text = src.read_text(encoding="utf-8")
|
text = BAD_XML.sub("", src.read_text(encoding="utf-8"))
|
||||||
css = re.search(r"<style>(.*?)</style>", text, re.S)
|
css = re.search(r"<style>(.*?)</style>", text, re.S)
|
||||||
css = css.group(1) if css else ""
|
css = css.group(1) if css else ""
|
||||||
body = re.search(r"<body>(.*)</body>", text, re.S)
|
body = re.search(r"<body>(.*)</body>", text, re.S)
|
||||||
@@ -60,7 +83,10 @@ def build(src, out, meta):
|
|||||||
|
|
||||||
imgdir = src.parent / "images"
|
imgdir = src.parent / "images"
|
||||||
used, files = [], []
|
used, files = [], []
|
||||||
with zipfile.ZipFile(out, "w") as z:
|
# ZIP_DEFLATED обязателен явно: у zipfile умолчание — ZIP_STORED, и книга
|
||||||
|
# выходит втрое толще (у Брикмана 1.42 МБ против 0.43 при том же тексте).
|
||||||
|
# Первым файлом всё равно идёт несжатый mimetype — этого требует формат.
|
||||||
|
with zipfile.ZipFile(out, "w", zipfile.ZIP_DEFLATED) as z:
|
||||||
# mimetype обязан идти первым и без сжатия — иначе часть читалок
|
# mimetype обязан идти первым и без сжатия — иначе часть читалок
|
||||||
# не опознаёт архив как EPUB
|
# не опознаёт архив как EPUB
|
||||||
z.writestr(zipfile.ZipInfo("mimetype"), "application/epub+zip",
|
z.writestr(zipfile.ZipInfo("mimetype"), "application/epub+zip",
|
||||||
@@ -82,8 +108,11 @@ def build(src, out, meta):
|
|||||||
CHAPTER % (NS, html.escape(title), chunk))
|
CHAPTER % (NS, html.escape(title), chunk))
|
||||||
for img in used:
|
for img in used:
|
||||||
z.write(imgdir / img, "OEBPS/images/" + img)
|
z.write(imgdir / img, "OEBPS/images/" + img)
|
||||||
if meta["cover"] and (imgdir / meta["cover"]).exists():
|
# Обложка бывает и среди картинок текста — второй раз её класть нельзя,
|
||||||
z.write(imgdir / meta["cover"], "OEBPS/images/" + meta["cover"])
|
# zip примет дубликат имени, а читалки на такой архив ругаются.
|
||||||
|
cover = meta["cover"] if meta["cover"] not in used else ""
|
||||||
|
if cover and (imgdir / cover).exists():
|
||||||
|
z.write(imgdir / cover, "OEBPS/images/" + cover)
|
||||||
|
|
||||||
items = ['<item id="css" href="style.css" media-type="text/css"/>',
|
items = ['<item id="css" href="style.css" media-type="text/css"/>',
|
||||||
'<item id="ncx" href="toc.ncx" media-type="application/x-dtbncx+xml"/>']
|
'<item id="ncx" href="toc.ncx" media-type="application/x-dtbncx+xml"/>']
|
||||||
@@ -94,7 +123,7 @@ def build(src, out, meta):
|
|||||||
spine.append('<itemref idref="c%d"/>' % i)
|
spine.append('<itemref idref="c%d"/>' % i)
|
||||||
mime = {".png": "image/png", ".jpg": "image/jpeg", ".jpeg": "image/jpeg",
|
mime = {".png": "image/png", ".jpg": "image/jpeg", ".jpeg": "image/jpeg",
|
||||||
".gif": "image/gif", ".svg": "image/svg+xml"}
|
".gif": "image/gif", ".svg": "image/svg+xml"}
|
||||||
for i, img in enumerate(used + ([meta["cover"]] if meta["cover"] else [])):
|
for i, img in enumerate(used + ([cover] if cover else [])):
|
||||||
items.append('<item id="i%d" href="images/%s" media-type="%s"%s/>'
|
items.append('<item id="i%d" href="images/%s" media-type="%s"%s/>'
|
||||||
% (i, img, mime.get(Path(img).suffix.lower(), "image/jpeg"),
|
% (i, img, mime.get(Path(img).suffix.lower(), "image/jpeg"),
|
||||||
' properties="cover-image"' if img == meta["cover"] else ""))
|
' properties="cover-image"' if img == meta["cover"] else ""))
|
||||||
@@ -104,12 +133,12 @@ def build(src, out, meta):
|
|||||||
<dc:identifier id="bid">%s</dc:identifier>
|
<dc:identifier id="bid">%s</dc:identifier>
|
||||||
<dc:title>%s</dc:title>
|
<dc:title>%s</dc:title>
|
||||||
<dc:creator>%s</dc:creator>
|
<dc:creator>%s</dc:creator>
|
||||||
<dc:language>%s</dc:language>
|
<dc:language>%s</dc:language>%s
|
||||||
</metadata>
|
</metadata>
|
||||||
<manifest>%s</manifest>
|
<manifest>%s</manifest>
|
||||||
<spine toc="ncx">%s</spine>
|
<spine toc="ncx">%s</spine>
|
||||||
</package>""" % (html.escape(meta["id"]), html.escape(meta["title"]),
|
</package>""" % (html.escape(meta["id"]), html.escape(meta["title"]),
|
||||||
html.escape(meta["author"]), meta["lang"],
|
html.escape(meta["author"]), meta["lang"], series_meta(meta),
|
||||||
"\n ".join(items), "\n ".join(spine)))
|
"\n ".join(items), "\n ".join(spine)))
|
||||||
|
|
||||||
nav = "\n".join(
|
nav = "\n".join(
|
||||||
@@ -133,9 +162,13 @@ def main():
|
|||||||
ap.add_argument("--author", default="")
|
ap.add_argument("--author", default="")
|
||||||
ap.add_argument("--lang", default="ru")
|
ap.add_argument("--lang", default="ru")
|
||||||
ap.add_argument("--cover", default="cover.jpg", help="имя файла в images/")
|
ap.add_argument("--cover", default="cover.jpg", help="имя файла в images/")
|
||||||
|
ap.add_argument("--series", default="", help="название цикла")
|
||||||
|
ap.add_argument("--series-index", default="", help="номер в цикле")
|
||||||
args = ap.parse_args()
|
args = ap.parse_args()
|
||||||
meta = {"title": args.title, "author": args.author, "lang": args.lang,
|
meta = {"title": args.title, "author": args.author, "lang": args.lang,
|
||||||
"cover": args.cover, "id": "urn:uuid:" + args.output.stem.replace(" ", "-")}
|
"cover": args.cover, "series": args.series,
|
||||||
|
"series_index": args.series_index,
|
||||||
|
"id": "urn:uuid:" + args.output.stem.replace(" ", "-")}
|
||||||
n, imgs = build(args.source, args.output, meta)
|
n, imgs = build(args.source, args.output, meta)
|
||||||
size = args.output.stat().st_size / 2**20
|
size = args.output.stat().st_size / 2**20
|
||||||
print("%s: глав %d, картинок %d, %.1f МБ" % (args.output.name, n, imgs, size))
|
print("%s: глав %d, картинок %d, %.1f МБ" % (args.output.name, n, imgs, size))
|
||||||
|
|||||||
+199
-9
@@ -39,6 +39,30 @@ MAX_IMG_SHARE = 0.6 # доля площади страницы: больше
|
|||||||
# Запасное опознание листинга, когда издатель вычистил и имена шрифтов, и флаги:
|
# Запасное опознание листинга, когда издатель вычистил и имена шрифтов, и флаги:
|
||||||
# кегль мельче основного текста плюс пунктуация, которой в прозе не бывает.
|
# кегль мельче основного текста плюс пунктуация, которой в прозе не бывает.
|
||||||
CODEY = re.compile(r"[{}();\[\]<>=*/#]|::|->")
|
CODEY = re.compile(r"[{}();\[\]<>=*/#]|::|->")
|
||||||
|
ROTATED = 0.1 # синус угла строки: больше — текст повёрнут, а не набран по горизонтали
|
||||||
|
HEAD_BAND = 0.07 # доля высоты полосы сверху, где лежит колонтитул, а не текст
|
||||||
|
SOFT_HYPHEN = "\u00ad"
|
||||||
|
OCR = False
|
||||||
|
|
||||||
|
|
||||||
|
def _image_png(raw):
|
||||||
|
"""Байты картинки блока -> настоящий PNG.
|
||||||
|
|
||||||
|
b["ext"] врёт: на «Запускаем Ansible» 3-го изд. блоки с ext="png" и
|
||||||
|
colorspace=4 на деле — CMYK JPEG (magic FFD8, подтверждено `file`), и
|
||||||
|
PNG-инструменты (pngquant) отказываются их декодировать с именем .png.
|
||||||
|
Вместо доверия заявленному формату декодируем через Pixmap (MuPDF сам
|
||||||
|
распознаёт реальный формат по данным) и всегда отдаём RGB PNG — тогда
|
||||||
|
имя файла (.png) не расходится с содержимым и CMYK не ловит инверсию
|
||||||
|
цвета в читалках, которые не ждут Adobe-JPEG с четырьмя каналами.
|
||||||
|
"""
|
||||||
|
try:
|
||||||
|
pix = pymupdf.Pixmap(raw)
|
||||||
|
if pix.colorspace is not None and pix.colorspace.n not in (1, 3):
|
||||||
|
pix = pymupdf.Pixmap(pymupdf.csRGB, pix)
|
||||||
|
return pix.tobytes("png")
|
||||||
|
except Exception:
|
||||||
|
return raw # не смогли декодировать — пишем как есть, лучше так, чем ничего
|
||||||
|
|
||||||
|
|
||||||
def style(font, flags=0):
|
def style(font, flags=0):
|
||||||
@@ -50,10 +74,40 @@ def style(font, flags=0):
|
|||||||
|
|
||||||
def is_mono(s):
|
def is_mono(s):
|
||||||
"""Моноширинный шрифт = листинг кода. Признак из флагов, а не из имени:
|
"""Моноширинный шрифт = листинг кода. Признак из флагов, а не из имени:
|
||||||
имена у каждого издательства свои, флаг одинаковый."""
|
имена у каждого издательства свои, флаг одинаковый.
|
||||||
|
|
||||||
|
Исключение — невидимый слой Tesseract: он весь набран GlyphLessFont с
|
||||||
|
выставленным моно-битом, и без этой проверки книга целиком опознаётся как
|
||||||
|
один сплошной листинг (измерено на Ansible после OCR: pre 10498, p 0).
|
||||||
|
В таком слое листинги ловятся кеглем и пунктуацией, см. block_kind().
|
||||||
|
"""
|
||||||
|
if "glyphless" in s["font"].lower():
|
||||||
|
return False
|
||||||
return bool(s["flags"] & MONO)
|
return bool(s["flags"] & MONO)
|
||||||
|
|
||||||
|
|
||||||
|
def ocr_layer(doc, step=23):
|
||||||
|
"""Текст сделан OCR-ом: у всего слоя один невидимый шрифт. Кегль в таком
|
||||||
|
слое подгоняется под рамку слова и гуляет (9.0, 9.2, 9.4 в одном абзаце),
|
||||||
|
поэтому дальше он округляется до целого — иначе гистограмма размазывается
|
||||||
|
и основной кегль книги определяется неверно."""
|
||||||
|
glyphless = total = 0
|
||||||
|
for pno in range(1, doc.page_count, step):
|
||||||
|
for b in doc[pno].get_text("dict")["blocks"]:
|
||||||
|
for l in b.get("lines", []):
|
||||||
|
for s in l["spans"]:
|
||||||
|
n = len(s["text"])
|
||||||
|
total += n
|
||||||
|
if "glyphless" in s["font"].lower():
|
||||||
|
glyphless += n
|
||||||
|
return total > 0 and glyphless / total > 0.9
|
||||||
|
|
||||||
|
|
||||||
|
def size_of(span):
|
||||||
|
"""Кегль спана. В OCR-слое округляется до целого: см. ocr_layer()."""
|
||||||
|
return round(span["size"]) if OCR else round(span["size"], 1)
|
||||||
|
|
||||||
|
|
||||||
def profile(doc, step=7):
|
def profile(doc, step=7):
|
||||||
"""(кегль текста, x0 красной строки, кегль глав) по выборке страниц.
|
"""(кегль текста, x0 красной строки, кегль глав) по выборке страниц.
|
||||||
|
|
||||||
@@ -63,6 +117,7 @@ def profile(doc, step=7):
|
|||||||
"""
|
"""
|
||||||
size_chars = collections.Counter()
|
size_chars = collections.Counter()
|
||||||
head_blocks = collections.Counter()
|
head_blocks = collections.Counter()
|
||||||
|
head_len = {}
|
||||||
x0 = collections.Counter()
|
x0 = collections.Counter()
|
||||||
for pno in range(1, doc.page_count, step):
|
for pno in range(1, doc.page_count, step):
|
||||||
for b in doc[pno].get_text("dict")["blocks"]:
|
for b in doc[pno].get_text("dict")["blocks"]:
|
||||||
@@ -73,16 +128,28 @@ def profile(doc, step=7):
|
|||||||
for l in b["lines"]:
|
for l in b["lines"]:
|
||||||
x0[round(l["bbox"][0])] += 1
|
x0[round(l["bbox"][0])] += 1
|
||||||
for sp in spans:
|
for sp in spans:
|
||||||
size_chars[round(sp["size"], 1)] += len(sp["text"])
|
size_chars[size_of(sp)] += len(sp["text"])
|
||||||
top = max(spans, key=lambda sp: len(sp["text"]))
|
top = max(spans, key=lambda sp: len(sp["text"]))
|
||||||
head_blocks[round(top["size"], 1)] += 1
|
head_blocks[size_of(top)] += 1
|
||||||
|
head_len.setdefault(size_of(top), []).append(
|
||||||
|
"".join(sp["text"] for sp in spans).strip())
|
||||||
if not size_chars:
|
if not size_chars:
|
||||||
sys.exit("в PDF нет извлекаемого текста — нужен OCR, этот скрипт не поможет")
|
sys.exit("в PDF нет извлекаемого текста — нужен OCR, этот скрипт не поможет")
|
||||||
body = size_chars.most_common(1)[0][0]
|
body = size_chars.most_common(1)[0][0]
|
||||||
# кандидаты в главы: крупнее текста и встречаются не единожды (единичный
|
# кандидаты в главы: крупнее текста и встречаются не единожды (единичный
|
||||||
# размер — это титул, а не уровень заголовка)
|
# размер — это титул, а не уровень заголовка)
|
||||||
|
# Ещё условие: у кандидата должны быть слова, а не цифры. У Нейгарда номер
|
||||||
|
# главы набран кеглем 100 отдельным блоком («1», «2», …), и без проверки
|
||||||
|
# главой объявлялся он, а настоящие заголовки в 30 пунктов уезжали в
|
||||||
|
# разделы — оглавление выходило пустым. Отбор идёт по наличию букв, а не по
|
||||||
|
# длине: «Preface» — семь знаков, и порог по длине отсекал бы и его.
|
||||||
|
def _wordy(sz):
|
||||||
|
txts = head_len.get(sz) or []
|
||||||
|
letters = sum(1 for t in txts if re.search(r"[^\W\d_]", t))
|
||||||
|
return bool(txts) and letters * 2 >= len(txts)
|
||||||
|
|
||||||
cand = sorted((sz for sz, n in head_blocks.items()
|
cand = sorted((sz for sz, n in head_blocks.items()
|
||||||
if sz >= body * 1.15 and n >= 3), reverse=True)
|
if sz >= body * 1.15 and n >= 3 and _wordy(sz)), reverse=True)
|
||||||
h1 = cand[0] if cand else body * 1.55
|
h1 = cand[0] if cand else body * 1.55
|
||||||
left = min(x for x, n in x0.most_common(4))
|
left = min(x for x, n in x0.most_common(4))
|
||||||
return body, left + max(4, body * 0.6), h1
|
return body, left + max(4, body * 0.6), h1
|
||||||
@@ -111,13 +178,18 @@ def looks_like_code(spans):
|
|||||||
return len(text) > 8 and len(CODEY.findall(text)) / len(text) > 0.03
|
return len(text) > 8 and len(CODEY.findall(text)) / len(text) > 0.03
|
||||||
|
|
||||||
|
|
||||||
|
def is_rotated(b):
|
||||||
|
"""Строки блока набраны не по горизонтали (боковая врезка, подпись к таблице)."""
|
||||||
|
return bool(b.get("lines")) and any(abs(l["dir"][1]) > ROTATED for l in b["lines"])
|
||||||
|
|
||||||
|
|
||||||
def block_kind(b, body, h1_size):
|
def block_kind(b, body, h1_size):
|
||||||
"""(tag, css class) по кеглю блока относительно основного текста книги."""
|
"""(tag, css class) по кеглю блока относительно основного текста книги."""
|
||||||
spans = [s for l in b["lines"] for s in l["spans"] if s["text"].strip()]
|
spans = [s for l in b["lines"] for s in l["spans"] if s["text"].strip()]
|
||||||
if not spans:
|
if not spans:
|
||||||
return None
|
return None
|
||||||
top = max(spans, key=lambda s: len(s["text"]))
|
top = max(spans, key=lambda s: len(s["text"]))
|
||||||
sz = top["size"]
|
sz = size_of(top)
|
||||||
if all(is_mono(sp) for sp in spans):
|
if all(is_mono(sp) for sp in spans):
|
||||||
return ("pre", None)
|
return ("pre", None)
|
||||||
if sz <= body * 0.93 and looks_like_code(spans):
|
if sz <= body * 0.93 and looks_like_code(spans):
|
||||||
@@ -126,7 +198,11 @@ def block_kind(b, body, h1_size):
|
|||||||
return ("h1", "title")
|
return ("h1", "title")
|
||||||
if sz >= h1_size * 0.97:
|
if sz >= h1_size * 0.97:
|
||||||
return ("h1", None)
|
return ("h1", None)
|
||||||
if sz >= body * 1.15:
|
# 1.10, а не 1.15: у Вернона подзаголовок набран 17.2 при тексте 15.0, то
|
||||||
|
# есть ровно на пять сотых ниже прежнего порога — и все 85 разделов книги
|
||||||
|
# уезжали в прозу. Тот же множитель, что у split_by_size, иначе разрез и
|
||||||
|
# классификация расходятся.
|
||||||
|
if sz >= body * 1.10:
|
||||||
return ("h2", None)
|
return ("h2", None)
|
||||||
return ("p", None)
|
return ("p", None)
|
||||||
|
|
||||||
@@ -137,7 +213,10 @@ def block_text(b, pre=False):
|
|||||||
# слова. Ни склейки строк, ни де-дефисации здесь быть не должно.
|
# слова. Ни склейки строк, ни де-дефисации здесь быть не должно.
|
||||||
rows = ["".join(span_html(s, plain=True) for s in l["spans"])
|
rows = ["".join(span_html(s, plain=True) for s in l["spans"])
|
||||||
for l in b["lines"]]
|
for l in b["lines"]]
|
||||||
return "\n".join(rows).rstrip()
|
# В настоящем коде мягкого переноса не бывает: если он тут есть, блок
|
||||||
|
# опознан как листинг ошибочно (обычно это таблица опций), и слово надо
|
||||||
|
# склеить, а не оставить разорванным переводом строки.
|
||||||
|
return "\n".join(rows).replace(SOFT_HYPHEN + "\n", "").rstrip()
|
||||||
out = []
|
out = []
|
||||||
for i, l in enumerate(b["lines"]):
|
for i, l in enumerate(b["lines"]):
|
||||||
line = "".join(span_html(s) for s in l["spans"])
|
line = "".join(span_html(s) for s in l["spans"])
|
||||||
@@ -145,7 +224,12 @@ def block_text(b, pre=False):
|
|||||||
continue
|
continue
|
||||||
if out:
|
if out:
|
||||||
prev = out[-1]
|
prev = out[-1]
|
||||||
if prev.endswith("-") and not prev.endswith("--"):
|
# Русские издания переносят мягким дефисом U+00AD, а не обычным:
|
||||||
|
# без этой ветки строки склеиваются через пробел и слово рвётся
|
||||||
|
# пополам («введен ные»). Измерено на Шоттсе (Питер, 2020).
|
||||||
|
if prev.endswith(SOFT_HYPHEN):
|
||||||
|
out[-1] = prev[:-1]
|
||||||
|
elif prev.endswith("-") and not prev.endswith("--"):
|
||||||
out[-1] = prev[:-1] # de-hyphenate
|
out[-1] = prev[:-1] # de-hyphenate
|
||||||
else:
|
else:
|
||||||
out.append(" ")
|
out.append(" ")
|
||||||
@@ -154,11 +238,33 @@ def block_text(b, pre=False):
|
|||||||
txt = re.sub(r"</i>(\s*)<i>", r"\1", txt)
|
txt = re.sub(r"</i>(\s*)<i>", r"\1", txt)
|
||||||
txt = re.sub(r"</b>(\s*)<b>", r"\1", txt)
|
txt = re.sub(r"</b>(\s*)<b>", r"\1", txt)
|
||||||
txt = re.sub(r"(?<=\S) {2,}(?=\S)", " ", txt)
|
txt = re.sub(r"(?<=\S) {2,}(?=\S)", " ", txt)
|
||||||
|
# Мягкие переносы внутри строки читалке не нужны и ломают поиск по тексту.
|
||||||
|
txt = txt.replace(SOFT_HYPHEN, "")
|
||||||
return txt.strip()
|
return txt.strip()
|
||||||
|
|
||||||
|
|
||||||
|
if SRC.name == "--selftest": # python3 pdf2html.py --selftest --selftest
|
||||||
|
# Повёрнутый текст ловится не по буквам, а по направлению строки: у
|
||||||
|
# «aoHegoduenueArdng» из боковой врезки Брикмана структура слова приличная.
|
||||||
|
assert is_rotated({"lines": [{"dir": (0.0, 1.0)}]})
|
||||||
|
assert is_rotated({"lines": [{"dir": (0.0, -1.0)}]})
|
||||||
|
assert not is_rotated({"lines": [{"dir": (1.0, 0.0)}, {"dir": (1.0, 0.01)}]})
|
||||||
|
for bad in ("30HWod1Q", "o|qisuy", "1 2 3 4", "|", "0Q1 v9 8|", "В",
|
||||||
|
"9120905", "LF. FF. FT. TIT", "ee. лов. о", "лов", "TIT. ITITIT", "TT",
|
||||||
|
"‘ая. ®No TERRAFORM", "юн"):
|
||||||
|
assert bookhtml.looks_garbled(bad), bad
|
||||||
|
for good in ("Глава 1. Зачем нужен Terraform", "C++11 и лямбды", "Python3",
|
||||||
|
"Модули", "Часть II. Основы", "Что дальше?", "Часть III", "API",
|
||||||
|
"AI", "Резюме", "06 авторе", "systemd в деталях"):
|
||||||
|
assert not bookhtml.looks_garbled(good), good
|
||||||
|
print("selftest OK")
|
||||||
|
sys.exit()
|
||||||
|
|
||||||
doc = pymupdf.open(SRC)
|
doc = pymupdf.open(SRC)
|
||||||
|
OCR = ocr_layer(doc)
|
||||||
BODY, INDENT_X, H1_SIZE = profile(doc)
|
BODY, INDENT_X, H1_SIZE = profile(doc)
|
||||||
|
if OCR:
|
||||||
|
print("слой распознан OCR-ом: моно-признак отключён, кегль округляется")
|
||||||
print("кегль текста %.1f, глав %.1f, красная строка от x0=%.0f"
|
print("кегль текста %.1f, глав %.1f, красная строка от x0=%.0f"
|
||||||
% (BODY, H1_SIZE, INDENT_X))
|
% (BODY, H1_SIZE, INDENT_X))
|
||||||
parts = []
|
parts = []
|
||||||
@@ -168,12 +274,73 @@ open_para = None # (tag, cls, text) still collecting
|
|||||||
pix = doc[0].get_pixmap(dpi=150)
|
pix = doc[0].get_pixmap(dpi=150)
|
||||||
pix.save(IMGDIR / "cover.jpg")
|
pix.save(IMGDIR / "cover.jpg")
|
||||||
|
|
||||||
|
def running_heads(doc, band, limit=80, min_pages=5):
|
||||||
|
"""Тексты, повторяющиеся в верхнем поле на многих полосах.
|
||||||
|
|
||||||
|
Позиционный признак в одиночку рубит и настоящие заголовки: у книг,
|
||||||
|
свёрстанных calibre, глава начинается ровно с верха полосы и короче 80
|
||||||
|
знаков. У Вернона так пропали 14 заголовков из 15. Повтор по десяткам
|
||||||
|
страниц — то, чем колонтитул отличается от заголовка.
|
||||||
|
"""
|
||||||
|
seen = collections.Counter()
|
||||||
|
for page in doc:
|
||||||
|
for b in page.get_text("dict")["blocks"]:
|
||||||
|
if b.get("type") == 1 or "lines" not in b:
|
||||||
|
continue
|
||||||
|
if b["bbox"][1] >= page.rect.height * band:
|
||||||
|
continue
|
||||||
|
txt = "".join(sp["text"] for l in b["lines"] for sp in l["spans"])
|
||||||
|
txt = re.sub(r"\d+", "", txt).strip().lower()
|
||||||
|
if txt and len(txt) < limit:
|
||||||
|
seen[txt] += 1
|
||||||
|
return {t for t, n in seen.items() if n >= min_pages}
|
||||||
|
|
||||||
|
|
||||||
|
def split_by_size(b, body):
|
||||||
|
"""Разрезать блок, где шапка набрана крупнее следующего за ней текста.
|
||||||
|
|
||||||
|
У Вернона 85 подзаголовков лежат в одном блоке с первым абзацем раздела:
|
||||||
|
строка 17.2 и сразу за ней 15.0. Без разреза они становятся частью абзаца,
|
||||||
|
оглавление теряет второй уровень, а текст начинается с заголовка без точки.
|
||||||
|
"""
|
||||||
|
lines = [l for l in b.get("lines", [])
|
||||||
|
if any(sp["text"].strip() for sp in l["spans"])]
|
||||||
|
if len(lines) < 2:
|
||||||
|
return [b]
|
||||||
|
big = [max(size_of(sp) for sp in l["spans"] if sp["text"].strip())
|
||||||
|
> body * 1.12 for l in lines]
|
||||||
|
if not big[0] or all(big):
|
||||||
|
return [b]
|
||||||
|
cut = big.index(False)
|
||||||
|
if not any(big[:cut]) or any(big[cut:]):
|
||||||
|
return [b] # разрез только когда крупное строго сверху
|
||||||
|
head = dict(b, lines=lines[:cut],
|
||||||
|
bbox=(b["bbox"][0], b["bbox"][1], b["bbox"][2],
|
||||||
|
lines[cut - 1]["bbox"][3]))
|
||||||
|
rest = dict(b, lines=lines[cut:],
|
||||||
|
bbox=(b["bbox"][0], lines[cut]["bbox"][1], b["bbox"][2],
|
||||||
|
b["bbox"][3]))
|
||||||
|
return [head, rest]
|
||||||
|
|
||||||
|
|
||||||
|
HEADS = running_heads(doc, HEAD_BAND)
|
||||||
|
print("колонтитулов опознано по повтору:", len(HEADS))
|
||||||
|
|
||||||
img_n = 0
|
img_n = 0
|
||||||
for pno, page in enumerate(doc):
|
for pno, page in enumerate(doc):
|
||||||
if pno == 0:
|
if pno == 0:
|
||||||
continue
|
continue
|
||||||
blocks = sorted(page.get_text("dict")["blocks"], key=lambda b: b["bbox"][1])
|
blocks = sorted(page.get_text("dict")["blocks"], key=lambda b: b["bbox"][1])
|
||||||
|
blocks = [part for b in blocks
|
||||||
|
for part in (split_by_size(b, BODY) if b.get("type") != 1 else [b])]
|
||||||
for bi, b in enumerate(blocks):
|
for bi, b in enumerate(blocks):
|
||||||
|
# Повёрнутый на 90° текст: боковые врезки и подписи к таблицам. В книге
|
||||||
|
# их единицы, а распознаются они всегда в кашу («aoHegoduenueArdng») и
|
||||||
|
# лезут в заголовки, потому что кегль у них крупный. Замерено у Брикмана:
|
||||||
|
# 105 строк из 17 300, все до одной — брак. Настоящая повёрнутая таблица
|
||||||
|
# тоже отсеется, но её всё равно нечем показать в потоковой вёрстке.
|
||||||
|
if is_rotated(b):
|
||||||
|
continue
|
||||||
if b["type"] == 1:
|
if b["type"] == 1:
|
||||||
x0, y0, x1, y1 = b["bbox"]
|
x0, y0, x1, y1 = b["bbox"]
|
||||||
if x1 - x0 < MIN_IMG or y1 - y0 < MIN_IMG:
|
if x1 - x0 < MIN_IMG or y1 - y0 < MIN_IMG:
|
||||||
@@ -183,7 +350,7 @@ for pno, page in enumerate(doc):
|
|||||||
continue # подложка отсканированной полосы: текст берём из OCR-слоя
|
continue # подложка отсканированной полосы: текст берём из OCR-слоя
|
||||||
img_n += 1
|
img_n += 1
|
||||||
name = "img%02d.png" % img_n
|
name = "img%02d.png" % img_n
|
||||||
(IMGDIR / name).write_bytes(b["image"])
|
(IMGDIR / name).write_bytes(_image_png(b["image"]))
|
||||||
if open_para:
|
if open_para:
|
||||||
parts.append(open_para)
|
parts.append(open_para)
|
||||||
open_para = None
|
open_para = None
|
||||||
@@ -200,6 +367,17 @@ for pno, page in enumerate(doc):
|
|||||||
txt = block_text(b, pre=(tag == "pre"))
|
txt = block_text(b, pre=(tag == "pre"))
|
||||||
if not txt:
|
if not txt:
|
||||||
continue
|
continue
|
||||||
|
# Колонтитул: верхние 7% полосы — поле, а не текст. У Брикмана так
|
||||||
|
# отсеивается 21 блок крупного кегля из 511, все до одного — шапки.
|
||||||
|
# ponytail: признак позиционный. Надёжнее — текст, повторяющийся на
|
||||||
|
# десятках полос, но за это платить вторым проходом по документу.
|
||||||
|
if (b["bbox"][1] < page.rect.height * HEAD_BAND and len(txt) < 80
|
||||||
|
and re.sub(r"\d+", "", bookhtml.strip_tags(txt)
|
||||||
|
if hasattr(bookhtml, "strip_tags")
|
||||||
|
else re.sub(r"<[^>]+>", "", txt)).strip().lower() in HEADS):
|
||||||
|
continue
|
||||||
|
if tag in ("h1", "h2") and bookhtml.looks_garbled(txt):
|
||||||
|
tag, cls = "p", None
|
||||||
indented = b["lines"][0]["bbox"][0] >= INDENT_X
|
indented = b["lines"][0]["bbox"][0] >= INDENT_X
|
||||||
cont = (
|
cont = (
|
||||||
open_para
|
open_para
|
||||||
@@ -213,6 +391,13 @@ for pno, page in enumerate(doc):
|
|||||||
join = "" if open_para[2].endswith("-") else " "
|
join = "" if open_para[2].endswith("-") else " "
|
||||||
open_para = (tag, cls, open_para[2] + join + txt, open_para[3])
|
open_para = (tag, cls, open_para[2] + join + txt, open_para[3])
|
||||||
continue
|
continue
|
||||||
|
# Соседние листинги склеиваются в один блок: OCR-слой отдаёт каждую
|
||||||
|
# строку кода отдельным блоком, и без склейки листинг рассыпается на
|
||||||
|
# десяток однострочных <pre> подряд. Строки внутри блока значимы,
|
||||||
|
# поэтому соединяются переводом строки, а не пробелом.
|
||||||
|
if OCR and tag == "pre" and open_para and open_para[0] == "pre":
|
||||||
|
open_para = (tag, cls, open_para[2] + "\n" + txt, open_para[3])
|
||||||
|
continue
|
||||||
if open_para:
|
if open_para:
|
||||||
parts.append(open_para)
|
parts.append(open_para)
|
||||||
open_para = (tag, cls, txt, at_top)
|
open_para = (tag, cls, txt, at_top)
|
||||||
@@ -243,6 +428,11 @@ for part in parts:
|
|||||||
merged.append(part)
|
merged.append(part)
|
||||||
parts = merged
|
parts = merged
|
||||||
|
|
||||||
|
# Проверка повторяется после склейки: по отдельности «LF», «FF» и «TIT» проходят
|
||||||
|
# как аббревиатуры, а склеенные в один заголовок — это обрывок таблицы.
|
||||||
|
parts = [(("p", None, x) if t in ("h1", "h2") and bookhtml.looks_garbled(x)
|
||||||
|
else (t, c, x)) for t, c, x in parts]
|
||||||
|
|
||||||
doc_html = bookhtml.document(doc.metadata.get("title") or SRC.stem, parts)
|
doc_html = bookhtml.document(doc.metadata.get("title") or SRC.stem, parts)
|
||||||
# merge tag runs split at page/line boundaries. Пробелы здесь уже не трогаем:
|
# merge tag runs split at page/line boundaries. Пробелы здесь уже не трогаем:
|
||||||
# внутри <pre> они значимы, а в прозе схлопнуты в block_text.
|
# внутри <pre> они значимы, а в прозе схлопнуты в block_text.
|
||||||
|
|||||||
@@ -57,6 +57,47 @@ def test_pack(tmp: Path):
|
|||||||
assert "Текст второй главы" not in ch1, "главы не должны склеиваться"
|
assert "Текст второй главы" not in ch1, "главы не должны склеиваться"
|
||||||
|
|
||||||
|
|
||||||
|
def test_cover_used_in_text_is_not_duplicated(tmp: Path):
|
||||||
|
"""Обложка, встречающаяся и в тексте, не должна попадать в архив дважды."""
|
||||||
|
src = tmp / "book.html"
|
||||||
|
src.write_text("<h1>Глава</h1>\n"
|
||||||
|
'<p class="figure"><img src="images/cover.jpg"/></p>\n'
|
||||||
|
"<p>Текст главы.</p>", encoding="utf-8")
|
||||||
|
(tmp / "images").mkdir()
|
||||||
|
(tmp / "images" / "cover.jpg").write_bytes(b"\xff\xd8\xff")
|
||||||
|
out = tmp / "book.epub"
|
||||||
|
pack_epub.build(src, out, {"title": "t", "author": "a", "lang": "ru",
|
||||||
|
"cover": "cover.jpg", "id": "i"})
|
||||||
|
names = zipfile.ZipFile(out).namelist()
|
||||||
|
assert names.count("OEBPS/images/cover.jpg") == 1, names
|
||||||
|
opf = zipfile.ZipFile(out).read("OEBPS/content.opf").decode()
|
||||||
|
assert opf.count('href="images/cover.jpg"') == 1, opf
|
||||||
|
|
||||||
|
|
||||||
|
def test_control_chars_stripped(tmp: Path):
|
||||||
|
"""Управляющий знак из PDF делает главу неразбираемой как XML.
|
||||||
|
|
||||||
|
У Бейера («Site Reliability Engineering») так падали 33 главы из 45:
|
||||||
|
один \x02 в начале абзаца, и читалка молча спотыкается на файле.
|
||||||
|
"""
|
||||||
|
import xml.etree.ElementTree as ET
|
||||||
|
|
||||||
|
src = tmp / "book.html"
|
||||||
|
src.write_text(BOOK.replace("<p>Текст первой главы.</p>",
|
||||||
|
"<p>\x02Текст первой главы.</p>"),
|
||||||
|
encoding="utf-8")
|
||||||
|
(tmp / "images").mkdir()
|
||||||
|
(tmp / "images" / "fig.png").write_bytes(b"\x89PNG\r\n\x1a\n")
|
||||||
|
(tmp / "images" / "cover.jpg").write_bytes(b"\xff\xd8\xff")
|
||||||
|
out = tmp / "book.epub"
|
||||||
|
pack_epub.build(src, out, {"title": "T", "author": "A", "lang": "ru",
|
||||||
|
"cover": "cover.jpg", "id": "urn:uuid:test"})
|
||||||
|
z = zipfile.ZipFile(out)
|
||||||
|
for name in z.namelist():
|
||||||
|
if name.endswith((".xhtml", ".opf", ".ncx")):
|
||||||
|
ET.fromstring(z.read(name)) # падает, если знак остался
|
||||||
|
|
||||||
|
|
||||||
def test_no_content(tmp: Path):
|
def test_no_content(tmp: Path):
|
||||||
src = tmp / "empty.html"
|
src = tmp / "empty.html"
|
||||||
src.write_text("<html><body>\n</body></html>", encoding="utf-8")
|
src.write_text("<html><body>\n</body></html>", encoding="utf-8")
|
||||||
@@ -69,7 +110,8 @@ def test_no_content(tmp: Path):
|
|||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
for case in (test_pack, test_no_content):
|
for case in (test_pack, test_cover_used_in_text_is_not_duplicated,
|
||||||
|
test_no_content):
|
||||||
with tempfile.TemporaryDirectory() as d:
|
with tempfile.TemporaryDirectory() as d:
|
||||||
case(Path(d))
|
case(Path(d))
|
||||||
print("OK")
|
print("OK")
|
||||||
|
|||||||
@@ -105,6 +105,59 @@ def test_code_listings_are_left_alone(tmp: Path):
|
|||||||
assert t.verify_output(src, "ru"), "код не должен заваливать приёмку"
|
assert t.verify_output(src, "ru"), "код не должен заваливать приёмку"
|
||||||
|
|
||||||
|
|
||||||
|
def test_junk_translation_falls_back_to_original(tmp: Path):
|
||||||
|
"""Модель иногда возвращает на длинный абзац огрызок «, v», а число абзацев
|
||||||
|
при этом сходится — прежние ворота такое пропускали."""
|
||||||
|
src = tmp / "book.html"
|
||||||
|
long_en = ("A vector is simply a sequence of elements that you can access by "
|
||||||
|
"an index, and it is the workhorse of the standard library. ") * 2
|
||||||
|
src.write_text("\n".join([
|
||||||
|
"<h1>Chapter</h1>",
|
||||||
|
"<p>%s</p>" % long_en,
|
||||||
|
"<p>Second paragraph of the very same chapter, also reasonably long.</p>",
|
||||||
|
]), encoding="utf-8")
|
||||||
|
chapters = t.split_chapters(t.parse_blocks(src))
|
||||||
|
trans = tmp / "translations"
|
||||||
|
trans.mkdir()
|
||||||
|
(trans / "chapter_000_translated.json").write_text(json.dumps(
|
||||||
|
{"number": 0, "paragraphs": ["Глава", ", v",
|
||||||
|
"Второй абзац той же самой главы, тоже достаточно длинный."]},
|
||||||
|
ensure_ascii=False), encoding="utf-8")
|
||||||
|
|
||||||
|
out = tmp / "out.html"
|
||||||
|
t.rebuild(src, chapters, tmp, out)
|
||||||
|
result = out.read_text(encoding="utf-8")
|
||||||
|
assert ", v" not in result, "огрызок не должен попадать в книгу"
|
||||||
|
assert "A vector is simply a sequence" in result, "вместо огрызка нужен оригинал"
|
||||||
|
assert "Второй абзац" in result, "нормальный перевод должен остаться"
|
||||||
|
|
||||||
|
|
||||||
|
def test_untouched_paragraphs_are_counted(tmp: Path):
|
||||||
|
"""Абзац, вернувшийся по-английски, доля кириллицы по документу не ловит:
|
||||||
|
у Страуструпа так осталось 182 упражнения из 10151 блока."""
|
||||||
|
src = tmp / "book.html"
|
||||||
|
en = ("Expanding on what you have learned, write a program that lists the "
|
||||||
|
"instructions for a computer to find the upstairs bedroom. ") * 2
|
||||||
|
ru_src = ("Этот абзац достаточно длинный, чтобы попасть под проверку языка "
|
||||||
|
"и быть переведённым как положено. ") * 2
|
||||||
|
src.write_text("\n".join(["<h1>Chapter</h1>", "<p>%s</p>" % en, "<p>%s</p>" % en]),
|
||||||
|
encoding="utf-8")
|
||||||
|
chapters = t.split_chapters(t.parse_blocks(src))
|
||||||
|
trans = tmp / "translations"
|
||||||
|
trans.mkdir()
|
||||||
|
(trans / "chapter_000_translated.json").write_text(json.dumps(
|
||||||
|
{"number": 0, "paragraphs": ["Глава", en, ru_src]}, ensure_ascii=False),
|
||||||
|
encoding="utf-8")
|
||||||
|
|
||||||
|
out = tmp / "out.html"
|
||||||
|
import io, contextlib
|
||||||
|
buf = io.StringIO()
|
||||||
|
with contextlib.redirect_stdout(buf):
|
||||||
|
t.rebuild(src, chapters, tmp, out)
|
||||||
|
report = buf.getvalue()
|
||||||
|
assert "осталось на языке оригинала: 1" in report, report
|
||||||
|
|
||||||
|
|
||||||
def test_workdir_belongs_to_one_book(tmp: Path):
|
def test_workdir_belongs_to_one_book(tmp: Path):
|
||||||
"""Имена chapter_NNN.json у всех книг одинаковы: чужой workdir склеит
|
"""Имена chapter_NNN.json у всех книг одинаковы: чужой workdir склеит
|
||||||
перевод одной книги с текстом другой."""
|
перевод одной книги с текстом другой."""
|
||||||
@@ -155,6 +208,8 @@ if __name__ == "__main__":
|
|||||||
test_broken_marks_drop_tags()
|
test_broken_marks_drop_tags()
|
||||||
test_language_detection()
|
test_language_detection()
|
||||||
for case in (test_split_and_rebuild, test_code_listings_are_left_alone,
|
for case in (test_split_and_rebuild, test_code_listings_are_left_alone,
|
||||||
|
test_junk_translation_falls_back_to_original,
|
||||||
|
test_untouched_paragraphs_are_counted,
|
||||||
test_workdir_belongs_to_one_book, test_stale_chapters_removed,
|
test_workdir_belongs_to_one_book, test_stale_chapters_removed,
|
||||||
test_verify_output_catches_untranslated):
|
test_verify_output_catches_untranslated):
|
||||||
with tempfile.TemporaryDirectory() as d:
|
with tempfile.TemporaryDirectory() as d:
|
||||||
|
|||||||
+23
-5
@@ -148,9 +148,15 @@ def run_translator(repo, workdir, extracted, workers):
|
|||||||
sys.exit("book_translator завершился с кодом %d" % r.returncode)
|
sys.exit("book_translator завершился с кодом %d" % r.returncode)
|
||||||
|
|
||||||
|
|
||||||
|
# Доля от длины оригинала, ниже которой перевод считается мусором. Совпадение
|
||||||
|
# числа абзацев ничего не гарантирует: модель иногда возвращает на длинный абзац
|
||||||
|
# огрызок вида «, v», и прежние ворота такое пропускали.
|
||||||
|
MIN_LEN_SHARE = 0.25
|
||||||
|
|
||||||
|
|
||||||
def rebuild(src, chapters, workdir, out):
|
def rebuild(src, chapters, workdir, out):
|
||||||
lines = src.read_text(encoding="utf-8").splitlines()
|
lines = src.read_text(encoding="utf-8").splitlines()
|
||||||
broken = missing = translated = 0
|
broken = missing = translated = junk = untouched = 0
|
||||||
for n, ch in enumerate(chapters):
|
for n, ch in enumerate(chapters):
|
||||||
f = workdir / "translations" / ("chapter_%03d_translated.json" % n)
|
f = workdir / "translations" / ("chapter_%03d_translated.json" % n)
|
||||||
if not f.exists():
|
if not f.exists():
|
||||||
@@ -162,15 +168,27 @@ def rebuild(src, chapters, workdir, out):
|
|||||||
% (n, len(paragraphs), len(ch)))
|
% (n, len(paragraphs), len(ch)))
|
||||||
missing += len(ch)
|
missing += len(ch)
|
||||||
continue
|
continue
|
||||||
for (idx, tag, attrs, _), text in zip(ch, paragraphs):
|
for (idx, tag, attrs, original), text in zip(ch, paragraphs):
|
||||||
|
plain_src = re.sub("<[^>]+>", "", original)
|
||||||
|
if len(plain_src) > 120 and len(text) < max(20, len(plain_src) * MIN_LEN_SHARE):
|
||||||
|
junk += 1
|
||||||
|
lines[idx] = "<%s%s>%s</%s>" % (tag, attrs, original, tag)
|
||||||
|
continue # оставляем оригинал: английский абзац лучше огрызка
|
||||||
|
# Абзац, вернувшийся на языке оригинала. Доля кириллицы по всему
|
||||||
|
# документу такое не ловит: полтора процента в ней тонут, а на
|
||||||
|
# странице это заметный кусок английского текста.
|
||||||
|
if len(plain_src) > 150 and detect_language(re.sub("<[^>]+>", "", text))[0] \
|
||||||
|
== detect_language(plain_src)[0] != "unknown":
|
||||||
|
untouched += 1
|
||||||
body, ok = from_marks(text)
|
body, ok = from_marks(text)
|
||||||
broken += not ok
|
broken += not ok
|
||||||
translated += 1
|
translated += 1
|
||||||
lines[idx] = "<%s%s>%s</%s>" % (tag, attrs, body, tag)
|
lines[idx] = "<%s%s>%s</%s>" % (tag, attrs, body, tag)
|
||||||
out.write_text("\n".join(lines) + "\n", encoding="utf-8")
|
out.write_text("\n".join(lines) + "\n", encoding="utf-8")
|
||||||
print("переведено блоков: %d, без перевода: %d, разметка потеряна в %d"
|
print("переведено блоков: %d, без перевода: %d, разметка потеряна в %d, "
|
||||||
% (translated, missing, broken))
|
"огрызков заменено оригиналом: %d, осталось на языке оригинала: %d"
|
||||||
return missing == 0
|
% (translated, missing, broken, junk, untouched))
|
||||||
|
return missing == 0 and junk * 200 <= translated
|
||||||
|
|
||||||
|
|
||||||
def verify_output(out, target):
|
def verify_output(out, target):
|
||||||
|
|||||||
Reference in New Issue
Block a user