Compare commits

..

12 Commits

Author SHA1 Message Date
chesirecatt 5d957e7be9 pdf2html: CMYK-блоки с ext=png на деле JPEG — декодировать через Pixmap
На «Запускаем Ansible» 3-го изд. 10 из 593 картинок — блоки с
b["ext"]=="png" и colorspace=4, но реальные байты — CMYK JPEG (FFD8,
подтверждено file). Запись b["image"] как есть под именем .png давала файл,
который pngquant и часть читалок отказывались декодировать.

_image_png() декодирует сырые байты через pymupdf.Pixmap (MuPDF определяет
формат по содержимому, не по чужому ext), при colorspace не Gray/RGB
конвертирует в RGB, и всегда отдаёт настоящий PNG. На отказе декодирования —
откат к исходным байтам, не падать.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-10-05 12:01:05 +03:00
chesirecatt 33845bd47a Пять правок по итогам четырёх книг из to_read
Извлечение (pdf2html.py), каждая правка вызвана конкретной книгой:

* колонтитул опознаётся по повтору текста в верхнем поле на многих полосах, а
  не по одной позиции. Позиционный признак рубит и настоящие заголовки: у
  Вернона («DDD Distilled») глава начинается ровно с верха полосы и короче 80
  знаков, так пропали 14 заголовков из 15. Это ровно тот второй проход, о
  котором говорил ponytail-комментарий на прежнем фильтре;
* блок режется, если шапка набрана крупнее следующего за ней текста. У Вернона
  85 подзаголовков лежали в одном блоке с первым абзацем раздела: оглавление
  теряло второй уровень, а абзац начинался с заголовка без точки;
* порог раздела снижен с 1.15 до 1.10 кегля текста. У Вернона подзаголовок
  17.2 при тексте 15.0 — ровно на пять сотых ниже прежнего порога, и все 85
  разделов уезжали в прозу. Тот же множитель, что у разреза блока, иначе
  разрез и классификация расходятся;
* кандидатом в главы не может быть кегль, у которого в блоках нет букв. У
  Нейгарда («Release it!») номер главы набран кеглем 100 отдельным блоком
  («1», «2», …), и главой объявлялся он, а настоящие заголовки в 30 пунктов
  уезжали в разделы — оглавление выходило пустым. Отбор по наличию букв, а не
  по длине: «Preface» — семь знаков, порог по длине отсекал бы и его.

Упаковка (pack_epub.py): XML 1.0 запрещает управляющие символы, а из PDF они
приезжают регулярно. У Бейера («Site Reliability Engineering») 33 главы из 45
не разбирались как XML из-за одного такого знака в начале абзаца, и читалка
спотыкается на таком файле молча. Чистим на упаковке, а не надеемся на
извлечение. Тест на это добавлен и проверен: с обезвреженной правкой он падает.

SKILL.md приведён в соответствие с репозиторием. Шаг упаковки с 25.08 делает
pack_epub.py, а текст всё ещё требовал calibre на PATH — из-за этого проверка
инструментов давала ложный отказ. Замерено 05.09.2026: ebook-convert
отсутствует на всех трёх здешних машинах, и все четыре книги собрались без
него. За calibre остался только .azw3 для Kindle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XCiPWRi3Jqy4saD2LwU1Dp
2026-09-05 18:23:05 +03:00
chesirecatt af5c243687 pack_epub: EPUB собирался без сжатия
У zipfile.ZipFile умолчание ZIP_STORED, компрессия нигде не задавалась —
книги весили втрое больше при том же тексте (Брикман 1,42 МБ против 0,38).
Теперь ZIP_DEFLATED; mimetype по-прежнему пишется несжатым, этого требует
формат EPUB.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XCiPWRi3Jqy4saD2LwU1Dp
2026-08-25 19:39:37 +03:00
chesirecatt 0320d2d331 Мусор в оглавлении сканов: фильтр поворота, проверка заголовков, вход для DjVu
Причиной мусорных заголовков у Брикмана оказался не порог кегля, а повёрнутый
на 90° текст: боковые врезки и подписи к таблицам распознаются в кашу
(«aoHegoduenueArdng») и попадают в h1, потому что кегль у них крупный. Замер:
105 строк из 17 300, все до одной брак. Блоки с неgоризонтальным направлением
строки выбрасываются целиком (is_rotated).

Остаток — подписи внутри иллюстраций — понижается до <p>, а не удаляется:
looks_garbled() в bookhtml.py, шесть признаков структуры, любых двух хватает.
Проверка повторяется после склейки соседних заголовков: по отдельности «LF»,
«FF» и «TIT» проходят как аббревиатуры, склеенные — обрывок таблицы.
Вместе: 99 h1 у Брикмана против 48.

djvu2html.py — третий вход конвейера. Родной текстовый слой DjVu чище нашего
OCR (у Праты 10,9 против 2,8 слов со смесью алфавитов на 10 000), а через PDF
он не проходит: ddjvu -format=pdf кладёт страницы картинками.

Самопроверки: djvu2html.py --selftest, pdf2html.py --selftest --selftest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XCiPWRi3Jqy4saD2LwU1Dp
2026-08-25 19:03:08 +03:00
chesirecatt 356b9ddb22 OCR-слой: моно-признак не по флагу, кегль округляется, листинги склеиваются
Невидимый слой tesseract набран GlyphLessFont с выставленным моно-битом, и
is_mono() опознавала книгу как один сплошной листинг: измерено на Ansible после
OCR — pre 10498, p 0. Кегль в таком слое подгоняется под рамку слова и гуляет
внутри абзаца (9.0, 9.2, 9.4), из-за чего гистограмма размазывалась и основной
кегль книги определялся неверно; в OCR-слое он округляется до целого. Каждая
строка кода приходит отдельным блоком — соседние pre склеиваются переводом
строки, иначе листинг рассыпается на десяток однострочных.

Признак OCR-слоя считается по выборке страниц, обычные PDF идут прежним путём.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 19:42:14 +03:00
Илья Поляков 1a2c03f999 Мягкий перенос U+00AD: русские PDF рвали слова пополам
Питер и ДМК переносят строку мягким дефисом, а не обычным. Ветка склейки
проверяла только "-", поэтому строки соединялись через пробел: «введен ные»,
«Практиче ски». Теперь U+00AD снимается при склейке, остатки внутри строки
удаляются, а в ошибочно опознанных листингах перенос со строкой схлопывается.

Измерено на Шоттсе (Питер, 2020, 544 стр.): было 8 мягких переносов на
проверенной выборке, стало ноль.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L2PNJCFFXvizoNjTFuTaqi
2026-08-24 11:50:26 +03:00
chesirecatt f19410d9ce Вход fb2 и метаданные цикла в EPUB
fb2html.py — третий вход в конвейер. FB2 в домашней библиотеке основной формат
(704 тысячи файлов), и без него конвейер до них не дотягивался. Разметка fb2
семантическая, поэтому задача та же, что у epub2html.py: свести чужие теги к
нашим. Уровень заголовка берётся из вложенности section.

pack_epub.py: --series и --series-index. Пишутся в двух видах — calibre-совместимый
meta понимают почти все читалки, belongs-to-collection требует EPUB 3. Без них
читалка не выстраивает книги цикла по порядку.
2026-08-21 19:45:37 +03:00
chesirecatt 43d6cfc456 glossary.py: двуязычный глоссарий из параллельных изданий цикла 2026-08-21 18:41:45 +03:00
chesirecatt 14c7fad763 Счётчик абзацев, вернувшихся на языке оригинала: упражнения вида [2] модель не переводит 2026-08-21 17:47:48 +03:00
chesirecatt 7ed6e500ca Перевод не возобновляем: --all переводит все главы заново, механизм progress не работает 2026-08-21 17:36:10 +03:00
chesirecatt a21a4d713b Автоматический подбор порогов вместо ручного + упаковка EPUB без calibre
Ручная настройка block_kind() под каждую книгу оказалась неприменима: у части
изданий имена шрифтов вычищены до «Fd820825», а флаги стилей у всех спанов
равны 4, то есть ни курсива, ни моноширинности по ним не видно. Пороги теперь
считаются из самого файла:

- кегль основного текста — по гистограмме символов;
- кегль глав — по частоте блочных кеглей, а не множителем от текста: иначе
  подзаголовки разделов попадают в главы и оглавление пухнет до трёхсот пунктов;
- глава отличается от раздела положением в верхней трети полосы, и правило
  применяется, только если глав так набирается хотя бы пять;
- красная строка — по распределению x0.

Листинг опознаётся флагом моноширинности, а при его отсутствии — кеглем мельче
основного текста плюс плотностью кодовой пунктуации.

Отсев картинок: мельче 24 pt (линейки и буллеты) и крупнее 60% площади полосы
(подложка отсканированной страницы). На «Программисте-прагматике» второе
правило убрало 370 сканов полос и уменьшило EPUB со 124 МБ до 17 МБ.

pack_epub.py собирает EPUB на стандартной библиотеке: calibre для упаковки не
нужен, а его установка требует root. Книга режется по h1, оглавление NCX
строится само. Тест test_pack_epub.py проверяет несжатый mimetype первым файлом,
экранирование разметки в метаданных, сохранение переносов в листинге и то, что
преамбула до первой главы не теряется.
2026-08-20 20:11:03 +03:00
chesirecatt 1bacd38292 Восстановление структуры EPUB по CSS для дампов из PDF
Дамп, собранный конвертером из PDF, теряет всю семантику: ни заголовков, ни
листингов, только абзацы с обфусцированными классами. Такие книги уезжали в
перевод одним куском, а код — вместе с текстом.

Теперь при нулевом числе заголовков или листингов разметка восстанавливается из
таблицы стилей книги: три самых крупных кегля становятся уровнями заголовков,
моноширинная гарнитура — листингом, соседние строки листинга склеиваются в один
блок. Книга со своей разметкой в этот путь не попадает.

Проверено на двух настоящих книгах: дамп APoSD — 1833 плоских абзаца стали 29
главами и 58 листингами; Страуструп с собственной разметкой не изменился.
2026-08-20 19:11:51 +03:00
13 changed files with 1619 additions and 72 deletions
+35
View File
@@ -26,9 +26,44 @@
к абзацу с предыдущей страницы; к абзацу с предыдущей страницы;
- переносы на конце строки снимаются, соседние `</i><i>` схлопываются; - переносы на конце строки снимаются, соседние `</i><i>` схлопываются;
- иллюстрации выгружаются в `images/`, обложка рендерится со страницы 1 в 150 dpi; - иллюстрации выгружаются в `images/`, обложка рендерится со страницы 1 в 150 dpi;
- повёрнутый на 90° текст выбрасывается целиком: боковые врезки и подписи к
таблицам распознаются в кашу («aoHegoduenueArdng») и лезут в заголовки, потому
что кегль у них крупный. У Брикмана это 105 строк из 17 300, и все до одной —
брак;
- мусорный заголовок понижается до `<p>`, а не удаляется: оглавление чистится,
текст остаётся. Проверка — `looks_garbled()` в `scripts/bookhtml.py`, шесть
признаков структуры, любых двух хватает. Вместе с фильтром поворота это увело
Брикмана с 99 `<h1>` до 48;
- позиционирование не сохраняется намеренно — текст должен течь под любой - позиционирование не сохраняется намеренно — текст должен течь под любой
размер шрифта на читалке. размер шрифта на читалке.
## DjVu со своим текстовым слоем: `scripts/djvu2html.py`
У сканов в DjVu слой распознавания обычно уже есть, и он лучше нашего прогона
через `ocrmypdf`. Замер на Прате (C++ 6-е рус. изд., 1244 полосы): слов со
смесью алфавитов внутри слова было 433 на 391 тысячу слов (10,9 на 10 000),
после сборки из родного слоя — 109 (2,8).
⚠️ **Через PDF этот слой не проходит.** `ddjvu -format=pdf` кладёт страницы
картинками, `pdftotext` после этого отдаёт пустоту. Текст берётся напрямую из
`djvutxt --detail=line`, скрипт разбирает его сам.
Чего в слое DjVu нет вовсе — курсива и полужирного: в скане их и не было.
Заголовки опознаются высотой строки, листинги — пунктуацией.
⚠️ **Высота строки в DjVu — это габарит с выносными элементами, а не кегль.**
Строка с «Ц» или «р» выше соседней на те же 10%, поэтому пороги взяты по
измеренным разрывам, а не «чуть выше основного текста». У Праты при основной
строке 78: колонтитулы 85–90, разделы 105–125, названия глав 170–185.
Строки заголовка склеиваются подряд: название главы занимает две-три строки, и
без склейки «Класс string и стандартная библиотека шаблонов» разваливается на
четыре пункта оглавления.
Самопроверки без сети: `python3 scripts/djvu2html.py --selftest` и
`python3 scripts/pdf2html.py --selftest --selftest` (флаг занимает оба
обязательных аргумента).
## Подключение как скил Claude Code ## Подключение как скил Claude Code
Репозиторий одновременно является скилом (`SKILL.md` в корне) и клонируется Репозиторий одновременно является скилом (`SKILL.md` в корне) и клонируется
+105 -21
View File
@@ -15,23 +15,47 @@ tag, and 2817 body paragraphs became `<h2>`.
properties, and emits semantic XHTML. calibre is then used only as the properties, and emits semantic XHTML. calibre is then used only as the
XHTML→EPUB packer. XHTML→EPUB packer.
## Two entry points ## Three entry points
`scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB. Both emit the `scripts/pdf2html.py` for PDF, `scripts/epub2html.py` for EPUB,
same normalized XHTML — one block per line, `<pre>` for code listings — and `scripts/djvu2html.py` for DjVu. All three emit the same normalized XHTML — one
everything downstream (translation, packing, audiobook) is identical. Shared block per line, `<pre>` for code listings — and everything downstream
document skeleton and CSS live in `scripts/bookhtml.py`. (translation, packing, audiobook) is identical. The shared document skeleton,
CSS and the junk-heading check `looks_garbled()` live in `scripts/bookhtml.py`.
**A DjVu with its own text layer must not be re-OCR'd.** The layer that is
already in the file beats a fresh `ocrmypdf` run — measured on Prata's C++ 6th
Russian edition, 1244 pages: words mixing Latin and Cyrillic inside one word
dropped from 433 to 109 per 391k words (10.9 → 2.8 per 10 000). That layer does
not survive a trip through PDF: `ddjvu -format=pdf` writes the pages as images
and `pdftotext` then returns nothing at all. `djvu2html.py` reads `djvutxt
--detail=line` directly.
Styling is the price: a DjVu text layer carries no italic or bold at all (the
scan never had them). Headings come from line height, listings from punctuation.
Line height there is the glyph bounding box, not the type size — a line holding
a descender is 10% taller than its neighbour — so the thresholds are set from
measured gaps (`H1`/`H2` in the script), not from "slightly above body text".
**Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws **Do not route an EPUB through the PDF path.** Converting EPUB→PDF→XHTML throws
away the semantic markup that is already there and re-derives it from font away the semantic markup that is already there and re-derives it from font
sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads sizes. `epub2html.py` needs no font profiling and no threshold tuning: it reads
the spine, drops the nav document, and maps the book's own headings. the spine, drops the nav document, and maps the book's own headings.
Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (only calibre is Steps 1, 3 and 4 below are PDF-only. For EPUB start at step 2 (no PyMuPDF
needed, not PyMuPDF) and convert with: needed) and convert with:
`python scripts/epub2html.py book.epub out/book.html` `python scripts/epub2html.py book.epub out/book.html`
**Dumps converted from PDF carry no semantics at all** — no headings, no `<pre>`,
just `<p class="class_s1e2">` with obfuscated names. The script detects this
(zero headings or zero listings) and rebuilds the structure from the book's own
stylesheet: the three largest font sizes become heading levels, a typewriter or
monospace family becomes `<pre>`, and adjacent listing paragraphs merge back into
one block. A book that carries its own markup never enters this path. Measured on
the dokumen.pub dump of Ousterhout's APoSD: 1833 flat paragraphs became 29
chapters and 58 listings.
It reports which heading level turned out to be the chapter level. In most It reports which heading level turned out to be the chapter level. In most
EPUBs `<h1>` is the book title and the parts, while chapters are `<h2>` — the EPUBs `<h1>` is the book title and the parts, while chapters are `<h2>` — the
script picks the deepest level that still gives a sane chapter count, because script picks the deepest level that still gives a sane chapter count, because
@@ -44,7 +68,10 @@ the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No 1. **Verify a text layer exists.** `pdfinfo` and `pdffonts` on the file. No
embedded fonts / no extractable text means a scan — stop and say OCR is embedded fonts / no extractable text means a scan — stop and say OCR is
needed; this skill does not apply. needed; this skill does not apply.
2. **Check tooling.** `ebook-convert` must be on PATH. For PyMuPDF, prefer an 2. **Check tooling.** Packing needs nothing but the standard library
(`scripts/pack_epub.py`); calibre is optional and only for `.azw3`. Measured
2026-09-05: `ebook-convert` was absent on all three machines at hand, and the
conversion went through end to end without it. For PyMuPDF, prefer an
existing interpreter that has it; otherwise build a throwaway venv in the existing interpreter that has it; otherwise build a throwaway venv in the
scratchpad — do not install into the system Python: scratchpad — do not install into the system Python:
@@ -76,21 +103,29 @@ the translation bridge splits the book on `<h1>`. Measured on Stroustrup's PPP
means the mono flag never fired and every listing is about to be reflowed as means the mono flag never fired and every listing is about to be reflowed as
prose. Fix thresholds and re-run until prose. Fix thresholds and re-run until
the counts are sane. Cheap to iterate; do not skip to packing. the counts are sane. Cheap to iterate; do not skip to packing.
7. **Pack with calibre**, always passing explicit TOC XPaths — without them 7. **Pack with `scripts/pack_epub.py`** — no calibre, no root, splits on `<h1>`
calibre applies its own heuristics and re-breaks the chapters: and writes the conservative EPUB 2 + NCX that PocketBook and KOReader both
read:
```bash ```bash
ebook-convert out/book.html "Title.epub" \ python scripts/pack_epub.py out/book.html "Title.epub" \
--title="Title" --authors="Author" --language=en \ --title="Title" --author="Author" --lang=en --cover=cover.jpg
--cover=out/images/cover.jpg \
--level1-toc='//h:h1' --level2-toc='//h:h2' \
--page-breaks-before='//h:h1' \
--no-default-epub-cover
ebook-convert "Title.epub" "Title.azw3"
``` ```
Take title/author from the user or the book's own title page — PDF metadata Take title/author from the user or the book's own title page — PDF metadata
is often an ASIN or a filename. is often an ASIN or a filename.
calibre is needed only for `.azw3` (Kindle over USB), and only if it happens
to be installed:
```bash
ebook-convert "Title.epub" "Title.azw3"
```
⚠️ Do **not** pack the XHTML with `ebook-convert` directly: without explicit
`--level1-toc='//h:h1' --page-breaks-before='//h:h1'` it applies its own
heuristics and re-breaks the chapters. `pack_epub.py` has no such mode to
get wrong.
8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle 8. **Deliver both files** and say what each is for: `.epub` for Send to Kindle
(Amazon converts server-side), `.azw3` for USB copy into `documents/`. (Amazon converts server-side), `.azw3` for USB copy into `documents/`.
Delete intermediate artifacts left in the user's directories. Delete intermediate artifacts left in the user's directories.
@@ -114,17 +149,51 @@ EOF
is being promoted; is being promoted;
- many double spaces means line joining is off; - many double spaces means line joining is off;
- list the extracted headings (`grep -o '<h1[^>]*>.\{0,60\}'`) and read them — - list the extracted headings (`grep -o '<h1[^>]*>.\{0,60\}'`) and read them —
they become the TOC, and a wrong one is obvious at a glance; they become the TOC, and a wrong one is obvious at a glance. On a scan most of
the junk there comes from **rotated** text, not from bad thresholds: sideways
captions and table stubs OCR into mush (`aoHegoduenueArdng`) and land in
headings because their type is large. `pdf2html.py` drops any block whose line
direction is not horizontal (`is_rotated()`, measured on Brikman: 105 lines out
of 17 300, every one of them garbage), and demotes what is left of the mush to
`<p>` rather than deleting it. That pair took Brikman from 99 `<h1>` to 48;
- read one full page of body text and confirm paragraphs merge across page - read one full page of body text and confirm paragraphs merge across page
breaks and hyphenated words are rejoined. breaks and hyphenated words are rejoined.
After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx` After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx`
navPoint count for the TOC size. navPoint count for the TOC size.
## Glossary from existing translations
For a book in a series that already has published translations, `scripts/glossary.py`
mines a bilingual glossary so the machine translation does not invent new spellings
for names the reader already knows. Feed it pairs of editions of the *same* volume:
```
python scripts/glossary.py --en vol12.fb2 --ru vol12.ru.fb2 \
--en vol15.epub --ru vol15.ru.fb2 --score 0.6
```
Two signals, and both are needed. Position: paragraph indices do not line up
(Russian editions split dialogue, giving 2–3× more paragraphs), so offsets are
measured as a **share of characters**, where the texts track each other closely.
Transliteration: a proper name in Russian is nearly always a transliteration, so
the Cyrillic candidate is romanized and compared to the English term — this is
what turns the output from noise into a usable list.
Two mirrored filters remove the rest of the junk: a candidate whose head word
also appears lowercase in the same text is a sentence-initial common word, not a
name — applied on both sides. Measured on four Dresden Files volumes: 71 pairs,
of which two were wrong.
⚠️ Concept terms (`White Council` → `Белый Совет`, `Spire` → `Копьё`) do **not**
come out of the transliteration path and the positional one alone is too noisy
for them. Extract those by hand and verify by grepping the existing translation.
## Optional stage: translation ## Optional stage: translation
Only when the user asks for a translated book. It slots between step 6 and Only when the user asks for a translated book. It slots between step 6 and
step 7 — translate the XHTML, then pack the translated file with calibre. step 7 — translate the XHTML, then pack the translated file with
`pack_epub.py`.
**Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>` **Code listings are never translated.** The bridge only recognizes `<p>`, `<h1>`
and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into and `<h2>`; a `<pre>` block is not parsed, and anything unparsed is carried into
@@ -171,8 +240,23 @@ is the bridge: it writes the repo's input format, shells out to
`python scripts/translate.py book.html book.ru.html --repo <clone> --workdir <dir> --workers 12` `python scripts/translate.py book.html book.ru.html --repo <clone> --workdir <dir> --workers 12`
Resumable: the external repo tracks completed chapters and skips them on a ⚠️ **Not resumable, despite what the external repo claims.** Its
re-run, so an interrupted run costs nothing to restart. `progress/translation_progress.json` stays `{"chapters": {}}` and `--all`
re-translates every chapter, including ones already sitting in
`translations/`. Measured 2026-08-21 on Stroustrup: a re-run to repair 4
failed chapters re-did all 29. Budget a full book on every restart, and
prefer getting one clean run over patching a partial one.
The one thing that *does* skip work: deleting a chapter from `extracted/`
before the run. It is then never sent, and `rebuild` keeps the original
English for it — the right treatment for an index.
**Numbered exercise items come back untranslated.** Measured on Stroustrup
2026-08-21: 182 paragraphs of 10151 (1.8%) shaped `[2] Expanding on what you
have learned…` were echoed back in English. The document-wide Cyrillic check
cannot see this — 1.8% drowns in it — so `rebuild` counts them separately as
"осталось на языке оригинала". If the count is high, add an explicit line to the
prompt that numbered items are prose and must be translated too.
4. **Read the reported counts.** "без перевода" above zero means a chapter came 4. **Read the reported counts.** "без перевода" above zero means a chapter came
back with a different paragraph count and kept its original text; "разметка back with a different paragraph count and kept its original text; "разметка
потеряна" counts paragraphs where the model mangled the inline-tag markers потеряна" counts paragraphs where the model mangled the inline-tag markers
+55
View File
@@ -6,6 +6,14 @@
строки, а всё, чего не узнал, доносит до результата нетронутым. строки, а всё, чего не узнал, доносит до результата нетронутым.
""" """
import html import html
import re
# Брак OCR-а внутри слова: цифра вплотную к букве («30HWod1Q»), смесь алфавитов
# в одном слове («aoHegoduenue»), заглавная посреди слова.
DIGIT_WORD = re.compile(r"[^\W\d_][\d]|[\d][^\W\d_]")
WORD = re.compile(r"\w+")
LAT, CYR = re.compile(r"[A-Za-z]"), re.compile(r"[А-Яа-яЁё]")
MIDCAP = re.compile(r"[a-zа-яё][A-ZА-ЯЁ]")
CSS = """ CSS = """
body { margin: 0 1em; } body { margin: 0 1em; }
@@ -28,6 +36,53 @@ code { font-family: monospace; font-size: 0.9em; }
""" """
def looks_garbled(t):
"""Заголовок ли это вообще, или подпись из схемы, распознанная по буквам.
Основную часть брака снимает фильтр поворота в главном цикле — здесь остаётся
то, что набрано горизонтально: подписи внутри иллюстраций и обрывки таблиц
(«LF. FF. FT. TIT», «0|91239123», «В»).
Безусловный брак — три случая, каждый сам по себе: в заголовке нет ни одной
буквы; заголовок короче трёх знаков; в нём есть палка (в наборе её не бывает,
в распознанном скане она попадается постоянно).
Дальше шесть признаков структуры, любых двух хватает. Одного мало: «C++11 и
лямбды» даёт «не-букв больше трети», «Python3» — цифру вплотную к букве,
и оба заголовка настоящие.
"""
t = t.strip()
words = WORD.findall(t)
if not words or not any(c.isalpha() for c in t):
return True
if "|" in t:
return True
# Заголовок короче четырёх знаков — обрывок («В», «лов», «юн»). Аббревиатуры
# целиком заглавными («API», «AI») настоящими заголовками бывают, их щадим —
# но только если букв в них не одна и та же: «TT» это обрывок таблицы.
if len(t) < 4 and not (t.isupper() and len(set(t)) >= 2):
return True
# Слова из двух букв по кругу: «TIT. ITITIT», «LF. FF. FT». Нужны минимум два
# таких слова подряд и ни одного нормального, иначе под нож попадёт «Часть III».
long_words = [w for w in words if len(w) >= 3]
if len(long_words) >= 2 and all(len(set(w.lower())) <= 2 for w in long_words):
return True
letters = sum(c.isalpha() for c in t)
toks = t.split()
# Обрывок — короткое слово без цифр внутри: «LF.», «оо». Цифры исключены
# намеренно, иначе «C++11» и «Qt 6» считались бы обрывками.
frags = [w for w in toks if not any(c.isdigit() for c in w)
and sum(c.isalpha() for c in w) < 3]
signals = (
len(frags) * 2 > len(toks), # обрывки слов
letters * 3 < len(t) * 2, # не-букв больше трети
bool(DIGIT_WORD.search(t)),
any(LAT.search(w) and CYR.search(w) for w in words), # смесь алфавитов в слове
any(MIDCAP.search(w) for w in words), # заглавная посреди слова
next((c.islower() for c in t if c.isalpha()), False), # начинается со строчной
)
return sum(signals) >= 2
def document(title, parts): def document(title, parts):
"""parts: [(tag, cls, inner_html)]; tag 'figure' — псевдотег для картинки.""" """parts: [(tag, cls, inner_html)]; tag 'figure' — псевдотег для картинки."""
buf = ['<?xml version="1.0" encoding="utf-8"?>', buf = ['<?xml version="1.0" encoding="utf-8"?>',
+195
View File
@@ -0,0 +1,195 @@
#!/usr/bin/env python3
"""DjVu со своим текстовым слоем -> semantic XHTML.
Для сканов, у которых слой распознавания уже есть в файле: он почти всегда
лучше нашего OCR-а. Замерено на Прате 6-го изд.: родной слой 302 523 слова
при смеси алфавитов 0.8 на 10 000 знаков, наш прогон через ocrmypdf —
296 954 слова при 12.6 и на 26 страниц меньше.
Через PDF этот слой не проходит: ddjvu -format=pdf кладёт страницы картинками
и текст теряется целиком (проверено, pdftotext даёт пустоту). Поэтому слой
берётся напрямую из djvutxt.
Стиля (курсив, полужирный) в слое DjVu нет вовсе — в скане его и не было.
Заголовки опознаются высотой строки, листинги — пунктуацией, как в pdf2html.py.
djvu2html.py book.djvu out.html
djvu2html.py --selftest
"""
import collections
import html
import re
import subprocess
import sys
from pathlib import Path
import bookhtml
CODEY = re.compile(r"[{}();\[\]<>=*/#]|::|->")
SOFT_HYPHEN = "­"
HEAD_BAND = 0.07 # доля высоты полосы: выше — колонтитул, а не текст
FOOT_BAND = 0.05 # то же снизу — колонцифра
# Высота строки в DjVu — это габарит с выносными элементами, а не кегль: строка
# с «Ц» или «р» выше соседней на те же 10%. Поэтому пороги взяты не «чуть выше
# основного текста», а по измеренным разрывам у Праты (основная строка 78):
# 85–90 — колонтитулы, 105–125 — разделы, 170–185 — названия глав.
H1 = 2.0
H2 = 1.30
def parse(txt):
"""djvutxt --detail=line -> [[(x0, y0, x1, y1, text), ...] по страницам].
Свой разбор, а не sexpdata: формат простой, а зависимость ради него —
лишняя. Строка может переноситься между скобкой и текстом, поэтому идём
по знакам, а не построчно.
"""
pages, cur, i, n = [], None, 0, len(txt)
while i < n:
c = txt[i]
if c == '"': # строка: до неэкранированной закрывающей кавычки
j, buf = i + 1, []
while j < n and txt[j] != '"':
if txt[j] == "\\" and j + 1 < n:
buf.append({"n": "\n", "t": "\t"}.get(txt[j + 1], txt[j + 1]))
j += 2
continue
buf.append(txt[j])
j += 1
if cur is not None and cur["box"]:
cur["lines"].append(tuple(cur["box"]) + ("".join(buf),))
cur["box"] = None
i = j + 1
continue
m = re.match(r"\((page|line)((?:\s+-?\d+){4})", txt[i:])
if m:
box = [int(v) for v in m.group(2).split()]
if m.group(1) == "page":
cur = {"h": box[3] - box[1], "lines": [], "box": None}
pages.append(cur)
elif cur is not None:
cur["box"] = box
i += m.end()
continue
i += 1
return [(p["h"], p["lines"]) for p in pages]
def profile(pages):
"""Высота основной строки и левое поле — по самым частым значениям."""
heights, lefts = collections.Counter(), collections.Counter()
for _, lines in pages:
for x0, y0, x1, y1, _ in lines:
heights[y1 - y0] += 1
lefts[x0] += 1
if not heights:
sys.exit("в DjVu нет текстового слоя — этот скрипт не поможет, нужен OCR")
# Высота строки гуляет на пиксель-другой, поэтому берётся не голая мода, а
# та высота, у которой вместе с соседями ±1 набирается больше всего строк.
body = max(heights, key=lambda h: sum(heights[h + d] for d in (-1, 0, 1)))
left = min(x for x, _ in lefts.most_common(4))
return body, left
def looks_like_code(t):
return len(t) > 8 and len(CODEY.findall(t)) / len(t) > 0.03
def kind(height, body):
if height >= body * H1:
return "h1"
if height >= body * H2:
return "h2"
return "p"
def convert(pages, body, left):
parts, open_para = [], None # open_para: [tag, cls, text]
for page_h, lines in pages:
for x0, y0, x1, y1, raw in lines:
txt = raw.strip()
if not txt:
continue
# Координаты DjVu считаются снизу вверх.
top = (page_h - y1) / page_h if page_h else 0.5
bottom = y0 / page_h if page_h else 0.5
if (top < HEAD_BAND or bottom < FOOT_BAND) and len(txt) < 80:
continue # колонтитул и колонцифра
tag = kind(y1 - y0, body)
if tag == "p" and looks_like_code(txt):
tag = "pre"
if tag in ("h1", "h2") and bookhtml.looks_garbled(txt):
tag = "p"
# Абзац продолжается, пока строки идут от левого поля: красная
# строка и смена тега начинают новый. Заголовок склеивается всегда:
# название главы занимает две-три строки («Класс string» / «и
# стандартная» / «библиотека» / «шаблонов» — это один заголовок),
# а два разных заголовка подряд без текста между ними не встречаются.
cont = open_para and open_para[0] == tag and (
tag in ("h1", "h2") or (tag == "p" and x0 <= left + (y1 - y0)))
if tag == "pre" and open_para and open_para[0] == "pre":
open_para[2] += "\n" + txt
continue
if cont:
prev = open_para[2]
if prev.endswith(SOFT_HYPHEN):
open_para[2] = prev[:-1] + txt
elif prev.endswith("-") and not prev.endswith("--"):
open_para[2] = prev[:-1] + txt
else:
open_para[2] = prev + " " + txt
continue
if open_para:
parts.append(tuple(open_para))
open_para = [tag, None, txt]
if open_para:
parts.append(tuple(open_para))
return [(t, c, html.escape(x).replace(SOFT_HYPHEN, "")) for t, c, x in parts]
# Кегли взяты с настоящих полос Праты: колонтитул 64, текст 77, раздел 108,
# название главы 180. Начало координат в DjVu внизу полосы.
SAMPLE = """(page 0 0 100 1000 (line 10 950 90 985 "300 Глава 6 ")
(line 10 720 90 900 "Упражнения по программированию ")
(line 10 590 90 698 "Подраздел про циклы ")
(line 10 500 90 577 "Напишите программу, которая читает ввод до сим-")
(line 10 420 90 497 "вола @ и повторяет его. ")
(line 10 320 90 397 "int main() { return 0; }")
(line 10 20 90 60 "300 "))"""
def selftest():
pages = parse(SAMPLE)
assert len(pages) == 1, pages
assert len(pages[0][1]) == 7, pages[0][1]
body, left = profile(pages)
assert body == 77, body
parts = convert(pages, body, left)
tags = [p[0] for p in parts]
assert tags == ["h1", "h2", "p", "pre"], tags
# Колонтитул сверху и колонцифра снизу выброшены, перенос склеен без пробела.
assert "символа" in parts[2][2], parts[2][2]
assert "300" not in " ".join(p[2] for p in parts), parts
assert parse('(page 0 0 10 10 (line 1 1 2 2 "a \\"b\\" c"))')[0][1][0][4] == 'a "b" c'
print("selftest OK")
if len(sys.argv) == 2 and sys.argv[1] == "--selftest":
selftest()
sys.exit()
if len(sys.argv) < 3:
sys.exit("usage: djvu2html.py book.djvu out.html | djvu2html.py --selftest")
SRC, OUT = Path(sys.argv[1]), Path(sys.argv[2])
raw = subprocess.run(["djvutxt", "--detail=line", str(SRC)],
capture_output=True, text=True, check=True).stdout
pages = parse(raw)
BODY, LEFT = profile(pages)
print("страниц %d, высота строки %d, левое поле %d" % (len(pages), BODY, LEFT))
parts = convert(pages, BODY, LEFT)
OUT.write_text(bookhtml.document(SRC.stem, parts), encoding="utf-8")
print("blocks:", len(parts),
"h1:", sum(1 for p in parts if p[0] == "h1"),
"h2:", sum(1 for p in parts if p[0] == "h2"),
"p:", sum(1 for p in parts if p[0] == "p"),
"pre:", sum(1 for p in parts if p[0] == "pre"))
+89 -6
View File
@@ -39,16 +39,57 @@ INLINE_MAP = {"i": "i", "em": "i", "cite": "i", "b": "b", "strong": "b",
"code": "code", "kbd": "code", "samp": "code", "tt": "code", "code": "code", "kbd": "code", "samp": "code", "tt": "code",
"var": "code"} "var": "code"}
DROP = {"script", "style", "head", "title", "nav"} DROP = {"script", "style", "head", "title", "nav"}
# Дампы, собранные конвертером из PDF, теряют всю семантику: ни заголовков, ни
# <pre>, только <p class="class_s1e2"> с обфусцированными именами. Разметку
# приходится восстанавливать из таблицы стилей — по кеглю и по гарнитуре.
MONO_FONT = re.compile(r"mono|courier|consol|typewriter|menlo|inconsolata|"
r"liberationmono|andalemono|lucidasanstype", re.I)
CSS_RULE = re.compile(r"\.([\w-]+)\s*\{([^}]*)\}")
IMG_EXT = {"image/jpeg": ".jpg", "image/png": ".png", "image/gif": ".gif", IMG_EXT = {"image/jpeg": ".jpg", "image/png": ".png", "image/gif": ".gif",
"image/svg+xml": ".svg", "image/webp": ".webp"} "image/svg+xml": ".svg", "image/webp": ".webp"}
def css_styles(zf):
"""class -> (кегль в em, моноширинный ли) из всех таблиц стилей книги."""
out = {}
for name in zf.namelist():
if not name.lower().endswith(".css"):
continue
for cls, body in CSS_RULE.findall(zf.read(name).decode("utf-8", "replace")):
size = out.get(cls, (0.0, False))[0]
m = re.search(r"font-size:\s*([\d.]+)\s*(em|rem|pt|px|%)", body)
if m:
v = float(m.group(1))
size = {"em": v, "rem": v, "pt": v / 12, "px": v / 16,
"%": v / 100}[m.group(2)]
fam = re.search(r"font-family:\s*([^;]+)", body)
mono = bool(fam and MONO_FONT.search(fam.group(1)))
was = out.get(cls, (0.0, False))
out[cls] = (size or was[0], mono or was[1])
return out
def derive_from_css(styles):
"""(класс -> уровень заголовка, множество моноширинных классов).
Заголовки — три самых крупных кегля выше основного текста. Больше трёх
уровней брать нельзя: дальше начинаются подписи и колонтитулы.
"""
sizes = sorted({s for s, _ in styles.values() if s > 1.05}, reverse=True)[:3]
level = {s: i + 1 for i, s in enumerate(sizes)}
head = {c: level[s] for c, (s, _) in styles.items() if s in level}
mono = {c for c, (_, m) in styles.items() if m}
return head, mono
class Reader(HTMLParser): class Reader(HTMLParser):
"""Собирает блоки из одного XHTML-документа книги.""" """Собирает блоки из одного XHTML-документа книги."""
def __init__(self, on_image): def __init__(self, on_image, head_cls=None, mono_cls=None):
super().__init__(convert_charrefs=True) super().__init__(convert_charrefs=True)
self.on_image = on_image self.on_image = on_image
self.head_cls = head_cls or {}
self.mono_cls = mono_cls or set()
self.blocks = [] self.blocks = []
self.cur = None # (tag, cls, [куски]) self.cur = None # (tag, cls, [куски])
self.open_inline = [] # незакрытые инлайновые теги текущего блока self.open_inline = [] # незакрытые инлайновые теги текущего блока
@@ -69,7 +110,12 @@ class Reader(HTMLParser):
else: else:
txt = re.sub(r"\s+", " ", txt).strip() txt = re.sub(r"\s+", " ", txt).strip()
txt = re.sub(r"<(i|b|code)>(\s*)</\1>", r"\2", txt) # пустая разметка txt = re.sub(r"<(i|b|code)>(\s*)</\1>", r"\2", txt) # пустая разметка
if txt.strip(): if not txt.strip():
return
if tag == "pre" and self.blocks and self.blocks[-1][0] == "pre":
prev = self.blocks[-1]
self.blocks[-1] = (prev[0], prev[1], prev[2] + "\n" + txt)
return
self.blocks.append((tag, cls, txt)) self.blocks.append((tag, cls, txt))
def start(self, tag, cls): def start(self, tag, cls):
@@ -98,7 +144,16 @@ class Reader(HTMLParser):
self.cur[2].append("\n" if self.cur[0] == "pre" else " ") self.cur[2].append("\n" if self.cur[0] == "pre" else " ")
return return
if tag in BLOCK_MAP: if tag in BLOCK_MAP:
self.start(*BLOCK_MAP[tag]) out_tag, out_cls = BLOCK_MAP[tag]
if out_tag in ("p", "pre"):
for c in (a.get("class") or "").split():
if c in self.head_cls:
out_tag, out_cls = "h%d" % self.head_cls[c], None
break
if c in self.mono_cls:
out_tag, out_cls = "pre", None
break
self.start(out_tag, out_cls)
return return
if tag in ("td", "th") and self.cur and self.cur[0] == "p" \ if tag in ("td", "th") and self.cur and self.cur[0] == "p" \
and self.cur[1] == "row" and self.cur[2]: and self.cur[1] == "row" and self.cur[2]:
@@ -220,6 +275,8 @@ def convert(src, out):
if cover: if cover:
extract(cover, stem="cover") extract(cover, stem="cover")
def parse_all(head_cls=None, mono_cls=None):
got = []
for doc in docs: for doc in docs:
if doc not in names: if doc not in names:
print("нет в архиве, пропущен: %s" % doc, file=sys.stderr) print("нет в архиве, пропущен: %s" % doc, file=sys.stderr)
@@ -227,19 +284,45 @@ def convert(src, out):
here = posixpath.dirname(doc) here = posixpath.dirname(doc)
def on_image(src_attr, here=here): def on_image(src_attr, here=here):
target = posixpath.normpath(posixpath.join(here, src_attr.split("#")[0])) target = posixpath.normpath(
posixpath.join(here, src_attr.split("#")[0]))
return extract(target) return extract(target)
r = Reader(on_image) r = Reader(on_image, head_cls, mono_cls)
r.feed(zf.read(doc).decode("utf-8", "replace")) r.feed(zf.read(doc).decode("utf-8", "replace"))
r.close() r.close()
parts.extend(r.blocks) got.extend(r.blocks)
return got
parts = parse_all()
has_head = any(re.fullmatch(r"h[1-6]", t) for t, _, _ in parts)
has_pre = any(t == "pre" for t, _, _ in parts)
if not (has_head and has_pre):
head_cls, mono_cls = derive_from_css(css_styles(zf))
if (not has_head and head_cls) or (not has_pre and mono_cls):
print("семантики в книге нет (заголовки: %s, листинги: %s) — "
"восстанавливаю по CSS" % (has_head, has_pre))
seen.clear()
counter[0] = 0
parts = parse_all(head_cls if not has_head else None,
mono_cls if not has_pre else None)
title = out.stem title = out.stem
lvl = chapter_level(parts) lvl = chapter_level(parts)
parts = [(("h1" if int(t[1]) <= lvl else "h2", c, x) parts = [(("h1" if int(t[1]) <= lvl else "h2", c, x)
if re.fullmatch(r"h[1-6]", t) else (t, c, x)) if re.fullmatch(r"h[1-6]", t) else (t, c, x))
for t, c, x in parts] for t, c, x in parts]
# «Chapter 1» и «Introduction» приезжают отдельными заголовками: в дампе это
# две строки. Иначе каждая вторая глава состоит из одного блока.
merged = []
for part in parts:
if (part[0] == "h1" and merged and merged[-1][0] == "h1"
and len(merged[-1][2]) < 60):
merged[-1] = ("h1", None, merged[-1][2] + ". " + part[2])
continue
merged.append(part)
parts = merged
h1 = sum(1 for p in parts if p[0] == "h1") h1 = sum(1 for p in parts if p[0] == "h1")
print("главы размечены h%d, глав получилось: %d" % (lvl, h1)) print("главы размечены h%d, глав получилось: %d" % (lvl, h1))
if h1 < 3: if h1 < 3:
+163
View File
@@ -0,0 +1,163 @@
#!/usr/bin/env python3
"""FB2 -> тот же XHTML, что выдают pdf2html.py и epub2html.py.
FB2 в этой библиотеке основной формат (704 тысячи файлов), и без этого входа
конвейер до них не дотягивается. Разметка у fb2 уже семантическая, поэтому
задача та же, что у epub2html.py: привести чужие теги к нашим.
Только стандартная библиотека.
"""
import argparse
import html
import re
import sys
import zipfile
from html.parser import HTMLParser
from pathlib import Path
import bookhtml
BLOCK = {"p": ("p", None), "v": ("p", "li"), "subtitle": ("h2", None),
"text-author": ("p", "note"), "th": ("p", "row"), "td": ("p", "row")}
INLINE = {"emphasis": "i", "strong": "b", "code": "code", "sub": "i", "sup": "i"}
DROP = {"description", "binary", "stylesheet"}
class Reader(HTMLParser):
def __init__(self):
super().__init__(convert_charrefs=True)
self.blocks, self.cur, self.open_i, self.drop = [], None, [], 0
self.depth = 0 # вложенность section: даёт уровень заголовка
self.in_title = False
def flush(self):
if not self.cur:
return
tag, cls, parts = self.cur
self.cur = None
for t in reversed(self.open_i):
parts.append("</%s>" % t)
self.open_i = []
txt = re.sub(r"\s+", " ", "".join(parts)).strip()
if txt:
self.blocks.append((tag, cls, txt))
def handle_starttag(self, tag, attrs):
if tag in DROP:
self.drop += 1
return
if self.drop:
return
if tag == "section":
self.depth += 1
return
if tag == "title":
self.flush()
self.in_title = True
return
if tag == "empty-line":
self.flush()
return
if tag == "image":
return # картинки fb2 лежат в base64, пропускаем
if self.in_title and tag == "p":
# заголовок раздела: уровень по вложенности section
self.flush()
self.cur = ("h1" if self.depth <= 2 else "h2", None, [])
return
if tag in BLOCK:
self.flush()
self.cur = (BLOCK[tag][0], BLOCK[tag][1], [])
return
if tag in INLINE and self.cur:
out = INLINE[tag]
self.open_i.append(out)
self.cur[2].append("<%s>" % out)
def handle_endtag(self, tag):
if tag in DROP:
self.drop = max(0, self.drop - 1)
return
if self.drop:
return
if tag == "section":
self.flush()
self.depth = max(0, self.depth - 1)
return
if tag == "title":
self.flush()
self.in_title = False
return
if tag in BLOCK:
self.flush()
return
if tag in INLINE and self.cur:
out = INLINE[tag]
if out in self.open_i:
self.open_i.remove(out)
self.cur[2].append("</%s>" % out)
def handle_data(self, data):
if self.drop or not self.cur:
return
self.cur[2].append(html.escape(data, quote=False))
def close(self):
super().close()
self.flush()
def read_fb2(path):
p = Path(path)
if p.suffix.lower() == ".zip" or zipfile.is_zipfile(p):
with zipfile.ZipFile(p) as z:
name = next(n for n in z.namelist() if n.lower().endswith(".fb2"))
raw = z.read(name)
else:
raw = p.read_bytes()
enc = "utf-8"
m = re.search(rb'encoding="([\w-]+)"', raw[:200])
if m:
enc = m.group(1).decode("ascii", "ignore")
return raw.decode(enc, "replace")
def meta(text):
"""(автор, название) из description — для метаданных EPUB."""
def tag(name):
m = re.search(r"<%s>(.*?)</%s>" % (name, name), text, re.S)
return re.sub(r"<[^>]+>", " ", m.group(1)).strip() if m else ""
ti = re.search(r"<title-info>(.*?)</title-info>", text, re.S)
block = ti.group(1) if ti else text
first = re.search(r"<first-name>(.*?)</first-name>", block, re.S)
last = re.search(r"<last-name>(.*?)</last-name>", block, re.S)
author = " ".join(re.sub(r"<[^>]+>", "", x.group(1)).strip()
for x in (first, last) if x).strip()
book = re.search(r"<book-title>(.*?)</book-title>", block, re.S)
return author, (re.sub(r"<[^>]+>", "", book.group(1)).strip() if book else "")
def main():
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument("source", type=Path)
ap.add_argument("output", type=Path)
args = ap.parse_args()
text = read_fb2(args.source)
body = re.search(r"<body[^>]*>(.*)</body>", text, re.S)
r = Reader()
r.feed(body.group(1) if body else text)
r.close()
if not r.blocks:
sys.exit("в %s не нашлось текста" % args.source)
author, title = meta(text)
args.output.parent.mkdir(parents=True, exist_ok=True)
args.output.write_text(bookhtml.document(title or args.output.stem, r.blocks),
encoding="utf-8")
print("blocks: %d, h1: %d, p: %d" % (len(r.blocks),
sum(1 for t, _, _ in r.blocks if t == "h1"),
sum(1 for t, _, _ in r.blocks if t == "p")))
print("метаданные: %s — %s" % (author or "?", title or "?"))
if __name__ == "__main__":
main()
+219
View File
@@ -0,0 +1,219 @@
#!/usr/bin/env python3
"""Двуязычный глоссарий из параллельных изданий одной книги.
Задача: машинный перевод очередного тома цикла не должен расходиться с уже
изданными переводами в именах и реалиях. Частотный список даёт только имена;
пары вида White Council → Белый Совет так не получить.
Метод — выравнивание по относительной позиции в тексте. Абзацы английского и
русского изданий не совпадают ни числом, ни границами, но идут в одном порядке,
поэтому термин, встречающийся в английском тексте на 12%, 34% и 78% длины,
в переводе окажется примерно там же. Кандидат, чьи позиции совпали с позициями
термина лучше, чем с текстом вообще, и есть перевод.
Работает с fb2, epub и голым текстом. Только стандартная библиотека.
"""
import argparse
import collections
import difflib
import html
import re
import sys
import zipfile
from pathlib import Path
# Слова, с которых начинается предложение, — не имена собственные.
RU_STOP = {"Когда", "Если", "Что", "Как", "Это", "Она", "Они", "Мы", "Вы", "Так",
"Потом", "Затем", "Даже", "Его", "Её", "Их", "Может", "Все", "Теперь",
"Нет", "Да", "Меня", "Мне", "Тогда", "Только", "После", "Пока", "Там",
"Здесь", "Однако", "Впрочем", "Просто", "Один", "Вот", "Или", "Почему",
"Мой", "Моя", "Кто", "Конечно", "Хорошо", "Ага", "Думаю", "Возможно",
"Глава", "Никто", "Ничего", "Что-то", "Кажется", "Значит", "Сейчас"}
EN_STOP = {"The", "A", "An", "And", "But", "I", "It", "He", "She", "They", "We",
"You", "This", "That", "There", "Then", "When", "If", "So", "My", "His",
"Her", "Chapter", "What", "Why", "How", "No", "Yes", "Not", "For", "Of",
"In", "On", "At", "To", "With", "As", "Was", "Were", "Had", "Have",
"Would", "Could", "Should", "One", "All", "Just", "Like", "Now", "Well",
"Okay", "Oh", "Maybe", "Something", "Nothing", "Someone"}
def read_text(path):
"""Текст книги из fb2, epub или txt — разметка выкидывается."""
p = Path(path)
if p.suffix.lower() == ".epub":
with zipfile.ZipFile(p) as z:
parts = [z.read(n).decode("utf-8", "replace")
for n in z.namelist() if n.lower().endswith((".xhtml", ".html"))]
raw = "\n".join(parts)
elif p.suffix.lower() in (".fb2", ".xml"):
raw = p.read_bytes().decode("utf-8", "replace")
else:
raw = p.read_text(encoding="utf-8", errors="replace")
raw = re.sub(r"<(script|style|head)[^>]*>.*?</\1>", " ", raw, flags=re.S | re.I)
raw = re.sub(r"</p>|</section>|<br\s*/?>", "\n", raw, flags=re.I)
return html.unescape(re.sub(r"<[^>]+>", " ", raw))
def paragraphs(text):
return [p.strip() for p in text.split("\n") if len(p.strip()) > 40]
def candidates(paras, pattern, stop, min_count):
"""{термин: [относительные позиции вхождений]}
Позиция считается по символам, а не по номеру абзаца: русские издания
разбивают диалоги построчно, и абзацев там втрое больше — по индексу
тексты расходятся, по доле знаков идут почти вровень.
"""
pos = collections.defaultdict(list)
total = sum(len(p) for p in paras) or 1
acc = 0
for p in paras:
rel = acc / total
acc += len(p)
for m in set(pattern.findall(p)):
head = m.split()[0]
if head in stop or len(m) < 3:
continue
pos[m].append(rel)
return {t: v for t, v in pos.items() if len(v) >= min_count}
# Кириллица -> латиница для сравнения имён. Имена собственные в переводе почти
# всегда транслитерация, и это признак куда надёжнее позиционного совпадения.
LAT = {"а":"a","б":"b","в":"v","г":"g","д":"d","е":"e","ё":"e","ж":"j","з":"z",
"и":"i","й":"i","к":"k","л":"l","м":"m","н":"n","о":"o","п":"p","р":"r",
"с":"s","т":"t","у":"u","ф":"f","х":"h","ц":"c","ч":"c","ш":"s","щ":"s",
"ъ":"","ы":"i","ь":"","э":"e","ю":"u","я":"a"}
def translit(s):
return "".join(LAT.get(c, c) for c in s.lower())
def name_similarity(en, ru):
"""Похожесть после огрубления: латиница обеих сторон без удвоений и гласных
на конце. Nicodemus/Никодимус дают почти единицу, Harry/Молли — ноль."""
a = re.sub(r"[^a-z]", "", en.lower())
b = re.sub(r"[^a-z]", "", translit(ru))
a = re.sub(r"(.)\1+", r"\1", a)
b = re.sub(r"(.)\1+", r"\1", b)
if not a or not b:
return 0.0
return difflib.SequenceMatcher(None, a, b[:len(a) + 3]).ratio()
def overlap(a, b, tol):
"""Доля вхождений a, у которых нашлось вхождение b поблизости."""
if not a or not b:
return 0.0
b = sorted(b)
hit = 0
for x in a:
lo, hi = 0, len(b) - 1
best = 1.0
while lo <= hi:
mid = (lo + hi) // 2
best = min(best, abs(b[mid] - x))
if b[mid] < x:
lo = mid + 1
else:
hi = mid - 1
if best <= tol:
hit += 1
return hit / len(a)
def main():
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument("--en", action="append", required=True, help="английское издание")
ap.add_argument("--ru", action="append", required=True, help="русское издание того же тома")
ap.add_argument("--min-count", type=int, default=4)
ap.add_argument("--tol", type=float, default=0.02, help="допуск по относительной позиции")
ap.add_argument("--score", type=float, default=0.6, help="порог уверенности")
ap.add_argument("--top", type=int, default=80)
args = ap.parse_args()
if len(args.en) != len(args.ru):
sys.exit("нужно одинаковое число --en и --ru: это должны быть пары изданий")
EN = re.compile(r"\b[A-Z][a-z]+(?:\s+(?:of\s+|the\s+)?[A-Z][a-z]+){0,2}")
lower_seen = collections.Counter()
ru_lower = collections.Counter()
RU = re.compile(r"\b[А-ЯЁ][а-яё]+(?:\s+[А-ЯЁ][а-яё]+){0,2}")
en_pos, ru_pos = collections.defaultdict(list), collections.defaultdict(list)
for i, (fe, fr) in enumerate(zip(args.en, args.ru)):
pe, pr = paragraphs(read_text(fe)), paragraphs(read_text(fr))
# Настоящее имя собственное почти не встречается со строчной буквы:
# так отсеиваются After, Once, More и прочие начала предложений.
for w in re.findall(r"\b[a-z]{3,}\b", " ".join(pe)):
lower_seen[w] += 1
for w in re.findall(r"\b[а-яё]{3,}\b", " ".join(pr)):
ru_lower[w] += 1
print("пара %d: %d абзацев EN, %d RU" % (i + 1, len(pe), len(pr)), file=sys.stderr)
for t, v in candidates(pe, EN, EN_STOP, 1).items():
en_pos[t] += [(i, x) for x in v]
for t, v in candidates(pr, RU, RU_STOP, 1).items():
ru_pos[t] += [(i, x) for x in v]
en_pos = {t: v for t, v in en_pos.items()
if len(v) >= args.min_count
and lower_seen[t.split()[0].lower()] < max(3, len(v) * 0.2)}
# Зеркальный отсев: «Надеюсь», «Зачем», «Парень» — обычные слова, они
# встречаются со строчной буквы и именами собственными быть не могут.
ru_pos = {t: v for t, v in ru_pos.items()
if len(v) >= args.min_count
and ru_lower[t.split()[0].lower()] < max(3, len(v) * 0.2)}
print("кандидатов: %d EN, %d RU" % (len(en_pos), len(ru_pos)), file=sys.stderr)
by_book_ru = collections.defaultdict(dict)
for t, v in ru_pos.items():
for b, x in v:
by_book_ru[b].setdefault(t, []).append(x)
out = []
for term, occ in sorted(en_pos.items(), key=lambda kv: -len(kv[1])):
mine = collections.defaultdict(list)
for b, x in occ:
mine[b].append(x)
best, best_score = None, 0.0
for cand, cocc in ru_pos.items():
# Частота кандидата должна быть сопоставима: «Гарри» встречается на
# каждой странице и по одностороннему совпадению побеждает всех.
if not (0.25 <= len(cocc) / len(occ) <= 4.0):
continue
fwd = back = tot_f = tot_b = 0.0
for b, xs in mine.items():
cand_pos = by_book_ru[b].get(cand)
if not cand_pos:
continue
fwd += overlap(xs, cand_pos, args.tol) * len(xs)
back += overlap(cand_pos, xs, args.tol) * len(cand_pos)
tot_f += len(xs); tot_b += len(cand_pos)
if tot_f < len(occ) * 0.5 or not tot_b:
continue
# Симметричная мера: термин должен находить перевод, а перевод —
# термин. Иначе частотное слово выигрывает у настоящего соответствия.
score = (fwd / tot_f) * (back / tot_b)
# Транслитерация — сильный самостоятельный довод: если написание
# совпадает, позиционного подтверждения нужно куда меньше.
sim = name_similarity(term, cand)
if sim >= 0.72:
score = max(score, sim) + 0.25
elif sim < 0.3 and " " not in term:
score *= 0.35 # разные имена, позиции совпали случайно
if score > best_score:
best, best_score = cand, score
if best and best_score >= args.score:
out.append((len(occ), best_score, term, best))
print("# частота | уверенность | оригинал -> перевод")
seen = set()
for n, sc, en, ru in out[:args.top]:
if (en, ru) in seen:
continue
seen.add((en, ru))
print("%5d %.2f %-30s -> %s" % (n, sc, en, ru))
if __name__ == "__main__":
main()
+178
View File
@@ -0,0 +1,178 @@
#!/usr/bin/env python3
"""XHTML от pdf2html.py/epub2html.py -> EPUB без сторонних зависимостей.
calibre для упаковки не нужен: EPUB — это zip с манифестом, а установка calibre
требует root, которого на чужой машине может не быть. Формат намеренно
консервативный (EPUB 2 + NCX): его одинаково понимают KOReader, штатные читалки
PocketBook и calibre, если он всё-таки понадобится.
Книга режется по <h1>: одна глава — один файл. Так читалка не разбирает
мегабайтный документ на каждом перелистывании и получает оглавление даром.
"""
import argparse
import html
import re
import shutil
import sys
import zipfile
from pathlib import Path
# XML 1.0 запрещает управляющие символы, а из PDF они приезжают регулярно:
# у Бейера («Site Reliability Engineering») 33 главы из 45 не разбирались как
# XML из-за одного такого знака в начале абзаца. Читалка спотыкается молча,
# поэтому чистим на упаковке, а не надеемся на извлечение.
BAD_XML = re.compile(r"[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]")
H1 = re.compile(r"^<h1[^>]*>(.*?)</h1>$")
IMG = re.compile(r'<img src="images/([^"]+)"')
NS = 'xmlns="http://www.w3.org/1999/xhtml"'
CHAPTER = """<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE html>
<html %s><head><meta charset="utf-8"/><title>%s</title>
<link rel="stylesheet" type="text/css" href="style.css"/></head>
<body>
%s
</body></html>
"""
def split_chapters(lines):
"""[(заголовок, [строки])] — режем по <h1>, преамбулу оставляем главой."""
chapters, title, buf = [], "Начало", []
for line in lines:
m = H1.match(line.strip())
if m and buf:
chapters.append((title, buf))
title, buf = re.sub("<[^>]+>", "", m.group(1)).strip() or "Глава", [line]
continue
if m:
title = re.sub("<[^>]+>", "", m.group(1)).strip() or "Глава"
buf.append(line)
if buf:
chapters.append((title, buf))
return [(t, b) for t, b in chapters if any(x.strip() for x in b)]
def series_meta(meta):
"""Цикл и номер в нём. Пишем в двух видах: calibre-совместимый meta читают
почти все читалки, belongs-to-collection — требование EPUB 3."""
name = meta.get("series")
if not name:
return ""
idx = meta.get("series_index") or ""
out = ['\n <meta name="calibre:series" content="%s"/>' % html.escape(name)]
if idx:
out.append('<meta name="calibre:series_index" content="%s"/>' % html.escape(str(idx)))
out.append('<meta property="belongs-to-collection" id="c01">%s</meta>' % html.escape(name))
out.append('<meta refines="#c01" property="collection-type">series</meta>')
if idx:
out.append('<meta refines="#c01" property="group-position">%s</meta>' % html.escape(str(idx)))
return "\n ".join(out)
def build(src, out, meta):
text = BAD_XML.sub("", src.read_text(encoding="utf-8"))
css = re.search(r"<style>(.*?)</style>", text, re.S)
css = css.group(1) if css else ""
body = re.search(r"<body>(.*)</body>", text, re.S)
lines = (body.group(1) if body else text).splitlines()
chapters = split_chapters(lines)
if not chapters:
sys.exit("в %s нет содержимого" % src)
imgdir = src.parent / "images"
used, files = [], []
# ZIP_DEFLATED обязателен явно: у zipfile умолчание — ZIP_STORED, и книга
# выходит втрое толще (у Брикмана 1.42 МБ против 0.43 при том же тексте).
# Первым файлом всё равно идёт несжатый mimetype — этого требует формат.
with zipfile.ZipFile(out, "w", zipfile.ZIP_DEFLATED) as z:
# mimetype обязан идти первым и без сжатия — иначе часть читалок
# не опознаёт архив как EPUB
z.writestr(zipfile.ZipInfo("mimetype"), "application/epub+zip",
compress_type=zipfile.ZIP_STORED)
z.writestr("META-INF/container.xml",
'<?xml version="1.0"?>\n<container version="1.0" '
'xmlns="urn:oasis:names:tc:opendocument:xmlns:container">'
'<rootfiles><rootfile full-path="OEBPS/content.opf" '
'media-type="application/oebps-package+xml"/></rootfiles></container>')
z.writestr("OEBPS/style.css", css)
for n, (title, buf) in enumerate(chapters):
name = "ch%03d.xhtml" % n
files.append((name, title))
chunk = "\n".join(buf)
for img in IMG.findall(chunk):
if img not in used and (imgdir / img).exists():
used.append(img)
z.writestr("OEBPS/" + name,
CHAPTER % (NS, html.escape(title), chunk))
for img in used:
z.write(imgdir / img, "OEBPS/images/" + img)
# Обложка бывает и среди картинок текста — второй раз её класть нельзя,
# zip примет дубликат имени, а читалки на такой архив ругаются.
cover = meta["cover"] if meta["cover"] not in used else ""
if cover and (imgdir / cover).exists():
z.write(imgdir / cover, "OEBPS/images/" + cover)
items = ['<item id="css" href="style.css" media-type="text/css"/>',
'<item id="ncx" href="toc.ncx" media-type="application/x-dtbncx+xml"/>']
spine = []
for i, (name, _) in enumerate(files):
items.append('<item id="c%d" href="%s" media-type="application/xhtml+xml"/>'
% (i, name))
spine.append('<itemref idref="c%d"/>' % i)
mime = {".png": "image/png", ".jpg": "image/jpeg", ".jpeg": "image/jpeg",
".gif": "image/gif", ".svg": "image/svg+xml"}
for i, img in enumerate(used + ([cover] if cover else [])):
items.append('<item id="i%d" href="images/%s" media-type="%s"%s/>'
% (i, img, mime.get(Path(img).suffix.lower(), "image/jpeg"),
' properties="cover-image"' if img == meta["cover"] else ""))
z.writestr("OEBPS/content.opf", """<?xml version="1.0" encoding="utf-8"?>
<package xmlns="http://www.idpf.org/2007/opf" version="3.0" unique-identifier="bid">
<metadata xmlns:dc="http://purl.org/dc/elements/1.1/">
<dc:identifier id="bid">%s</dc:identifier>
<dc:title>%s</dc:title>
<dc:creator>%s</dc:creator>
<dc:language>%s</dc:language>%s
</metadata>
<manifest>%s</manifest>
<spine toc="ncx">%s</spine>
</package>""" % (html.escape(meta["id"]), html.escape(meta["title"]),
html.escape(meta["author"]), meta["lang"], series_meta(meta),
"\n ".join(items), "\n ".join(spine)))
nav = "\n".join(
'<navPoint id="n%d" playOrder="%d"><navLabel><text>%s</text></navLabel>'
'<content src="%s"/></navPoint>' % (i, i + 1, html.escape(t), n)
for i, (n, t) in enumerate(files))
z.writestr("OEBPS/toc.ncx", """<?xml version="1.0" encoding="utf-8"?>
<ncx xmlns="http://www.daisy.org/z3986/2005/ncx/" version="2005-1">
<head><meta name="dtb:uid" content="%s"/></head>
<docTitle><text>%s</text></docTitle>
<navMap>%s</navMap>
</ncx>""" % (html.escape(meta["id"]), html.escape(meta["title"]), nav))
return len(chapters), len(used)
def main():
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument("source", type=Path, help="XHTML от pdf2html.py/epub2html.py")
ap.add_argument("output", type=Path, help="куда писать .epub")
ap.add_argument("--title", required=True)
ap.add_argument("--author", default="")
ap.add_argument("--lang", default="ru")
ap.add_argument("--cover", default="cover.jpg", help="имя файла в images/")
ap.add_argument("--series", default="", help="название цикла")
ap.add_argument("--series-index", default="", help="номер в цикле")
args = ap.parse_args()
meta = {"title": args.title, "author": args.author, "lang": args.lang,
"cover": args.cover, "series": args.series,
"series_index": args.series_index,
"id": "urn:uuid:" + args.output.stem.replace(" ", "-")}
n, imgs = build(args.source, args.output, meta)
size = args.output.stat().st_size / 2**20
print("%s: глав %d, картинок %d, %.1f МБ" % (args.output.name, n, imgs, size))
if __name__ == "__main__":
main()
+302 -28
View File
@@ -1,5 +1,6 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""PDF (calibre-generated) -> semantic XHTML, preserving italic/bold/headings.""" """PDF (calibre-generated) -> semantic XHTML, preserving italic/bold/headings."""
import collections
import html import html
import re import re
import sys import sys
@@ -13,8 +14,6 @@ if len(sys.argv) < 3 or sys.argv[1] in ("-h", "--help"):
sys.exit("usage: pdf2html.py book.pdf out.html | pdf2html.py --fonts book.pdf") sys.exit("usage: pdf2html.py book.pdf out.html | pdf2html.py --fonts book.pdf")
if sys.argv[1] == "--fonts": # разведка: какие шрифты/кегли в PDF if sys.argv[1] == "--fonts": # разведка: какие шрифты/кегли в PDF
import collections
c = collections.Counter() c = collections.Counter()
d = pymupdf.open(sys.argv[2]) d = pymupdf.open(sys.argv[2])
for p in d: for p in d:
@@ -31,23 +30,131 @@ OUT = Path(sys.argv[2])
IMGDIR = OUT.parent / "images" IMGDIR = OUT.parent / "images"
IMGDIR.mkdir(parents=True, exist_ok=True) IMGDIR.mkdir(parents=True, exist_ok=True)
# Пороги подобраны под вёрстку 15pt/letter. Для другой книги сначала # Пороги не зашиты: имена шрифтов у многих издательств вычищены до вида
# посмотреть реальные шрифты и кегли: pdf2html.py --fonts file.pdf # «Fd820825», а кегль основного текста у каждой книги свой. Профиль считается по
INDENT_X = 88 # x0 первой строки: больше — абзац с красной строки # самому файлу, увидеть его: pdf2html.py --fonts file.pdf
MONO = 8 # бит моноширинного шрифта в span["flags"] (PyMuPDF) ITALIC, MONO, BOLD = 2, 8, 16 # биты span["flags"] в PyMuPDF
MIN_IMG = 24 # pt: всё мельче — линейки, буллеты и прочая вёрстка, не иллюстрации
MAX_IMG_SHARE = 0.6 # доля площади страницы: больше — это скан полосы, а не рисунок
# Запасное опознание листинга, когда издатель вычистил и имена шрифтов, и флаги:
# кегль мельче основного текста плюс пунктуация, которой в прозе не бывает.
CODEY = re.compile(r"[{}();\[\]<>=*/#]|::|->")
ROTATED = 0.1 # синус угла строки: больше — текст повёрнут, а не набран по горизонтали
HEAD_BAND = 0.07 # доля высоты полосы сверху, где лежит колонтитул, а не текст
SOFT_HYPHEN = "\u00ad"
OCR = False
def style(font): def _image_png(raw):
"""Байты картинки блока -> настоящий PNG.
b["ext"] врёт: на «Запускаем Ansible» 3-го изд. блоки с ext="png" и
colorspace=4 на деле — CMYK JPEG (magic FFD8, подтверждено `file`), и
PNG-инструменты (pngquant) отказываются их декодировать с именем .png.
Вместо доверия заявленному формату декодируем через Pixmap (MuPDF сам
распознаёт реальный формат по данным) и всегда отдаём RGB PNG — тогда
имя файла (.png) не расходится с содержимым и CMYK не ловит инверсию
цвета в читалках, которые не ждут Adobe-JPEG с четырьмя каналами.
"""
try:
pix = pymupdf.Pixmap(raw)
if pix.colorspace is not None and pix.colorspace.n not in (1, 3):
pix = pymupdf.Pixmap(pymupdf.csRGB, pix)
return pix.tobytes("png")
except Exception:
return raw # не смогли декодировать — пишем как есть, лучше так, чем ничего
def style(font, flags=0):
"""Полужирный и курсив: сначала флаги, затем имя — имена бывают пустыми."""
f = font.lower() f = font.lower()
return ("bold" in f or "semibold" in f, "-it" in f or "italic" in f) return (bool(flags & BOLD) or "bold" in f or "semibold" in f,
bool(flags & ITALIC) or "-it" in f or "italic" in f)
def is_mono(s): def is_mono(s):
"""Моноширинный шрифт = листинг кода. Признак берётся из флагов PyMuPDF, а """Моноширинный шрифт = листинг кода. Признак из флагов, а не из имени:
не из имени шрифта: имена у каждого издательства свои, флаг одинаковый.""" имена у каждого издательства свои, флаг одинаковый.
Исключение — невидимый слой Tesseract: он весь набран GlyphLessFont с
выставленным моно-битом, и без этой проверки книга целиком опознаётся как
один сплошной листинг (измерено на Ansible после OCR: pre 10498, p 0).
В таком слое листинги ловятся кеглем и пунктуацией, см. block_kind().
"""
if "glyphless" in s["font"].lower():
return False
return bool(s["flags"] & MONO) return bool(s["flags"] & MONO)
def ocr_layer(doc, step=23):
"""Текст сделан OCR-ом: у всего слоя один невидимый шрифт. Кегль в таком
слое подгоняется под рамку слова и гуляет (9.0, 9.2, 9.4 в одном абзаце),
поэтому дальше он округляется до целого — иначе гистограмма размазывается
и основной кегль книги определяется неверно."""
glyphless = total = 0
for pno in range(1, doc.page_count, step):
for b in doc[pno].get_text("dict")["blocks"]:
for l in b.get("lines", []):
for s in l["spans"]:
n = len(s["text"])
total += n
if "glyphless" in s["font"].lower():
glyphless += n
return total > 0 and glyphless / total > 0.9
def size_of(span):
"""Кегль спана. В OCR-слое округляется до целого: см. ocr_layer()."""
return round(span["size"]) if OCR else round(span["size"], 1)
def profile(doc, step=7):
"""(кегль текста, x0 красной строки, кегль глав) по выборке страниц.
Кегль глав берётся из гистограммы, а не как «в полтора раза больше текста»:
у книги обычно два-три размера заголовков, и если объявить главой каждый из
них, оглавление на читалке распухает до трёхсот пунктов вместо двадцати.
"""
size_chars = collections.Counter()
head_blocks = collections.Counter()
head_len = {}
x0 = collections.Counter()
for pno in range(1, doc.page_count, step):
for b in doc[pno].get_text("dict")["blocks"]:
spans = [sp for l in b.get("lines", []) for sp in l["spans"]
if sp["text"].strip()]
if not spans:
continue
for l in b["lines"]:
x0[round(l["bbox"][0])] += 1
for sp in spans:
size_chars[size_of(sp)] += len(sp["text"])
top = max(spans, key=lambda sp: len(sp["text"]))
head_blocks[size_of(top)] += 1
head_len.setdefault(size_of(top), []).append(
"".join(sp["text"] for sp in spans).strip())
if not size_chars:
sys.exit("в PDF нет извлекаемого текста — нужен OCR, этот скрипт не поможет")
body = size_chars.most_common(1)[0][0]
# кандидаты в главы: крупнее текста и встречаются не единожды (единичный
# размер — это титул, а не уровень заголовка)
# Ещё условие: у кандидата должны быть слова, а не цифры. У Нейгарда номер
# главы набран кеглем 100 отдельным блоком («1», «2», …), и без проверки
# главой объявлялся он, а настоящие заголовки в 30 пунктов уезжали в
# разделы — оглавление выходило пустым. Отбор идёт по наличию букв, а не по
# длине: «Preface» — семь знаков, и порог по длине отсекал бы и его.
def _wordy(sz):
txts = head_len.get(sz) or []
letters = sum(1 for t in txts if re.search(r"[^\W\d_]", t))
return bool(txts) and letters * 2 >= len(txts)
cand = sorted((sz for sz, n in head_blocks.items()
if sz >= body * 1.15 and n >= 3 and _wordy(sz)), reverse=True)
h1 = cand[0] if cand else body * 1.55
left = min(x for x, n in x0.most_common(4))
return body, left + max(4, body * 0.6), h1
def span_html(s, plain=False): def span_html(s, plain=False):
t = html.escape(s["text"]) t = html.escape(s["text"])
if not t: if not t:
@@ -56,7 +163,7 @@ def span_html(s, plain=False):
return t return t
if is_mono(s): if is_mono(s):
return "<code>%s</code>" % t return "<code>%s</code>" % t
bold, ital = style(s["font"]) bold, ital = style(s["font"], s["flags"])
if ital: if ital:
t = "<i>%s</i>" % t t = "<i>%s</i>" % t
if bold: if bold:
@@ -64,23 +171,39 @@ def span_html(s, plain=False):
return t return t
def block_kind(b): def looks_like_code(spans):
"""(tag, css class) from dominant span font/size.""" """Доля кодовой пунктуации. Порог низкий: в прозе на 200 знаков не набирается
и десятка скобок со звёздочками, в листинге — набирается всегда."""
text = "".join(s["text"] for s in spans)
return len(text) > 8 and len(CODEY.findall(text)) / len(text) > 0.03
def is_rotated(b):
"""Строки блока набраны не по горизонтали (боковая врезка, подпись к таблице)."""
return bool(b.get("lines")) and any(abs(l["dir"][1]) > ROTATED for l in b["lines"])
def block_kind(b, body, h1_size):
"""(tag, css class) по кеглю блока относительно основного текста книги."""
spans = [s for l in b["lines"] for s in l["spans"] if s["text"].strip()] spans = [s for l in b["lines"] for s in l["spans"] if s["text"].strip()]
if not spans: if not spans:
return None return None
top = max(spans, key=lambda s: len(s["text"])) top = max(spans, key=lambda s: len(s["text"]))
f, sz = top["font"], top["size"] sz = size_of(top)
if sz >= 28 or f.startswith("Arvo") or f.startswith("CloisterBlack"):
return ("h1", "title")
if sz >= 24:
return ("h1", None)
if all(is_mono(sp) for sp in spans): if all(is_mono(sp) for sp in spans):
return ("pre", None) return ("pre", None)
if sz >= 17: if sz <= body * 0.93 and looks_like_code(spans):
return ("pre", None)
if sz >= h1_size * 1.35:
return ("h1", "title")
if sz >= h1_size * 0.97:
return ("h1", None)
# 1.10, а не 1.15: у Вернона подзаголовок набран 17.2 при тексте 15.0, то
# есть ровно на пять сотых ниже прежнего порога — и все 85 разделов книги
# уезжали в прозу. Тот же множитель, что у split_by_size, иначе разрез и
# классификация расходятся.
if sz >= body * 1.10:
return ("h2", None) return ("h2", None)
if f.startswith("MyriadPro"):
return ("p", "note") # Maxine's journal / chat / slides
return ("p", None) return ("p", None)
@@ -90,7 +213,10 @@ def block_text(b, pre=False):
# слова. Ни склейки строк, ни де-дефисации здесь быть не должно. # слова. Ни склейки строк, ни де-дефисации здесь быть не должно.
rows = ["".join(span_html(s, plain=True) for s in l["spans"]) rows = ["".join(span_html(s, plain=True) for s in l["spans"])
for l in b["lines"]] for l in b["lines"]]
return "\n".join(rows).rstrip() # В настоящем коде мягкого переноса не бывает: если он тут есть, блок
# опознан как листинг ошибочно (обычно это таблица опций), и слово надо
# склеить, а не оставить разорванным переводом строки.
return "\n".join(rows).replace(SOFT_HYPHEN + "\n", "").rstrip()
out = [] out = []
for i, l in enumerate(b["lines"]): for i, l in enumerate(b["lines"]):
line = "".join(span_html(s) for s in l["spans"]) line = "".join(span_html(s) for s in l["spans"])
@@ -98,7 +224,12 @@ def block_text(b, pre=False):
continue continue
if out: if out:
prev = out[-1] prev = out[-1]
if prev.endswith("-") and not prev.endswith("--"): # Русские издания переносят мягким дефисом U+00AD, а не обычным:
# без этой ветки строки склеиваются через пробел и слово рвётся
# пополам («введен ные»). Измерено на Шоттсе (Питер, 2020).
if prev.endswith(SOFT_HYPHEN):
out[-1] = prev[:-1]
elif prev.endswith("-") and not prev.endswith("--"):
out[-1] = prev[:-1] # de-hyphenate out[-1] = prev[:-1] # de-hyphenate
else: else:
out.append(" ") out.append(" ")
@@ -107,10 +238,35 @@ def block_text(b, pre=False):
txt = re.sub(r"</i>(\s*)<i>", r"\1", txt) txt = re.sub(r"</i>(\s*)<i>", r"\1", txt)
txt = re.sub(r"</b>(\s*)<b>", r"\1", txt) txt = re.sub(r"</b>(\s*)<b>", r"\1", txt)
txt = re.sub(r"(?<=\S) {2,}(?=\S)", " ", txt) txt = re.sub(r"(?<=\S) {2,}(?=\S)", " ", txt)
# Мягкие переносы внутри строки читалке не нужны и ломают поиск по тексту.
txt = txt.replace(SOFT_HYPHEN, "")
return txt.strip() return txt.strip()
if SRC.name == "--selftest": # python3 pdf2html.py --selftest --selftest
# Повёрнутый текст ловится не по буквам, а по направлению строки: у
# «aoHegoduenueArdng» из боковой врезки Брикмана структура слова приличная.
assert is_rotated({"lines": [{"dir": (0.0, 1.0)}]})
assert is_rotated({"lines": [{"dir": (0.0, -1.0)}]})
assert not is_rotated({"lines": [{"dir": (1.0, 0.0)}, {"dir": (1.0, 0.01)}]})
for bad in ("30HWod1Q", "o|qisuy", "1 2 3 4", "|", "0Q1 v9 8|", "В",
"9120905", "LF. FF. FT. TIT", "ee. лов. о", "лов", "TIT. ITITIT", "TT",
"‘ая. ®No TERRAFORM", "юн"):
assert bookhtml.looks_garbled(bad), bad
for good in ("Глава 1. Зачем нужен Terraform", "C++11 и лямбды", "Python3",
"Модули", "Часть II. Основы", "Что дальше?", "Часть III", "API",
"AI", "Резюме", "06 авторе", "systemd в деталях"):
assert not bookhtml.looks_garbled(good), good
print("selftest OK")
sys.exit()
doc = pymupdf.open(SRC) doc = pymupdf.open(SRC)
OCR = ocr_layer(doc)
BODY, INDENT_X, H1_SIZE = profile(doc)
if OCR:
print("слой распознан OCR-ом: моно-признак отключён, кегль округляется")
print("кегль текста %.1f, глав %.1f, красная строка от x0=%.0f"
% (BODY, H1_SIZE, INDENT_X))
parts = [] parts = []
open_para = None # (tag, cls, text) still collecting open_para = None # (tag, cls, text) still collecting
@@ -118,28 +274,110 @@ open_para = None # (tag, cls, text) still collecting
pix = doc[0].get_pixmap(dpi=150) pix = doc[0].get_pixmap(dpi=150)
pix.save(IMGDIR / "cover.jpg") pix.save(IMGDIR / "cover.jpg")
def running_heads(doc, band, limit=80, min_pages=5):
"""Тексты, повторяющиеся в верхнем поле на многих полосах.
Позиционный признак в одиночку рубит и настоящие заголовки: у книг,
свёрстанных calibre, глава начинается ровно с верха полосы и короче 80
знаков. У Вернона так пропали 14 заголовков из 15. Повтор по десяткам
страниц — то, чем колонтитул отличается от заголовка.
"""
seen = collections.Counter()
for page in doc:
for b in page.get_text("dict")["blocks"]:
if b.get("type") == 1 or "lines" not in b:
continue
if b["bbox"][1] >= page.rect.height * band:
continue
txt = "".join(sp["text"] for l in b["lines"] for sp in l["spans"])
txt = re.sub(r"\d+", "", txt).strip().lower()
if txt and len(txt) < limit:
seen[txt] += 1
return {t for t, n in seen.items() if n >= min_pages}
def split_by_size(b, body):
"""Разрезать блок, где шапка набрана крупнее следующего за ней текста.
У Вернона 85 подзаголовков лежат в одном блоке с первым абзацем раздела:
строка 17.2 и сразу за ней 15.0. Без разреза они становятся частью абзаца,
оглавление теряет второй уровень, а текст начинается с заголовка без точки.
"""
lines = [l for l in b.get("lines", [])
if any(sp["text"].strip() for sp in l["spans"])]
if len(lines) < 2:
return [b]
big = [max(size_of(sp) for sp in l["spans"] if sp["text"].strip())
> body * 1.12 for l in lines]
if not big[0] or all(big):
return [b]
cut = big.index(False)
if not any(big[:cut]) or any(big[cut:]):
return [b] # разрез только когда крупное строго сверху
head = dict(b, lines=lines[:cut],
bbox=(b["bbox"][0], b["bbox"][1], b["bbox"][2],
lines[cut - 1]["bbox"][3]))
rest = dict(b, lines=lines[cut:],
bbox=(b["bbox"][0], lines[cut]["bbox"][1], b["bbox"][2],
b["bbox"][3]))
return [head, rest]
HEADS = running_heads(doc, HEAD_BAND)
print("колонтитулов опознано по повтору:", len(HEADS))
img_n = 0 img_n = 0
for pno, page in enumerate(doc): for pno, page in enumerate(doc):
if pno == 0: if pno == 0:
continue continue
blocks = sorted(page.get_text("dict")["blocks"], key=lambda b: b["bbox"][1]) blocks = sorted(page.get_text("dict")["blocks"], key=lambda b: b["bbox"][1])
blocks = [part for b in blocks
for part in (split_by_size(b, BODY) if b.get("type") != 1 else [b])]
for bi, b in enumerate(blocks): for bi, b in enumerate(blocks):
# Повёрнутый на 90° текст: боковые врезки и подписи к таблицам. В книге
# их единицы, а распознаются они всегда в кашу («aoHegoduenueArdng») и
# лезут в заголовки, потому что кегль у них крупный. Замерено у Брикмана:
# 105 строк из 17 300, все до одной — брак. Настоящая повёрнутая таблица
# тоже отсеется, но её всё равно нечем показать в потоковой вёрстке.
if is_rotated(b):
continue
if b["type"] == 1: if b["type"] == 1:
x0, y0, x1, y1 = b["bbox"]
if x1 - x0 < MIN_IMG or y1 - y0 < MIN_IMG:
continue # линейка или буллет, а не иллюстрация
page_area = page.rect.width * page.rect.height
if page_area and (x1 - x0) * (y1 - y0) / page_area > MAX_IMG_SHARE:
continue # подложка отсканированной полосы: текст берём из OCR-слоя
img_n += 1 img_n += 1
name = "img%02d.png" % img_n name = "img%02d.png" % img_n
(IMGDIR / name).write_bytes(b["image"]) (IMGDIR / name).write_bytes(_image_png(b["image"]))
if open_para: if open_para:
parts.append(open_para) parts.append(open_para)
open_para = None open_para = None
parts.append(("figure", None, '<img src="images/%s"/>' % name)) parts.append(("figure", None, '<img src="images/%s"/>' % name, False))
continue continue
kind = block_kind(b) kind = block_kind(b, BODY, H1_SIZE)
if not kind: if not kind:
continue continue
tag, cls = kind tag, cls = kind
# Глава открывает полосу, раздел встречается по тексту. Признак —
# верхняя треть страницы, а не «первый блок»: сверху нередко идёт
# колонтитул или плашка, и тогда первым заголовок не бывает никогда.
at_top = b["bbox"][1] < page.rect.height * 0.35
txt = block_text(b, pre=(tag == "pre")) txt = block_text(b, pre=(tag == "pre"))
if not txt: if not txt:
continue continue
# Колонтитул: верхние 7% полосы — поле, а не текст. У Брикмана так
# отсеивается 21 блок крупного кегля из 511, все до одного — шапки.
# ponytail: признак позиционный. Надёжнее — текст, повторяющийся на
# десятках полос, но за это платить вторым проходом по документу.
if (b["bbox"][1] < page.rect.height * HEAD_BAND and len(txt) < 80
and re.sub(r"\d+", "", bookhtml.strip_tags(txt)
if hasattr(bookhtml, "strip_tags")
else re.sub(r"<[^>]+>", "", txt)).strip().lower() in HEADS):
continue
if tag in ("h1", "h2") and bookhtml.looks_garbled(txt):
tag, cls = "p", None
indented = b["lines"][0]["bbox"][0] >= INDENT_X indented = b["lines"][0]["bbox"][0] >= INDENT_X
cont = ( cont = (
open_para open_para
@@ -151,14 +389,50 @@ for pno, page in enumerate(doc):
) )
if cont: if cont:
join = "" if open_para[2].endswith("-") else " " join = "" if open_para[2].endswith("-") else " "
open_para = (tag, cls, open_para[2] + join + txt) open_para = (tag, cls, open_para[2] + join + txt, open_para[3])
continue
# Соседние листинги склеиваются в один блок: OCR-слой отдаёт каждую
# строку кода отдельным блоком, и без склейки листинг рассыпается на
# десяток однострочных <pre> подряд. Строки внутри блока значимы,
# поэтому соединяются переводом строки, а не пробелом.
if OCR and tag == "pre" and open_para and open_para[0] == "pre":
open_para = (tag, cls, open_para[2] + "\n" + txt, open_para[3])
continue continue
if open_para: if open_para:
parts.append(open_para) parts.append(open_para)
open_para = (tag, cls, txt) open_para = (tag, cls, txt, at_top)
if open_para: if open_para:
parts.append(open_para) parts.append(open_para)
# Понижаем заголовки не с верха полосы до раздела — но только если глав после
# этого остаётся достаточно: у части книг главы начинаются в середине страницы,
# и тогда правило обнулило бы оглавление целиком.
top_h1 = sum(1 for t, c, x, top in parts if t == "h1" and c is None and top)
if top_h1 >= 5:
parts = [(("h2" if (t == "h1" and c is None and not top) else t), c, x, top)
for t, c, x, top in parts]
parts = [(t, c, x) for t, c, x, _ in parts]
# «1» и «Списки» приезжают отдельными блоками: номер главы набран крупно, её
# название — отдельной строкой. В оглавлении читалки нужен один пункт.
merged = []
for part in parts:
if (part[0] == "h1" and merged and merged[-1][0] == "h1"
and len(merged[-1][2]) < 60):
merged[-1] = ("h1", merged[-1][1], merged[-1][2] + ". " + part[2])
continue
if (part[0] == "h2" and merged and merged[-1][0] == "h1"
and len(merged[-1][2]) < 12):
merged[-1] = ("h1", merged[-1][1], merged[-1][2] + ". " + part[2])
continue
merged.append(part)
parts = merged
# Проверка повторяется после склейки: по отдельности «LF», «FF» и «TIT» проходят
# как аббревиатуры, а склеенные в один заголовок — это обрывок таблицы.
parts = [(("p", None, x) if t in ("h1", "h2") and bookhtml.looks_garbled(x)
else (t, c, x)) for t, c, x in parts]
doc_html = bookhtml.document(doc.metadata.get("title") or SRC.stem, parts) doc_html = bookhtml.document(doc.metadata.get("title") or SRC.stem, parts)
# merge tag runs split at page/line boundaries. Пробелы здесь уже не трогаем: # merge tag runs split at page/line boundaries. Пробелы здесь уже не трогаем:
# внутри <pre> они значимы, а в прозе схлопнуты в block_text. # внутри <pre> они значимы, а в прозе схлопнуты в block_text.
+72 -1
View File
@@ -113,7 +113,78 @@ def test_convert(tmp: Path):
assert text.count("<h1>") == 3 and "<h2>" not in text assert text.count("<h1>") == 3 and "<h2>" not in text
# Дамп, собранный конвертером из PDF: ни одного заголовка, ни одного <pre>,
# только <p> с обфусцированными классами и таблица стилей.
FLAT_CSS = """
.cls_head { font-size: 1.66667em; font-family: Reg; }
.cls_body { font-size: 1em; font-family: Reg; }
.cls_code { font-size: 0.9em; font-family: lucidasanstypewriter; }
"""
FLAT_OPF = """<?xml version="1.0"?>
<package xmlns="http://www.idpf.org/2007/opf" version="2.0">
<metadata/>
<manifest>
<item id="s" href="style.css" media-type="text/css"/>
<item id="a" href="p1.xhtml" media-type="application/xhtml+xml"/>
<item id="b" href="p2.xhtml" media-type="application/xhtml+xml"/>
<item id="c" href="p3.xhtml" media-type="application/xhtml+xml"/>
</manifest>
<spine><itemref idref="a"/><itemref idref="b"/><itemref idref="c"/></spine>
</package>"""
def flat_doc(n):
return """<?xml version="1.0" encoding="utf-8"?>
<html xmlns="http://www.w3.org/1999/xhtml"><body>
<p class="cls_head">Chapter %d</p>
<p class="cls_head">Title Of Chapter %d</p>
<p class="cls_body">Body text of the chapter, long enough to be a paragraph.</p>
<p class="cls_code">def f(x):</p>
<p class="cls_code"> return x + 1</p>
<p class="cls_body">Closing paragraph of the chapter goes here.</p>
</body></html>""" % (n, n)
def test_flat_dump_recovered(tmp: Path):
src = tmp / "flat.epub"
with zipfile.ZipFile(src, "w") as z:
z.writestr("mimetype", "application/epub+zip")
z.writestr("META-INF/container.xml", CONTAINER)
z.writestr("OEBPS/content.opf", FLAT_OPF)
z.writestr("OEBPS/style.css", FLAT_CSS)
for i in (1, 2, 3):
z.writestr("OEBPS/p%d.xhtml" % i, flat_doc(i))
out = tmp / "flat.html"
epub2html.convert(src, out)
text = out.read_text(encoding="utf-8")
# заголовки восстановлены по кеглю и склеены с номером главы
assert "<h1>Chapter 2. Title Of Chapter 2</h1>" in text, text
assert text.count("<h1>") == 3
# листинг восстановлен по гарнитуре, соседние строки склеены в один блок
assert "<pre>def f(x):\n return x + 1</pre>" in text, text
assert text.count("<pre>") == 3
# обычный текст остался абзацем
assert "<p>Body text of the chapter" in text
def test_real_markup_wins_over_css(tmp: Path):
"""Книга со своей разметкой не должна уезжать в запасной путь."""
src = tmp / "book.epub"
build_epub(src)
out = tmp / "book.html"
epub2html.convert(src, out)
text = out.read_text(encoding="utf-8")
assert text.count("<h1>") == 3, "заголовки книги должны остаться её собственными"
assert "def main():" in text
if __name__ == "__main__": if __name__ == "__main__":
for case in (test_convert, test_flat_dump_recovered,
test_real_markup_wins_over_css):
with tempfile.TemporaryDirectory() as d: with tempfile.TemporaryDirectory() as d:
test_convert(Path(d)) case(Path(d))
print("OK") print("OK")
+117
View File
@@ -0,0 +1,117 @@
#!/usr/bin/env python3
"""Самопроверка упаковщика без сети: python3 test_pack_epub.py"""
import re
import tempfile
import zipfile
from pathlib import Path
import pack_epub
BOOK = """<?xml version="1.0" encoding="utf-8"?>
<html xmlns="http://www.w3.org/1999/xhtml"><head><title>T</title>
<style>p { margin: 0; }</style></head><body>
<p>Предисловие до первой главы.</p>
<h1>Глава 1. Начало</h1>
<p>Текст первой главы.</p>
<pre>def f():
return 1</pre>
<p class="figure"><img src="images/fig.png"/></p>
<h1>Глава 2. Продолжение</h1>
<p>Текст второй главы.</p>
</body></html>"""
def test_pack(tmp: Path):
src = tmp / "book.html"
src.write_text(BOOK, encoding="utf-8")
(tmp / "images").mkdir()
(tmp / "images" / "fig.png").write_bytes(b"\x89PNG\r\n\x1a\n")
(tmp / "images" / "cover.jpg").write_bytes(b"\xff\xd8\xff")
out = tmp / "book.epub"
chapters, imgs = pack_epub.build(src, out, {
"title": "Книга & книга", "author": "Автор", "lang": "ru",
"cover": "cover.jpg", "id": "urn:uuid:test"})
# преамбула до первой главы не теряется, поэтому глав три, а не две
assert (chapters, imgs) == (3, 1), (chapters, imgs)
z = zipfile.ZipFile(out)
names = z.namelist()
# без этого часть читалок не опознаёт файл как EPUB
assert names[0] == "mimetype"
assert z.getinfo("mimetype").compress_type == zipfile.ZIP_STORED
assert "OEBPS/images/fig.png" in names and "OEBPS/images/cover.jpg" in names
opf = z.read("OEBPS/content.opf").decode("utf-8")
assert "<dc:title>Книга &amp; книга</dc:title>" in opf, "разметку надо экранировать"
assert opf.count("<itemref") == 3
assert 'properties="cover-image"' in opf
ncx = z.read("OEBPS/toc.ncx").decode("utf-8")
assert re.findall(r"<text>([^<]*)</text>", ncx)[1:] == [
"Начало", "Глава 1. Начало", "Глава 2. Продолжение"]
ch1 = z.read("OEBPS/ch001.xhtml").decode("utf-8")
assert "<pre>def f():\n return 1</pre>" in ch1, "переносы в листинге значимы"
assert 'href="style.css"' in ch1
assert "Текст второй главы" not in ch1, "главы не должны склеиваться"
def test_cover_used_in_text_is_not_duplicated(tmp: Path):
"""Обложка, встречающаяся и в тексте, не должна попадать в архив дважды."""
src = tmp / "book.html"
src.write_text("<h1>Глава</h1>\n"
'<p class="figure"><img src="images/cover.jpg"/></p>\n'
"<p>Текст главы.</p>", encoding="utf-8")
(tmp / "images").mkdir()
(tmp / "images" / "cover.jpg").write_bytes(b"\xff\xd8\xff")
out = tmp / "book.epub"
pack_epub.build(src, out, {"title": "t", "author": "a", "lang": "ru",
"cover": "cover.jpg", "id": "i"})
names = zipfile.ZipFile(out).namelist()
assert names.count("OEBPS/images/cover.jpg") == 1, names
opf = zipfile.ZipFile(out).read("OEBPS/content.opf").decode()
assert opf.count('href="images/cover.jpg"') == 1, opf
def test_control_chars_stripped(tmp: Path):
"""Управляющий знак из PDF делает главу неразбираемой как XML.
У Бейера («Site Reliability Engineering») так падали 33 главы из 45:
один \x02 в начале абзаца, и читалка молча спотыкается на файле.
"""
import xml.etree.ElementTree as ET
src = tmp / "book.html"
src.write_text(BOOK.replace("<p>Текст первой главы.</p>",
"<p>\x02Текст первой главы.</p>"),
encoding="utf-8")
(tmp / "images").mkdir()
(tmp / "images" / "fig.png").write_bytes(b"\x89PNG\r\n\x1a\n")
(tmp / "images" / "cover.jpg").write_bytes(b"\xff\xd8\xff")
out = tmp / "book.epub"
pack_epub.build(src, out, {"title": "T", "author": "A", "lang": "ru",
"cover": "cover.jpg", "id": "urn:uuid:test"})
z = zipfile.ZipFile(out)
for name in z.namelist():
if name.endswith((".xhtml", ".opf", ".ncx")):
ET.fromstring(z.read(name)) # падает, если знак остался
def test_no_content(tmp: Path):
src = tmp / "empty.html"
src.write_text("<html><body>\n</body></html>", encoding="utf-8")
try:
pack_epub.build(src, tmp / "x.epub", {"title": "t", "author": "", "lang": "ru",
"cover": "", "id": "i"})
except SystemExit:
return
raise AssertionError("пустой источник должен останавливать упаковку")
if __name__ == "__main__":
for case in (test_pack, test_cover_used_in_text_is_not_duplicated,
test_no_content):
with tempfile.TemporaryDirectory() as d:
case(Path(d))
print("OK")
+55
View File
@@ -105,6 +105,59 @@ def test_code_listings_are_left_alone(tmp: Path):
assert t.verify_output(src, "ru"), "код не должен заваливать приёмку" assert t.verify_output(src, "ru"), "код не должен заваливать приёмку"
def test_junk_translation_falls_back_to_original(tmp: Path):
"""Модель иногда возвращает на длинный абзац огрызок «, v», а число абзацев
при этом сходится — прежние ворота такое пропускали."""
src = tmp / "book.html"
long_en = ("A vector is simply a sequence of elements that you can access by "
"an index, and it is the workhorse of the standard library. ") * 2
src.write_text("\n".join([
"<h1>Chapter</h1>",
"<p>%s</p>" % long_en,
"<p>Second paragraph of the very same chapter, also reasonably long.</p>",
]), encoding="utf-8")
chapters = t.split_chapters(t.parse_blocks(src))
trans = tmp / "translations"
trans.mkdir()
(trans / "chapter_000_translated.json").write_text(json.dumps(
{"number": 0, "paragraphs": ["Глава", ", v",
"Второй абзац той же самой главы, тоже достаточно длинный."]},
ensure_ascii=False), encoding="utf-8")
out = tmp / "out.html"
t.rebuild(src, chapters, tmp, out)
result = out.read_text(encoding="utf-8")
assert ", v" not in result, "огрызок не должен попадать в книгу"
assert "A vector is simply a sequence" in result, "вместо огрызка нужен оригинал"
assert "Второй абзац" in result, "нормальный перевод должен остаться"
def test_untouched_paragraphs_are_counted(tmp: Path):
"""Абзац, вернувшийся по-английски, доля кириллицы по документу не ловит:
у Страуструпа так осталось 182 упражнения из 10151 блока."""
src = tmp / "book.html"
en = ("Expanding on what you have learned, write a program that lists the "
"instructions for a computer to find the upstairs bedroom. ") * 2
ru_src = ("Этот абзац достаточно длинный, чтобы попасть под проверку языка "
"и быть переведённым как положено. ") * 2
src.write_text("\n".join(["<h1>Chapter</h1>", "<p>%s</p>" % en, "<p>%s</p>" % en]),
encoding="utf-8")
chapters = t.split_chapters(t.parse_blocks(src))
trans = tmp / "translations"
trans.mkdir()
(trans / "chapter_000_translated.json").write_text(json.dumps(
{"number": 0, "paragraphs": ["Глава", en, ru_src]}, ensure_ascii=False),
encoding="utf-8")
out = tmp / "out.html"
import io, contextlib
buf = io.StringIO()
with contextlib.redirect_stdout(buf):
t.rebuild(src, chapters, tmp, out)
report = buf.getvalue()
assert "осталось на языке оригинала: 1" in report, report
def test_workdir_belongs_to_one_book(tmp: Path): def test_workdir_belongs_to_one_book(tmp: Path):
"""Имена chapter_NNN.json у всех книг одинаковы: чужой workdir склеит """Имена chapter_NNN.json у всех книг одинаковы: чужой workdir склеит
перевод одной книги с текстом другой.""" перевод одной книги с текстом другой."""
@@ -155,6 +208,8 @@ if __name__ == "__main__":
test_broken_marks_drop_tags() test_broken_marks_drop_tags()
test_language_detection() test_language_detection()
for case in (test_split_and_rebuild, test_code_listings_are_left_alone, for case in (test_split_and_rebuild, test_code_listings_are_left_alone,
test_junk_translation_falls_back_to_original,
test_untouched_paragraphs_are_counted,
test_workdir_belongs_to_one_book, test_stale_chapters_removed, test_workdir_belongs_to_one_book, test_stale_chapters_removed,
test_verify_output_catches_untranslated): test_verify_output_catches_untranslated):
with tempfile.TemporaryDirectory() as d: with tempfile.TemporaryDirectory() as d:
+23 -5
View File
@@ -148,9 +148,15 @@ def run_translator(repo, workdir, extracted, workers):
sys.exit("book_translator завершился с кодом %d" % r.returncode) sys.exit("book_translator завершился с кодом %d" % r.returncode)
# Доля от длины оригинала, ниже которой перевод считается мусором. Совпадение
# числа абзацев ничего не гарантирует: модель иногда возвращает на длинный абзац
# огрызок вида «, v», и прежние ворота такое пропускали.
MIN_LEN_SHARE = 0.25
def rebuild(src, chapters, workdir, out): def rebuild(src, chapters, workdir, out):
lines = src.read_text(encoding="utf-8").splitlines() lines = src.read_text(encoding="utf-8").splitlines()
broken = missing = translated = 0 broken = missing = translated = junk = untouched = 0
for n, ch in enumerate(chapters): for n, ch in enumerate(chapters):
f = workdir / "translations" / ("chapter_%03d_translated.json" % n) f = workdir / "translations" / ("chapter_%03d_translated.json" % n)
if not f.exists(): if not f.exists():
@@ -162,15 +168,27 @@ def rebuild(src, chapters, workdir, out):
% (n, len(paragraphs), len(ch))) % (n, len(paragraphs), len(ch)))
missing += len(ch) missing += len(ch)
continue continue
for (idx, tag, attrs, _), text in zip(ch, paragraphs): for (idx, tag, attrs, original), text in zip(ch, paragraphs):
plain_src = re.sub("<[^>]+>", "", original)
if len(plain_src) > 120 and len(text) < max(20, len(plain_src) * MIN_LEN_SHARE):
junk += 1
lines[idx] = "<%s%s>%s</%s>" % (tag, attrs, original, tag)
continue # оставляем оригинал: английский абзац лучше огрызка
# Абзац, вернувшийся на языке оригинала. Доля кириллицы по всему
# документу такое не ловит: полтора процента в ней тонут, а на
# странице это заметный кусок английского текста.
if len(plain_src) > 150 and detect_language(re.sub("<[^>]+>", "", text))[0] \
== detect_language(plain_src)[0] != "unknown":
untouched += 1
body, ok = from_marks(text) body, ok = from_marks(text)
broken += not ok broken += not ok
translated += 1 translated += 1
lines[idx] = "<%s%s>%s</%s>" % (tag, attrs, body, tag) lines[idx] = "<%s%s>%s</%s>" % (tag, attrs, body, tag)
out.write_text("\n".join(lines) + "\n", encoding="utf-8") out.write_text("\n".join(lines) + "\n", encoding="utf-8")
print("переведено блоков: %d, без перевода: %d, разметка потеряна в %d" print("переведено блоков: %d, без перевода: %d, разметка потеряна в %d, "
% (translated, missing, broken)) "огрызков заменено оригиналом: %d, осталось на языке оригинала: %d"
return missing == 0 % (translated, missing, broken, junk, untouched))
return missing == 0 and junk * 200 <= translated
def verify_output(out, target): def verify_output(out, target):