glossary.py: двуязычный глоссарий из параллельных изданий цикла

This commit is contained in:
chesirecatt
2026-08-21 18:41:45 +03:00
parent 14c7fad763
commit 43d6cfc456
4 changed files with 271 additions and 4 deletions
+27
View File
@@ -130,6 +130,33 @@ EOF
After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx`
navPoint count for the TOC size.
## Glossary from existing translations
For a book in a series that already has published translations, `scripts/glossary.py`
mines a bilingual glossary so the machine translation does not invent new spellings
for names the reader already knows. Feed it pairs of editions of the *same* volume:
```
python scripts/glossary.py --en vol12.fb2 --ru vol12.ru.fb2 \
--en vol15.epub --ru vol15.ru.fb2 --score 0.6
```
Two signals, and both are needed. Position: paragraph indices do not line up
(Russian editions split dialogue, giving 2–3× more paragraphs), so offsets are
measured as a **share of characters**, where the texts track each other closely.
Transliteration: a proper name in Russian is nearly always a transliteration, so
the Cyrillic candidate is romanized and compared to the English term — this is
what turns the output from noise into a usable list.
Two mirrored filters remove the rest of the junk: a candidate whose head word
also appears lowercase in the same text is a sentence-initial common word, not a
name — applied on both sides. Measured on four Dresden Files volumes: 71 pairs,
of which two were wrong.
⚠️ Concept terms (`White Council` → `Белый Совет`, `Spire` → `Копьё`) do **not**
come out of the transliteration path and the positional one alone is too noisy
for them. Extract those by hand and verify by grepping the existing translation.
## Optional stage: translation
Only when the user asks for a translated book. It slots between step 6 and