glossary.py: двуязычный глоссарий из параллельных изданий цикла
This commit is contained in:
@@ -130,6 +130,33 @@ EOF
|
||||
After packing, confirm the EPUB: `ebook-meta` for metadata, and the `.ncx`
|
||||
navPoint count for the TOC size.
|
||||
|
||||
## Glossary from existing translations
|
||||
|
||||
For a book in a series that already has published translations, `scripts/glossary.py`
|
||||
mines a bilingual glossary so the machine translation does not invent new spellings
|
||||
for names the reader already knows. Feed it pairs of editions of the *same* volume:
|
||||
|
||||
```
|
||||
python scripts/glossary.py --en vol12.fb2 --ru vol12.ru.fb2 \
|
||||
--en vol15.epub --ru vol15.ru.fb2 --score 0.6
|
||||
```
|
||||
|
||||
Two signals, and both are needed. Position: paragraph indices do not line up
|
||||
(Russian editions split dialogue, giving 2–3× more paragraphs), so offsets are
|
||||
measured as a **share of characters**, where the texts track each other closely.
|
||||
Transliteration: a proper name in Russian is nearly always a transliteration, so
|
||||
the Cyrillic candidate is romanized and compared to the English term — this is
|
||||
what turns the output from noise into a usable list.
|
||||
|
||||
Two mirrored filters remove the rest of the junk: a candidate whose head word
|
||||
also appears lowercase in the same text is a sentence-initial common word, not a
|
||||
name — applied on both sides. Measured on four Dresden Files volumes: 71 pairs,
|
||||
of which two were wrong.
|
||||
|
||||
⚠️ Concept terms (`White Council` → `Белый Совет`, `Spire` → `Копьё`) do **not**
|
||||
come out of the transliteration path and the positional one alone is too noisy
|
||||
for them. Extract those by hand and verify by grepping the existing translation.
|
||||
|
||||
## Optional stage: translation
|
||||
|
||||
Only when the user asks for a translated book. It slots between step 6 and
|
||||
|
||||
Reference in New Issue
Block a user