Локальная озвучка через Silero: нормализатор русского текста под TTS, рендер по главам, замечание про AVX2

This commit is contained in:
Илья Поляков
2026-08-03 14:58:54 +03:00
parent 3bbd75b1c3
commit 6c81673b4f
3 changed files with 346 additions and 2 deletions
+34 -2
View File
@@ -202,9 +202,41 @@ the external repo is still never edited.
stays viable. Hand the result to the `prepare-audiobooks` skill for covers
and library layout.
### Prefer local Silero over edge-tts
edge-tts drops fragments silently under load — measured 1 of 49 on one run and
7 of 49 on the next, at *fewer* workers, so it is volume- not concurrency-bound.
`scripts/silero_render.py` replaces the external synthesis stage entirely with a
local model: no network, no dropouts, no `repair` cycle, free.
```bash
pip install torch numpy --index-url https://download.pytorch.org/whl/cpu
curl -O https://models.silero.ai/models/tts/ru/v5_5_ru.pt # 145 МБ
python scripts/silero_render.py v5_5_ru.pt <workdir>/translations_tts <out> \
--album "Название" --author "Автор" --speaker eugene --tempo 0.87 \
--skip 0 1 2 3 37
```
**Check `lscpu | grep avx2` before choosing the host.** PyTorch needs AVX2;
without it inference is ~30x slower — measured 3.4x realtime on a Celeron N5095
(SSE4 only) versus 97x on a Ryzen 7 5800H. Use 8 threads, not 16: hyperthreading
loses (97x vs 83x).
`scripts/tts_normalize.py` prepares the text — Latin script is what makes a
Russian voice sound worst, and a full book carries far more of it than a sample
chapter suggests (37 unique tokens in one chapter, 728 across the book). It maps
named entities and acronyms by hand with stress marks, transliterates the rest
by rule, spells out numbers, and drops URLs. It also generates Russian chapter
titles, since translated headings often stay English and the intro fragment
would otherwise read them aloud in Latin.
Voice tempo: SSML `<prosody rate>` quantizes to named levels, so percentages
cluster instead of stepping evenly. For fine control use `--tempo`, which
time-stretches with ffmpeg and preserves pitch.
The phonetics stage (`07_extract_terms.py` + `08_generate_phonetics.py`) is
worth running first for a technical book, or the Russian voice will mangle
every English term.
worth running first for a technical book **when using edge-tts**; with
`tts_normalize.py` it is redundant.
**Ordering constraint for the whole stage:** `verify` and `split` both read
`audiobook/temp_audio/`, and the external script's `cleanup_temp_files()`