Локальная озвучка через Silero: нормализатор русского текста под TTS, рендер по главам, замечание про AVX2
This commit is contained in:
@@ -202,9 +202,41 @@ the external repo is still never edited.
|
||||
stays viable. Hand the result to the `prepare-audiobooks` skill for covers
|
||||
and library layout.
|
||||
|
||||
### Prefer local Silero over edge-tts
|
||||
|
||||
edge-tts drops fragments silently under load — measured 1 of 49 on one run and
|
||||
7 of 49 on the next, at *fewer* workers, so it is volume- not concurrency-bound.
|
||||
`scripts/silero_render.py` replaces the external synthesis stage entirely with a
|
||||
local model: no network, no dropouts, no `repair` cycle, free.
|
||||
|
||||
```bash
|
||||
pip install torch numpy --index-url https://download.pytorch.org/whl/cpu
|
||||
curl -O https://models.silero.ai/models/tts/ru/v5_5_ru.pt # 145 МБ
|
||||
python scripts/silero_render.py v5_5_ru.pt <workdir>/translations_tts <out> \
|
||||
--album "Название" --author "Автор" --speaker eugene --tempo 0.87 \
|
||||
--skip 0 1 2 3 37
|
||||
```
|
||||
|
||||
**Check `lscpu | grep avx2` before choosing the host.** PyTorch needs AVX2;
|
||||
without it inference is ~30x slower — measured 3.4x realtime on a Celeron N5095
|
||||
(SSE4 only) versus 97x on a Ryzen 7 5800H. Use 8 threads, not 16: hyperthreading
|
||||
loses (97x vs 83x).
|
||||
|
||||
`scripts/tts_normalize.py` prepares the text — Latin script is what makes a
|
||||
Russian voice sound worst, and a full book carries far more of it than a sample
|
||||
chapter suggests (37 unique tokens in one chapter, 728 across the book). It maps
|
||||
named entities and acronyms by hand with stress marks, transliterates the rest
|
||||
by rule, spells out numbers, and drops URLs. It also generates Russian chapter
|
||||
titles, since translated headings often stay English and the intro fragment
|
||||
would otherwise read them aloud in Latin.
|
||||
|
||||
Voice tempo: SSML `<prosody rate>` quantizes to named levels, so percentages
|
||||
cluster instead of stepping evenly. For fine control use `--tempo`, which
|
||||
time-stretches with ffmpeg and preserves pitch.
|
||||
|
||||
The phonetics stage (`07_extract_terms.py` + `08_generate_phonetics.py`) is
|
||||
worth running first for a technical book, or the Russian voice will mangle
|
||||
every English term.
|
||||
worth running first for a technical book **when using edge-tts**; with
|
||||
`tts_normalize.py` it is redundant.
|
||||
|
||||
**Ordering constraint for the whole stage:** `verify` and `split` both read
|
||||
`audiobook/temp_audio/`, and the external script's `cleanup_temp_files()`
|
||||
|
||||
Reference in New Issue
Block a user