Локальная озвучка через Silero: нормализатор русского текста под TTS, рендер по главам, замечание про AVX2

This commit is contained in:
Илья Поляков
2026-08-03 14:58:54 +03:00
parent 3bbd75b1c3
commit 6c81673b4f
3 changed files with 346 additions and 2 deletions
+34 -2
View File
@@ -202,9 +202,41 @@ the external repo is still never edited.
stays viable. Hand the result to the `prepare-audiobooks` skill for covers stays viable. Hand the result to the `prepare-audiobooks` skill for covers
and library layout. and library layout.
### Prefer local Silero over edge-tts
edge-tts drops fragments silently under load — measured 1 of 49 on one run and
7 of 49 on the next, at *fewer* workers, so it is volume- not concurrency-bound.
`scripts/silero_render.py` replaces the external synthesis stage entirely with a
local model: no network, no dropouts, no `repair` cycle, free.
```bash
pip install torch numpy --index-url https://download.pytorch.org/whl/cpu
curl -O https://models.silero.ai/models/tts/ru/v5_5_ru.pt # 145 МБ
python scripts/silero_render.py v5_5_ru.pt <workdir>/translations_tts <out> \
--album "Название" --author "Автор" --speaker eugene --tempo 0.87 \
--skip 0 1 2 3 37
```
**Check `lscpu | grep avx2` before choosing the host.** PyTorch needs AVX2;
without it inference is ~30x slower — measured 3.4x realtime on a Celeron N5095
(SSE4 only) versus 97x on a Ryzen 7 5800H. Use 8 threads, not 16: hyperthreading
loses (97x vs 83x).
`scripts/tts_normalize.py` prepares the text — Latin script is what makes a
Russian voice sound worst, and a full book carries far more of it than a sample
chapter suggests (37 unique tokens in one chapter, 728 across the book). It maps
named entities and acronyms by hand with stress marks, transliterates the rest
by rule, spells out numbers, and drops URLs. It also generates Russian chapter
titles, since translated headings often stay English and the intro fragment
would otherwise read them aloud in Latin.
Voice tempo: SSML `<prosody rate>` quantizes to named levels, so percentages
cluster instead of stepping evenly. For fine control use `--tempo`, which
time-stretches with ffmpeg and preserves pitch.
The phonetics stage (`07_extract_terms.py` + `08_generate_phonetics.py`) is The phonetics stage (`07_extract_terms.py` + `08_generate_phonetics.py`) is
worth running first for a technical book, or the Russian voice will mangle worth running first for a technical book **when using edge-tts**; with
every English term. `tts_normalize.py` it is redundant.
**Ordering constraint for the whole stage:** `verify` and `split` both read **Ordering constraint for the whole stage:** `verify` and `split` both read
`audiobook/temp_audio/`, and the external script's `cleanup_temp_files()` `audiobook/temp_audio/`, and the external script's `cleanup_temp_files()`
+150
View File
@@ -0,0 +1,150 @@
#!/usr/bin/env python3
"""Озвучка книги локальной моделью Silero — треки по главам, без сети.
Зачем не edge-tts: тот бесплатен, но на объёме молча теряет фрагменты
(на одной главе выпадало от 1 до 7 из 49). Silero считает локально,
детерминированно и без пропусков.
ВАЖНО про железо: PyTorch считает нейросеть инструкциями AVX2. На процессоре
без них (проверять `lscpu | grep avx2`) скорость падает примерно в тридцать
раз — на Celeron N5095 вышло 3.4x реального времени против 97x на Ryzen 5800H.
Гипертрединг вредит: 8 потоков быстрее 16.
Модель: https://models.silero.ai/models/tts/ru/v5_5_ru.pt (145 МБ)
Нужны: torch (CPU-сборка), numpy, ffmpeg.
"""
import argparse
import html
import json
import pathlib
import re
import subprocess
import sys
import time
import wave
import torch
sys.path.insert(0, str(pathlib.Path(__file__).parent))
import tts_normalize as norm
SR = 48000
def ssml(text):
"""Абзац как <p> из <s>: модель сама строит интонацию конца фразы."""
sents = [s.strip() for s in re.split(r'(?<=[.!?…])\s+', text) if s.strip()]
sents = [s for s in sents if re.search(r'\w', s)]
if not sents:
return None
return '<speak><p>%s</p></speak>' % "".join(
"<s>%s</s>" % html.escape(s, quote=False) for s in sents)
def safe(s):
return re.sub(r'\s+', ' ', re.sub(r'[/\\\x00-\x1f]', ' ', s)).strip(' .')[:70]
def main():
ap = argparse.ArgumentParser(
description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("model", type=pathlib.Path, help="v5_5_ru.pt")
ap.add_argument("texts", type=pathlib.Path,
help="каталог с chapter_*_translated*.json (вывод audiobook.py prep)")
ap.add_argument("out", type=pathlib.Path, help="куда складывать треки")
ap.add_argument("--album", required=True, help="название книги для тегов")
ap.add_argument("--author", required=True, help="автор для тегов")
ap.add_argument("--speaker", default="eugene",
help="aidar, baya, kseniya, eugene, xenia")
ap.add_argument("--tempo", type=float, default=1.0,
help="растяжение времени: 0.87 = на 13%% медленнее, высота сохраняется")
ap.add_argument("--threads", type=int, default=8)
ap.add_argument("--skip", type=int, nargs="*", default=[],
help="номера разделов, которые не озвучивать (титулы, библиография)")
args = ap.parse_args()
torch.set_num_threads(args.threads)
model = torch.package.PackageImporter(str(args.model)).load_pickle("tts_models", "model")
model.to(torch.device('cpu'))
args.out.mkdir(parents=True, exist_ok=True)
kw = dict(speaker=args.speaker, sample_rate=SR, put_accent=True, put_yo=True,
put_stress_homo=True, put_yo_homo=True)
skip = set(args.skip)
todo = []
for f in sorted(args.texts.glob('chapter_*_translated*.json')):
n = int(re.search(r'chapter_(\d+)', f.name).group(1))
if n in skip:
continue
j = json.loads(f.read_text(encoding="utf-8"))
title = norm.chapter_title(n, j.get('title', '')) or ('Раздел %03d' % n)
todo.append((n, title, j['paragraphs']))
if not todo:
sys.exit("в %s нет глав" % args.texts)
print("глав к озвучке: %d (пропущено %d)" % (len(todo), len(skip)), flush=True)
started = time.time()
for idx, (n, title, paras) in enumerate(todo, 1):
dest = args.out / ('%03d - %s.mp3' % (n, safe(title)))
if dest.exists() and dest.stat().st_size:
print(" %03d %-24s уже готово" % (n, title), flush=True)
continue
t0, chunks, failed = time.time(), [], 0
intro = ssml(title + '.')
if intro:
chunks += [model.apply_tts(ssml_text=intro, **kw), torch.zeros(int(SR * 0.6))]
for p in paras:
src = norm.normalize(p)
if not src.strip():
continue
s = ssml(src)
try:
a = model.apply_tts(ssml_text=s, **kw) if s else None
except Exception:
try: # парсер SSML спотыкается на редких абзацах
a = model.apply_tts(text=src, **kw)
except Exception:
failed += 1
continue
if a is not None:
chunks += [a, torch.zeros(int(SR * 0.30))]
if not chunks:
print(" %03d %-24s пусто, пропуск" % (n, title), flush=True)
continue
full = torch.cat(chunks)
wav = args.out / ('.render_%03d.wav' % n)
with wave.open(str(wav), 'wb') as w:
w.setnchannels(1); w.setsampwidth(2); w.setframerate(SR)
w.writeframes((full.clamp(-1, 1) * 32767).to(torch.int16).numpy().tobytes())
cmd = ['ffmpeg', '-v', 'error', '-y', '-i', str(wav)]
if abs(args.tempo - 1.0) > 1e-6:
cmd += ['-filter:a', 'atempo=%.3f' % args.tempo]
cmd += ['-b:a', '64k', '-ac', '1',
'-metadata', 'title=%s' % title,
'-metadata', 'album=%s' % args.album,
'-metadata', 'artist=%s' % args.author,
'-metadata', 'album_artist=%s' % args.author,
'-metadata', 'track=%d/%d' % (idx, len(todo)),
'-metadata', 'genre=Audiobook', str(dest)]
subprocess.run(cmd, check=True)
wav.unlink()
raw, el = len(full) / SR, time.time() - t0
print(" %03d %-24s %5.1f мин (синтез %4.1f мин, %.1fx)%s"
% (n, title, raw / args.tempo / 60, el / 60, raw / el,
' СБОЕВ: %d' % failed if failed else ''), flush=True)
files = sorted(args.out.glob('*.mp3'))
total = sum(float(subprocess.run(
['ffprobe', '-v', 'error', '-show_entries', 'format=duration',
'-of', 'default=nw=1:nk=1', str(f)],
capture_output=True, text=True).stdout or 0) for f in files)
print("\nготово: %d треков, %.1f ч, %.0f МБ, заняло %.1f ч"
% (len(files), total / 3600,
sum(f.stat().st_size for f in files) / 1024 / 1024,
(time.time() - started) / 3600))
print("каталог:", args.out)
if __name__ == '__main__':
main()
+162
View File
@@ -0,0 +1,162 @@
#!/usr/bin/env python3
"""Нормализация русского текста под русский TTS.
Латиницу русский голос читает плохо, поэтому она вся уходит в кириллицу:
именованные сущности по ручному словарю, остальное — общими правилами.
"""
import re
# Имена собственные и термины, где побуквенная транслитерация даёт мимо.
# Ударение помечается «+» перед гласной.
NAMES = {
'Parts Unlimited': 'П+артс Анлим+итед', 'Wayne-Yokohama': 'У+эйн-Йокох+ама',
'Equity Partners': '+Эквити П+артнерс', 'Data Hub': 'Д+ата Хаб',
'Phoenix': 'Ф+еникс', 'Unicorn': 'Ю+никорн', 'Narwhal': 'Н+арвал',
'Panther': 'П+антер', 'Horizon': 'Хор+айзон', 'Summit': 'С+аммит',
'Kumquat': 'К+амкват', 'Inversion': 'Инв+ерсия', 'Rebellion': 'Ребелл+ион',
'Orca': '+Орка', 'Shamu': 'Шам+у', 'Unikitty': 'Юник+итти',
'Dockside': 'Д+оксайд', 'Tomcat': 'Т+омкэт', 'Makefile': 'М+ейкфайл',
'Microsoft': 'Майкрос+офт', 'Google': 'Г+угл', 'Amazon': 'Ам+азон',
'Apple': '+Эппл', 'Netflix': 'Н+етфликс', 'Facebook': 'Ф+ейсбук',
'Twitter': 'Тв+иттер', 'YouTube': 'Ють+юб', 'Toyota': 'Той+ота',
'Nokia': 'Н+окиа', 'Alcoa': 'Алк+оа', 'Walmart': 'Уолм+арт',
'Tesla': 'Т+есла', 'Uber': '+Убер', 'Lyft': 'Лифт', 'Blockbuster': 'Блокб+астер',
'Compuware': 'Компь+юуэр', 'Symbian': 'С+имбиан', 'Windows': 'В+индоус',
'Linux': 'Л+инукс', 'Android': 'Андр+оид', 'iPhone': 'Айф+он',
'Python': 'П+айтон', 'Java': 'Дж+ава', 'Clojure': 'Кл+ожур',
'Docker': 'Д+окер', 'Git': 'Гит', 'Excel': '+Эксель', 'Word': 'Ворд',
'PowerPoint': 'П+ауэр-П+ойнт', 'SharePoint': 'Шер-П+ойнт',
'Visio': 'В+изио', 'OmniGraffle': 'Омни-Гр+аффл', 'Office': '+Офис',
'DevOps': 'Дев-+Опс', 'NoSQL': 'Ноу-эс-кью-+эль', 'Agile': '+Эджайл',
'San Francisco': 'Сан-Франц+иско', 'Las Vegas': 'Лас-В+егас',
'London': 'Л+ондон', 'Elkhart Grove': '+Элкхарт-Гр+оув',
'Gene Kim': 'Джин Ким', 'Maxine': 'Макс+ин',
}
# Аббревиатуры: как их произносят по-русски
ACRONYMS = {
'IT': 'айт+и', 'API': 'эй-пи-+ай', 'QA': 'кь+ю-эй', 'HR': 'эйч+ар',
'VP': 'вип+и', 'CIO': 'си-ай-+о', 'CEO': 'си-и-+о', 'CTO': 'си-ти-+о',
'MRP': 'эм-эр-п+и', 'ERP': 'и-ар-п+и', 'CRM': 'си-эр-+эм',
'USB': 'ю-эс-б+и', 'SKU': 'эс-кей-+ю', 'BOM': 'би-о-+эм',
'VIN': 'вин', 'ETL': 'и-ти-+эль', 'PII': 'пи-ай-+ай',
'ITIL': '+айтил', 'CI': 'си-+ай', 'CD': 'си-д+и', 'DSL': 'ди-эс-+эль',
'HTML': 'эйч-ти-эм-+эль', 'IP': 'ай-п+и', 'SQL': 'эс-кью-+эль',
'PhD': 'пи-эйч-д+и', 'ISBN': 'и-эс-би-+эн', 'OEM': 'о-и-+эм',
'TEP': 'тэп', 'LARB': 'ларб', 'CSG': 'си-эс-дж+и', 'SPSS': 'эс-пи-эс-+эс',
'PUL': 'пи-ю-+эль', 'DEVP': 'дев-п+и', 'R': 'эр', 'v': 'в+ерсия',
}
DIGITS = ['ноль', 'один', 'два', 'три', 'четыре', 'пять',
'шесть', 'семь', 'восемь', 'девять']
TENS = {10: 'десять', 11: 'одиннадцать', 12: 'двенадцать', 13: 'тринадцать',
14: 'четырнадцать', 15: 'пятнадцать', 16: 'шестнадцать',
17: 'семнадцать', 18: 'восемнадцать', 19: 'девятнадцать',
20: 'двадцать', 30: 'тридцать', 40: 'сорок', 50: 'пятьдесят',
60: 'шестьдесят', 70: 'семьдесят', 80: 'восемьдесят', 90: 'девяносто'}
HUNDREDS = {100: 'сто', 200: 'двести', 300: 'триста', 400: 'четыреста',
500: 'пятьсот', 600: 'шестьсот', 700: 'семьсот',
800: 'восемьсот', 900: 'девятьсот'}
# английская орфография -> русское звучание, длинные сочетания раньше коротких
TRANSLIT = [
('tch', 'ч'), ('sch', 'ш'), ('sh', 'ш'), ('ch', 'ч'), ('th', 'т'),
('ph', 'ф'), ('ck', 'к'), ('qu', 'кв'), ('oo', 'у'), ('ee', 'и'),
('ea', 'и'), ('ou', 'ау'), ('ow', 'оу'), ('ay', 'ей'), ('ai', 'ей'),
('ey', 'и'), ('oy', 'ой'), ('au', 'о'), ('aw', 'о'), ('igh', 'ай'),
('a', 'а'), ('b', 'б'), ('c', 'к'), ('d', 'д'), ('e', 'е'), ('f', 'ф'),
('g', 'г'), ('h', 'х'), ('i', 'и'), ('j', 'дж'), ('k', 'к'), ('l', 'л'),
('m', 'м'), ('n', 'н'), ('o', 'о'), ('p', 'п'), ('q', 'к'), ('r', 'р'),
('s', 'с'), ('t', 'т'), ('u', 'у'), ('v', 'в'), ('w', 'у'), ('x', 'кс'),
('y', 'й'), ('z', 'з'),
]
def number_to_words(n):
n = int(n)
if n < 10:
return DIGITS[n]
if n in TENS:
return TENS[n]
if n < 100:
return '%s %s' % (TENS[n // 10 * 10], DIGITS[n % 10])
if n in HUNDREDS:
return HUNDREDS[n]
if n < 1000:
return '%s %s' % (HUNDREDS[n // 100 * 100], number_to_words(n % 100))
if n < 10000 and n % 1000 == 0:
return '%s тысяч' % DIGITS[n // 1000]
return ' '.join(DIGITS[int(c)] for c in str(n))
def translit_word(w):
"""Незнакомое латинское слово — общими правилами чтения."""
low = w.lower()
out = low
for src, dst in TRANSLIT:
out = out.replace(src, dst)
out = re.sub(r'[^а-яё]', '', out)
return out.capitalize() if w[:1].isupper() else out
def normalize(text):
# ссылки читать бессмысленно
text = re.sub(r'https?://\S+', ' ссылка ', text)
text = re.sub(r'\bwww\.\S+', ' ссылка ', text)
# именованные сущности (сначала многословные)
for name in sorted(NAMES, key=len, reverse=True):
text = re.sub(re.escape(name), NAMES[name], text)
# время 6:07
text = re.sub(r'\b(\d{1,2}):(\d{2})\b',
lambda m: '%s %s' % (number_to_words(m.group(1)),
' '.join(DIGITS[int(c)] for c in m.group(2))),
text)
# идентификаторы: MRP-8, DEVP-101
text = re.sub(r'\b([A-Za-z]{2,6})-(\d+)\b',
lambda m: '%s %s' % (ACRONYMS.get(m.group(1).upper(),
translit_word(m.group(1))),
number_to_words(m.group(2))), text)
# аббревиатуры
for a in sorted(ACRONYMS, key=len, reverse=True):
text = re.sub(r'\b%s\b' % re.escape(a), ACRONYMS[a], text)
# оставшаяся латиница — общими правилами
text = re.sub(r'\b[A-Za-z][A-Za-z\'-]*\b', lambda m: translit_word(m.group(0)), text)
# числа
text = re.sub(r'\b\d+\b', lambda m: number_to_words(m.group(0)), text)
# мусор после замен
text = re.sub(r'[ \t]{2,}', ' ', text)
return text.strip()
TITLES = {
7: 'Пролог', 34: 'Эпилог',
4: 'Посвящение', 5: 'Обращение к читателю',
6: 'Действующие лица', 8: 'Газетная заметка',
9: 'Часть первая', 18: 'Часть вторая', 26: 'Часть третья',
35: 'Должностная инструкция', 36: 'Пять идеалов',
38: 'Благодарности', 39: 'Об авторе',
}
ORDINAL = ['первая', 'вторая', 'третья', 'четвёртая', 'пятая', 'шестая',
'седьмая', 'восьмая', 'девятая', 'десятая', 'одиннадцатая',
'двенадцатая', 'тринадцатая', 'четырнадцатая', 'пятнадцатая',
'шестнадцатая', 'семнадцатая', 'восемнадцатая', 'девятнадцатая']
def chapter_title(num, raw):
if num in TITLES:
return TITLES[num]
m = re.search(r'CHAPTER\s+(\d+)', raw, re.I)
if m:
i = int(m.group(1))
return 'Глава %s' % (ORDINAL[i - 1] if i <= len(ORDINAL) else number_to_words(i))
return None # титульные страницы и прочее без озвучиваемого заголовка
if __name__ == '__main__':
for s in ['Максин открыла Phoenix в SharePoint через API за 20 минут.',
'Сервер DEVP-101 и MRP-8 в Parts Unlimited, встреча в 6:07.',
'Смотри https://example.com/page и репозиторий Git.',
'Кофе с CIO и VP по вопросам IT и QA.']:
print(' ', s, '\n ->', normalize(s), '\n')