Обвязка озвучки: снятие маркеров разметки для TTS и проверка целостности синтеза

This commit is contained in:
Илья Поляков
2026-08-03 10:48:04 +03:00
parent b92a0a39c0
commit 93c40c76bf
4 changed files with 373 additions and 1 deletions
+44 -1
View File
@@ -1,6 +1,6 @@
---
name: pdf-to-kindle
description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. Also covers translating such a book into another language before packing it, delegated to the external book_translator project.
description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. Also covers translating such a book into another language before packing it, delegated to the external book_translator project, and generating a TTS audiobook from it with an integrity check for dropped fragments.
---
# PDF to Kindle
@@ -153,6 +153,49 @@ of each block round-trips, and blocks are reassembled by index.
fallback, language detection, and a full split/rebuild against a faked
translator response — no network, no API key. Run it after touching the bridge.
## Optional stage: audiobook
Also delegated to `book_translator` (`05_create_audiobook.py`, Microsoft
edge-tts — free, needs internet). Two wrappers live in `scripts/audiobook.py`;
the external repo is still never edited.
1. **Strip markup markers first.** The translated JSON still holds the
`⟦i⟧` markers from the translation stage — TTS would read them aloud:
`python scripts/audiobook.py prep <workdir>/translations <workdir>/translations_tts`
Point the external script at the *stripped* copy; the original keeps its
italics for the EPUB.
2. **Synthesize:** `05_create_audiobook.py --translations-dir <…>/translations_tts
--voice dmitry --rate '+0%'`. Fragments are **one per paragraph** (the
`--paragraphs-per-group` flag is not used by the loop), named
`chapter_NNN_intro.mp3` / `chapter_NNN_para_NNNN.mp3` under
`audiobook/temp_audio/`.
3. **Verify before the temp files are deleted** — `cleanup_temp_files()` wipes
`temp_audio/`, and after that only the merged file can be checked:
`python scripts/audiobook.py verify <workdir>/translations_tts <workdir>/audiobook`
It flags chapters missing fragments, zero-byte/undecodable mp3s, and
chapters whose duration falls short of what their character count predicts.
The seconds-per-character baseline is the median across chapters, so it
self-calibrates to whatever voice and `--rate` were used. Exit code is
non-zero when anything is wrong.
4. **Re-running fills gaps cheaply** — the external script skips any fragment
file that already exists *and is non-empty*, so delete the bad ones `verify`
named and run it again. It never re-checks that an existing file is sane,
which is exactly why step 3 exists.
Known limits of the external stage: it merges everything into a single
`audiobook_complete.mp3` with no per-chapter files and no chapter marks —
poor for players; splitting is not implemented here. The phonetics stage
(`07_extract_terms.py` + `08_generate_phonetics.py`) is worth running first for
a technical book, or the Russian voice will mangle every English term.
`scripts/test_audiobook.py` builds real silent mp3s with ffmpeg and checks that
`verify` catches a missing fragment, an empty file, and passes a clean book.
Needs ffmpeg; no network.
## What the script does
- style from span font names (`-It`, `Italic`, `Bold`, `Semibold`) → `<i>`/`<b>`;