Обвязка озвучки: снятие маркеров разметки для TTS и проверка целостности синтеза
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
---
|
||||
name: pdf-to-kindle
|
||||
description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. Also covers translating such a book into another language before packing it, delegated to the external book_translator project.
|
||||
description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. Also covers translating such a book into another language before packing it, delegated to the external book_translator project, and generating a TTS audiobook from it with an integrity check for dropped fragments.
|
||||
---
|
||||
|
||||
# PDF to Kindle
|
||||
@@ -153,6 +153,49 @@ of each block round-trips, and blocks are reassembled by index.
|
||||
fallback, language detection, and a full split/rebuild against a faked
|
||||
translator response — no network, no API key. Run it after touching the bridge.
|
||||
|
||||
## Optional stage: audiobook
|
||||
|
||||
Also delegated to `book_translator` (`05_create_audiobook.py`, Microsoft
|
||||
edge-tts — free, needs internet). Two wrappers live in `scripts/audiobook.py`;
|
||||
the external repo is still never edited.
|
||||
|
||||
1. **Strip markup markers first.** The translated JSON still holds the
|
||||
`⟦i⟧` markers from the translation stage — TTS would read them aloud:
|
||||
|
||||
`python scripts/audiobook.py prep <workdir>/translations <workdir>/translations_tts`
|
||||
|
||||
Point the external script at the *stripped* copy; the original keeps its
|
||||
italics for the EPUB.
|
||||
2. **Synthesize:** `05_create_audiobook.py --translations-dir <…>/translations_tts
|
||||
--voice dmitry --rate '+0%'`. Fragments are **one per paragraph** (the
|
||||
`--paragraphs-per-group` flag is not used by the loop), named
|
||||
`chapter_NNN_intro.mp3` / `chapter_NNN_para_NNNN.mp3` under
|
||||
`audiobook/temp_audio/`.
|
||||
3. **Verify before the temp files are deleted** — `cleanup_temp_files()` wipes
|
||||
`temp_audio/`, and after that only the merged file can be checked:
|
||||
|
||||
`python scripts/audiobook.py verify <workdir>/translations_tts <workdir>/audiobook`
|
||||
|
||||
It flags chapters missing fragments, zero-byte/undecodable mp3s, and
|
||||
chapters whose duration falls short of what their character count predicts.
|
||||
The seconds-per-character baseline is the median across chapters, so it
|
||||
self-calibrates to whatever voice and `--rate` were used. Exit code is
|
||||
non-zero when anything is wrong.
|
||||
4. **Re-running fills gaps cheaply** — the external script skips any fragment
|
||||
file that already exists *and is non-empty*, so delete the bad ones `verify`
|
||||
named and run it again. It never re-checks that an existing file is sane,
|
||||
which is exactly why step 3 exists.
|
||||
|
||||
Known limits of the external stage: it merges everything into a single
|
||||
`audiobook_complete.mp3` with no per-chapter files and no chapter marks —
|
||||
poor for players; splitting is not implemented here. The phonetics stage
|
||||
(`07_extract_terms.py` + `08_generate_phonetics.py`) is worth running first for
|
||||
a technical book, or the Russian voice will mangle every English term.
|
||||
|
||||
`scripts/test_audiobook.py` builds real silent mp3s with ffmpeg and checks that
|
||||
`verify` catches a missing fragment, an empty file, and passes a clean book.
|
||||
Needs ffmpeg; no network.
|
||||
|
||||
## What the script does
|
||||
|
||||
- style from span font names (`-It`, `Italic`, `Bold`, `Semibold`) → `<i>`/`<b>`;
|
||||
|
||||
Reference in New Issue
Block a user