Нарезка аудиокниги по главам: треки с ID3-тегами вместо одного слитного файла

This commit is contained in:
Илья Поляков
2026-08-03 11:02:32 +03:00
parent 93c40c76bf
commit a0c26b41bf
4 changed files with 227 additions and 7 deletions
+25 -6
View File
@@ -1,6 +1,6 @@
---
name: pdf-to-kindle
description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. Also covers translating such a book into another language before packing it, delegated to the external book_translator project, and generating a TTS audiobook from it with an integrity check for dropped fragments.
description: Convert a text-layer PDF (usually one generated by calibre from an ebook) into EPUB/AZW3 for Kindle while preserving italics, bold, chapter headings, sidebars/journal blocks, images and cross-page paragraphs. Use when the user asks to read a PDF book on Kindle or another e-reader, to convert a PDF to EPUB/AZW3/MOBI, or complains that a converted book lost its italics, merged its chapters, or reflows badly on the device. Not for scanned PDFs (no text layer) — those need OCR first. Also covers translating such a book into another language before packing it, delegated to the external book_translator project, and generating a TTS audiobook from it with an integrity check for dropped fragments and a split into per-chapter tagged tracks.
---
# PDF to Kindle
@@ -185,12 +185,31 @@ the external repo is still never edited.
file that already exists *and is non-empty*, so delete the bad ones `verify`
named and run it again. It never re-checks that an existing file is sane,
which is exactly why step 3 exists.
5. **Split into chapter tracks**, also before the temp files go:
Known limits of the external stage: it merges everything into a single
`audiobook_complete.mp3` with no per-chapter files and no chapter marks —
poor for players; splitting is not implemented here. The phonetics stage
(`07_extract_terms.py` + `08_generate_phonetics.py`) is worth running first for
a technical book, or the Russian voice will mangle every English term.
```bash
python scripts/audiobook.py split <workdir>/translations_tts <workdir>/audiobook \
<workdir>/tracks --album "Название" --author "Автор" --gap 0.3
```
The external stage only ever produces one merged `audiobook_complete.mp3`
with no chapter marks, which is bad for players. `split` rebuilds per-chapter
`NNN - Title.mp3` from the same fragments, with ID3 album/artist/track/title
tags, and refuses a chapter whose fragments are incomplete (`--force`
overrides). Concatenation is stream-copy, falling back to a re-encode only if
the mp3 streams don't line up; `--gap` inserts silence between paragraphs,
generated to match the fragments' own codec parameters so the copy path
stays viable. Hand the result to the `prepare-audiobooks` skill for covers
and library layout.
The phonetics stage (`07_extract_terms.py` + `08_generate_phonetics.py`) is
worth running first for a technical book, or the Russian voice will mangle
every English term.
**Ordering constraint for the whole stage:** `verify` and `split` both read
`audiobook/temp_audio/`, and the external script's `cleanup_temp_files()`
deletes it right after merging. Run both before that, or the fragments are gone
and only the merged file's total duration can be checked.
`scripts/test_audiobook.py` builds real silent mp3s with ffmpeg and checks that
`verify` catches a missing fragment, an empty file, and passes a clean book.