Corpus design for Uzbek text-to-speech

The data side of an Uzbek text-to-speech system, at Openbank.

The corpus pipeline chains diarization, ECAPA speaker verification, speech-boundary segmentation, selective source separation and 2-of-3 consensus ASR filtering, with an immutable provenance ledger and audit gates a clip has to clear before it can enter the training manifest.

The design is complete: corpus construction, curation thresholds, the evaluation protocol and the manifest gate. It started from a survey of diffusion and flow-matching TTS, multilingual speech corpora including FLEURS, and TTS scaling studies.