Text-to-Speech (TTS) / Chuyển Văn bản thành Giọng nói
📖 Text-to-Speech (TTS)
Definition (English):
TTS synthesizes natural-sounding speech from text. Modern TTS uses neural networks: Tacotron 2 (attention-based), FastSpeech (non-autoregressive), VITS (end-to-end variational), and VALL-E (neural codec language model). TTS handles prosody, emotion, and multi-speaker generation. VALL-E and similar models enable few-shot voice cloning from 3 seconds of audio.
📖 Chuyển Văn bản thành Giọng nói
Định nghĩa (Tiếng Việt):
TTS tổng hợp giọng nói tự nhiên từ văn bản. TTS hiện đại sử dụng mạng nơ-ron: Tacotron 2 (dựa trên attention), FastSpeech (không autoregressive), VITS (biến thể đầu cuối) và VALL-E (mô hình ngôn ngữ codec nơ-ron). TTS xử lý prosody, cảm xúc và tạo đa người nói. VALL-E và mô hình tương tự cho phép voice cloning few-shot từ 3 giây âm thanh.
📂 Phân loại: Speech & Audio
HỆ SINH THÁI CiCC