Text-to-Speech (TTS) / Chuyển Văn bản thành Giọng nói

✍️
Từ điển và thuật ngữ trí tuệ nhân tạo

📖 Text-to-Speech (TTS)

Definition (English):

TTS synthesizes natural-sounding speech from text. Modern TTS uses neural networks: Tacotron 2 (attention-based), FastSpeech (non-autoregressive), VITS (end-to-end variational), and VALL-E (neural codec language model). TTS handles prosody, emotion, and multi-speaker generation. VALL-E and similar models enable few-shot voice cloning from 3 seconds of audio.


📖 Chuyển Văn bản thành Giọng nói

Định nghĩa (Tiếng Việt):

TTS tổng hợp giọng nói tự nhiên từ văn bản. TTS hiện đại sử dụng mạng nơ-ron: Tacotron 2 (dựa trên attention), FastSpeech (không autoregressive), VITS (biến thể đầu cuối) và VALL-E (mô hình ngôn ngữ codec nơ-ron). TTS xử lý prosody, cảm xúc và tạo đa người nói. VALL-E và mô hình tương tự cho phép voice cloning few-shot từ 3 giây âm thanh.


📂 Phân loại: Speech & Audio

HỆ SINH THÁI CiCC

Đi tiếp cùng nội dung này

Hỏi AI Edu?