Speaker Diarization / Phân biệt Người nói

✍️

📖 Speaker Diarization

Definition (English):

Speaker diarization answers "who spoke when" in audio — segmenting a recording by speaker identity. It combines voice activity detection, speaker embedding extraction (x-vectors, ECAPA-TDNN), and clustering. Diarization Error Rate (DER) is the primary metric. Applications: meeting transcription, call center analytics, podcast indexing, and surveillance.


📖 Phân biệt Người nói

Định nghĩa (Tiếng Việt):

Phân biệt người nói trả lời "ai nói khi nào" trong âm thanh — phân đoạn bản ghi theo danh tính người nói. Nó kết hợp phát hiện hoạt động giọng nói, trích xuất nhúng người nói (x-vectors, ECAPA-TDNN) và phân cụm. DER là chỉ số chính. Ứng dụng: phiên âm cuộc họp, phân tích tổng đài, lập chỉ mục podcast và giám sát.


📂 Phân loại: Speech & Audio

Hỏi AI Edu?