Speaker Diarization / Phân biệt Người nói
📖 Speaker Diarization
Definition (English):
Speaker diarization answers "who spoke when" in audio — segmenting a recording by speaker identity. It combines voice activity detection, speaker embedding extraction (x-vectors, ECAPA-TDNN), and clustering. Diarization Error Rate (DER) is the primary metric. Applications: meeting transcription, call center analytics, podcast indexing, and surveillance.
📖 Phân biệt Người nói
Định nghĩa (Tiếng Việt):
Phân biệt người nói trả lời "ai nói khi nào" trong âm thanh — phân đoạn bản ghi theo danh tính người nói. Nó kết hợp phát hiện hoạt động giọng nói, trích xuất nhúng người nói (x-vectors, ECAPA-TDNN) và phân cụm. DER là chỉ số chính. Ứng dụng: phiên âm cuộc họp, phân tích tổng đài, lập chỉ mục podcast và giám sát.
📂 Phân loại: Speech & Audio