The TNO speaker diarization system for NIST RT05s meeting data
conference paper
The TNO speaker speaker diarization system is based on a standard BIC segmentation and clustering algorithm. Since for the NIST Rich Transcription speaker dizarization evaluation measure correct speech detection appears to be essential, we have developed a speech activity detector (SAD) as well. This is based on decoding the speech signal using two Gaussian Mixture Models trained on silence and speech. The SAD was trained on only AMI development test data, and performed quite well in the evaluation on all 5 meeting locations, with a SAD error rate of 5.0 %. For the speaker clustering algorithm we optimized the BIC penalty parameter ? to 14, which is quite high with respect to the theoretical value of 1. The final speaker diarization error rate was evaluated at 35.1 %.
Topics
TNO Identifier
16861
Source title
2nd International Workshop on Machine Learning for Multimodal Interaction, MLMI 2005, 11 July 2005 through 13 July 2005, Edinburgh,
Pages
440-449
Files
To receive the publication files, please send an e-mail request to TNO Repository.