This open-source dataset consists of 4.25 hours of transcribed Guangzhou Cantonese conversational speech on certain topics, where ten conversations between ten pairs of speakers were contained.
This dataset features a rich array of audio recordings for wake word detection in Indian English. It encompasses diverse accents, regional dialects, and recording conditions to enhance the performance