relate
收藏资源简介:
该数据集包含2,846个训练样本,总大小约为8.99MB。每个样本包含8个特征字段:文本内容(text)、音频类型(audio_type)、标注数量(num_annotations)、波形文件名(wave_filename)、持续时间(duration)、文本相关度评分(text_relevance_score)、相关度推理(text_relevance_reasoning)以及文本相关思考(text_relevance_thoughts,为字符串列表)。数据集仅提供训练集划分,下载尺寸为4.26MB。数据文件路径为'train-*'格式,适用于音频-文本相关性分析等任务。
This dataset comprises 2,846 training samples with an overall size of approximately 8.99 MB. Each sample contains 8 feature fields: text (text content), audio_type (audio type), num_annotations (number of annotations), wave_filename (waveform file name), duration, text_relevance_score (text relevance score), text_relevance_reasoning (text relevance reasoning), and text_relevance_thoughts (a list of strings). The dataset only offers a training set split, with a download size of 4.26 MB. The data files follow the 'train-*' naming format, and are applicable to tasks including audio-text correlation analysis.




