EduViQA
收藏资源简介:
EduVQA-Alpha是一个多语言教育视频问答(VideoQA)数据集,包含学术视频和合成的问答对,支持英语和波斯语。数据集采用CLIP-SSIM自适应分块技术进行视频分割,确保高质量的多模态AI系统语义对齐。数据集结构包括视频分块、视频转录文件和问答对,涵盖了多种学术主题和教学风格。数据集的创建过程包括视频来源、分块和注释,确保伦理合规性。数据集适用于多模态视频问答、RAG管道训练和视觉语言模型基准测试。
EduVQA-Alpha is a multilingual educational video question answering (VideoQA) dataset that consists of academic videos and synthesized question-answer pairs, supporting both English and Persian. This dataset adopts the CLIP-SSIM adaptive chunking technique for video segmentation, ensuring high-quality semantic alignment for multimodal AI systems. The dataset structure includes video chunks, video transcript files, and question-answer pairs, covering a wide range of academic topics and teaching styles. The dataset's creation process covers video sourcing, chunking, and annotation, ensuring ethical compliance. This dataset is applicable to multimodal video question answering, RAG pipeline training, and vision-language model benchmark testing.




