VT-SSum
收藏资源简介:
VT-SSum是由微软亚洲研究院创建的一个视频转录分割与摘要的基准数据集,包含125,004对转录-摘要数据,来源于9,616个视频。该数据集利用VideoLectures.NET的视频及其配套幻灯片内容,通过弱监督方法生成摘要。创建过程中,数据集通过精确的视频与幻灯片时间线对齐,分割音频并转换为文本,提取幻灯片文本,并进行转录分割。VT-SSum主要应用于视频理解领域,旨在解决视频转录的摘要问题,提高模型在口语文本摘要任务上的性能。
VT-SSum is a benchmark dataset for video transcript segmentation and summarization created by Microsoft Research Asia. It contains 125,004 transcript-summary pairs sourced from 9,616 videos. The dataset leverages videos and their accompanying slide content from VideoLectures.NET, and generates summaries via weakly-supervised methods. During its development, the dataset achieves precise timeline alignment between videos and their corresponding slides, segments audio and converts it into text, extracts slide text, and performs transcript segmentation. VT-SSum is primarily applied in the field of video understanding, aiming to address the problem of video transcript summarization and improve model performance on spoken text summarization tasks.

- 1VT-SSum: A Benchmark Dataset for Video Transcript Segmentation and Summarization微软亚洲研究院 · 2021年



