SEVENLLM-Instruct
收藏资源简介:
SEVENLLM-Instruct是由北京航空航天大学国家重点实验室创建的高质量双语指令数据集,包含85000个样本,用于训练大型语言模型以增强网络安全事件分析能力。该数据集通过爬取网络安全网站的原始文本构建,采用Select-Instruct方法生成监督学习数据。SEVENLLM-Instruct旨在解决网络安全领域数据稀缺问题,通过多任务学习提升模型在威胁识别和响应方面的性能,广泛应用于网络安全事件的自动化和智能化处理。
SEVENLLM-Instruct is a high-quality bilingual instruction dataset developed by the State Key Laboratory of Beihang University, which contains 85,000 samples. It is designed to train large language models (LLMs) to enhance their capabilities in cybersecurity incident analysis. This dataset is constructed by crawling raw text from cybersecurity websites, and adopts the Select-Instruct method to generate supervised learning data. SEVENLLM-Instruct aims to address the data scarcity issue in the cybersecurity domain, improve the model's performance in threat identification and response via multi-task learning, and is widely applied to the automated and intelligent processing of cybersecurity incidents.

- 1SEvenLLM: Benchmarking, Eliciting, and Enhancing Abilities of Large Language Models in Cyber Threat Intelligence北京航空航天大学复杂与关键软件环境国家重点实验室 · 2024年



