RefactorCoderQA
收藏资源简介:
RefactorCoderQA是一个全面且多样化的基准数据集,旨在评估和提升大型语言模型(LLMs)在编码任务上的性能。该数据集涵盖了四个关键技术领域:软件工程(SE)、数据科学(DS)、机器学习(ML)和自然语言处理(NLP),并使用来自Stack Overflow的真实世界编码问题构建而成。每个问题都包括详细的问题描述和经过验证的解决方案,并已重新格式化为一致的输入-输出格式,以支持结构化提示和客观评估。数据集的开发经过系统性的数据收集、清理和组织过程。通过使用实际开发场景中的真实问题和答案,RefactorCoderQA提供了一个更现实和有意义的方式来评估LLMs在广泛领域和编码任务中的能力。
RefactorCoderQA is a comprehensive and diverse benchmark dataset designed to evaluate and enhance the performance of Large Language Models (LLMs) on coding tasks. This dataset covers four key technical domains: Software Engineering (SE), Data Science (DS), Machine Learning (ML), and Natural Language Processing (NLP), and is constructed using real-world coding problems sourced from Stack Overflow. Each problem included in the dataset features a detailed problem description and a validated solution, and has been reformatted into a consistent input-output format to support structured prompting and objective evaluation. The development of the dataset follows a systematic process of data collection, cleaning, and organization. By leveraging real problems and answers from actual development scenarios, RefactorCoderQA provides a more realistic and meaningful approach to evaluating the capabilities of LLMs across a wide range of domains and coding tasks.




