OpenDCAI/FlipVQA
收藏资源简介:
--- license: apache-2.0 configs: - config_name: Multimodal data_files: - split: train path: Multimodal/train-* - split: test path: Multimodal/test-* - config_name: Text_base data_files: - split: train path: Text_base/train-* task_categories: - question-answering language: - zh - en size_categories: - 10K<n<100K dataset_info: config_name: Text_base features: - name: question dtype: string - name: answer dtype: string - name: solution dtype: string - name: type dtype: string - name: subject dtype: string - name: images dtype: 'null' - name: generated_cot dtype: string - name: answer_match_result dtype: bool splits: - name: train num_bytes: 662976117 num_examples: 72126 download_size: 289871745 dataset_size: 662976117 --- # FlipVQA-85K FlipVQA-85k is a high-fidelity reasoning benchmark curated from a corpus of 544 college-level educational PDF documents, including expert-authored textbooks and exercise sets. The collection spans 11 academic disciplines, primarily in STEM domains where problems typically involve rigorous and verifiable reasoning processes. Paper link: https://huggingface.co/papers/2511.16216 Please see https://github.com/OpenDCAI/DataFlow-VQA for detailed data curation method. **Important**: - The `solution` is extracted from the original textbooks, and might not be suitable to directly used for LLM training. - The `generated_cot` consists of detailed thinking processes generated by Qwen3-235B-A22B-Thinking and Qwen3-VL-235B-A22B-Thinking, which have proven effective for LLM training. We retain **both** correct and incorrect thinking processes in the dataset, with the `answer_match_result` key indicating whether the thinking process yielded a correct answer.




