遇见数据集

MM-MT-Bench

收藏
魔搭社区2026-07-30 更新2026-08-02 收录
官方服务:

资源简介:

# MM-MT-Bench MM-MT-Bench is a multi-turn LLM-as-a-judge evaluation benchmark similar to the text MT-Bench for testing multimodal instruction-tuned models. While existing benchmarks like MMMU, MathVista, ChartQA and so on are focused on closed-ended questions with short responses, they do not evaluate model's ability to follow user instructions in multi-turn dialogues and answer open-ended questions in a zero-shot manner. MM MT-Bench is designed to overcome this limitation. The [mistral-evals](https://github.com/mistralai/mistral-evals) repository provides an example script on how to evaluate a model on this benchmark. [Paper link](https://arxiv.org/abs/2410.07073)

提供机构:
maas
创建时间:
2026-06-15
二维码
社区交流群
二维码
科研交流群
商业服务