vuduylinh150804/vietmed-rag-dataset
收藏资源简介:
该数据集包含越南语医疗数据以及为检索增强生成(RAG)实验处理的LightRAG存储工件。具体内容包括:爬取的越南医疗源文档(vietmed_crawled/)、排队等待LightRAG摄入的源文档(__enqueued__/)、用于文档、块、实体和关系的LightRAG键值存储(kv_store_*.json)、块、实体和关系的向量数据库导出(vdb_*.json)、稀疏检索的BM25索引(bm25_*.pkl)、链接块、实体和关系的知识图谱(graph_chunk_entity_relation.graphml),以及实体名称查找和规范化名称矩阵工件(name_*.json和name_matrix_normed.npy)。数据集旨在用于越南医疗RAG、检索评估、知识图谱探索,以及使用LightRAG兼容存储文件的下游研究。
This dataset contains Vietnamese medical data and processed LightRAG storage artifacts for retrieval-augmented generation (RAG) experiments. It includes crawled Vietnamese medical source documents (vietmed_crawled/), source documents queued for LightRAG ingestion (__enqueued__/), LightRAG key-value stores for documents, chunks, entities, and relationships (kv_store_*.json), vector database exports for chunks, entities, and relationships (vdb_*.json), BM25 indexes for sparse retrieval (bm25_*.pkl), a knowledge graph linking chunks, entities, and relations (graph_chunk_entity_relation.graphml), and entity name lookup and normalized name matrix artifacts (name_*.json and name_matrix_normed.npy). The dataset is intended for Vietnamese medical RAG, retrieval evaluation, knowledge graph exploration, and downstream research using LightRAG-compatible storage files.




