Synthetic Medicare‑Like Inpatient Claims and Beneficiary Data Conforming to ResDAC FTS Layouts
收藏资源简介:
This dataset provides a fully synthetic Medicare‑like database that mimics the raw fixed‑width files and reproduces the structure of ResDAC File Transfer Summary (FTS) layouts for beneficiary and claims files. It can be ingested by the same tooling that processes real CMS Medicare data. Because the data are fully synthetic, it can be used to test such tooling without exposing any protected health information (PHI) or personally identifiable information (PII). In particular, the dataset can be used to demonstrate how the Dorieh Medicare Data Processing Pipeline works. Synthetic records are generated from scratch using AI‑generated FTS specifications (produced by the ChatGPT‑4.1 model) together with public reference data (e.g., U.S. Census ZIP‑code–level population counts and SSA–FIPS crosswalks) and CMS’s DE‑SynPUF synthetic Medicare files. An internal “beneficiary” table ensures that identifiers and core demographics are consistent across all files and across years. The generator also introduces low levels of realistic data‑quality issues (such as missing identifiers, miscoded race, and small shifts in dates of birth) to support testing of error‑handling and quality‑control workflows. The Zenodo archive contains: Fixed‑width data files conforming to selected ResDAC Medicare FTS layouts for calendar years 2011–2016 Machine‑readable schema information derived from the FTS‑style specifications Pointers to the code and documentation used to generate the data, including the Dorieh platform and the Medicare pipeline description This artificial Medicare database is intended for education, methods development, and testing of ETL and analysis pipelines in population‑health research. It is not suitable for drawing substantive conclusions about real patients, providers, or health‑care utilization. Ethics and privacy This dataset consists entirely of artificially generated Medicare‑like beneficiary and claims records. No individual‑level CMS ResDAC data or any other real patient‑level datasets were used as inputs. The only external sources are publicly available aggregate statistics (e.g., U.S. Census data, SSA–FIPS crosswalks) and CMS’s DE‑SynPUF synthetic Medicare files, which are themselves non‑identifiable. The generation process does not attempt to reconstruct records for any actual person, and the resulting files contain no PHI or PII as defined under U.S. regulations. Because the data are fully synthetic and non‑identifiable, use of this resource is not expected to constitute human‑subjects research. Users remain responsible for ensuring that their own use complies with applicable laws, institutional policies, and ethics/IRB requirements.



