A 5.0T Ultra-High-Field fMRI Dataset for Naturalistic Visual Scene Processing
收藏资源简介:
Modeling neural responses under naturalistic visual stimulation is an important goal in computational neuroscience and brain-computer interface research. Progress in this area depends on neuroimaging datasets that combine repeated measurements, shared stimulus anchors, and sufficient stimulus diversity for evaluating encoding and decoding models. Here, we present the Natural Vision Dataset (NVD), a publicly available 5.0 T fMRI dataset designed for static natural-image viewing. The field strength is reported as part of the acquisition context, and the dataset was not designed to isolate field-strength effects or to compare 5.0 T performance with 3 T or 7 T acquisitions. Twenty healthy participants viewed a shared set of 1,268 natural images and 500 participant-specific images per participant, yielding 10,000 participant-specific images across the dataset. Each image was presented three times across separate sessions. This hybrid shared-unique stimulus design supports assessment of response reliability, cross-participant alignment, and model generalization beyond a fixed shared image set. Functional data were acquired at 1.8 mm isotropic resolution with a TR of 1 s and are organized according to the Brain Imaging Data Structure, with both volumetric and surface-based derivatives provided. Image-level response estimates were derived using GLMsingle toolbox. Data quality and benchmark utility were characterized using motion, vigilance, tSNR, noise-ceiling, and brain-to-CLIP decoding analyses. NVD provides a standardized resource for investigating human visual representations and evaluating computational models of natural-image processing. Stimulus images The stimulus images are organized in the dataset-level stimuli/ directory. The full stimulus inventory is provided in stimuli/images.tsv, which contains 11,268 image entries. Each row corresponds to one stimulus image and includes the image identifier (image_id), image file name (image_name), COCO image identifier (coco_id), stimulus assignment set (stimulus_set), participant viewing scope (viewed_by), task identifier (task_id), Hard Subset membership (hard_subset), and Hard Subset cluster identifier (hard_subset_cluster_id). The stimulus_set column indicates whether the image belongs to the Shared or Unique stimulus set. The viewed_by column specifies whether the image was presented to all participants (all) or only to a specific participant, denoted by values of the form sub-<label>. The task_id column should be interpreted as a BIDS task label and stimulus-subset identifier rather than as a separate behavioral task. Participants performed the same viewing/vigilance paradigm throughout the experiment; the labels s1–s11 distinguish stimulus subsets used for BIDS organization and file naming. The hard_subset column indicates whether the image belongs to the Hard Subset selected from semantically dense CLIP embedding clusters. Missing values are encoded as n/a; specifically, hard_subset_cluster_id = n/a indicates that the image is not included in the Hard Subset and therefore has no Hard Subset cluster assignment. The accompanying metadata file, stimuli/images.json, provides column-level descriptions for the stimulus inventory. Raw MRI data The folder for each participant consists of several session folders. The session folder in turn includes two or three folders, named “anat”, “func” or “fmap”, for corresponding modality data. In the “func” folder, the “sub-<subID>_ses-<sesID>_task-<taskID>_run-<runID>_events.tsv” file contains task events of each run. The stimulus information for each trial is listed in the last column of this events file, which is named “condition” for the experiments. As a result, the specific stimulus image for each trial can be located in the “stimuli” folder according to the stimuli information listed in the last column of the events file. Preprocessed volume and surface data from fMRIPrep Functional MRI data were preprocessed using fMRIPrep with outputs in volumetric (T1w and MNI) and surface (fsLR) spaces. For each functional run, preprocessed BOLD time series in native anatomical space (T1w) are provided as: “sub-<subID>_ses-<sesID>_task-<taskID>_space-T1w_desc-preproc_bold.nii.gz”. Spatially normalized volumes in MNI space are available as: “sub-<subID>_ses-<sesID>_task-<taskID>_space-MNI152NLin2009cAsym_desc-preproc_bold.nii.gz”. Nuisance regressors estimated during preprocessing are stored in: “sub-<subID>_ses-<sesID>_task-<taskID>_desc-confounds_timeseries.tsv”. Surface-based outputs were generated in fsLR standard space. Combined cortical and subcortical time series are provided in CIFTI-2 format (91k grayordinates): “sub-<subID>_ses-<sesID>_task-<taskID>_space-fsLR_den-91k_bold.dtseries.nii”. Hemisphere-specific cortical surface time series at 32k vertex density are provided: “sub-<subID>_ses-<sesID>_task-<taskID>_hemi-L_space-fsLR_den-32k_bold.func.gii”, “sub-<subID>_ses-<sesID>_task-<taskID>_hemi-R_space-fsLR_den-32k_bold.func.gii”. Brain activation data from surface-based analysis Brain activation data were derived from GLMsingle analyses and reorganized under “derivatives/glmsingle/”. For each subject and run, trial-wise GLMsingle beta estimates in the standard fsLR 91k grayordinate space are stored as CIFTI-2 dense scalar files: “derivatives/glmsingle/sub-<subID>/ses-<sesID>/func/sub-<subID>_ses-<sesID>_task-<taskID>_space-fsLR_den-91k_desc-rawBetasmd_betas.dscalar.nii”. Each run-level beta file is accompanied by a paired .tsv file with the same basename, in which each row corresponds to one trial/presentation represented in the CIFTI scalar axis. A sidecar .json file provides additional metadata, including the source files, image space, density, beta type, and row description. Subject-level condition-averaged beta maps, obtained by averaging GLMsingle betasmd estimates across all runs/repeats for the same condition, are stored separately as: “derivatives/glmsingle/sub-<subID>/avg/sub-<subID>_space-fsLR_den-91k_desc-avgBetasmd_betas.dscalar.nii”. The corresponding paired .tsv file contains the condition-level information, with each row matching one averaged condition in the CIFTI file. In addition to the full fsLR 91k outputs, visual-ROI-restricted beta maps are also provided. These files use the same organization but include VisualROIs in the description field, for example: “desc-rawBetasmdVisualROIs_betas.dscalar.nii”, ”desc-avgBetasmdVisualROIs_betas.dscalar.nii”. The visual ROI definitions and grayordinate indices are provided under “derivatives/glmsingle/metadata/”, including the ROI table, visual grayordinate columns, visual ROI indices, and visual ROI roimap files.



