BrainWorld

A Structural-Prior-Conditioned Generative Model for Whole-Brain 4D fMRI Dynamics

Junfeng Xia, Wenhao Ye, Junxiang Zhang, Xuanye Pan, Mo Wang, Quanying Liu

NeurIPS 2026 · Oral

NeurIPS 2026 · OralBrainWorld original paper figure
Original paper figure View full figure ↗
At a glance

What this study adds

Summary of the linked study; findings refer to that study’s own comparisons.

Question
Can anatomy guide long-range whole-brain fMRI generation?
Approach
Condition latent diffusion on structural MRI and prior functional context.
Evidence
Evaluation across 22 datasets and diverse brain states.
Finding
The paper reports stable generated trajectories up to 400 frames, plus gains from generated-example augmentation and transferable features.

Abstract

Whole-brain 4D fMRI generation is valuable for modeling functional brain dynamics, yet existing fMRI foundation models mainly target representation learning and downstream prediction rather than conditional predictive generation. We introduce BrainWorld, a structural-prior-conditioned generative model for whole-brain 4D fMRI dynamics. BrainWorld uses sMRI as subject-level anatomical context to guide future fMRI generation, integrating structural information into the denoising process rather than treating it as a parallel modality. Evaluated on 22 datasets spanning diverse cohorts and brain states, BrainWorld generates stable 4D fMRI trajectories up to 400 frames, improves downstream performance through generated-example augmentation, and learns transferable multimodal representations that outperform baselines. Together, these results establish BrainWorld as a condition-aware generative framework for long-horizon brain dynamics modeling and multimodal representation learning. Code is available at the code repository.

Method

BrainWorld standardizes whole-brain fMRI and compresses it into a pretrained VAE latent space. A conditional Diffusion Transformer models brain dynamics using structural MRI, past functional connectivity, and optional visual or audio context. The VAE decoder reconstructs voxel-level 4D fMRI, while intermediate DiT features support downstream representation learning.

Animated BrainWorld framework: horizontal fMRI preprocessing followed by structural-prior-conditioned latent diffusion and downstream representation extraction.
Preprocessing & latent diffusionConditioning & representationsView full-size animation ↗

BibTeX

@misc{brainworld2026,
  title={BrainWorld: A Structural-Prior-Conditioned Generative Model for Whole-Brain 4D fMRI Dynamics},
  author={Xia, Junfeng and Ye, Wenhao and Zhang, Junxiang and Pan, Xuanye and Wang, Mo and Liu, Quanying},
  year={2026},
  eprint={2606.17742},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2606.17742}
}
Continue exploring

BrainTaskonomy

Next project