Abstract
Recent fMRI foundation models differ substantially in the spatial scale at which they represent brain activity. ROI- and connectivity-based models are efficient but coarse, whereas voxel-level models preserve fine-grained spatial structure but require specialized 3D/4D architectures and costly fMRI-specific pretraining. We ask whether part of this performance gap reflects the importance of preserving cortical geometry. Motivated by evidence that macroscale brain activity is strongly constrained by brain geometry, we introduce FlatClip, a frozen-encoder surface-level baseline that renders cortical activity as geometry-aware flatmap sequences and reuses a frozen SigLIP2 image encoder with only a lightweight downstream probe. Across resting-state benchmarks, FlatClip serves as a competitive middle-ground representation, outperforming ROI-level baselines on HCP and ADNI tasks while remaining weaker on PPMI and below the strongest voxel-level models overall. On visual-fMRI decoding, performance improves when the input is restricted from whole cortex to visual or NSD-provided task-active cortex, suggesting that geometry is most useful when it is aligned with the prediction target. We further probe this interpretation with geometry controls: disrupting flatmap spatial organization reduces performance, mapping ROI-level signals back into atlas-defined 4D volumes partially recovers performance, and adaptive patching in highly activated regions improves several voxel-based prediction metrics. Together, these results position surface-level flatmap sequences as a practical middle-ground baseline between ROI and voxel models, and suggest that geometry-preserving spatial organization can be useful when designing fMRI models. Code is available.

