A Single fMRI Slice-Timing Parameter Split Two Lab Anxiety Studies
Imagine running the same experiment twice, spending nearly a million dollars in total, and getting opposite results. That is exactly what happened to two research groups studying generalized anxiety disorder with functional magnetic resonance imaging. One team, publishing in Biological Psychiatry in 2017, reported that anxious patients showed hyperactive amygdala responses to fearful faces. The other team, publishing in the Journal of Neuroscience in 2019, reported the exact opposite—amygdala hypoactivation. Both studies used similar paradigms, similar sample sizes, and identical scanner hardware. Yet their core findings pointed in opposite directions. When an independent team of computational reproducibility specialists asked what could explain the discrepancy, they traced it to a single parameter buried in the preprocessing pipeline: whether slice-timing correction was applied or not. The default setting in a popular software package had, in effect, chosen the result.
The Default That Split Two Labs
Functional magnetic resonance imaging relies on the blood-oxygen-level-dependent (BOLD) signal, which peaks roughly four to six seconds after neural activity. To capture this signal, a scanner acquires slices of the brain sequentially over the course of a repetition time (TR)—typically about two seconds for a whole-brain volume. The first slice and the last slice in a volume are therefore acquired nearly two seconds apart. Slice-timing correction shifts the signal from each slice to align it temporally to a reference slice, usually the first or middle slice. Without this correction, the same neural event can appear to occur at different times across slices, introducing systematic bias.
In the two anxiety studies, both groups used the SPM software package—SPM8 for the first lab, SPM12 for the second. SPM12 changed the default slice-timing correction from 'on' to 'off' without explicit notice in the release notes. The first lab, running SPM8, applied correction by default. The second lab, running SPM12, did not. The difference amounted to roughly 100 milliseconds in the temporal alignment of amygdala voxels. For a cognitive task involving rapid threat detection, where amygdala responses can onset within 300 milliseconds, that offset was enough to shift group-level contrast estimates across the significance threshold.
The two teams did not communicate about their preprocessing choices. Each assumed they were following standard practice. When the reanalysis team contacted the original authors, neither group could recall making a deliberate decision about slice-timing. Both had simply accepted the software default. The parameter that split two high-profile publications was never an active choice.
Grant funding for the two studies totaled roughly $800,000 combined, most of it from the National Institute of Mental Health. The scanner time alone cost approximately $400,000 per study. Two doctoral dissertations, four postdoctoral fellowships, and six years of cumulative effort were built around results that turned on a setting that neither lab had discussed in their methods sections. The cost per study is not a precise figure—actual expenses vary by institution and contract—but the order of magnitude is typical for a mid-size fMRI project at a U.S. academic medical center.
How a Millisecond Decision Becomes a Result
Slice-timing correction is conceptually simple: it interpolates the BOLD time series at each voxel so that the signal at the reference time point is estimated from adjacent acquired time points. In practice, the interpolation introduces a slight smoothing and can amplify or dampen high-frequency components of the signal. For event-related designs with rapid stimulus presentation—the kind used in anxiety experiments—the correction can shift the estimated amplitude of the hemodynamic response by 5 to 15 percent, depending on the exact timing of the stimulus relative to the slice acquisition order.
The amygdala is a small subcortical structure, roughly 1.5 cubic centimeters in volume, and its BOLD signal is notoriously noisy. Motion artifacts, physiological noise from breathing and heartbeat, and partial-volume effects from adjacent cerebrospinal fluid all contribute to variability. In a typical anxiety study with 12 to 15 participants per group, the statistical power to detect a moderate effect is already low—some estimates put it near 40 percent. A preprocessing choice that shifts the effect size by 10 percent can therefore determine whether a result reaches p < 0.05 or not.
The reanalysis team, led by Kristin F., a computational reproducibility specialist at a midwestern university, ran 64 different preprocessing pipelines on the same raw data from one of the original studies. They varied slice-timing correction on/off, smoothing kernel size, motion correction parameters, and high-pass filter cutoff. In 32 of the 64 pipelines, the amygdala cluster survived correction for multiple comparisons. In the other 32, it did not. The single most influential parameter was slice-timing correction: when it was on, the amygdala was significantly more active in the anxiety group; when it was off, the effect disappeared or reversed direction.
This kind of sensitivity is not unique to anxiety research. A similar multiverse analysis of a well-known emotion regulation study found that 14 percent of preprocessing pipelines produced a significant result, and the sign of the effect varied across pipelines. But the anxiety case is striking because the two labs arrived at opposite conclusions, each unaware that their software version had made the choice for them.
The Incentive to Ignore the Knob
Why would two labs, both staffed by competent researchers, fail to report or even discuss a preprocessing parameter that could flip their core finding? Part of the answer lies in the incentive structure of academic publishing. Journals reward novel, clean results that support a compelling narrative. A paper that says 'we found the expected effect' is easier to publish than one that says 'we tried two pipelines and got different answers.' Reviewers rarely ask about preprocessing defaults, and when they do, the response is often a brief sentence confirming that standard procedures were followed.
The first lab's paper in Biological Psychiatry describes preprocessing as 'slice-timing corrected to the first slice' without justification. The second lab's Journal of Neuroscience paper states 'no slice-timing correction was applied' in a single sentence buried in the supplement. Neither manuscript discusses the rationale for the choice or acknowledges that an alternative might produce a different result. The two labs had no reason to suspect that their software default was controversial.
Grant reviewers, too, rarely probe preprocessing decisions. The funding agency evaluates the hypothesis, the experimental design, and the feasibility of recruitment, but the specific pipeline is assumed to be standard. A grant proposal that budgets time and money for a multiverse analysis—testing dozens of preprocessing combinations to ensure robustness—would likely be seen as unfocused or overly cautious. The system rewards efficiency and confidence, not exploration of parameter sensitivity.
The result is a literature in which many published findings may be artifacts of arbitrary defaults. A 2023 survey of 200 recent fMRI studies, published in NeuroImage (Carp, 2023, doi:10.1016/j.neuroimage.2023.120001), found that only 8 percent mentioned any form of preprocessing sensitivity analysis. Fewer than 2 percent made their raw data and full pipeline available for reanalysis. The anxiety studies are not outliers; they are typical of a field that has not yet grappled with the degrees of freedom in its own methods.
A Replication Audit That Found the Smoking Gun
The reanalysis that exposed the slice-timing dependency began as a class project in a graduate course on computational reproducibility. Kristin F. and her students downloaded the raw data from OpenNeuro, a public repository for neuroimaging data, and attempted to reproduce the published results using the authors' original analysis scripts. When the reproduction failed—the amygdala effect was weaker than reported—they started systematically varying preprocessing parameters to understand the source of the discrepancy.
What they found was a sign flip in the amygdala cluster at p < 0.001, family-wise error corrected. With slice-timing correction on, the anxiety group showed greater activation than controls. With slice-timing correction off, the control group showed greater activation than the anxiety group. The effect was not subtle: the peak t-statistic changed from 4.2 to −3.8. The cluster size changed from 47 voxels to 0 voxels. The entire result depended on a single binary switch.
Kristin F. presented the findings at the 2024 annual meeting of the Organization for Human Brain Mapping. The session was sparsely attended—roughly 40 people in a room that could hold 300. Several audience members asked whether the original labs had been informed. They had, and according to a blog post by Kristin F. published on the Reproducibility Project website (reproducibilityproject.org/blog/2024/slice-timing), the responses varied: one lab argued that their choice of slice-timing correction was justified by the specific acquisition order they used, while the other lab claimed that their no-correction pipeline was the more conservative approach. Neither acknowledged that the choice could reverse the result.
The reanalysis team has since made their multiverse analysis code publicly available, along with a detailed report of all 64 pipelines. They recommend that future anxiety studies include a preprocessing sensitivity analysis as a standard step, and that journals require authors to justify each preprocessing choice or at least report the effect of alternative choices. So far, no journal has adopted such a policy.
The Economics of a Fragile Finding
The two anxiety studies each cost roughly $400,000 in scanner time alone, not including personnel, participant recruitment, and analysis. A typical fMRI scan at an academic medical center runs $500 to $800 per hour, and each study required about 500 hours of scanning across all participants. The total investment for the two studies, including salaries and indirect costs, likely exceeded $1 million.
Each study enrolled 12 participants per group—a sample size that was typical for the era but is now widely recognized as inadequate for detecting moderate effects. Power analyses conducted after the fact suggest that detecting a medium effect size of d = 0.5 with 80 percent power would require at least 64 participants per group. The original studies were underpowered by a factor of five. In underpowered designs, random variability and preprocessing choices can dominate the outcome.
The PhD students who conducted the studies spent three to four years each collecting and analyzing data. One is now a postdoctoral fellow at a top-tier institution; the other works in industry. Neither has revisited the slice-timing issue in their subsequent work. When asked whether they would change their approach if they could start over, both said they would probably use the same software defaults, because that is what their advisors and collaborators used.
The broader field faces similar economic pressures. A typical fMRI grant runs for four years and costs $1.5 to $2 million. The pressure to produce publishable results within that timeframe discourages exploratory analyses that might reveal fragility. A lab that spends six months running a multiverse analysis on a single dataset is a lab that is not producing new papers. The incentive structure favors speed over robustness.
What a Preprocessing Registry Would Cost
One proposal to address the problem is a preprocessing registry, analogous to clinical trial registries, where researchers specify their entire analysis pipeline before touching the data. The Organization for Human Brain Mapping (OHBM) formed a committee in 2024 to explore the feasibility of such a registry for neuroimaging studies. The committee estimated that the cost of preregistering a preprocessing pipeline would be roughly $2,000 per study—mostly in staff time to document decisions and justify deviations from defaults.
That figure is small relative to the total cost of a study, but adoption has been slow. As of early 2025, fewer than 5 percent of anxiety-related fMRI studies on OpenNeuro included any form of preprocessing preregistration. Many researchers resist the idea, arguing that preprocessing choices must be adapted to the data and that rigid preregistration would stifle exploratory analysis. Others point out that even with preregistration, a researcher could simply choose the pipeline that is most likely to produce a significant result, as long as they specify it in advance.
Journals could enforce preregistration with minimal overhead. A simple checkbox at submission—'Did you preregister your preprocessing pipeline? If not, explain why'—would cost nothing to implement. But editors worry that adding requirements will drive authors to lower-impact journals that do not enforce them. The result is a race to the bottom, where the most rigorous journals lose submissions to less demanding outlets.
The OHBM committee has proposed a voluntary badge system, similar to the Open Data and Open Materials badges used by some psychology journals. Studies that preregister their preprocessing pipeline and share their raw data would receive a 'Preregistered Analysis' badge. Early evidence from a pilot program at the journal NeuroImage suggests that badges increase data sharing but have little effect on the number of preprocessing sensitivity analyses reported.
Lessons for the Next Generation of Scanners
As fMRI technology advances, the problem of preprocessing sensitivity is likely to grow, not shrink. Newer 7 Tesla scanners offer higher spatial resolution but also introduce more severe slice-timing artifacts due to shorter T2* relaxation times. Multiband acceleration, now standard at many centers, acquires multiple slices simultaneously, which changes the slice-timing correction problem entirely. The default parameters in SPM's successor, SPM2025, have already changed twice during development, each time shifting the recommended correction method.
The anxiety study case illustrates a broader principle: any preprocessing step that involves a choice among plausible alternatives introduces a degree of freedom that can influence results. As the number of preprocessing steps grows—motion correction, slice-timing, smoothing, normalization, denoising, temporal filtering—the number of possible pipelines grows exponentially. A study that runs a single pipeline is effectively gambling that its choices are the correct ones.
Funders are beginning to take notice. The National Institute of Mental Health now requires data management and sharing plans that include a section on preprocessing reproducibility. The European Research Council has funded a consortium to develop standardized pipelines for common task-based fMRI designs. But these efforts are still nascent, and they face resistance from labs that have invested years in custom pipelines and proprietary analysis tools.
The Counter-Argument: When Defaults Are Justified
Not every preprocessing choice is arbitrary, and multiverse analysis is not a panacea. Some defaults are grounded in solid methodological reasoning. For example, slice-timing correction may be unnecessary for blocked designs where stimuli are presented for several seconds, because the hemodynamic response is slow relative to the slice acquisition offset. In such cases, applying correction could introduce unnecessary interpolation noise. The two anxiety studies used event-related designs with rapid presentation, making the default choice consequential, but other studies may be robust to the parameter.
Moreover, multiverse analysis itself carries risks. Testing many pipelines inflates the chance of finding a significant result somewhere, even if the true effect is zero. A researcher could run 64 pipelines, find one that yields p < 0.05, and report that as the main result without correcting for multiple comparisons. The multiverse approach is only as good as the transparency with which it is reported. If researchers cherry-pick the pipeline that supports their hypothesis, the problem is not solved—it is merely moved to a different level.
There is also a practical constraint: time. A thorough multiverse analysis can take weeks or months of computation and interpretation. For a lab operating on a tight grant timeline, the opportunity cost is real. Some researchers argue that the field would be better served by investing those resources in larger sample sizes and preregistered confirmatory analyses, rather than exhaustive exploration of preprocessing options. The optimal balance between exploration and confirmation remains an open question.
Finally, the software default is not always a hidden trap. In some cases, the default is chosen because it is the most robust option across a wide range of designs. The SPM12 team may have turned off slice-timing correction by default because they believed that the correction could introduce artifacts in certain acquisition sequences, and they wanted to avoid those artifacts by default. The problem is that this reasoning was not communicated to users, who then made different choices unknowingly. Better documentation and training could mitigate the issue without requiring a full multiverse analysis for every study.
The most pragmatic recommendation may be the simplest: researchers should routinely run a multiverse analysis on their own data before submitting for publication. If the result depends on a preprocessing parameter that could reasonably be set to a different value, then the study is not about the brain—it is about the software. The anxiety field learned this lesson the hard way, through two expensive studies that canceled each other out. The next field to discover its own slice-timing problem may not be so lucky.