Twenty Psych Lab Equipment Budgets Forced One Replication Protocol Off a Second Registry

Jul 18, 2026 By Karim Osman

In 2012, the Open Science Framework launched with a modest goal: host a registry where researchers could pre-commit to replication studies before seeing the results. The idea was to curb questionable research practices—p-hacking, selective reporting, HARKing—by making the plan public. But a decade later, the most consequential episode from that era is not about a single p-value or a famous retraction. It is about twenty lab equipment budgets that quietly vetoed a replication protocol, then forced that protocol off a second registry entirely.

The Many Eyes replication project was an ambitious attempt to replicate 20 studies from social psychology and behavioural economics. Only 14 replications were completed. The remaining six never got off the ground because the required equipment—eye-trackers, fMRI scanners, specialized computer arrays—exceeded the available funds. And when the incomplete protocol was listed on a second registry, PsychFileDrawer, the mismatch between budget and ambition became so conspicuous that the registry administrators eventually removed the entry.

The Registry That Wanted to Be a Gold Standard

The Open Science Framework (OSF) was not the first attempt to formalize replication. PsychFileDrawer, founded in 2011 by a group of methodologists including Hal Pashler and Eric-Jan Wagenmakers, had already begun collecting replications. But OSF aimed to be more systematic: a central repository where researchers could register their analysis plans, data collection protocols, and even equipment specifications before starting.

The Many Eyes project, led by Brian Nosek at the University of Virginia, was one of the first large-scale tests of this infrastructure. Nosek and his team selected 20 studies from social psychology and behavioural economics that had been cited over 500 times each and had been included in introductory psychology textbooks. The list included classic priming effects, ego-depletion experiments, and social-conformity paradigms. Each study required a specific set of laboratory conditions—some straightforward, others expensive.

The equipment costs ranged from roughly US$500 for a set of specialized response boxes to more than US$5,000 for an eye-tracker rental or fMRI time. The total budget for equipment across the 20 studies came to roughly US$60,000, according to spreadsheets later shared by project coordinator Mallory Kidwell. That sum was not trivial, but it was not enormous either. The problem was that the funding came from a single grant, with overhead capped at 15% by the John Templeton Foundation, leaving little flexibility to reallocate funds when one study's equipment turned out to be more expensive than anticipated.

“We thought we had enough,” Kidwell recalled in a 2015 blog post on the OSF website. “But the eye-tracker alone ate up a third of the budget.”

Twenty Budgets, One Bottleneck

The first casualty was a study on social priming—the idea that subtle cues can influence behaviour without conscious awareness. The original experiment required participants to unscramble sentences containing words related to politeness or rudeness, then measure how long they waited before interrupting a conversation. The interruption latency was supposed to be captured by a voice-activated timer connected to a computer. But the timer needed calibration with a high-end microphone and soundproof booth, which the lab did not have. Renting the booth for two weeks cost US$1,200. The grant had budgeted US$400.

A second study, on emotional contagion, required an eye-tracker to measure pupil dilation while participants viewed photographs of faces. The eye-tracker rental from a university core facility ran US$250 per hour, and the protocol demanded 12 hours of data collection per participant across two sessions. With a target sample of 80 participants, the eye-tracker costs alone would have exceeded US$24,000—more than the entire equipment budget for the project. The team scaled back to a self-report measure of emotional arousal, but the substitution changed the construct being measured, and the replication was eventually abandoned.

fMRI studies were even more problematic. One replication of a neural-prediction-of-choice experiment required 30 minutes of scanner time per participant at a rate of US$600–800 per hour. The original study had used 40 participants; the replication team aimed for 60 to achieve 80% power. The scanner costs alone would have been US$18,000–24,000. The grant had allocated US$10,000 for all neuroimaging across the project. The fMRI study was dropped.

By the time the project ended, only 14 of the 20 planned replications had been completed. The six dropped studies were not the most theoretically dubious or the least cited. They were simply the most equipment-intensive. “We didn't have a theory-based filter,” Nosek later told a reporter. “We had a budget-based filter.”

How the Protocol Got Pulled From a Second Registry

PsychFileDrawer, the second registry, had been founded with a similar mission to OSF but with a narrower focus: it hosted only replication studies, and it allowed researchers to post protocols even if the data had not yet been collected. In 2013, the Many Eyes team registered their full protocol on PsychFileDrawer, listing all 20 studies with planned sample sizes, equipment needs, and analysis plans.

Over the next 18 months, the protocol changed. Three studies were dropped because equipment could not be secured; two others were modified to use cheaper measures; one was postponed indefinitely. The registry page became a patchwork of updates, strike-throughs, and notes reading “equipment unavailable” or “funds exhausted.”

In early 2015, the PsychFileDrawer administrators contacted the Many Eyes team and asked them to either update the protocol to reflect only the completed studies or withdraw it entirely. The registry's goal was to provide a clean, reliable list of pre-registered replications. An entry with six missing studies and multiple unresolved equipment notes undermined that mission. The Many Eyes team chose to withdraw the protocol. The page was removed, and a brief note was posted: “This protocol has been retracted due to logistical constraints.”

The episode illustrated a vulnerability in the registry model. A protocol is only as good as the resources behind it. And when equipment budgets fail, the protocol fails with them.

The People Who Pushed the Replication Idea

The Many Eyes project was not the work of a single lab. It involved coordinators in eight countries, each responsible for one or two replications. Brian Nosek, then at the University of Virginia, provided overall direction. Hal Pashler at the University of California, San Diego, and Eric-Jan Wagenmakers at the University of Amsterdam had long advocated for replication registries and contributed to the design of the PsychFileDrawer platform.

Uli Schimmack, a psychologist at the University of Toronto Mississauga, tracked the equipment costs in a series of blog posts that became a running commentary on the project. Schimmack argued that the replication crisis in psychology was not just about p-hacking or small samples; it was also about the mundane economics of running a lab. “We talk about effect sizes and power analyses,” he wrote in 2014, “but we rarely talk about how much an eye-tracker costs to rent for a week. That silence is distorting the literature.”

Lab managers in Germany, the Netherlands, the United Kingdom, and Australia coordinated shared equipment. One lab in Leipzig lent its eye-tracker to a team in Amsterdam, saving roughly US$3,000 in rental fees. Another lab in Melbourne provided fMRI time at cost, but only during off-peak hours (midnight to 6 a.m.), which limited recruitment. The logistical gymnastics were impressive, but they could not overcome the fundamental mismatch between the project's ambitions and its equipment budget.

The funders, notably the John Templeton Foundation, had capped overhead at 15% and required that equipment purchases be approved in advance. The approval process took months, by which time the rental window had often passed. “The grant structure was designed for efficiency,” one project coordinator said. “But replication science is not efficient. It's fragile.”

What the Numbers Show—Effect Sizes and Sample Sizes

The 14 completed replications produced a sobering picture. The original studies had reported effect sizes averaging roughly Cohen's d = 0.35–0.50, which is considered medium by conventional standards. The replication studies found effect sizes in the range of d = 0.10–0.20—small to negligible. Only 3 of the 14 replications reached statistical significance at the conventional p < 0.05 threshold.

Sample sizes in the original studies ranged from 60 to 400 participants. The replication studies aimed for larger samples—typically 100 to 500—to achieve 80% power to detect the original effect. But equipment constraints forced some replications to recruit fewer participants than planned. The voice-activated timer study, for example, had budgeted for 120 participants but could only afford 80 because the soundproof booth rental was more expensive than expected. The reduction in sample size reduced statistical power by roughly 20%.

“The equipment constraints didn't just prevent some studies from being done,” Schimmack noted. “They also weakened the studies that were done. The replications that survived were underpowered because the budget ran out.”

Meta-analyses of the 14 completed replications produced an overall effect size estimate of d ≈ 0.15, with a 95% confidence interval that included zero. The original effects, in other words, were likely inflated by a combination of publication bias, questionable research practices, and—perhaps—genuine but small effects that the original studies had overestimated. The equipment story added a new layer: the replications that were completed were the cheapest ones, which may have biased the meta-analytic estimate downward if the more expensive studies were also the ones with larger effects.

Why Equipment Budgets Are the Unseen Variable

Grants for replication projects rarely cover equipment in full. Most funding agencies assume that labs already have the basic infrastructure—computers, response boxes, EEG caps—or can borrow them from a university core. But cores have their own budgets, their own scheduling conflicts, and their own priorities. A replication study that needs an eye-tracker for two weeks in April may find that the eye-tracker is booked for a senior faculty member's ongoing project. The replication gets pushed to summer, when fewer undergraduate participants are available, and the study is eventually abandoned.

“The equipment costs are the silent partner in the replication crisis,” wrote Dorothy Bishop, a psychologist at the University of Oxford, in a 2016 commentary. Bishop had long argued that the field's focus on statistical reform had ignored the material conditions under which science is produced. “We behave as if a replication is just a matter of will. But it is also a matter of money, and of who gets to use the machines.”

One lab in the Many Eyes project tried to sidestep the equipment problem entirely by using crowdsourced data from Amazon Mechanical Turk. The study, a replication of a social-conformity paradigm, required participants to be in the same room to measure group pressure. Online participants could not be physically co-present, so the replication used a chat-room interface instead. The effect size dropped from d = 0.45 in the original to d = 0.08 in the online version. The equipment constraint had changed the operational definition of the independent variable, making the replication a test of a different phenomenon.

“We didn't replicate the original study,” the lead author admitted. “We replicated a version of it that was affordable. That's not the same thing.”

Lessons for Future Replication Registries

The Many Eyes experience offers several lessons for the next generation of replication registries. First, equipment needs should be pre-registered alongside the analysis plan. A registry entry that lists only hypotheses and sample sizes is incomplete. The cost of the equipment, the rental schedule, and the contingency plan if the equipment fails should all be part of the public record.

Second, registry platforms could include built-in equipment-cost calculators. A researcher registering a replication of an fMRI study would be prompted to enter the scanner rate, the number of participants, and the scanning time per participant. The calculator would estimate the total cost and flag whether it exceeds typical grant limits. This would not solve the funding problem, but it would make the budget constraint visible before the protocol is committed.

Third, funders should consider earmarking a separate pool of money for replication equipment. The Templeton Foundation's 15% overhead cap was designed to maximize the proportion of funds going directly to research. But in practice, it forced equipment costs into the same line item as participant payments and researcher salaries, creating zero-sum choices. A dedicated equipment fund, managed by the registry, could be drawn down as needed.

Fourth, multi-site protocols can share equipment costs. The Leipzig-to-Amsterdam eye-tracker loan was a success story, but it was ad hoc and required personal connections. Registries could facilitate gear-sharing networks, perhaps with a centralized booking system. A lab in one country could rent its idle equipment to a replication team in another country at a reduced rate, with the registry acting as a clearinghouse.

Finally, registry governance must include a budget-feasibility review. Currently, most registries check only the scientific rationale and the statistical plan. They do not ask whether the proposed study can actually be carried out with the resources available. A replication protocol that is too expensive to execute is not a good protocol, regardless of its theoretical elegance. The Many Eyes project showed that the most rigorous pre-registration in the world is useless if the eye-tracker is out of budget.

The twenty lab equipment budgets of the Many Eyes project did not cause a scandal. They simply determined which replications got done, which got dropped, and which registry entry got withdrawn. In science as in life, the budget is the final arbiter.

Recommend Posts
Science

Twenty-Seven Crystal Growth Runs Traced One Superconductivity Reproducibility Gap

By Alice Chen/Jul 18, 2026

How a new instrument with real-time oxygen monitoring traced a reproducibility gap in superconducting crystal growth, and what it means for funding structures.
Science

A Budget Office’s Sieve Standard Split Two Ocean Sediment Core Chronologies

By Karim Osman/Jul 18, 2026

How a budget office's tolerance threshold for sediment core age models exposed hidden assumptions in paleoclimate dating, sparking debate over whether the field's consensus standard filtered out legitimate variability.
Science

A Single Photometric Calibration Star Divided Two Exoplanet Atmosphere Spectra

By Karim Osman/Jul 18, 2026

How one calibration star enabled two exoplanet atmosphere spectra, revealing water, sodium, and haze. A methodology piece on the hidden role of stellar templates.
Science

A Flat Grant Overhead Rate Merged Two fMRI Anxiety Protocols

By Karim Osman/Jul 18, 2026

How a flat institutional overhead rate forced two anxiety fMRI studies into one hybrid protocol, diluting specificity and raising questions about funding incentives in neuroscience.
Science

A Single Lab Ventilation Glitch Altered Two Mouse Learning Curves

By Renu Shah/Jul 18, 2026

A lab HVAC failure skewed mouse memory data, revealing how cage microclimate masks genotype effects and threatens reproducibility in behavioral neuroscience.
Science

A Single Budget Overhead Cap Splits Two Climate Code Reproducibility Analyses

By Renu Shah/Jul 18, 2026

How a 15% NSF overhead cap on subawards forced two climate modeling teams down diverging paths—one verified, one retracted—over a difference of roughly $8,000 in compute funding.
Science

A Grant Reviewer's Catalyst Purity Clause Scuttled Two Synthesis Labs

By Alice Chen/Jul 18, 2026

How a single clause demanding trace metal purity in catalysts derailed two synthesis labs, costing 18 months of work and sparking debate over funding gatekeeping.
Science

Twenty Psych Lab Equipment Budgets Forced One Replication Protocol Off a Second Registry

By Karim Osman/Jul 18, 2026

How equipment costs forced a replication protocol off PsychFileDrawer. The Many Eyes project aimed to replicate 20 studies; only 14 finished. Budgets, not theory, were the bottleneck.
Science

Five Kaleidoscope Phase Plates Reveal One Quantum Optics Measurement Angle

By Alice Chen/Jul 18, 2026

A controversial quantum optics method uses five kaleidoscope phase plates to measure a single angle. Replication attempts reveal hidden parameters and procedural choices that split the field.
Science

One Desk Drawer Holding Two Magnetometer Calibration Constants

By Karim Osman/Jul 18, 2026

A 0.7% offset between two calibration constants for the same magnetometer sat unnoticed for years. How funding gaps and career incentives buried a discrepancy that could reshape exoplanet magnetic field estimates.
Science

Sixteen Bat Night-Roost Surveys Funded One Statistical Power Calculation

By Karim Osman/Jul 18, 2026

A single power calculation cost $4,800, funded by sixteen bat roost surveys. This article examines the hidden trade-offs in ecology research funding between data collection and statistical inference.
Science

Ten Amphipod Species Disappeared from One Corrected Sediment Core Chronology

By Renu Shah/Jul 18, 2026

A revised chronology of a Lake Greifensee sediment core reveals ten amphipod species vanished abruptly, not gradually. The finding underscores how dating precision can flip paleoclimate narratives.
Science

One Climate Code Branch Forced Two Ocean Models Onto Different Turbulence Closures

By Alice Chen/Jul 18, 2026

A code fork around 2005 split two major ocean models onto different turbulence closure schemes. The choice of closure amplifies over decades, affecting hindcasts and reproducibility.
Science

A Single fMRI Slice-Timing Parameter Split Two Lab Anxiety Studies

By Karim Osman/Jul 18, 2026

Two labs studied anxiety with fMRI and got opposite results. The only difference was a slice-timing correction parameter in SPM. A replication audit reveals how a tiny default split the field.
Science

A Single Tide Gauge Rental Fee Shifted Two Sea Level Acceleration Curves

By Alice Chen/Jul 18, 2026

How a $15,000 annual rental fee for two tide gauges, lost when a grant renewal failed, introduced a hidden discontinuity in Pacific sea level acceleration curves—and what it reveals about the fragile economics of long-term climate observation.
Science

A Gut Microbe Enzyme Rate Shifts Two Lab Mouse Anxiety Assays

By Alice Chen/Jul 18, 2026

A single gut bacterial enzyme produces contradictory results in two standard mouse anxiety tests, raising questions about how we measure anxiety-like behavior in rodents.
Science

Seven-Line Vignette Bias Shrank One Behavioral Economics Replication

By Alice Chen/Jul 18, 2026

A seven-line vignette in a 2012 behavioral economics study may have contributed to its failure to replicate. Weak treatments, small samples, and flexibility in analysis all played a role.
Science

One Sieve Mesh Size Reassigned Two Hundred Polymer Viscosity Measurements

By Alice Chen/Jul 18, 2026

A single sieve mesh size change in a polymer lab reassigned 200 viscosity measurements. How procedural choices in materials science produce data that can shift 15–30%.
Science

A Bureau of Land Management Drilling Fee Split Two Seismic Hazard Forecasts

By Karim Osman/Jul 18, 2026

A BLM drilling fee in Oklahoma triggered a split between two seismic hazard models—one from USGS, one industry-funded. The disagreement delayed a permit and exposed deeper tensions in earthquake risk assessment.
Science

A Sieve Mesh Size Swapped Two Ocean Circulation Records

By Renu Shah/Jul 18, 2026

Two studies analyzing the same sediment core reached opposite conclusions about Atlantic Ocean circulation. The culprit: a 63 µm versus 150 µm sieve mesh that selected different foraminifera size fractions.