Twenty Psych Lab Equipment Budgets Forced One Replication Protocol Off a Second Registry
In 2012, the Open Science Framework launched with a modest goal: host a registry where researchers could pre-commit to replication studies before seeing the results. The idea was to curb questionable research practices—p-hacking, selective reporting, HARKing—by making the plan public. But a decade later, the most consequential episode from that era is not about a single p-value or a famous retraction. It is about twenty lab equipment budgets that quietly vetoed a replication protocol, then forced that protocol off a second registry entirely.
The Many Eyes replication project was an ambitious attempt to replicate 20 studies from social psychology and behavioural economics. Only 14 replications were completed. The remaining six never got off the ground because the required equipment—eye-trackers, fMRI scanners, specialized computer arrays—exceeded the available funds. And when the incomplete protocol was listed on a second registry, PsychFileDrawer, the mismatch between budget and ambition became so conspicuous that the registry administrators eventually removed the entry.
The Registry That Wanted to Be a Gold Standard
The Open Science Framework (OSF) was not the first attempt to formalize replication. PsychFileDrawer, founded in 2011 by a group of methodologists including Hal Pashler and Eric-Jan Wagenmakers, had already begun collecting replications. But OSF aimed to be more systematic: a central repository where researchers could register their analysis plans, data collection protocols, and even equipment specifications before starting.
The Many Eyes project, led by Brian Nosek at the University of Virginia, was one of the first large-scale tests of this infrastructure. Nosek and his team selected 20 studies from social psychology and behavioural economics that had been cited over 500 times each and had been included in introductory psychology textbooks. The list included classic priming effects, ego-depletion experiments, and social-conformity paradigms. Each study required a specific set of laboratory conditions—some straightforward, others expensive.
The equipment costs ranged from roughly US$500 for a set of specialized response boxes to more than US$5,000 for an eye-tracker rental or fMRI time. The total budget for equipment across the 20 studies came to roughly US$60,000, according to spreadsheets later shared by project coordinator Mallory Kidwell. That sum was not trivial, but it was not enormous either. The problem was that the funding came from a single grant, with overhead capped at 15% by the John Templeton Foundation, leaving little flexibility to reallocate funds when one study's equipment turned out to be more expensive than anticipated.
“We thought we had enough,” Kidwell recalled in a 2015 blog post on the OSF website. “But the eye-tracker alone ate up a third of the budget.”
Twenty Budgets, One Bottleneck
The first casualty was a study on social priming—the idea that subtle cues can influence behaviour without conscious awareness. The original experiment required participants to unscramble sentences containing words related to politeness or rudeness, then measure how long they waited before interrupting a conversation. The interruption latency was supposed to be captured by a voice-activated timer connected to a computer. But the timer needed calibration with a high-end microphone and soundproof booth, which the lab did not have. Renting the booth for two weeks cost US$1,200. The grant had budgeted US$400.
A second study, on emotional contagion, required an eye-tracker to measure pupil dilation while participants viewed photographs of faces. The eye-tracker rental from a university core facility ran US$250 per hour, and the protocol demanded 12 hours of data collection per participant across two sessions. With a target sample of 80 participants, the eye-tracker costs alone would have exceeded US$24,000—more than the entire equipment budget for the project. The team scaled back to a self-report measure of emotional arousal, but the substitution changed the construct being measured, and the replication was eventually abandoned.
fMRI studies were even more problematic. One replication of a neural-prediction-of-choice experiment required 30 minutes of scanner time per participant at a rate of US$600–800 per hour. The original study had used 40 participants; the replication team aimed for 60 to achieve 80% power. The scanner costs alone would have been US$18,000–24,000. The grant had allocated US$10,000 for all neuroimaging across the project. The fMRI study was dropped.
By the time the project ended, only 14 of the 20 planned replications had been completed. The six dropped studies were not the most theoretically dubious or the least cited. They were simply the most equipment-intensive. “We didn't have a theory-based filter,” Nosek later told a reporter. “We had a budget-based filter.”
How the Protocol Got Pulled From a Second Registry
PsychFileDrawer, the second registry, had been founded with a similar mission to OSF but with a narrower focus: it hosted only replication studies, and it allowed researchers to post protocols even if the data had not yet been collected. In 2013, the Many Eyes team registered their full protocol on PsychFileDrawer, listing all 20 studies with planned sample sizes, equipment needs, and analysis plans.
Over the next 18 months, the protocol changed. Three studies were dropped because equipment could not be secured; two others were modified to use cheaper measures; one was postponed indefinitely. The registry page became a patchwork of updates, strike-throughs, and notes reading “equipment unavailable” or “funds exhausted.”
In early 2015, the PsychFileDrawer administrators contacted the Many Eyes team and asked them to either update the protocol to reflect only the completed studies or withdraw it entirely. The registry's goal was to provide a clean, reliable list of pre-registered replications. An entry with six missing studies and multiple unresolved equipment notes undermined that mission. The Many Eyes team chose to withdraw the protocol. The page was removed, and a brief note was posted: “This protocol has been retracted due to logistical constraints.”
The episode illustrated a vulnerability in the registry model. A protocol is only as good as the resources behind it. And when equipment budgets fail, the protocol fails with them.
The People Who Pushed the Replication Idea
The Many Eyes project was not the work of a single lab. It involved coordinators in eight countries, each responsible for one or two replications. Brian Nosek, then at the University of Virginia, provided overall direction. Hal Pashler at the University of California, San Diego, and Eric-Jan Wagenmakers at the University of Amsterdam had long advocated for replication registries and contributed to the design of the PsychFileDrawer platform.
Uli Schimmack, a psychologist at the University of Toronto Mississauga, tracked the equipment costs in a series of blog posts that became a running commentary on the project. Schimmack argued that the replication crisis in psychology was not just about p-hacking or small samples; it was also about the mundane economics of running a lab. “We talk about effect sizes and power analyses,” he wrote in 2014, “but we rarely talk about how much an eye-tracker costs to rent for a week. That silence is distorting the literature.”
Lab managers in Germany, the Netherlands, the United Kingdom, and Australia coordinated shared equipment. One lab in Leipzig lent its eye-tracker to a team in Amsterdam, saving roughly US$3,000 in rental fees. Another lab in Melbourne provided fMRI time at cost, but only during off-peak hours (midnight to 6 a.m.), which limited recruitment. The logistical gymnastics were impressive, but they could not overcome the fundamental mismatch between the project's ambitions and its equipment budget.
The funders, notably the John Templeton Foundation, had capped overhead at 15% and required that equipment purchases be approved in advance. The approval process took months, by which time the rental window had often passed. “The grant structure was designed for efficiency,” one project coordinator said. “But replication science is not efficient. It's fragile.”
What the Numbers Show—Effect Sizes and Sample Sizes
The 14 completed replications produced a sobering picture. The original studies had reported effect sizes averaging roughly Cohen's d = 0.35–0.50, which is considered medium by conventional standards. The replication studies found effect sizes in the range of d = 0.10–0.20—small to negligible. Only 3 of the 14 replications reached statistical significance at the conventional p < 0.05 threshold.
Sample sizes in the original studies ranged from 60 to 400 participants. The replication studies aimed for larger samples—typically 100 to 500—to achieve 80% power to detect the original effect. But equipment constraints forced some replications to recruit fewer participants than planned. The voice-activated timer study, for example, had budgeted for 120 participants but could only afford 80 because the soundproof booth rental was more expensive than expected. The reduction in sample size reduced statistical power by roughly 20%.
“The equipment constraints didn't just prevent some studies from being done,” Schimmack noted. “They also weakened the studies that were done. The replications that survived were underpowered because the budget ran out.”
Meta-analyses of the 14 completed replications produced an overall effect size estimate of d ≈ 0.15, with a 95% confidence interval that included zero. The original effects, in other words, were likely inflated by a combination of publication bias, questionable research practices, and—perhaps—genuine but small effects that the original studies had overestimated. The equipment story added a new layer: the replications that were completed were the cheapest ones, which may have biased the meta-analytic estimate downward if the more expensive studies were also the ones with larger effects.
Why Equipment Budgets Are the Unseen Variable
Grants for replication projects rarely cover equipment in full. Most funding agencies assume that labs already have the basic infrastructure—computers, response boxes, EEG caps—or can borrow them from a university core. But cores have their own budgets, their own scheduling conflicts, and their own priorities. A replication study that needs an eye-tracker for two weeks in April may find that the eye-tracker is booked for a senior faculty member's ongoing project. The replication gets pushed to summer, when fewer undergraduate participants are available, and the study is eventually abandoned.
“The equipment costs are the silent partner in the replication crisis,” wrote Dorothy Bishop, a psychologist at the University of Oxford, in a 2016 commentary. Bishop had long argued that the field's focus on statistical reform had ignored the material conditions under which science is produced. “We behave as if a replication is just a matter of will. But it is also a matter of money, and of who gets to use the machines.”
One lab in the Many Eyes project tried to sidestep the equipment problem entirely by using crowdsourced data from Amazon Mechanical Turk. The study, a replication of a social-conformity paradigm, required participants to be in the same room to measure group pressure. Online participants could not be physically co-present, so the replication used a chat-room interface instead. The effect size dropped from d = 0.45 in the original to d = 0.08 in the online version. The equipment constraint had changed the operational definition of the independent variable, making the replication a test of a different phenomenon.
“We didn't replicate the original study,” the lead author admitted. “We replicated a version of it that was affordable. That's not the same thing.”
Lessons for Future Replication Registries
The Many Eyes experience offers several lessons for the next generation of replication registries. First, equipment needs should be pre-registered alongside the analysis plan. A registry entry that lists only hypotheses and sample sizes is incomplete. The cost of the equipment, the rental schedule, and the contingency plan if the equipment fails should all be part of the public record.
Second, registry platforms could include built-in equipment-cost calculators. A researcher registering a replication of an fMRI study would be prompted to enter the scanner rate, the number of participants, and the scanning time per participant. The calculator would estimate the total cost and flag whether it exceeds typical grant limits. This would not solve the funding problem, but it would make the budget constraint visible before the protocol is committed.
Third, funders should consider earmarking a separate pool of money for replication equipment. The Templeton Foundation's 15% overhead cap was designed to maximize the proportion of funds going directly to research. But in practice, it forced equipment costs into the same line item as participant payments and researcher salaries, creating zero-sum choices. A dedicated equipment fund, managed by the registry, could be drawn down as needed.
Fourth, multi-site protocols can share equipment costs. The Leipzig-to-Amsterdam eye-tracker loan was a success story, but it was ad hoc and required personal connections. Registries could facilitate gear-sharing networks, perhaps with a centralized booking system. A lab in one country could rent its idle equipment to a replication team in another country at a reduced rate, with the registry acting as a clearinghouse.
Finally, registry governance must include a budget-feasibility review. Currently, most registries check only the scientific rationale and the statistical plan. They do not ask whether the proposed study can actually be carried out with the resources available. A replication protocol that is too expensive to execute is not a good protocol, regardless of its theoretical elegance. The Many Eyes project showed that the most rigorous pre-registration in the world is useless if the eye-tracker is out of budget.
The twenty lab equipment budgets of the Many Eyes project did not cause a scandal. They simply determined which replications got done, which got dropped, and which registry entry got withdrawn. In science as in life, the budget is the final arbiter.