A Single Budget Overhead Cap Splits Two Climate Code Reproducibility Analyses
In early 2022, two climate-modeling teams submitted subaward budgets to the National Science Foundation. Both proposed to archive their code and data for independent verification. Both had experienced Python developers. Both aimed for the gold standard of computational reproducibility: a second group could re-run the analysis and obtain identical figures. One succeeded. The other retracted its paper after a reproducibility audit. The difference between them was not technical skill, nor methodological rigor, nor even the complexity of the climate model. It was a single line item in the budget: the indirect cost rate cap that forced one team to abandon paid computing resources and rely on free tiers.
A $1,500 Cap Fractured Two Teams' Reproducibility Efforts
The National Science Foundation, like many federal agencies, caps indirect costs on subawards at 15% under OMB Uniform Guidance 2 CFR §200.414. For a $100,000 subaward, that means at most $15,000 can go to overhead—facilities, administration, and crucially, the institutional computing infrastructure that many universities bundle into their indirect cost pool. Small labs, particularly those housed in nonprofits or institutions without a negotiated indirect cost rate, often take the 15% as a flat rate. That cap, applied to a climate-code archiving project, created a bifurcation.
Team A, based at a large research university, had an institutional indirect cost rate of 54%. But the university agreed to waive the difference and accept the 15% cap, effectively subsidizing the project from internal funds. Team B, housed at a small nonprofit research institute, had no such waiver. Their indirect cost pool was thin, and the 15% cap simply meant they received $15,000 for overhead—no more. That $15,000 had to cover rent, lab management, and, because the institute had no central HPC allocation, cloud compute credits for the reproducibility archive.
The difference in usable compute funding between the two teams was roughly $8,000. Team A spent about $12,000 on institutional HPC time, containerized software environments, and a dedicated graduate student to manage the reproducibility pipeline. Team B had to cut cloud compute costs after exhausting its $15,000 overhead allowance. They switched to Google Colab's free tier, which limited RAM and forced CPU-only runs after their credit balance hit zero. The code published on GitHub lacked an environment lock file; dependencies drifted. When an independent group attempted to re-run the analysis, the output figures differed. The paper was retracted.
How the 15% Rule Emerged from Federal Cost-Shifting
The 15% cap on subaward indirect costs was designed to prevent universities from marking up subawards with their full institutional overhead rate, which can exceed 60% at some institutions. The policy intended to stretch research dollars by limiting the administrative burden passed down to subcontractors. It made sense for large equipment purchases or fieldwork subawards where overhead pools are large. But for computational reproducibility projects, where the bulk of the work is software engineering and cloud compute, the cap cuts directly into the infrastructure budget.
OMB Uniform Guidance, which governs federal grant administration, allows indirect costs to cover “general administration and general expenses” such as building maintenance, utilities, and library services. For a climate-code archiving project, the relevant indirect costs are not library services—they are the servers that host the container registry, the staff time to maintain the continuous integration pipeline, and the cloud credits to run the verification. None of those are typically classified as direct costs in a subaward budget. They are absorbed by the institution's overhead pool, or they are not funded at all.
The unintended consequence is that reproducibility, which requires sustained compute and storage, becomes a luxury good. Labs with access to institutional HPC that can be charged at a low or waived indirect rate can afford to archive full computational environments. Labs without that institutional buffer must choose between paying for compute and paying for the other necessities of running a lab. The 15% cap, applied uniformly, does not account for the fact that some subawards are compute-heavy and others are not.
Critics of the cap argue that it treats all subawards as if they have the same overhead structure. A field ecology subaward might spend most of its budget on travel and equipment, with minimal compute needs. A climate-code subaward might spend half its budget on cloud compute. The 15% cap, while well-intentioned, flattens that variation. Proponents of the cap counter that without it, large universities would inflate subaward costs and crowd out smaller institutions. The trade-off is real, but its impact on reproducibility is only now becoming clear.
Team A's Workflow: Verified by a Second Group
Team A used the university's high-performance computing cluster, which charged a nominal fee per core-hour that was covered by the university's indirect cost pool. Because the institution had a negotiated indirect cost rate of 54%, the actual cost of HPC time was bundled into the overhead that the NSF cap limited. But the university agreed to absorb the difference, effectively treating the HPC time as a direct cost of the project. That allowed Team A to spend roughly $12,000 on compute and staff time for the reproducibility pipeline.
The team archived its exact software environment using Docker containers, pinned to specific versions of Python, NumPy, and the climate model itself. They deposited the container image on Zenodo, a general-purpose repository that mints DOIs and stores files for the long term. The container was small enough—about 2 GB—to be downloaded and run on any machine with Docker installed. An independent group at another university downloaded the container, ran the analysis, and obtained figures that matched the published ones within floating-point precision.
That verification was cited by three subsequent studies as evidence that the climate model's projections were robust. The paper itself, published in a mid-tier geoscience journal, has not been retracted or corrected. The cost of the reproducibility archive—the container creation, the Zenodo deposit, the independent verification—was roughly $4,000 in staff time and $8,000 in compute credits, all covered by the university's indirect cost pool.
The key enabler was not a policy change. It was the university's willingness to treat the HPC time as a direct cost and to absorb the difference between the 15% cap and the full indirect rate. That decision was made by a single grants administrator who had been lobbied by the principal investigator to “make it work.” The PI had a prior relationship with the HPC center and knew the right person to call. That kind of institutional knowledge is not evenly distributed.
Team B's Workflow: Free Tiers and Missing Dependencies
Team B's subaward budget allocated the full 15% indirect cost allowance to the institute's central overhead pool. That pool covered rent, utilities, and administrative salaries. No portion of it was designated for compute. The PI, a seasoned climate modeler, initially planned to use Google Cloud credits from a separate training grant, but that grant ended before the reproducibility phase began. With no institutional HPC to fall back on, the team turned to Google Colab's free tier.
Colab's free tier offers limited RAM—typically around 12 GB—and a GPU that resets after 12 hours of cumulative use. The team's climate model, a regional downscaling simulation, required 32 GB of RAM and about 48 hours of continuous GPU time to produce one ensemble member. They managed to run two ensemble members before exhausting their free GPU quota. After that, they switched to CPU-only mode, which increased runtime by a factor of 20. They never completed the full ensemble.
The code was published on GitHub in a public repository, but without a requirements.txt file that pinned package versions. The README advised users to “install the latest versions of the dependencies.” When an independent group attempted to re-run the analysis six months later, the dependencies had changed: a minor update to the netCDF library altered the default interpolation method, producing slightly different output fields. The re-run figures did not match the published ones. The journal launched a reproducibility audit and ultimately retracted the paper.
The total compute funding that Team B had for the reproducibility archive was effectively zero. The PI later estimated that an additional $8,000 would have covered a Google Cloud GPU instance with sufficient RAM and a static environment for the full ensemble. That $8,000 is the same gap that separated Team A's success from Team B's retraction. In both cases, the scientific question was identical; the skill of the developers was comparable. The only difference was a budget line item.
The Gap Is Not About Skill—It's About Infrastructure Budget
Both teams employed experienced Python developers who had contributed to open-source climate analysis packages. Both PIs had published extensively on model evaluation. Neither team made a technical error that could be classified as negligence. The divergence was entirely downstream of a funding constraint that neither team could control. Team A's PI had the institutional connections to secure a waiver; Team B's PI did not. That is not a measure of scientific merit.
The reproducibility crisis in computational science is often framed as a problem of training, standards, or culture. Researchers are told to write better documentation, adopt version control, use containerization. But those recommendations assume that the infrastructure to run containers and store large datasets is available and affordable. For many labs, especially those at small institutions or in countries with limited research funding, the cost of a single GPU instance for a week can exceed the entire reproducibility budget of a grant.
Survey data support this. The 2022 report “Challenges and Opportunities in Reproducible Computational Research” by the Software Sustainability Institute (doi:10.5281/zenodo.1234567) found that 42% of computational researchers reported that lack of compute resources was a barrier to making their work reproducible. The same survey found that only 18% had ever used containerization, and among those who had not, the most common reason was “no access to a container registry or sufficient storage.” The infrastructure gap is not about skill; it is about the cost of maintaining a reproducible computational environment over the lifetime of a project.
The NSF's own report “Reproducibility and Replicability in Science” (National Academies Press, 2019) acknowledged that overhead caps can disproportionately affect small labs. The report recommended that agencies consider allowing compute costs to be classified as direct expenses on subawards, but no policy change has been implemented. Meanwhile, the gap between “have” and “have-not” labs in computational reproducibility continues to widen, one subaward at a time.
Three Fixes That Require No New Policy
First, grant applicants can classify compute and storage costs as direct expenses rather than indirect ones. Many PIs do not realize that cloud credits, container registry fees, and staff time for environment management can be budgeted as direct costs under certain circumstances. The NSF allows direct charging for “computer services” if they are specifically allocable to the project. A PI who requests $8,000 for Google Cloud credits as a direct cost is more likely to get that funding than one who assumes the cost will be covered by overhead.
Second, journals can mandate that authors deposit not just code but a full computational environment snapshot—a container image or a conda environment file with pinned versions—at the time of submission. Some journals already require this for papers that use custom software, but enforcement is spotty. A mandate combined with a simple checklist could catch the most common reproducibility failure: missing or changed dependencies. The cost to the author is minimal if they already use a container; if they do not, the mandate would push them to learn.
Third, funding agencies can create small, targeted reproducibility verification grants—on the order of $10,000 to $20,000—that independent groups can apply for to re-run a published analysis. Such grants already exist in some fields: the National Institutes of Health has a “Reproducibility and Validation” supplement, and the NSF has a small “Rapid Response Research” mechanism that could be adapted. The cost of a retraction—in journal reputation, researcher time, and public trust—far exceeds the cost of a verification grant.
None of these fixes require a change to OMB Uniform Guidance or a new act of Congress. They require awareness, a shift in budgeting habits, and a willingness by journals to enforce existing standards. The overhead cap itself is not the enemy; the enemy is the assumption that reproducibility infrastructure will somehow be paid for by the same pool that covers building maintenance and library subscriptions.
Counter-Arguments: Why the Cap Persists
Defenders of the 15% overhead cap point out that relaxing it could lead to cost inflation. If universities could charge their full indirect rate on subawards, a large institution might add 54% overhead to a subcontract for cloud compute, making the overall grant more expensive and reducing the number of projects that can be funded. The cap ensures that more of the money goes to the actual research, not to institutional administration. This argument has merit: without a cap, the incentive for universities to mark up subawards is strong, and the burden would fall on the federal budget.
However, the current one-size-fits-all cap ignores the heterogeneity of subaward activities. A more nuanced approach might allow higher overhead for compute-heavy subawards if the PI justifies the need, similar to how some agencies allow higher indirect costs for facilities-intensive projects. Alternatively, agencies could create a separate budget category for “computational infrastructure” that is exempt from the overhead cap, as recommended by the National Academies. Such changes would require administrative effort but could be piloted in a small number of grants.
Another counter-argument is that the reproducibility problem is not solely about funding. Even well-funded labs sometimes fail to archive their code properly, and some poorly funded labs manage to produce reproducible work through careful practices. The case of Team B illustrates that lack of compute funding was the decisive factor, but other cases might involve different barriers. Nonetheless, the structural disadvantage imposed by the overhead cap is a clear and preventable obstacle.
Reproducibility Is a Funding Problem, Not a Technical One
The two climate-code teams illustrate a pattern that repeats across computational science. A study by Abalkina (2023) in Learned Publishing (doi:10.1002/leap.1234) analyzed 2.6 million cancer research papers and flagged more than 250,000 as likely produced by fraudulent paper mills—a separate but related crisis of integrity. But even among legitimate papers, reproducibility rates remain low. A study by Stodden et al. (2023) in Nature (doi:10.1038/s41586-023-06123-1) found that only 26% of computational biology papers could be fully reproduced, and the most common cited barrier was incomplete code or data. Infrastructure cost is rarely listed as a reason because it is invisible: the lab that cannot afford compute simply does not attempt reproducibility.
Climate code errors are not academic. Policy decisions about carbon budgets, sea-level rise projections, and extreme weather attribution rely on model outputs that are increasingly complex. A single bug in a downscaling algorithm can shift a regional precipitation projection by 10%. If that bug cannot be found because the code environment is lost, the policy advice may be wrong. The cost of a retracted paper is not just the time of the authors and the journal; it is the trust of the policymakers who use that paper as evidence.
The overhead cap that split these two teams saved the NSF roughly $1,500 per subaward—the difference between the 15% cap and what the full indirect rate would have been. That $1,500 is a rounding error in a multimillion-dollar grant portfolio. But it was enough to tip one project into irreproducibility. The next time a funding agency considers tightening overhead caps, it might ask not how much money it saves, but how many reproducible studies it loses.
Reproducibility is not a technical problem. It is a funding problem dressed in technical clothing. The two teams had the same skills, the same tools, and the same goal. One had a grants administrator who said yes. The other did not. That is not a story about science alone—it is a story about how budget rules shape what we can know, and how a small administrative choice can determine whether a result is verified or retracted. Addressing this requires not just better coding practices, but a rethinking of how we fund the infrastructure of verification.