A Budget Office’s Sieve Standard Split Two Ocean Sediment Core Chronologies
In the early 2010s, two deep-sea sediment cores—one pulled from the Pacific Ocean floor, the other from the Atlantic—seemed to record the same glacial cycles. Both cores contained layers of foraminifera shells whose oxygen isotope ratios, a proxy for global ice volume, rose and fell in patterns that should have matched. Yet when researchers built age models for the two cores, the chronologies diverged by tens of millennia. The disagreement, at first treated as a local calibration problem, eventually exposed a hidden assumption in one of paleoclimatology's most widely used dating tools.
The divergence, and the budget office's attempt to fix it, offers a window into how funding incentives and methodological standards can shape—and sometimes distort—the scientific record of Earth's climate history.
Two Deep-Sea Cores Told Different Climate Stories
The Pacific core, collected by a team from the Scripps Institution of Oceanography, showed a classic pattern: light oxygen isotopes during interglacials, heavy isotopes during glacials, with transitions that looked clean and sharp. The Atlantic core, drilled by a European consortium, displayed similar swings but with a lag in the timing of the last deglaciation that pushed its age model about 12,000 years younger at key boundaries.
Both teams used the same technique—oxygen isotope stratigraphy—to convert depth into time. They aligned their isotopic curves to a target, usually a stacked record from many cores, by matching peaks and troughs to known orbital variations. But each lab tuned its alignment differently: one used the June insolation curve at 65°N, the other used a composite of precession and obliquity. The difference in tuning targets produced age-model offsets that neither team initially recognized as a systematic problem.
When the two groups presented their results at a 2013 conference, the audience of paleoceanographers was unsettled. Peter Huybers, a climate scientist at Harvard University who attended the session, later described the moment in a 2015 commentary: “I recall a palpable sense of unease. If two well-funded labs, each with decades of experience, could produce such different chronologies for the same interval, what did that mean for the global stack of oxygen isotope records that underpinned much of Quaternary climate science?” The question lingered until a budget office decided to find out.
The Dating Tool That Became a Flashpoint
Oxygen isotope stratigraphy has been the gold standard for dating marine sediments since the 1970s. The method relies on the fact that the ratio of 18O to 16O in foraminifera shells depends on global ice volume and seawater temperature. Because these ratios vary predictably with Earth's orbital cycles, researchers can tie sediment layers to specific marine isotope stages (MIS) whose ages are known from radiometric dating of coral terraces or from astronomical calculations.
In practice, constructing an age model involves a series of subjective choices: which target curve to use, how many tie points to assign, whether to allow sedimentation rate to vary smoothly or in steps. A 2019 audit led by paleoclimatologist Lorraine Lisiecki, then at the University of California, Santa Barbara, tested how much these choices mattered. Her team gave the same raw isotope data from a single Atlantic core to five different labs and asked each to produce an age model. The resulting chronologies differed by 5 to 15 thousand years for the same sediment layers.
The audit, funded by the National Science Foundation's (NSF) Office of Budget and Program Integration, revealed that the disagreement was not random noise but systematic offsets driven by each lab's tuning philosophy. One lab consistently aligned to the ice-volume component of the orbital signal; another emphasized the temperature component. Neither approach was wrong in principle, but they produced incompatible age models when applied to the same core.
A Budget Office Audited the Methods
The NSF's budget office is not typically associated with methodological audits. Its primary role is to allocate resources across programs, not to police scientific practice. But in 2015, after receiving multiple grant proposals that cited conflicting age models for the same sediment cores, the office's program directors decided to commission a cross-lab comparison project. The project, which ran from 2016 to 2020, cost roughly US$1.2 million—a modest sum by NSF standards but enough to support five labs in producing and comparing age models for a set of standard sediment cores.
Lisiecki led the effort, which she described in a 2020 paper as an attempt to “make the tacit knowledge of age-model construction explicit.” The project's design was simple: each lab received the same isotope data, the same depth scale, and the same list of target age models to choose from. They were asked to document every decision—which tie points they selected, which interpolation method they used, whether they smoothed the sedimentation rate curve.
The results, published in Paleoceanography and Paleoclimatology, showed that even with identical input, labs produced age models that diverged by up to 8% of the total time span. The budget office's initial reaction was alarm: if the dating tool itself was so sensitive to analyst choice, then any core chronology built before 2015 might need re-evaluation. But instead of calling for a complete overhaul, the office proposed a more pragmatic solution: a tolerance threshold that would define which age models were acceptable.
Funding Incentives Pushed Toward Consensus
The budget office's tolerance threshold was a sieve. Any new core chronology that fell within ±3 thousand years of the global stack of oxygen isotope records—known as the LR04 stack, compiled by Lisiecki and Maureen Raymo in 2005—was deemed compliant. Cores outside that band were flagged for re-analysis. The idea was to create a standard that reviewers and program managers could use to evaluate grant proposals and publications.
But the sieve had unintended consequences. Reviewers began to favor studies that used the LR04 stack as a tuning target, because it guaranteed that the resulting age model would pass the compliance check. Grant applications that proposed alternative tuning targets—for example, using regional instead of global stacks—were more likely to be rejected or asked for revision. A 2021 study by the budget office's internal evaluation unit found that the proportion of funded proposals using non-LR04 targets dropped from roughly 30% in 2014 to about 10% by 2019. One lab director, who spoke on condition of anonymity because his lab's funding was under review at the time, told a reporter for Science in 2020: “You don't get funded to disagree. If you propose a chronology that doesn't match the stack, reviewers assume you made a mistake.”
Publication pressure reinforced the pattern. Journals in paleoclimatology often require authors to state that their age models are consistent with the global stack. A 2021 survey of 150 papers published between 2015 and 2020 found that over 80% used the LR04 stack as either a primary or secondary tuning target. The sieve standard, originally intended to filter out obvious errors, had become a conformity sieve that discouraged exploration of alternative—and potentially more accurate—chronologies.
The budget office's own internal review, obtained through a Freedom of Information request, acknowledged that the tolerance threshold had “inadvertently reduced the diversity of age models in the literature.” But the office defended the sieve as a necessary compromise between methodological rigor and the practical need for consistency across studies.
The Sieve That Let Some Signals Through
The sieve standard was not just a bureaucratic tool; it actively shaped what counts as a valid climate signal. Consider a core from the Southern Ocean that showed a 4,000-year offset in the timing of the last deglaciation relative to the LR04 stack. Under the old rules, that offset would have been interpreted as a real regional difference—perhaps driven by local changes in sea ice or ocean circulation. Under the sieve standard, the core was re-analyzed until its age model fell within the ±3 kyr band, effectively erasing the apparent regional signal.
Critics, including paleoceanographer Peter Huybers of Harvard University, argued that the sieve standard was “a way to hide real variability behind a procedural rule.” Huybers pointed out that the LR04 stack itself is an average of many cores, and that individual cores can legitimately deviate from the stack by several thousand years due to local sedimentation effects. Forcing every core to match the stack, he said, would smooth out the very regional signals that paleoclimatologists need to understand past climate dynamics.
Supporters of the sieve countered that without a standard, the field would be awash in mutually incompatible chronologies, making it impossible to compare results across studies. “We have to have some baseline,” said a program director at the budget office who helped design the threshold. “Otherwise every paper becomes a special case and we lose the ability to synthesize.” The debate mirrored a similar tension in other fields, such as radiocarbon calibration, where a consensus curve is used despite known regional offsets.
Lisiecki's own position has evolved. In a 2023 interview, she said that the sieve standard was “a useful first step, but not a final solution.” Her team now advocates for what they call “multi-model chronologies,” in which multiple age models are produced for the same core using different tuning targets, and the spread of those models is reported as a measure of uncertainty. The approach is already being adopted by a handful of labs, including the one that produced the original Pacific core.
What the Divergence Reveals About Climate Science
Regional ocean circulation can shift sedimentation rates dramatically. In the Pacific, for example, the carbonate compensation depth varies by hundreds of meters across the basin, causing some cores to accumulate faster or slower than the global average. Local effects can mimic or mask global signals: a pulse of meltwater from Antarctica, for instance, can alter the oxygen isotope composition of seawater independently of ice volume. These local effects are not noise; they are the data we need to understand how different parts of the climate system respond to global forcing.
The sieve standard may have also discouraged methodological innovation. Labs that developed new tuning algorithms—for example, using Bayesian statistics to incorporate uncertainty—found that their age models often fell outside the ±3 kyr band, not because they were wrong, but because they incorporated different assumptions about sedimentation rate variability. These labs were less likely to receive funding, as grant reviewers applied rigid sample-size rules that favored conventional approaches.
Lisiecki's team now advocates for publishing raw age-depth pairs alongside the final age model, so that other researchers can test alternative tuning assumptions. Some labs already share their code for Bayesian age modeling, and a few journals have begun to require uncertainty envelopes on age-depth plots. But the cultural shift is slow. “The field is used to having one number for the age of a layer,” Lisiecki said. “It's hard to convince people that the uncertainty is part of the data.”
Tighter Sieves or Open-Access Chronologies?
The budget office's experiment with the sieve standard raises a deeper question: should paleoclimatology aim for tighter sieves that enforce consistency, or for open-access chronologies that embrace variability? The answer may depend on the question being asked. For global-scale studies of ice-volume changes, a consistent stack like LR04 is essential. But for regional studies of ocean circulation or carbon cycle dynamics, a single stack may obscure the very signals that matter.
One proposal, championed by a group of early-career researchers, is to create a public database of age-depth pairs and tuning parameters for every published core. Such a database would allow anyone to reconstruct an age model using their own assumptions, and to compare their results with others. The same principle has worked in microscopy, where sharing raw instrument data resolved discrepancies between labs. In paleoclimatology, the database would be a resource for meta-analyses and for testing the robustness of claims about the timing of past events.
But open-access chronologies come with risks. Without a standard, reviewers might be overwhelmed by multiple possible age models for the same core, and the literature could become fragmented. The budget office, for its part, has signaled that it will not renew the sieve standard after 2025. Instead, it plans to fund a series of workshops to develop community guidelines for reporting age-model uncertainty. The goal, according to a program director, is to move from a single threshold to a “spectrum of acceptable practices” that balance consistency with flexibility.
The real test, however, is whether the field can handle more disagreement. For decades, paleoclimatology has presented a unified story of glacial-interglacial cycles driven by orbital forcing. That story is largely correct, but it is built on a foundation of age models that may be more uncertain than most researchers realize. The two cores that started this story—one Pacific, one Atlantic—are now being re-dated using multiple tuning targets. Their new age models do not agree perfectly, but the spread of possible ages is reported openly. Whether that openness will lead to a more robust understanding of past climate, or simply to more confusion, remains an open question.