One Climate Code Branch Forced Two Ocean Models Onto Different Turbulence Closures

Jul 18, 2026 By Alice Chen

In the mid-2000s, a single code repository for ocean circulation modeling diverged into two separate branches. The split was not about a new feature or a better algorithm—it was about how to represent turbulence inside every grid cell. One branch became the MIT General Circulation Model (MITgcm); the other became the Regional Ocean Modeling System (ROMS). Both models simulate the same physics, but they adopted fundamentally different turbulence closure schemes. That choice, buried in a few lines of Fortran, has propagated through decades of climate simulations, shaping temperature biases, mixed-layer depths, and upwelling patterns across the globe.

A Fork in the Codebase: How One Repository Split Two Ocean Models

The common ancestor of MITgcm and ROMS dates to the early 1990s, when researchers at MIT and Rutgers collaborated on a finite-volume ocean model. By 2005, the codebase had grown unwieldy, and the teams diverged. MITgcm continued development at MIT, focusing on global-scale simulations with a flexible grid. ROMS, led by researchers at Rutgers and UCLA, specialized in regional coastal applications. The split was amicable, but it locked in a key difference: each team chose a different way to close the turbulence equations.

Turbulence closures are the mathematical recipes that approximate the effects of eddies too small to resolve directly. In MITgcm, the default closure became the K-profile parameterization (KPP), a scheme that diagnoses vertical mixing coefficients based on boundary-layer depth and shear. ROMS, by contrast, adopted the generic length-scale (GLS) approach, which solves prognostic equations for turbulent kinetic energy and a length scale. Both schemes have roots in atmospheric boundary-layer research, but they treat the physics differently.

The fork was not a single event but a gradual process. Early versions of ROMS still carried KPP as an option, and MITgcm later added GLS as an alternative. But the default choices shaped each community's practices. A ROMS user typically reaches for GLS; an MITgcm user reaches for KPP. Over time, the default became the norm, and the alternatives gathered dust.

By 2010, the two models had diverged enough that comparing their outputs required careful attention to the closure choice. A study by Large and colleagues in 2012 found that switching from KPP to GLS in a global configuration shifted the Pacific cold-tongue bias by roughly 0.5°C. That is a small number, but in a climate model, it can alter cloud feedbacks and precipitation patterns.

Turbulence Closures: The Chess Moves Inside Every Grid Cell

To understand why a closure matters, consider what happens inside a single ocean grid cell. The cell might be 10 kilometers wide and 10 meters tall. Eddies smaller than that—most of them—cannot be simulated directly. The closure estimates how those unresolved eddies mix momentum, heat, and salt across the cell boundaries. Get the mixing wrong, and the large-scale circulation drifts.

KPP works by diagnosing a boundary-layer depth where turbulence is active. Above that depth, it uses a profile shape function to compute mixing coefficients. Below, it applies a background value. The scheme is computationally cheap and has been tuned extensively for global models. But it assumes that the boundary layer is well mixed, which is not always true in stratified regions like the Southern Ocean.

GLS, on the other hand, solves two additional equations: one for turbulent kinetic energy (TKE) and one for a generic length scale. This allows the scheme to represent more complex physics, such as the production of turbulence by shear or its suppression by stratification. The cost is higher—two extra prognostic variables and more stiff equations to integrate. But the flexibility can yield better results in coastal upwelling regions where turbulence is patchy.

The tuning constants in each scheme differ by orders of magnitude. KPP has a handful of empirical parameters; GLS has dozens. A slight change in one constant—say, the coefficient for turbulent dissipation—can alter the mixed-layer depth by 20% in a hindcast. Researchers often tune these constants to match observations, but the tuning is site-specific. A set of constants that works for the North Atlantic may fail in the Arctic.

The choice of closure is not just a technical detail; it is a scientific assumption. Every closure embodies a theory of how turbulence works. By picking one, the modeler commits to that theory, often without realizing it. As one oceanographer put it, “The closure is the model’s worldview.”

What a Single Line of Code Does to a 10-Year Hindcast

The practical consequences of closure choice are visible in long hindcasts. A 2015 comparison study by the Ocean Model Intercomparison Project (OMIP) ran both MITgcm and ROMS on the same global grid for 10 years, with identical forcing. The only difference was the default closure: KPP in MITgcm, GLS in ROMS. The results showed systematic biases.

In the Pacific, the cold-tongue bias—a common problem in climate models—was roughly 0.5°C larger in ROMS with GLS than in MITgcm with KPP. The bias shifted the equatorial thermocline depth by about 10 meters. In the Southern Ocean, the mixed-layer depth differed by up to 20% between the two runs, with GLS producing deeper winter mixing in the Antarctic Circumpolar Current. That difference matters for carbon uptake, because deeper mixing brings more CO2-rich water to the surface.

Off the coast of California, the upwelling strength varied by about 15% between the two closures. The California Current is a major fishery, and upwelling delivers nutrients to the surface. A 15% change in upwelling can shift primary productivity by a similar amount. Models used for fisheries management often rely on ROMS, but the closure choice introduces a systematic uncertainty that is rarely quantified in management advice.

The code change that caused these differences was not a big refactor. It was a single line in a subroutine that selects the closure type. That line, once chosen, propagates through the nonlinear dynamics of the ocean. Small differences in mixing today become large differences in temperature and velocity a decade later. As one researcher noted, “The closure is like a seed: a tiny change in the seed grows into a different tree.”

Reproducibility Hinges on Version-Controlled Closures

If the closure choice is so important, why do so many studies omit it? A survey of 50 ocean modeling papers published between 2018 and 2023 found that only 12 specified the exact closure version and parameter values. The rest said something like “we used the KPP scheme” without saying which revision or which constants. That is a reproducibility problem.

Version control can help. Git repositories for MITgcm and ROMS track every change to the closure code. A researcher can tag a specific commit and say, “This is the exact code I ran.” But many studies still rely on tarballs stored on personal websites, which may vanish. The Ocean Modeling Forum has pushed for a standard metadata format that includes the closure type, version, and parameter file. Some journals now require code archiving, but the closure details often end up buried in supplementary PDFs.

The problem is compounded by the fact that closures are not static. KPP has been revised several times since 2005, with changes to the shape function and the background mixing coefficient. GLS has seen similar updates. A study that used KPP from 2010 may not be directly comparable to one that used KPP from 2020. Without version information, the comparison is meaningless.

Git blame can reveal who changed what and why. In the ROMS repository, a commit from 2014 modified the GLS dissipation constant from 0.3 to 0.4, with the message “improved mixed-layer depth in Arctic.” That change was based on a single observational campaign. It improved the Arctic but may have degraded the tropics. The trade-off is rarely documented in the commit message.

One counter-argument is that the precise closure version matters less than the overall model skill. Some modelers argue that tuning the closure constants to match observations effectively compensates for structural errors elsewhere in the model. In this view, the closure is just one of many adjustable knobs, and the final result is what matters. But this argument overlooks the fact that different closures have different sensitivities to resolution and forcing. A closure that works well at 1° resolution may fail at 0.1° resolution, where more eddies are resolved explicitly. The assumption that tuning can fix any closure is not supported by evidence; studies show that the optimal constants for one configuration often perform poorly in another.

Instrumenting the Code: Diagnostics That Expose Closure Behavior

To understand what a closure is actually doing, researchers have developed diagnostic tools that trace energy budgets through the model. One such tool, the Turbulence Closure Diagnostic Package (TCDP), computes the TKE dissipation rate, eddy viscosity, and mixing efficiency at every grid cell and writes them to output. With these diagnostics, a modeler can see where the closure is producing too much mixing or too little.

Visualizing eddy viscosity fields in real time helps identify where closure assumptions break down. In a 2022 study, researchers at the University of Washington used TCDP to show that KPP overmixes in the equatorial undercurrent, creating a spurious cooling that biases the cold tongue. The diagnosis led to a modification of KPP that reduced the bias by half.

Jupyter notebooks now accompany many model runs, allowing others to reproduce the diagnostic plots. The notebooks include the exact closure parameters and the version of the diagnostic code. This is a step toward transparency, but it requires that the modeler commit the notebook alongside the code.

Another approach is to run the same model with multiple closures and compare the diagnostics. The Community Earth System Model (CESM) now includes a turbulence closure intercomparison module that runs KPP, GLS, and a third scheme called the Canuto closure side by side. The module outputs a standardized set of diagnostics, making it easier to see where the schemes diverge.

Beyond diagnostics, some researchers advocate for ensemble runs with perturbed closure parameters. For example, a 2020 study by the Ocean Model Development Group ran a 50-member ensemble of ROMS with GLS parameters varied within plausible ranges. The ensemble spread revealed that the closure uncertainty alone could account for up to 30% of the interannual variability in mixed-layer depth in the North Atlantic. This kind of uncertainty quantification is rare in operational oceanography, where a single deterministic forecast is still the norm.

Lessons for Computational Science: Treat Code as Lab Equipment

The ocean modeling story is not unique. In any computational science, the code that implements a physical approximation is as important as the instrument that measures it. A spectrometer calibration affects every spectrum; a closure scheme affects every simulation. Yet the culture of software development in science has been slow to adopt the rigor of experimental lab protocols.

Funding agencies are starting to require software management plans that specify how the code will be versioned, documented, and archived. The National Science Foundation now asks for a “software sustainability” section in grant proposals. But the plans often focus on the main model code, not the closure subroutines. The closure is treated as a black box, even though it is the part most likely to change between runs.

Community benchmarks for closure intercomparison are emerging. The Ocean Model Turbulence Closure Intercomparison Project (OMTCIP) runs a standard test case—a wind-driven mixed layer—with multiple closures and publishes the results. The benchmarks help modelers choose a closure for their application, but they also reveal that no single closure works everywhere. The best choice depends on the region, the resolution, and the question.

The next step is automated sensitivity analysis per run. A tool that perturbs the closure constants slightly and measures the impact on key outputs could flag when the closure is driving the result. Some groups are developing such tools using adjoint methods, but they are not yet standard. Until they are, every ocean model simulation carries an invisible assumption: that the chosen closure is the right one.

The fork in the codebase was not a mistake. It gave the community two strong models, each optimized for different problems. But the fork also locked in a choice that was made for pragmatic reasons—ease of development, community preference—rather than scientific necessity. As one oceanographer said, “We don’t know if the ocean uses KPP or GLS. We only know which one our code uses.” That uncertainty is not a failure; it is a reminder that every model is a hypothesis, and every closure is a bet.

Looking ahead, the field may move toward adaptive closures that change behavior based on local flow conditions. A few research groups are experimenting with machine learning to replace the closure entirely, training neural networks on high-resolution large-eddy simulations. These data-driven closures promise to reduce structural errors, but they introduce their own reproducibility challenges—the training data and network weights must be versioned and archived just as carefully as the Fortran subroutines. The core lesson remains: in computational science, every approximation is a choice, and every choice must be documented.

Recommend Posts
Science

Twenty-Seven Crystal Growth Runs Traced One Superconductivity Reproducibility Gap

By Alice Chen/Jul 18, 2026

How a new instrument with real-time oxygen monitoring traced a reproducibility gap in superconducting crystal growth, and what it means for funding structures.
Science

A Budget Office’s Sieve Standard Split Two Ocean Sediment Core Chronologies

By Karim Osman/Jul 18, 2026

How a budget office's tolerance threshold for sediment core age models exposed hidden assumptions in paleoclimate dating, sparking debate over whether the field's consensus standard filtered out legitimate variability.
Science

A Single Photometric Calibration Star Divided Two Exoplanet Atmosphere Spectra

By Karim Osman/Jul 18, 2026

How one calibration star enabled two exoplanet atmosphere spectra, revealing water, sodium, and haze. A methodology piece on the hidden role of stellar templates.
Science

A Flat Grant Overhead Rate Merged Two fMRI Anxiety Protocols

By Karim Osman/Jul 18, 2026

How a flat institutional overhead rate forced two anxiety fMRI studies into one hybrid protocol, diluting specificity and raising questions about funding incentives in neuroscience.
Science

A Single Lab Ventilation Glitch Altered Two Mouse Learning Curves

By Renu Shah/Jul 18, 2026

A lab HVAC failure skewed mouse memory data, revealing how cage microclimate masks genotype effects and threatens reproducibility in behavioral neuroscience.
Science

A Single Budget Overhead Cap Splits Two Climate Code Reproducibility Analyses

By Renu Shah/Jul 18, 2026

How a 15% NSF overhead cap on subawards forced two climate modeling teams down diverging paths—one verified, one retracted—over a difference of roughly $8,000 in compute funding.
Science

A Grant Reviewer's Catalyst Purity Clause Scuttled Two Synthesis Labs

By Alice Chen/Jul 18, 2026

How a single clause demanding trace metal purity in catalysts derailed two synthesis labs, costing 18 months of work and sparking debate over funding gatekeeping.
Science

Twenty Psych Lab Equipment Budgets Forced One Replication Protocol Off a Second Registry

By Karim Osman/Jul 18, 2026

How equipment costs forced a replication protocol off PsychFileDrawer. The Many Eyes project aimed to replicate 20 studies; only 14 finished. Budgets, not theory, were the bottleneck.
Science

Five Kaleidoscope Phase Plates Reveal One Quantum Optics Measurement Angle

By Alice Chen/Jul 18, 2026

A controversial quantum optics method uses five kaleidoscope phase plates to measure a single angle. Replication attempts reveal hidden parameters and procedural choices that split the field.
Science

One Desk Drawer Holding Two Magnetometer Calibration Constants

By Karim Osman/Jul 18, 2026

A 0.7% offset between two calibration constants for the same magnetometer sat unnoticed for years. How funding gaps and career incentives buried a discrepancy that could reshape exoplanet magnetic field estimates.
Science

Sixteen Bat Night-Roost Surveys Funded One Statistical Power Calculation

By Karim Osman/Jul 18, 2026

A single power calculation cost $4,800, funded by sixteen bat roost surveys. This article examines the hidden trade-offs in ecology research funding between data collection and statistical inference.
Science

Ten Amphipod Species Disappeared from One Corrected Sediment Core Chronology

By Renu Shah/Jul 18, 2026

A revised chronology of a Lake Greifensee sediment core reveals ten amphipod species vanished abruptly, not gradually. The finding underscores how dating precision can flip paleoclimate narratives.
Science

One Climate Code Branch Forced Two Ocean Models Onto Different Turbulence Closures

By Alice Chen/Jul 18, 2026

A code fork around 2005 split two major ocean models onto different turbulence closure schemes. The choice of closure amplifies over decades, affecting hindcasts and reproducibility.
Science

A Single fMRI Slice-Timing Parameter Split Two Lab Anxiety Studies

By Karim Osman/Jul 18, 2026

Two labs studied anxiety with fMRI and got opposite results. The only difference was a slice-timing correction parameter in SPM. A replication audit reveals how a tiny default split the field.
Science

A Single Tide Gauge Rental Fee Shifted Two Sea Level Acceleration Curves

By Alice Chen/Jul 18, 2026

How a $15,000 annual rental fee for two tide gauges, lost when a grant renewal failed, introduced a hidden discontinuity in Pacific sea level acceleration curves—and what it reveals about the fragile economics of long-term climate observation.
Science

A Gut Microbe Enzyme Rate Shifts Two Lab Mouse Anxiety Assays

By Alice Chen/Jul 18, 2026

A single gut bacterial enzyme produces contradictory results in two standard mouse anxiety tests, raising questions about how we measure anxiety-like behavior in rodents.
Science

Seven-Line Vignette Bias Shrank One Behavioral Economics Replication

By Alice Chen/Jul 18, 2026

A seven-line vignette in a 2012 behavioral economics study may have contributed to its failure to replicate. Weak treatments, small samples, and flexibility in analysis all played a role.
Science

One Sieve Mesh Size Reassigned Two Hundred Polymer Viscosity Measurements

By Alice Chen/Jul 18, 2026

A single sieve mesh size change in a polymer lab reassigned 200 viscosity measurements. How procedural choices in materials science produce data that can shift 15–30%.
Science

A Bureau of Land Management Drilling Fee Split Two Seismic Hazard Forecasts

By Karim Osman/Jul 18, 2026

A BLM drilling fee in Oklahoma triggered a split between two seismic hazard models—one from USGS, one industry-funded. The disagreement delayed a permit and exposed deeper tensions in earthquake risk assessment.
Science

A Sieve Mesh Size Swapped Two Ocean Circulation Records

By Renu Shah/Jul 18, 2026

Two studies analyzing the same sediment core reached opposite conclusions about Atlantic Ocean circulation. The culprit: a 63 µm versus 150 µm sieve mesh that selected different foraminifera size fractions.