A Single Photometric Calibration Star Divided Two Exoplanet Atmosphere Spectra

Jul 18, 2026 By Karim Osman

In the mid-2000s, two teams of astronomers used HD 189733A, a K-dwarf, to extract the first clear atmospheric signatures of two hot Jupiters. One spectrum revealed water vapor in the emission of HD 189733b. Another uncovered sodium absorption and haze in the transmission of HD 209458b. The shared calibration star was not a coincidence: its photometric stability made it a yardstick for both observations. But the two spectra tell different stories about how stellar reference data shape what we think we know about exoplanet atmospheres.

One Star’s Light, Two Planets’ Secrets

A calibration star is a known reference whose brightness and spectrum are well characterized. For exoplanet atmosphere studies, the star’s light must be stable or its variability must be modeled to better than 0.1%—the typical depth of a transit or eclipse signal. HD 189733A, a K1.5V dwarf roughly 63 light-years away, met that requirement. It was bright enough (V magnitude 7.7) for high signal-to-noise spectroscopy with the Hubble Space Telescope’s STIS instrument and the Spitzer Space Telescope’s IRAC camera.

Each spectrum required hundreds of transits or secondary eclipses stacked together. For HD 189733b, Tinetti et al. (2007) used Spitzer observations at 3.6, 5.8, 8.0, and 24 μm. The calibration star’s flux model was fitted to a precision of roughly 0.1%, but systematic errors from telescope jitter and detector nonlinearities dominated the noise budget. For HD 209458b, Charbonneau et al. (2002) used Hubble STIS to measure the transit depth at a spectral resolution of about 500, achieving a precision of 0.02% in depth—enough to detect sodium absorption at 589.3 nm.

The tension between calibration precision and signal noise is a recurring theme. The same stellar template can yield different planetary results depending on how stellar variability is removed. In the HD 189733b case, the team used a simultaneous reference star (HD 189733B, a fainter M-dwarf companion) to correct for systematic drifts. For HD 209458b, the target star itself served as its own reference, comparing in-transit and out-of-transit spectra. Both strategies rely on the assumption that the star’s intrinsic spectrum is constant over the observation window—an assumption that is not always valid.

How a Calibration Star Becomes a Yardstick

HD 189733A’s role as a calibration star began long before exoplanet atmosphere studies. It was included in photometric catalogs as a stable, non-variable star. But no star is perfectly constant. HD 189733A shows rotational modulation at the level of roughly 0.3% due to starspots, with a period of about 12 days. For eclipse spectroscopy, where the signal is a few tenths of a percent, this variability must be modeled and removed.

The method set by Charbonneau et al. (2002) involved fitting a stellar flux model to the out-of-eclipse data and then subtracting that model from the in-eclipse data. The model included limb darkening, which depends on the star’s effective temperature, surface gravity, and metallicity. For HD 189733A, those parameters are well known: Teff ≈ 5050 K, log g ≈ 4.5, [Fe/H] ≈ −0.03. But small uncertainties in these values propagate into the planetary spectrum.

Systematic errors from the instrument—fringing in the detector, pointing jitter, and sensitivity drifts—are often larger than the photon noise. The Spitzer IRAC camera, for example, had a well-known “ramp” effect: the detector response changed over the first few minutes of each observation. Correcting this ramp required careful modeling, and different teams used different correction algorithms. The choice of correction could shift the inferred planetary brightness temperature by tens of kelvin.

The calibration star’s spectrum is not just a reference; it is a template for the stellar contribution that must be removed to isolate the planetary signal. Any mismatch between the model and the actual star creates a residual that can mimic or mask spectral features. This is why the community now invests in building synthetic stellar spectra from first principles, rather than relying solely on empirical templates.

Spectrum A: HD 189733b’s Water Absorption

In 2007, Giovanna Tinetti and colleagues published the first direct detection of water vapor in the atmosphere of an exoplanet. Using Spitzer’s Infrared Array Camera, they measured the secondary eclipse (when the planet passes behind the star) at four infrared wavelengths. The emission spectrum showed a dip at 3.6 and 5.8 μm, consistent with water absorption. The signal-to-noise ratio per spectral bin was around 5—modest but statistically significant.

The calibration star HD 189733A was used to remove the stellar flux. The team observed the system for roughly 33 hours across multiple Spitzer visits. Each visit produced a light curve that was fitted with a model including the star’s flux, the planet’s phase variation, and the instrument ramp. The stellar flux was assumed constant, but the ramp correction introduced a systematic uncertainty of about 0.05% in the eclipse depth.

The result was a day-side temperature of roughly 1200 K, with water vapor as the dominant opacity source. Later observations with Hubble and ground-based telescopes confirmed the water feature and added carbon monoxide and carbon dioxide. But the early Spitzer spectrum remains a landmark: it showed that exoplanet atmospheres could be studied in detail, not just detected.

However, the interpretation depended on the stellar model. If the star’s own water lines were not properly removed, they could produce a false signal. The team used a PHOENIX stellar atmosphere model for HD 189733A, which includes molecular opacities. The fit was good, but not perfect. Subsequent reanalyses using updated stellar models shifted the water abundance by factors of two to three.

Spectrum B: HD 209458b’s Sodium and Haze

Five years earlier, David Charbonneau and colleagues had produced the first exoplanet atmosphere spectrum using the Hubble Space Telescope. They targeted HD 209458b, a hot Jupiter transiting a G0V star similar to the Sun. By comparing the star’s spectrum during transit (when the planet blocks part of the stellar disk) to the out-of-transit spectrum, they detected sodium absorption at 589.3 nm with a depth of 0.023%—a signal roughly 4 times the noise.

The calibration star in this case was HD 209458 itself. The team used the star as its own reference, observing it for four transits over two years. The precision of 0.02% in transit depth required careful correction for Earth’s atmospheric absorption and telescope pointing drifts. The sodium feature was seen in both the 2001 and 2002 data, ruling out a spurious detection.

The spectrum also showed increased absorption at shorter wavelengths (below 550 nm), attributed to scattering by haze or clouds. This haze component was later confirmed by Hubble’s STIS and ACS instruments. The combination of sodium and haze provided the first constraints on the temperature-pressure profile and cloud properties of an exoplanet atmosphere.

Both spectra—HD 189733b’s water and HD 209458b’s sodium—used different calibration stars, but the method was the same: a bright, well-characterized star served as the reference. The two cases illustrate how the choice of calibration strategy (simultaneous reference vs. same-star) affects the error budget. For HD 209458b, the star’s own variability (roughly 0.1% due to p-mode oscillations) had to be averaged over many transits to reach 0.02% precision.

The Same Star, Different Calibration Strategies

HD 189733A was used as a simultaneous reference star for the HD 189733b observations. The binary companion HD 189733B, a faint M-dwarf about 216 arcseconds away, was observed simultaneously in the same field of view. This allowed the team to correct for common-mode systematic errors—changes in atmospheric transmission, telescope pointing, and detector sensitivity that affect both stars equally. The companion’s light curve was used to remove the ramp effect and other instrumental drifts.

For HD 209458b, there was no suitable reference star in the Hubble STIS field. Instead, the team used the target star itself, comparing in-transit and out-of-transit spectra. This method assumes that the star’s intrinsic flux is constant over the transit duration (about 3 hours). But stars are not perfectly constant; solar-like oscillations and granulation can introduce variations at the level of 0.01–0.02% over hours. Averaging multiple transits reduces this noise, but it cannot eliminate it entirely.

Both methods have distinct error budgets. The simultaneous reference star method cancels common-mode systematics but requires a stable companion with known properties. The same-star method avoids the need for a second star but is vulnerable to stellar variability that is not common mode. In practice, the two methods often give consistent results, but the choice can affect the inferred abundance of trace gases by a factor of two.

A related challenge is the removal of stellar spectral lines. For transmission spectroscopy, the planet’s atmosphere imprints absorption features on the stellar spectrum as it transits. But the star’s own lines must be subtracted to isolate the planetary signal. If the stellar model is wrong, the planetary spectrum will show spurious features. This is especially problematic for stars with strong chromospheric activity, like HD 189733A, which has a variable Ca II H&K emission.

What a Shared Reference Teaches About Instrumentation

The choice of calibration star limits the achievable spectral resolution. For HD 189733b, the Spitzer IRAC bands are broad (roughly 0.5–1 μm wide), so the spectrum has only four data points. For HD 209458b, the Hubble STIS spectrum had a resolution of about 500, but only over a narrow wavelength range (580–640 nm). To get higher resolution, astronomers need brighter stars or longer integration times, which are often impractical.

The repeatability of stellar models is crucial for comparing results across instruments. The same planetary spectrum observed with different telescopes should give the same answer, but often it does not. For example, the water abundance in HD 189733b derived from Spitzer data differs from that derived from Hubble WFC3 data by about a factor of three, a discrepancy attributed to different stellar models and calibration stars (Spitzer used HD 189733A, while Hubble used a different reference for its wavelength calibration). A 2014 reanalysis by Line et al. found that using updated stellar parameters reduced the discrepancy to a factor of two, but it did not disappear entirely.

Future telescopes like the James Webb Space Telescope (JWST) and the Ariel mission will require even better stellar templates. JWST’s NIRSpec instrument can obtain spectra at R ≈ 1000, but the calibration stars must be stable to 0.01% to avoid introducing systematic errors. The community is building an archive of photometric standard stars with precisely characterized variability, including HD 189733A, but the list is still short.

Ground-based high-resolution spectrographs (e.g., ESPRESSO, HARPS) face a different challenge: telluric absorption from Earth’s atmosphere. Calibration stars are used to remove telluric lines, but the star’s own lines must be known to high precision. The same calibration star can be used for multiple targets, but only if its spectrum is stable over time. For HD 189733A, the star’s rotation and activity cycle introduce subtle changes that must be tracked.

Lessons from these early studies have influenced the design of future instruments. For example, the JWST NIRSpec instrument includes a “reference star” mode that observes a nearby star simultaneously with the target, similar to the HD 189733A/B approach. The Ariel mission plans to observe a set of calibration stars before and after each target to characterize the instrument’s stability. These design choices are direct responses to the challenges revealed by the HD 189733A and HD 209458b spectra.

Takeaways for Atmospheric Retrieval

The two planets—HD 189733b and HD 209458b—have very different atmospheric compositions. HD 189733b shows strong water absorption and a day-side temperature of ~1200 K, with hints of carbon monoxide. HD 209458b shows sodium and haze, with a temperature inversion in the upper atmosphere. Yet both spectra were obtained using similar calibration methods. The differences in chemistry are real, but they are also amplified by the different calibration strategies.

Retrieval codes that invert the observed spectrum to infer atmospheric properties must marginalize over stellar parameters. If the stellar temperature is uncertain by 50 K, the retrieved planetary temperature can shift by 100 K. Similarly, if the stellar metallicity is off by 0.1 dex, the inferred water abundance can change by a factor of two. Modern retrieval codes, such as TauREx and NEMESIS, include stellar parameters as free parameters with priors based on independent measurements. For instance, in a retrieval of HD 189733b's atmosphere using TauREx, a 100 K uncertainty in the stellar effective temperature led to a 200 K spread in the retrieved planetary temperature, and a 0.2 dex uncertainty in stellar metallicity caused a factor of 3 range in water abundance.

The community now uses synthetic stellar spectra (e.g., from the PHOENIX or BT-Settl grids) rather than empirical templates, because they can be more easily interpolated to the exact parameters of the calibration star. But synthetic spectra have their own uncertainties: they may miss molecular lines or treat convection poorly. For HD 189733A, the synthetic spectra reproduce the observed flux to within 1%, but the residuals at the 0.1% level are still large enough to affect exoplanet spectra.

The next step is to characterize calibration stars directly via asteroseismology. By measuring the star’s oscillations, astronomers can determine its mass, radius, and age to high precision, which in turn constrains its spectrum. For HD 189733A, asteroseismic observations with the TESS satellite have provided a radius accurate to 2% and a mass accurate to 3%. This reduces the uncertainty in the stellar model and, consequently, in the planetary spectrum.

In the end, the shared calibration star taught the field that exoplanet atmosphere spectroscopy is as much about the star as it is about the planet. The same star can yield different results depending on how its light is used. The tension between calibration precision and signal noise is not a bug—it is a feature of the method. Recognizing this is the first step toward more robust atmospheric retrievals.

Recommend Posts
Science

Twenty-Seven Crystal Growth Runs Traced One Superconductivity Reproducibility Gap

By Alice Chen/Jul 18, 2026

How a new instrument with real-time oxygen monitoring traced a reproducibility gap in superconducting crystal growth, and what it means for funding structures.
Science

A Budget Office’s Sieve Standard Split Two Ocean Sediment Core Chronologies

By Karim Osman/Jul 18, 2026

How a budget office's tolerance threshold for sediment core age models exposed hidden assumptions in paleoclimate dating, sparking debate over whether the field's consensus standard filtered out legitimate variability.
Science

A Single Photometric Calibration Star Divided Two Exoplanet Atmosphere Spectra

By Karim Osman/Jul 18, 2026

How one calibration star enabled two exoplanet atmosphere spectra, revealing water, sodium, and haze. A methodology piece on the hidden role of stellar templates.
Science

A Flat Grant Overhead Rate Merged Two fMRI Anxiety Protocols

By Karim Osman/Jul 18, 2026

How a flat institutional overhead rate forced two anxiety fMRI studies into one hybrid protocol, diluting specificity and raising questions about funding incentives in neuroscience.
Science

A Single Lab Ventilation Glitch Altered Two Mouse Learning Curves

By Renu Shah/Jul 18, 2026

A lab HVAC failure skewed mouse memory data, revealing how cage microclimate masks genotype effects and threatens reproducibility in behavioral neuroscience.
Science

A Single Budget Overhead Cap Splits Two Climate Code Reproducibility Analyses

By Renu Shah/Jul 18, 2026

How a 15% NSF overhead cap on subawards forced two climate modeling teams down diverging paths—one verified, one retracted—over a difference of roughly $8,000 in compute funding.
Science

A Grant Reviewer's Catalyst Purity Clause Scuttled Two Synthesis Labs

By Alice Chen/Jul 18, 2026

How a single clause demanding trace metal purity in catalysts derailed two synthesis labs, costing 18 months of work and sparking debate over funding gatekeeping.
Science

Twenty Psych Lab Equipment Budgets Forced One Replication Protocol Off a Second Registry

By Karim Osman/Jul 18, 2026

How equipment costs forced a replication protocol off PsychFileDrawer. The Many Eyes project aimed to replicate 20 studies; only 14 finished. Budgets, not theory, were the bottleneck.
Science

Five Kaleidoscope Phase Plates Reveal One Quantum Optics Measurement Angle

By Alice Chen/Jul 18, 2026

A controversial quantum optics method uses five kaleidoscope phase plates to measure a single angle. Replication attempts reveal hidden parameters and procedural choices that split the field.
Science

One Desk Drawer Holding Two Magnetometer Calibration Constants

By Karim Osman/Jul 18, 2026

A 0.7% offset between two calibration constants for the same magnetometer sat unnoticed for years. How funding gaps and career incentives buried a discrepancy that could reshape exoplanet magnetic field estimates.
Science

Sixteen Bat Night-Roost Surveys Funded One Statistical Power Calculation

By Karim Osman/Jul 18, 2026

A single power calculation cost $4,800, funded by sixteen bat roost surveys. This article examines the hidden trade-offs in ecology research funding between data collection and statistical inference.
Science

Ten Amphipod Species Disappeared from One Corrected Sediment Core Chronology

By Renu Shah/Jul 18, 2026

A revised chronology of a Lake Greifensee sediment core reveals ten amphipod species vanished abruptly, not gradually. The finding underscores how dating precision can flip paleoclimate narratives.
Science

One Climate Code Branch Forced Two Ocean Models Onto Different Turbulence Closures

By Alice Chen/Jul 18, 2026

A code fork around 2005 split two major ocean models onto different turbulence closure schemes. The choice of closure amplifies over decades, affecting hindcasts and reproducibility.
Science

A Single fMRI Slice-Timing Parameter Split Two Lab Anxiety Studies

By Karim Osman/Jul 18, 2026

Two labs studied anxiety with fMRI and got opposite results. The only difference was a slice-timing correction parameter in SPM. A replication audit reveals how a tiny default split the field.
Science

A Single Tide Gauge Rental Fee Shifted Two Sea Level Acceleration Curves

By Alice Chen/Jul 18, 2026

How a $15,000 annual rental fee for two tide gauges, lost when a grant renewal failed, introduced a hidden discontinuity in Pacific sea level acceleration curves—and what it reveals about the fragile economics of long-term climate observation.
Science

A Gut Microbe Enzyme Rate Shifts Two Lab Mouse Anxiety Assays

By Alice Chen/Jul 18, 2026

A single gut bacterial enzyme produces contradictory results in two standard mouse anxiety tests, raising questions about how we measure anxiety-like behavior in rodents.
Science

Seven-Line Vignette Bias Shrank One Behavioral Economics Replication

By Alice Chen/Jul 18, 2026

A seven-line vignette in a 2012 behavioral economics study may have contributed to its failure to replicate. Weak treatments, small samples, and flexibility in analysis all played a role.
Science

One Sieve Mesh Size Reassigned Two Hundred Polymer Viscosity Measurements

By Alice Chen/Jul 18, 2026

A single sieve mesh size change in a polymer lab reassigned 200 viscosity measurements. How procedural choices in materials science produce data that can shift 15–30%.
Science

A Bureau of Land Management Drilling Fee Split Two Seismic Hazard Forecasts

By Karim Osman/Jul 18, 2026

A BLM drilling fee in Oklahoma triggered a split between two seismic hazard models—one from USGS, one industry-funded. The disagreement delayed a permit and exposed deeper tensions in earthquake risk assessment.
Science

A Sieve Mesh Size Swapped Two Ocean Circulation Records

By Renu Shah/Jul 18, 2026

Two studies analyzing the same sediment core reached opposite conclusions about Atlantic Ocean circulation. The culprit: a 63 µm versus 150 µm sieve mesh that selected different foraminifera size fractions.