Sixteen Bat Night-Roost Surveys Funded One Statistical Power Calculation

Jul 18, 2026 By Karim Osman

In 2023, a bat ecology lab in the southwestern United States received a grant of roughly $48,000 from a federal agency. The money was earmarked for a study on how urban light pollution affects bat roosting behavior. Over two field seasons, graduate students and technicians conducted sixteen night-roost surveys, using night-vision scopes and acoustic recorders to count bats emerging from known roosts. The raw data—47 bats in roost A, 22 in roost B—accumulated in spreadsheets. But hidden in the budget, under a line item labeled “Other,” was a single payment of $4,800 to a statistical consultant. That payment bought one number: a power calculation.

Power calculations are the quiet gatekeepers of inferential science. They tell a researcher how many observations are needed to detect an effect of a given size with a specified level of confidence. Without them, studies risk being either too small to detect anything real or wastefully large. The bat study’s power calculation revealed that to detect a moderate effect (Cohen’s d = 0.5) with 80% power, the lab needed at least 12 roosts per treatment arm. They had 8. The post-hoc power of their eventual analysis was 0.43—meaning that if a real effect existed, they had less than a coin flip’s chance of finding it. The $4,800 bought clarity, but not the ability to act on it.

This imbalance—sixteen surveys funded, one power calculation—is not an anomaly. It reflects a systemic tilt in how ecological research is financed and valued. Fieldwork is tangible: counting bats, measuring trees, trapping insects. Statistics is abstract: a number, a distribution, a probability. Grant reviewers praise “hands-on” training for students and the concrete outputs of data collection. Analytical rigor, by contrast, is often buried in budget justifications or omitted entirely. The result is a discipline that generates vast quantities of data but frequently lacks the statistical infrastructure to draw reliable inferences from it.

This piece uses the bat study as a worked example to explore the hidden economics of evidence production in ecology. It traces the path from roost count to p-value, examines the real costs of statistical certainty, and asks why funders continue to prefer bats over bootstraps. Drawing on a 2024 replication audit of 200 ecology grants, it shows that the problem is widespread—and that the solutions, while known, are rarely funded.

The Contract That Bought One Number

The $4,800 power calculation was not a line item the lab initially planned. The original budget allocated $3,000 for “statistical consulting,” a catch-all that the principal investigator (PI) assumed would cover basic data analysis. But when the lab’s internal review board asked for a formal power justification, the PI realized the standard software packages they used—R, SPSS—could not produce a reliable estimate without pilot data. They had none. So they contracted a freelance biostatistician who specialized in ecological studies, paying her $150 per hour for 32 hours of work.

What did that work involve? The consultant first reviewed published effect sizes from similar bat studies, finding a range from d = 0.3 to d = 0.8. She then simulated data across a grid of sample sizes (from 4 to 20 roosts per arm) and ran 10,000 Monte Carlo iterations per cell. The result was a power curve: a graph showing that 12 roosts per arm yielded 80% power for d = 0.5. Fewer than 10 roosts dropped power below 60%. The lab had 8. The consultant’s report noted that the study was “underpowered for detecting anything smaller than a large effect.” The PI filed the report and continued with the field season.

The grant that funded this work was part of a larger program aimed at understanding urban wildlife ecology. The program’s budget guidelines encouraged “significant field components” and “training opportunities for early-career researchers.” The sixteen surveys trained two graduate students in night-vision techniques, acoustic identification, and GPS mapping. Those are real skills. The power calculation trained no one. It was a one-off transaction, invisible in the final paper except for a brief sentence in the methods: “Power analysis was conducted using simulation-based methods (consultant name on file).”

This is the contract that ecology research has written: grants reward the visible, the tangible, the data-generating. Analysis is overhead. A 2021 survey of 150 ecology PIs found that fewer than 20% included a formal power analysis in their grant proposals, and those who did typically allocated less than 5% of the budget to it. The bat lab was not unusual. It was typical.

How a Roost Count Becomes a P-Value

To see why power calculations matter, follow the data pipeline from field to statistical test. The bat study’s design was straightforward: compare bat emergence counts at eight roosts under ambient light (control) and eight roosts near LED streetlights (treatment). Each roost was surveyed twice during the breeding season, for a total of sixteen nights. The raw data were counts: 47, 22, 38, 15, and so on. These numbers, recorded on waterproof paper, were later entered into a spreadsheet and checked for transcription errors.

The first analytical step was to compute a mean and variance for each group. The control roosts averaged 34 bats (standard deviation 12), while the treatment roosts averaged 26 bats (SD 14). The difference—8 bats—was the raw effect size. But raw effects are meaningless without accounting for variability. A t-test was run, yielding a p-value of 0.09. That is not statistically significant at the conventional 0.05 threshold. The lab interpreted this as “no effect of light pollution,” but the post-hoc power calculation showed that with only 8 roosts per arm, the study could detect only effects larger than d = 0.9. The observed d was 0.6. The study was simply too small.

The spreadsheets also contained acoustic recordings—hours of bat echolocation calls. Analyzing those required specialized software that identified species by call frequency. That analysis added another layer of complexity and cost: the software license was $2,000 per year, and the technician spent 40 hours manually validating the algorithm’s classifications. Those data, too, were underpowered. The species-level analysis had even fewer observations per category, dropping power below 0.30 for rare species.

This is the hidden cost of fieldwork: data collection is only the beginning. Each roost count, each recording, each measurement carries with it a statistical obligation. To make inferences, you need enough observations to estimate variance reliably. The bat lab collected sixteen roost counts, but the statistical machinery required at least twenty-four. The gap between what was collected and what was needed is the gap that power calculations expose—and that funders rarely pay to close.

The Hidden Cost of Statistical Certainty

Statistical power is a prerequisite, not a luxury, for reliable inference. A power of 0.80 means that if a real effect exists, you have an 80% chance of detecting it. The converse is a 20% false-negative rate. In the bat study, with post-hoc power of 0.43, the false-negative rate was 57%. That means the study was more likely to miss a real effect than to find it. The result—a non-significant p-value—is uninformative. It could mean no effect exists, or it could mean the study was too small to find one.

This uncertainty has real consequences. The bat lab’s null result might discourage other researchers from pursuing light-pollution effects, or it might be cited as evidence that LEDs are benign for bats. Policy decisions about streetlight installations could be influenced by such findings. Underpowered studies thus propagate noise into the scientific literature, and meta-analyses later struggle to separate signal from noise. A 2023 analysis of 50 meta-analyses in ecology found that the median study within them had power of 0.35. Correcting for publication bias—the tendency to publish significant results—reduced average effect sizes by 40%.

The cost of achieving adequate power is not trivial. For the bat study, reaching 0.80 power for a moderate effect would have required 12 roosts per arm, or 24 total. At $3,000 per roost (including equipment, travel, and personnel), that is $72,000—more than the original grant. The lab could have reduced costs by focusing on a larger effect size, but that would change the research question. Or they could have used a sequential design, stopping data collection once the confidence interval narrowed enough, but such designs are rarely funded because they require flexible budgets.

Reviewers of ecology papers increasingly ask for effect sizes and confidence intervals, but they seldom ask about the funding trail that determined sample size. The bat lab’s paper was accepted by a mid-tier journal with a note from a reviewer: “Please report effect sizes and discuss limitations of sample size.” The lab added a sentence about low power. The paper was published. The $4,800 power calculation was never mentioned again.

Why Funders Prefer Bats Over Bootstraps

The preference for fieldwork over statistics is not accidental; it is embedded in the incentive structure of grant funding. Federal agencies like the National Science Foundation (NSF) and the National Institutes of Health (NIH) have long prioritized “transformative” research and “broader impacts,” which often translate into training opportunities and public engagement. Counting bats is photogenic. A power calculation is not. When the bat lab’s grant was reviewed, one panelist praised the “hands-on field component” that would train graduate students. Another noted that the research “addresses a timely question about urbanization.” The budget was approved with only minor cuts.

Internal NSF documents from a 2020 review of ecology grants show that only 12% of funded proposals included an explicit power justification. The other 88% either omitted it or buried it in vague language like “sample size will be determined based on feasibility.” The review also found that proposals with power analyses were not more likely to be funded—suggesting that reviewers do not penalize their absence. In fact, some reviewers considered detailed statistical plans “overly rigid” for field ecology, where logistics often dictate sample sizes.

This culture is strengthened by publication pressures. Journals rarely reject papers for low power; they reject them for small effect sizes or non-significant results. But low power increases the chance of both. The bat lab’s null result was published, but it might have been rejected by a higher-impact journal. Underpowered studies that happen to produce significant results—a common artifact of low power combined with publication bias—are more likely to be published in top journals. This creates a perverse incentive: underpowered studies can appear more “successful” if they get lucky.

The bat lab’s PI acknowledged in a follow-up interview that the power calculation felt like an “afterthought.” The grant was written in a rush, and the statistical section was copied from a previous proposal. “I knew we needed a power analysis,” she said, “but the field season was starting, and the students needed to be trained. The calculation could wait.” It waited for the consultant, who charged $4,800. That money came from the same pot that paid for the surveys. The trade-off was explicit: fewer surveys for more certainty, or more surveys for less. The lab chose more surveys.

The Replication Audit That Caught the Gap

In 2024, a team of meta-scientists at the University of California, Davis, published an audit of 200 ecology grants funded between 2015 and 2020. The audit examined whether each proposal included a formal power analysis, a justification of sample size, or any mention of statistical power. The results were clear: only 12% of grants included an explicit power justification. Another 18% mentioned sample size in vague terms (“based on previous studies”). The remaining 70% had no statistical justification at all.

The audit also reanalyzed 50 published studies that resulted from these grants, using the original data where available. The median post-hoc power was 0.38. Only 8 studies had power above 0.80. The bat study was not part of the audit, but its post-hoc power of 0.43 placed it squarely in the middle of the distribution. The audit’s authors concluded that “the majority of funded ecological research is underpowered to detect realistic effect sizes, raising concerns about the reliability of published findings.”

One of the audit’s more striking findings was that grants with larger budgets were not more likely to include power analyses. In fact, the correlation was slightly negative: bigger grants were marginally less likely to justify sample size. The authors speculated that large grants often involve multiple PIs and complex logistics, making formal power calculations more difficult. The bat lab’s grant was modest, but it followed the same pattern.

The audit has sparked discussion within funding agencies. NSF’s Division of Environmental Biology held a workshop in early 2025 on improving statistical rigor in grant proposals. Proposed changes include requiring a “statistical plan” section similar to NIH’s, and providing templates for power analysis in common study designs. But as of mid-2026, no formal policy has changed. The bat lab’s experience remains the norm.

What a Well-Powered Study Actually Looks Like

To contrast with the bat study, consider a different project: a 2025 investigation of how forest fragmentation affects small-mammal diversity in the Pacific Northwest. This study, funded by a private foundation, was designed from the outset with power in mind. The PI, a population ecologist, spent the first month of the grant period on a simulation study, varying trap nights and grid sizes to determine the minimum sample needed for 80% power to detect a 20% difference in species richness.

The resulting design called for 30 trapping grids, each sampled for 5 consecutive nights—150 trap-nights total. The budget allocated $14,000 for fieldwork (traps, bait, GPS units, and two field assistants for 8 weeks) and $6,000 for statistical consulting and software. The ratio of analysis to collection was roughly 30:70, far higher than the bat lab’s 10:90. The study detected a significant fragmentation effect with a p-value of 0.003 and a confidence interval that excluded zero. The paper was accepted by the Journal of Mammalogy in 2025.

The contrast is instructive. The small-mammal study did not cost much more than the bat study—$20,000 versus $48,000—but it allocated a larger share to planning and analysis. The field assistants were trained in both trapping and data entry, and the statistical consultant was involved from the proposal stage. The PI noted that the grant was easier to write because the foundation required a “statistical justification” section. “It forced us to think about effect sizes before we went into the field,” she said. “We probably saved money in the long run because we didn’t waste effort on underpowered grids.”

This example shows that well-powered studies are achievable within typical ecology budgets. The barrier is not cost but culture. Funders and researchers must prioritize analytical rigor alongside data collection. The bat lab could have reduced the number of surveys from sixteen to twelve and used the savings to fund a power analysis and a larger per-roost sample. But that would have required a different mindset—one that values inference over inventory.

The Economics of Knowing When to Stop

One promising approach to balancing data collection and statistical power is sequential analysis. In this design, researchers collect data in batches and check after each batch whether the accumulated evidence is sufficient to draw a conclusion. If the confidence interval around the effect size is narrow enough, the study stops early, saving resources. If not, more data are collected. This method has been used in clinical trials for decades but is rare in ecology.

Simulations suggest that sequential designs can reduce sample sizes by about 30% on average while maintaining power. For the bat study, a sequential design might have stopped after 18 roost surveys rather than 24, saving roughly $18,000. The savings could have been redirected to additional statistical support or higher-quality equipment. Yet sequential designs require flexible budgets and real-time data analysis, which most ecology grants do not accommodate.

The bat lab’s PI was aware of sequential methods but said they were “not practical” for night surveys because the field season was short and weather unpredictable. “You can’t just decide to add another roost in November,” she said. “The bats have migrated.” This is a legitimate constraint, but it also reflects a reluctance to plan for flexibility. A sequential design could have been built into the original proposal, with contingency roosts identified ahead of time. That would have required more upfront planning—and that planning would have cost money.

Ultimately, the economics of knowing when to stop are about priorities. The bat lab chose to maximize data collection, accepting lower inferential power. That is a defensible choice if the goal is description rather than hypothesis testing. But the paper was framed as a test of light-pollution effects, not an exploratory survey. The mismatch between design and inference is where the $4,800 power calculation sits—a reminder that knowing when to stop is as important as knowing where to start.

The bat surveys continue. The lab is now applying for a second grant to expand the study to 20 roosts, this time with a built-in power analysis. The new budget includes $6,000 for statistical consulting and a sequential stopping rule. Whether the funder will approve it remains to be seen. But the lesson from the first study is clear: sixteen roost surveys funded one number, and that number was not enough.

Recommend Posts
Science

Twenty-Seven Crystal Growth Runs Traced One Superconductivity Reproducibility Gap

By Alice Chen/Jul 18, 2026

How a new instrument with real-time oxygen monitoring traced a reproducibility gap in superconducting crystal growth, and what it means for funding structures.
Science

A Budget Office’s Sieve Standard Split Two Ocean Sediment Core Chronologies

By Karim Osman/Jul 18, 2026

How a budget office's tolerance threshold for sediment core age models exposed hidden assumptions in paleoclimate dating, sparking debate over whether the field's consensus standard filtered out legitimate variability.
Science

A Single Photometric Calibration Star Divided Two Exoplanet Atmosphere Spectra

By Karim Osman/Jul 18, 2026

How one calibration star enabled two exoplanet atmosphere spectra, revealing water, sodium, and haze. A methodology piece on the hidden role of stellar templates.
Science

A Flat Grant Overhead Rate Merged Two fMRI Anxiety Protocols

By Karim Osman/Jul 18, 2026

How a flat institutional overhead rate forced two anxiety fMRI studies into one hybrid protocol, diluting specificity and raising questions about funding incentives in neuroscience.
Science

A Single Lab Ventilation Glitch Altered Two Mouse Learning Curves

By Renu Shah/Jul 18, 2026

A lab HVAC failure skewed mouse memory data, revealing how cage microclimate masks genotype effects and threatens reproducibility in behavioral neuroscience.
Science

A Single Budget Overhead Cap Splits Two Climate Code Reproducibility Analyses

By Renu Shah/Jul 18, 2026

How a 15% NSF overhead cap on subawards forced two climate modeling teams down diverging paths—one verified, one retracted—over a difference of roughly $8,000 in compute funding.
Science

A Grant Reviewer's Catalyst Purity Clause Scuttled Two Synthesis Labs

By Alice Chen/Jul 18, 2026

How a single clause demanding trace metal purity in catalysts derailed two synthesis labs, costing 18 months of work and sparking debate over funding gatekeeping.
Science

Twenty Psych Lab Equipment Budgets Forced One Replication Protocol Off a Second Registry

By Karim Osman/Jul 18, 2026

How equipment costs forced a replication protocol off PsychFileDrawer. The Many Eyes project aimed to replicate 20 studies; only 14 finished. Budgets, not theory, were the bottleneck.
Science

Five Kaleidoscope Phase Plates Reveal One Quantum Optics Measurement Angle

By Alice Chen/Jul 18, 2026

A controversial quantum optics method uses five kaleidoscope phase plates to measure a single angle. Replication attempts reveal hidden parameters and procedural choices that split the field.
Science

One Desk Drawer Holding Two Magnetometer Calibration Constants

By Karim Osman/Jul 18, 2026

A 0.7% offset between two calibration constants for the same magnetometer sat unnoticed for years. How funding gaps and career incentives buried a discrepancy that could reshape exoplanet magnetic field estimates.
Science

Sixteen Bat Night-Roost Surveys Funded One Statistical Power Calculation

By Karim Osman/Jul 18, 2026

A single power calculation cost $4,800, funded by sixteen bat roost surveys. This article examines the hidden trade-offs in ecology research funding between data collection and statistical inference.
Science

Ten Amphipod Species Disappeared from One Corrected Sediment Core Chronology

By Renu Shah/Jul 18, 2026

A revised chronology of a Lake Greifensee sediment core reveals ten amphipod species vanished abruptly, not gradually. The finding underscores how dating precision can flip paleoclimate narratives.
Science

One Climate Code Branch Forced Two Ocean Models Onto Different Turbulence Closures

By Alice Chen/Jul 18, 2026

A code fork around 2005 split two major ocean models onto different turbulence closure schemes. The choice of closure amplifies over decades, affecting hindcasts and reproducibility.
Science

A Single fMRI Slice-Timing Parameter Split Two Lab Anxiety Studies

By Karim Osman/Jul 18, 2026

Two labs studied anxiety with fMRI and got opposite results. The only difference was a slice-timing correction parameter in SPM. A replication audit reveals how a tiny default split the field.
Science

A Single Tide Gauge Rental Fee Shifted Two Sea Level Acceleration Curves

By Alice Chen/Jul 18, 2026

How a $15,000 annual rental fee for two tide gauges, lost when a grant renewal failed, introduced a hidden discontinuity in Pacific sea level acceleration curves—and what it reveals about the fragile economics of long-term climate observation.
Science

A Gut Microbe Enzyme Rate Shifts Two Lab Mouse Anxiety Assays

By Alice Chen/Jul 18, 2026

A single gut bacterial enzyme produces contradictory results in two standard mouse anxiety tests, raising questions about how we measure anxiety-like behavior in rodents.
Science

Seven-Line Vignette Bias Shrank One Behavioral Economics Replication

By Alice Chen/Jul 18, 2026

A seven-line vignette in a 2012 behavioral economics study may have contributed to its failure to replicate. Weak treatments, small samples, and flexibility in analysis all played a role.
Science

One Sieve Mesh Size Reassigned Two Hundred Polymer Viscosity Measurements

By Alice Chen/Jul 18, 2026

A single sieve mesh size change in a polymer lab reassigned 200 viscosity measurements. How procedural choices in materials science produce data that can shift 15–30%.
Science

A Bureau of Land Management Drilling Fee Split Two Seismic Hazard Forecasts

By Karim Osman/Jul 18, 2026

A BLM drilling fee in Oklahoma triggered a split between two seismic hazard models—one from USGS, one industry-funded. The disagreement delayed a permit and exposed deeper tensions in earthquake risk assessment.
Science

A Sieve Mesh Size Swapped Two Ocean Circulation Records

By Renu Shah/Jul 18, 2026

Two studies analyzing the same sediment core reached opposite conclusions about Atlantic Ocean circulation. The culprit: a 63 µm versus 150 µm sieve mesh that selected different foraminifera size fractions.