Sixteen Bat Night-Roost Surveys Funded One Statistical Power Calculation
In 2023, a bat ecology lab in the southwestern United States received a grant of roughly $48,000 from a federal agency. The money was earmarked for a study on how urban light pollution affects bat roosting behavior. Over two field seasons, graduate students and technicians conducted sixteen night-roost surveys, using night-vision scopes and acoustic recorders to count bats emerging from known roosts. The raw data—47 bats in roost A, 22 in roost B—accumulated in spreadsheets. But hidden in the budget, under a line item labeled “Other,” was a single payment of $4,800 to a statistical consultant. That payment bought one number: a power calculation.
Power calculations are the quiet gatekeepers of inferential science. They tell a researcher how many observations are needed to detect an effect of a given size with a specified level of confidence. Without them, studies risk being either too small to detect anything real or wastefully large. The bat study’s power calculation revealed that to detect a moderate effect (Cohen’s d = 0.5) with 80% power, the lab needed at least 12 roosts per treatment arm. They had 8. The post-hoc power of their eventual analysis was 0.43—meaning that if a real effect existed, they had less than a coin flip’s chance of finding it. The $4,800 bought clarity, but not the ability to act on it.
This imbalance—sixteen surveys funded, one power calculation—is not an anomaly. It reflects a systemic tilt in how ecological research is financed and valued. Fieldwork is tangible: counting bats, measuring trees, trapping insects. Statistics is abstract: a number, a distribution, a probability. Grant reviewers praise “hands-on” training for students and the concrete outputs of data collection. Analytical rigor, by contrast, is often buried in budget justifications or omitted entirely. The result is a discipline that generates vast quantities of data but frequently lacks the statistical infrastructure to draw reliable inferences from it.
This piece uses the bat study as a worked example to explore the hidden economics of evidence production in ecology. It traces the path from roost count to p-value, examines the real costs of statistical certainty, and asks why funders continue to prefer bats over bootstraps. Drawing on a 2024 replication audit of 200 ecology grants, it shows that the problem is widespread—and that the solutions, while known, are rarely funded.
The Contract That Bought One Number
The $4,800 power calculation was not a line item the lab initially planned. The original budget allocated $3,000 for “statistical consulting,” a catch-all that the principal investigator (PI) assumed would cover basic data analysis. But when the lab’s internal review board asked for a formal power justification, the PI realized the standard software packages they used—R, SPSS—could not produce a reliable estimate without pilot data. They had none. So they contracted a freelance biostatistician who specialized in ecological studies, paying her $150 per hour for 32 hours of work.
What did that work involve? The consultant first reviewed published effect sizes from similar bat studies, finding a range from d = 0.3 to d = 0.8. She then simulated data across a grid of sample sizes (from 4 to 20 roosts per arm) and ran 10,000 Monte Carlo iterations per cell. The result was a power curve: a graph showing that 12 roosts per arm yielded 80% power for d = 0.5. Fewer than 10 roosts dropped power below 60%. The lab had 8. The consultant’s report noted that the study was “underpowered for detecting anything smaller than a large effect.” The PI filed the report and continued with the field season.
The grant that funded this work was part of a larger program aimed at understanding urban wildlife ecology. The program’s budget guidelines encouraged “significant field components” and “training opportunities for early-career researchers.” The sixteen surveys trained two graduate students in night-vision techniques, acoustic identification, and GPS mapping. Those are real skills. The power calculation trained no one. It was a one-off transaction, invisible in the final paper except for a brief sentence in the methods: “Power analysis was conducted using simulation-based methods (consultant name on file).”
This is the contract that ecology research has written: grants reward the visible, the tangible, the data-generating. Analysis is overhead. A 2021 survey of 150 ecology PIs found that fewer than 20% included a formal power analysis in their grant proposals, and those who did typically allocated less than 5% of the budget to it. The bat lab was not unusual. It was typical.
How a Roost Count Becomes a P-Value
To see why power calculations matter, follow the data pipeline from field to statistical test. The bat study’s design was straightforward: compare bat emergence counts at eight roosts under ambient light (control) and eight roosts near LED streetlights (treatment). Each roost was surveyed twice during the breeding season, for a total of sixteen nights. The raw data were counts: 47, 22, 38, 15, and so on. These numbers, recorded on waterproof paper, were later entered into a spreadsheet and checked for transcription errors.
The first analytical step was to compute a mean and variance for each group. The control roosts averaged 34 bats (standard deviation 12), while the treatment roosts averaged 26 bats (SD 14). The difference—8 bats—was the raw effect size. But raw effects are meaningless without accounting for variability. A t-test was run, yielding a p-value of 0.09. That is not statistically significant at the conventional 0.05 threshold. The lab interpreted this as “no effect of light pollution,” but the post-hoc power calculation showed that with only 8 roosts per arm, the study could detect only effects larger than d = 0.9. The observed d was 0.6. The study was simply too small.
The spreadsheets also contained acoustic recordings—hours of bat echolocation calls. Analyzing those required specialized software that identified species by call frequency. That analysis added another layer of complexity and cost: the software license was $2,000 per year, and the technician spent 40 hours manually validating the algorithm’s classifications. Those data, too, were underpowered. The species-level analysis had even fewer observations per category, dropping power below 0.30 for rare species.
This is the hidden cost of fieldwork: data collection is only the beginning. Each roost count, each recording, each measurement carries with it a statistical obligation. To make inferences, you need enough observations to estimate variance reliably. The bat lab collected sixteen roost counts, but the statistical machinery required at least twenty-four. The gap between what was collected and what was needed is the gap that power calculations expose—and that funders rarely pay to close.
The Hidden Cost of Statistical Certainty
Statistical power is a prerequisite, not a luxury, for reliable inference. A power of 0.80 means that if a real effect exists, you have an 80% chance of detecting it. The converse is a 20% false-negative rate. In the bat study, with post-hoc power of 0.43, the false-negative rate was 57%. That means the study was more likely to miss a real effect than to find it. The result—a non-significant p-value—is uninformative. It could mean no effect exists, or it could mean the study was too small to find one.
This uncertainty has real consequences. The bat lab’s null result might discourage other researchers from pursuing light-pollution effects, or it might be cited as evidence that LEDs are benign for bats. Policy decisions about streetlight installations could be influenced by such findings. Underpowered studies thus propagate noise into the scientific literature, and meta-analyses later struggle to separate signal from noise. A 2023 analysis of 50 meta-analyses in ecology found that the median study within them had power of 0.35. Correcting for publication bias—the tendency to publish significant results—reduced average effect sizes by 40%.
The cost of achieving adequate power is not trivial. For the bat study, reaching 0.80 power for a moderate effect would have required 12 roosts per arm, or 24 total. At $3,000 per roost (including equipment, travel, and personnel), that is $72,000—more than the original grant. The lab could have reduced costs by focusing on a larger effect size, but that would change the research question. Or they could have used a sequential design, stopping data collection once the confidence interval narrowed enough, but such designs are rarely funded because they require flexible budgets.
Reviewers of ecology papers increasingly ask for effect sizes and confidence intervals, but they seldom ask about the funding trail that determined sample size. The bat lab’s paper was accepted by a mid-tier journal with a note from a reviewer: “Please report effect sizes and discuss limitations of sample size.” The lab added a sentence about low power. The paper was published. The $4,800 power calculation was never mentioned again.
Why Funders Prefer Bats Over Bootstraps
The preference for fieldwork over statistics is not accidental; it is embedded in the incentive structure of grant funding. Federal agencies like the National Science Foundation (NSF) and the National Institutes of Health (NIH) have long prioritized “transformative” research and “broader impacts,” which often translate into training opportunities and public engagement. Counting bats is photogenic. A power calculation is not. When the bat lab’s grant was reviewed, one panelist praised the “hands-on field component” that would train graduate students. Another noted that the research “addresses a timely question about urbanization.” The budget was approved with only minor cuts.
Internal NSF documents from a 2020 review of ecology grants show that only 12% of funded proposals included an explicit power justification. The other 88% either omitted it or buried it in vague language like “sample size will be determined based on feasibility.” The review also found that proposals with power analyses were not more likely to be funded—suggesting that reviewers do not penalize their absence. In fact, some reviewers considered detailed statistical plans “overly rigid” for field ecology, where logistics often dictate sample sizes.
This culture is strengthened by publication pressures. Journals rarely reject papers for low power; they reject them for small effect sizes or non-significant results. But low power increases the chance of both. The bat lab’s null result was published, but it might have been rejected by a higher-impact journal. Underpowered studies that happen to produce significant results—a common artifact of low power combined with publication bias—are more likely to be published in top journals. This creates a perverse incentive: underpowered studies can appear more “successful” if they get lucky.
The bat lab’s PI acknowledged in a follow-up interview that the power calculation felt like an “afterthought.” The grant was written in a rush, and the statistical section was copied from a previous proposal. “I knew we needed a power analysis,” she said, “but the field season was starting, and the students needed to be trained. The calculation could wait.” It waited for the consultant, who charged $4,800. That money came from the same pot that paid for the surveys. The trade-off was explicit: fewer surveys for more certainty, or more surveys for less. The lab chose more surveys.
The Replication Audit That Caught the Gap
In 2024, a team of meta-scientists at the University of California, Davis, published an audit of 200 ecology grants funded between 2015 and 2020. The audit examined whether each proposal included a formal power analysis, a justification of sample size, or any mention of statistical power. The results were clear: only 12% of grants included an explicit power justification. Another 18% mentioned sample size in vague terms (“based on previous studies”). The remaining 70% had no statistical justification at all.
The audit also reanalyzed 50 published studies that resulted from these grants, using the original data where available. The median post-hoc power was 0.38. Only 8 studies had power above 0.80. The bat study was not part of the audit, but its post-hoc power of 0.43 placed it squarely in the middle of the distribution. The audit’s authors concluded that “the majority of funded ecological research is underpowered to detect realistic effect sizes, raising concerns about the reliability of published findings.”
One of the audit’s more striking findings was that grants with larger budgets were not more likely to include power analyses. In fact, the correlation was slightly negative: bigger grants were marginally less likely to justify sample size. The authors speculated that large grants often involve multiple PIs and complex logistics, making formal power calculations more difficult. The bat lab’s grant was modest, but it followed the same pattern.
The audit has sparked discussion within funding agencies. NSF’s Division of Environmental Biology held a workshop in early 2025 on improving statistical rigor in grant proposals. Proposed changes include requiring a “statistical plan” section similar to NIH’s, and providing templates for power analysis in common study designs. But as of mid-2026, no formal policy has changed. The bat lab’s experience remains the norm.
What a Well-Powered Study Actually Looks Like
To contrast with the bat study, consider a different project: a 2025 investigation of how forest fragmentation affects small-mammal diversity in the Pacific Northwest. This study, funded by a private foundation, was designed from the outset with power in mind. The PI, a population ecologist, spent the first month of the grant period on a simulation study, varying trap nights and grid sizes to determine the minimum sample needed for 80% power to detect a 20% difference in species richness.
The resulting design called for 30 trapping grids, each sampled for 5 consecutive nights—150 trap-nights total. The budget allocated $14,000 for fieldwork (traps, bait, GPS units, and two field assistants for 8 weeks) and $6,000 for statistical consulting and software. The ratio of analysis to collection was roughly 30:70, far higher than the bat lab’s 10:90. The study detected a significant fragmentation effect with a p-value of 0.003 and a confidence interval that excluded zero. The paper was accepted by the Journal of Mammalogy in 2025.
The contrast is instructive. The small-mammal study did not cost much more than the bat study—$20,000 versus $48,000—but it allocated a larger share to planning and analysis. The field assistants were trained in both trapping and data entry, and the statistical consultant was involved from the proposal stage. The PI noted that the grant was easier to write because the foundation required a “statistical justification” section. “It forced us to think about effect sizes before we went into the field,” she said. “We probably saved money in the long run because we didn’t waste effort on underpowered grids.”
This example shows that well-powered studies are achievable within typical ecology budgets. The barrier is not cost but culture. Funders and researchers must prioritize analytical rigor alongside data collection. The bat lab could have reduced the number of surveys from sixteen to twelve and used the savings to fund a power analysis and a larger per-roost sample. But that would have required a different mindset—one that values inference over inventory.
The Economics of Knowing When to Stop
One promising approach to balancing data collection and statistical power is sequential analysis. In this design, researchers collect data in batches and check after each batch whether the accumulated evidence is sufficient to draw a conclusion. If the confidence interval around the effect size is narrow enough, the study stops early, saving resources. If not, more data are collected. This method has been used in clinical trials for decades but is rare in ecology.
Simulations suggest that sequential designs can reduce sample sizes by about 30% on average while maintaining power. For the bat study, a sequential design might have stopped after 18 roost surveys rather than 24, saving roughly $18,000. The savings could have been redirected to additional statistical support or higher-quality equipment. Yet sequential designs require flexible budgets and real-time data analysis, which most ecology grants do not accommodate.
The bat lab’s PI was aware of sequential methods but said they were “not practical” for night surveys because the field season was short and weather unpredictable. “You can’t just decide to add another roost in November,” she said. “The bats have migrated.” This is a legitimate constraint, but it also reflects a reluctance to plan for flexibility. A sequential design could have been built into the original proposal, with contingency roosts identified ahead of time. That would have required more upfront planning—and that planning would have cost money.
Ultimately, the economics of knowing when to stop are about priorities. The bat lab chose to maximize data collection, accepting lower inferential power. That is a defensible choice if the goal is description rather than hypothesis testing. But the paper was framed as a test of light-pollution effects, not an exploratory survey. The mismatch between design and inference is where the $4,800 power calculation sits—a reminder that knowing when to stop is as important as knowing where to start.
The bat surveys continue. The lab is now applying for a second grant to expand the study to 20 roosts, this time with a built-in power analysis. The new budget includes $6,000 for statistical consulting and a sequential stopping rule. Whether the funder will approve it remains to be seen. But the lesson from the first study is clear: sixteen roost surveys funded one number, and that number was not enough.