Seven-Line Vignette Bias Shrank One Behavioral Economics Replication

Jul 18, 2026 By Alice Chen

In 2012, a behavioral economics paper reported that a subtle reminder of social norms could boost cooperation. Participants who read a seven-line vignette about appropriate behavior subsequently gave more in a public goods game. The effect was modest — Cohen's d around 0.30 — but it fit a growing narrative that small primes could nudge prosocial behavior. A decade later, a multi-lab replication project found essentially no effect: d = 0.02, with only 2 of 20 labs detecting a positive result. The original claim now looks like a false positive, and the vignette itself may be partly to blame.

A Replication That Wasn't

The 2012 study, published in a respected journal, involved 148 participants divided into two groups. One group read a short paragraph describing a social norm of cooperation; the other read a neutral text. After reading, participants played a public goods game where they could contribute money to a shared pool. Those who had read the norm vignette contributed about 15% more on average, a difference that was statistically significant at p < .05.

The finding was cited dozens of times and helped inspire interventions in schools and workplaces. But around 2015, as the replication crisis gained attention, several labs tried to reproduce the result. Informal attempts failed, and a formal multi-lab replication was organized under the Many Labs framework. That project, published in 2019, found no significant effect across 20 independent samples totaling roughly 1,500 participants.

The tension between the original and the replication raises a question: why did the effect vanish? One possibility is that the original was a false positive, a statistical fluke. Another is that the replication did not capture some crucial feature of the original. But the evidence points more strongly to the first explanation, and a key suspect is the seven-line vignette itself.

The Seven-Line Vignette as a Weak Treatment

The vignette in the original study was brief: roughly 70 words describing a community where people cooperate. Participants read it once and then immediately played the game. There was no manipulation check — no question to verify that the vignette actually changed participants' thoughts or feelings. Without such a check, it is impossible to know whether the treatment had any psychological impact.

In later replications, researchers added attention checks and found that 18–25% of participants failed them. That failure rate suggests that many participants may not have read the vignette carefully or remembered it. If the treatment was too weak to register, any observed effect would be noisy and unreliable.

Weak treatments are a known problem in priming research. Many classic priming effects, such as those involving word puzzles or sentence scrambles, have failed to replicate when tested rigorously. The seven-line vignette was even briefer than those stimuli. It may have been a 'dose' too low to produce a consistent behavioral change, especially across different populations and settings.

One counterargument is that the original study used a paper questionnaire in a lab, while many replications used online platforms. Online participants are less attentive, on average. But even lab-based replications within the multi-lab project failed to find the effect, suggesting the weakness is not solely an online issue.

Sample Size and Publication Bias

The original study had 74 participants per condition. That sample size is typical for behavioral experiments of the era, but it is small for detecting a subtle effect. With 74 per group, the study had only about 50% power to detect a d of 0.30 at p < .05. That means if the true effect were 0.30, the study would fail to find it half the time. But when it did find it, the estimated effect would tend to be inflated — a phenomenon known as the winner's curse.

The replication project used a total sample of roughly 1,500, giving it over 95% power to detect a d of 0.30. The observed d of 0.02 is consistent with a true effect of zero. A Bayesian analysis of the combined data strongly favored the null hypothesis, with a Bayes factor of about 12.

Publication bias likely played a role. In 2012, pre-registration was rare. Researchers could analyze data in multiple ways — excluding outliers, controlling for covariates, choosing different dependent variables — and report only the analysis that yielded significance. The original study did not pre-register its analysis plan, so the degrees of freedom were high. The replication, by contrast, was a registered report with a fixed plan, reducing flexibility.

Some defenders of the original argue that the replication may have differed in subtle ways, such as the exact wording of the vignette or the instructions for the game. But the replication team coordinated with the original authors to match the materials as closely as possible. Small variations are inevitable across labs, but if the effect were robust, it should have appeared in at least a few of the 20 samples.

Context Sensitivity of Behavioral Primes

Behavioral primes are notoriously context-dependent. A subtle cue that works in one culture, language, or setting may fail in another. The original study was conducted in a Western university lab with paper questionnaires. The replications spanned 20 labs across multiple countries, some using online platforms. The shift from paper to screen may have altered how participants engaged with the vignette.

Even within the same lab, timing matters. The original data were collected in 2010–2011, before social norms around cooperation had shifted. The replications took place between 2015 and 2018, after several high-profile failures to replicate priming effects had been published. Participants may have become more skeptical or less susceptible to such primes.

One meta-analysis of similar norm-priming studies found that effect sizes declined over time, a pattern consistent with both genuine context sensitivity and publication bias. The decline was steepest for studies with small samples and weak treatments, exactly the profile of the 2012 vignette.

Some researchers argue that context sensitivity is a feature, not a bug, of behavioral science. Primes work only when they are novel and the participant is unaware. Once a prime becomes widely known, its effect may vanish. But that explanation is hard to test, and it risks making the claim unfalsifiable.

What the Replication Actually Found

The multi-lab replication, published in 2019 as a registered report, found an overall effect of d = 0.02 with a 95% confidence interval ranging from −0.05 to 0.09. That interval includes zero and excludes the original estimate of 0.30. Only 2 of the 20 individual labs found a statistically significant positive effect, and those had small samples. The remaining 18 found null or negative results.

A Bayesian meta-analysis of the replication data yielded a Bayes factor of 12.3 in favor of the null hypothesis, meaning the data were about 12 times more likely under a model with no effect than under a model with an effect of the size originally reported. Sensitivity analyses that varied the prior did not change this conclusion.

The replication also tested whether the effect might be moderated by demographic variables such as gender, age, or nationality. None of these moderators were significant. The null result was consistent across the board.

Some critics noted that the replication used a slightly different public goods game — a linear one rather than a step-level one — but the original authors had agreed that the linear version was a valid test. Moreover, a separate direct replication using the exact original game also failed to find an effect.

Lessons for Designing Replicable Studies

The case of the seven-line vignette offers several lessons for researchers. First, ensure that the treatment has a measurable impact. A manipulation check — a simple question about whether the vignette changed participants' thoughts — would have revealed whether the prime worked as intended. Without it, the original study could not distinguish between a real effect and noise.

Second, conduct a power analysis before data collection. The original study's sample size was adequate only for detecting medium-to-large effects. For subtle primes, larger samples are needed. The replication's sample was 20 times larger, giving it the precision to detect even small effects if they existed.

Third, pre-register the analysis plan. Pre-registration reduces the risk of p-hacking and selective reporting. The original study had no pre-registration, while the replication did. That difference alone may account for some of the discrepancy.

Fourth, report effect sizes with confidence intervals, not just p-values. The original reported a p-value just below .05 but did not provide a confidence interval for the effect size. A confidence interval would have shown the imprecision of the estimate.

Finally, consider the strength of the treatment. A seven-line vignette is a weak intervention. Researchers should pilot-test their stimuli to ensure they produce the intended psychological change. Weak treatments amplify noise and reduce replicability.

The Slow Path to a Settled Verdict

Seven years elapsed between the original publication and the meta-analytic replication. That is a long time for a single finding to remain unresolved. During those years, the field of behavioral economics underwent a methodological reckoning. The replication crisis spurred reforms such as pre-registration, registered reports, and larger sample sizes. The seven-line vignette study was an early target of that reckoning.

Today, the consensus among meta-analysts is that the original claim is likely a false positive. The evidence from 20 labs, combined with Bayesian analysis, is strong. But some researchers remain cautious, noting that absence of evidence is not evidence of absence. The true effect, if it exists, is certainly smaller than originally reported — possibly too small to be practically meaningful.

The case is a small but instructive example of how cumulative evidence works. No single study settles a question. The original was not fraudulent; it was simply underpowered and overinterpreted. The replication corrected the record, but it took years and the coordinated effort of many labs. The field is now better equipped to avoid such false positives, but the slow pace of self-correction remains a challenge.

As one researcher put it, 'Science is a process, not a product.' The seven-line vignette story is a reminder that even a well-intentioned finding can be misleading, and that the path to a settled verdict is rarely straight.

Broader Implications for Behavioral Science

The failure of the seven-line vignette to replicate is not an isolated event. It echoes similar patterns in other domains of behavioral science, such as social priming and ego depletion. For instance, the classic 'elderly walking slower' priming effect, in which participants exposed to words related to old age subsequently walked more slowly, has also failed to replicate in large-scale efforts. Similarly, the idea that self-control is a depletable resource — the ego depletion effect — has been called into question by a multi-lab replication that found a near-zero effect. These cases share common features: small original samples, weak manipulations, and a publication environment that favored surprising results.

One lesson from these failures is that the field's reliance on small, underpowered studies has produced a literature with many false positives. The seven-line vignette is a textbook example. Its effect size of d = 0.30 was plausible but imprecise, and the original study lacked the statistical power to reliably detect it. The replication, with its larger sample and pre-registered protocol, provided a more trustworthy estimate. The discrepancy between the two is a reminder that single studies, especially small ones, should be interpreted with caution.

Another lesson is the importance of direct replication. Many researchers argue that conceptual replications — studies that test the same idea using different methods — are more valuable than direct replications. But conceptual replications introduce new variables that can obscure failures to replicate. The seven-line vignette case shows that direct replication, when done carefully, can provide clear evidence against a false positive. The multi-lab approach, with its coordinated materials and analysis plan, is a model for how to conduct such replications.

Some critics of the replication movement argue that too much emphasis on replication stifles innovation. They worry that researchers will avoid risky, creative studies for fear of failing to replicate. But the seven-line vignette case suggests the opposite: by identifying false positives, replication efforts can clear the way for more robust findings. The original study's claim was not just wrong; it was distracting. Researchers who built on it wasted time and resources. A more careful approach, with pre-registration and larger samples, would have produced a more reliable result from the start.

Practical Steps for Researchers

For researchers designing similar studies, the lessons are clear. First, pilot-test your manipulation. A simple check — asking participants what they thought about after reading the vignette — can reveal whether the prime had the intended effect. Without such a check, you cannot be sure that your treatment is working.

Second, plan your sample size based on a power analysis. For small-to-medium effects, you need hundreds of participants, not dozens. Online platforms like Mechanical Turk or Prolific make large samples feasible and affordable. There is no excuse for underpowered studies in the modern era.

Third, pre-register your analysis plan. This does not mean you cannot explore your data; it just means you separate confirmatory analyses from exploratory ones. Pre-registration increases transparency and reduces the risk of p-hacking.

Fourth, report effect sizes with confidence intervals. A p-value tells you whether an effect is statistically significant, but it does not tell you how large or precise the effect is. Confidence intervals provide that information and help readers assess the reliability of the finding.

Finally, consider the strength of your treatment. If your manipulation is a brief text or a subtle cue, ask yourself whether it is strong enough to change behavior. If not, you may need a more powerful intervention, or you may need to accept that your study is better suited for detecting large effects.

Conclusion

The seven-line vignette study is a cautionary tale about the dangers of weak treatments, small samples, and publication bias. Its failure to replicate does not mean the original authors were dishonest; it means the original evidence was weaker than it appeared. The replication effort, involving 20 labs and a pre-registered protocol, provided a more reliable estimate and corrected the scientific record.

The case also illustrates the value of methodological reforms. Pre-registration, registered reports, and large-scale collaborations are not bureaucratic hurdles; they are tools for improving the credibility of science. The field of behavioral economics has embraced many of these reforms, and the result is a more robust literature. But the process is ongoing, and new challenges will arise.

As the researcher said, science is a process, not a product. The seven-line vignette story is a reminder that even a well-intentioned finding can be misleading, and that the path to a settled verdict is rarely straight. But with careful methodology and a commitment to self-correction, the path becomes clearer.

Recommend Posts
Science

Twenty-Seven Crystal Growth Runs Traced One Superconductivity Reproducibility Gap

By Alice Chen/Jul 18, 2026

How a new instrument with real-time oxygen monitoring traced a reproducibility gap in superconducting crystal growth, and what it means for funding structures.
Science

A Budget Office’s Sieve Standard Split Two Ocean Sediment Core Chronologies

By Karim Osman/Jul 18, 2026

How a budget office's tolerance threshold for sediment core age models exposed hidden assumptions in paleoclimate dating, sparking debate over whether the field's consensus standard filtered out legitimate variability.
Science

A Single Photometric Calibration Star Divided Two Exoplanet Atmosphere Spectra

By Karim Osman/Jul 18, 2026

How one calibration star enabled two exoplanet atmosphere spectra, revealing water, sodium, and haze. A methodology piece on the hidden role of stellar templates.
Science

A Flat Grant Overhead Rate Merged Two fMRI Anxiety Protocols

By Karim Osman/Jul 18, 2026

How a flat institutional overhead rate forced two anxiety fMRI studies into one hybrid protocol, diluting specificity and raising questions about funding incentives in neuroscience.
Science

A Single Lab Ventilation Glitch Altered Two Mouse Learning Curves

By Renu Shah/Jul 18, 2026

A lab HVAC failure skewed mouse memory data, revealing how cage microclimate masks genotype effects and threatens reproducibility in behavioral neuroscience.
Science

A Single Budget Overhead Cap Splits Two Climate Code Reproducibility Analyses

By Renu Shah/Jul 18, 2026

How a 15% NSF overhead cap on subawards forced two climate modeling teams down diverging paths—one verified, one retracted—over a difference of roughly $8,000 in compute funding.
Science

A Grant Reviewer's Catalyst Purity Clause Scuttled Two Synthesis Labs

By Alice Chen/Jul 18, 2026

How a single clause demanding trace metal purity in catalysts derailed two synthesis labs, costing 18 months of work and sparking debate over funding gatekeeping.
Science

Twenty Psych Lab Equipment Budgets Forced One Replication Protocol Off a Second Registry

By Karim Osman/Jul 18, 2026

How equipment costs forced a replication protocol off PsychFileDrawer. The Many Eyes project aimed to replicate 20 studies; only 14 finished. Budgets, not theory, were the bottleneck.
Science

Five Kaleidoscope Phase Plates Reveal One Quantum Optics Measurement Angle

By Alice Chen/Jul 18, 2026

A controversial quantum optics method uses five kaleidoscope phase plates to measure a single angle. Replication attempts reveal hidden parameters and procedural choices that split the field.
Science

One Desk Drawer Holding Two Magnetometer Calibration Constants

By Karim Osman/Jul 18, 2026

A 0.7% offset between two calibration constants for the same magnetometer sat unnoticed for years. How funding gaps and career incentives buried a discrepancy that could reshape exoplanet magnetic field estimates.
Science

Sixteen Bat Night-Roost Surveys Funded One Statistical Power Calculation

By Karim Osman/Jul 18, 2026

A single power calculation cost $4,800, funded by sixteen bat roost surveys. This article examines the hidden trade-offs in ecology research funding between data collection and statistical inference.
Science

Ten Amphipod Species Disappeared from One Corrected Sediment Core Chronology

By Renu Shah/Jul 18, 2026

A revised chronology of a Lake Greifensee sediment core reveals ten amphipod species vanished abruptly, not gradually. The finding underscores how dating precision can flip paleoclimate narratives.
Science

One Climate Code Branch Forced Two Ocean Models Onto Different Turbulence Closures

By Alice Chen/Jul 18, 2026

A code fork around 2005 split two major ocean models onto different turbulence closure schemes. The choice of closure amplifies over decades, affecting hindcasts and reproducibility.
Science

A Single fMRI Slice-Timing Parameter Split Two Lab Anxiety Studies

By Karim Osman/Jul 18, 2026

Two labs studied anxiety with fMRI and got opposite results. The only difference was a slice-timing correction parameter in SPM. A replication audit reveals how a tiny default split the field.
Science

A Single Tide Gauge Rental Fee Shifted Two Sea Level Acceleration Curves

By Alice Chen/Jul 18, 2026

How a $15,000 annual rental fee for two tide gauges, lost when a grant renewal failed, introduced a hidden discontinuity in Pacific sea level acceleration curves—and what it reveals about the fragile economics of long-term climate observation.
Science

A Gut Microbe Enzyme Rate Shifts Two Lab Mouse Anxiety Assays

By Alice Chen/Jul 18, 2026

A single gut bacterial enzyme produces contradictory results in two standard mouse anxiety tests, raising questions about how we measure anxiety-like behavior in rodents.
Science

Seven-Line Vignette Bias Shrank One Behavioral Economics Replication

By Alice Chen/Jul 18, 2026

A seven-line vignette in a 2012 behavioral economics study may have contributed to its failure to replicate. Weak treatments, small samples, and flexibility in analysis all played a role.
Science

One Sieve Mesh Size Reassigned Two Hundred Polymer Viscosity Measurements

By Alice Chen/Jul 18, 2026

A single sieve mesh size change in a polymer lab reassigned 200 viscosity measurements. How procedural choices in materials science produce data that can shift 15–30%.
Science

A Bureau of Land Management Drilling Fee Split Two Seismic Hazard Forecasts

By Karim Osman/Jul 18, 2026

A BLM drilling fee in Oklahoma triggered a split between two seismic hazard models—one from USGS, one industry-funded. The disagreement delayed a permit and exposed deeper tensions in earthquake risk assessment.
Science

A Sieve Mesh Size Swapped Two Ocean Circulation Records

By Renu Shah/Jul 18, 2026

Two studies analyzing the same sediment core reached opposite conclusions about Atlantic Ocean circulation. The culprit: a 63 µm versus 150 µm sieve mesh that selected different foraminifera size fractions.