Pre-Registration as a Severe Testing Device

[Excerpts are taken from the preprint, “The Value of Preregistration for Psychological Science: A Conceptual Analysis” by Daniël Lakens, posted at PsyArXiv Preprints]
What is Preregistration For?
“If the only goal of a researcher is to prevent bias, it suffices to verbally agree upon the planned analysis with collaborators as long as everyone will perfectly remember the agreed upon analysis. In the conceptual analysis presented here, researchers preregister to allow future readers of the preregistration (which might include the researchers themselves) to evaluate whether the research question was tested in a way that could have falsified the prediction.”
“Mayo (1996) carefully develops arguments for the role that prediction plays in science and arrives at an error statistical philosophy based on a severity requirement.”
Severe Tests
“A test is severe when it is highly capable of demonstrating a claim is false.”
“Figure 1A visualizes a null hypothesis test, where only one specific state of the world (namely an effect of exactly zero) will falsify our prediction. All other possible states of the world are in line with our prediction.”
“Figure 1B represents a one-sided null-hypothesis test, where differences larger than zero are predicted, and the prediction is falsified when the difference is either equal to zero, or smaller than zero. This prediction is slightly riskier than a two-sided test, in that there are more ways in which our prediction could be wrong, because 50% of all possible outcomes falsify the prediction, and 50% corroborate it.”
“Finally, Figure 1C visualized a range prediction where only differences between 0.5 and 2.5 support the prediction. Since there are many more ways this prediction could be wrong, it is an even more severe test.”
“If we observe a difference of 1.5, with a 95% confidence interval from 1 to 2, all three predictions are confirmed with an alpha level of 0.05, but the prediction in Figure 1C has passed the most severe test since it was confirmed in a test that had a higher capac-ity of demonstrating the prediction is false. Note that the three tests differ in severity even when they are tested with the same Type 1 error rate.”Capture
“As far as I am aware, Mayo’s severity argument currently provides one of the few philosophies of science that allows for a coherent conceptual analysis of the value of preregistration.”
Examples of Practices that Reduce the Severity of Tests
“One example of such a practice is optional stopping, where researchers collect data, analyze their data, and continue the data collection only if the result is not statistically significant. In theory, a researcher who is willing to continue collecting data indefinitely will always observe a statistically significant result. By repeatedly looking at the data, the Type 1 error rate can inflate to 100%. In this extreme case the prediction can no longer be falsified, and the test has no severity.”
“The severity of a test can also be compromised by selecting a hypothesis based on the observed results. In this practice, known as Hypothesizing After the Results are Known (HARKing, Kerr, 1998) researchers look at their data, and then select a prediction. This reversal of the typical hypothesis testing procedure makes the test incapable of demonstrating the claim was false.”
“As a final example…think about the scenario [where a] researcher makes multitudes of observations and selects out of all these tests only those that support their prediction. Choosing to selectively report tests from among many tests that were performed strongly reduces the capability of a test to demonstrate the claim was false.”
“…a preregistration document should give us all the information that allows future readers to evaluate the severity of the test. This includes the theoretical and empirical basis for predictions, the experimental design, the materials, and the analysis code. Having access to this information should allow readers to see whether any choices were made during the research process that reduced the severity of a test.”
“Researchers should also specify when they will conclude their prediction is not supported. As De Groot (1969) writes: ‘The author of a theory should himself state…what potential outcomes would, if actually found, lead him to regard his theory as disproven.’”
Preregistration Makes it Possible to Evaluate the Severity of a Test
“The severity of a test could in theory be unrelated to whether it is preregistered. However, in practice there will almost always be a correlation between the ability to transparently evaluate the severity of a test and preregistration, both because researchers can often selectively report results, use optional stopping, or come up with a plausible hypothesis after the results are known, and because theories rarely completely constrain the test of predictions.”
“As this conceptual analysis of preregistration makes clear, the practice of specifying the design, data collection, and planned analyses in advance is based on a philosophy of science that values tests of predictions and puts more trust in claims that have passed severe tests (Lakatos, 1978; Mayo, 2018; Meehl, 1990; Platt, 1964; Popper, 1959).”
To read the article, click here.

REED: EiR* – What’s Supporting that Fixed Effects Estimate?

[* EiR = Econometrics in Replications, a feature of TRN that highlights useful econometrics procedures for re-analysing existing research.]
NOTE: All the data and code  necessary to produce the results in the tables below are available at Harvard’s Dataverse: click here.
Fixed effects estimators are often used when researchers are concerned about omitted variable bias due to unobserved, time-invariant variables. These can prove insightful if there is much within-variation to support the fixed effects estimate. However, they can be misleading when there is not.
Stata has several commands that can help the researcher gauge the extent of within-variation. In this example, we use the “wagepan” dataset that is bundled with Jeffrey Wooldridge’s text, “Introductory Econometrics: A Modern Approach, 6e”. The dataset consists of annual observations of 545 workers over the years 1980-1987. It is described here.
In this example we use fixed effects to regress log(wage) on education, labor market experience, labor market experience squared, dummy variables for marital and union status, and annual time dummies.
The table below reports the fixed effects (within-estimate) for the “married” variable. For the sake of comparision, it also reports the between-estimate for “married” (calculated used the Mundlak version of the Random Effects Within Between estimator (Bell, Fairbrother, and Jones, 2019).
TRN1(20200425)
The within-estimate of the marriage premium is smaller than the between-estimate. This is consistent with marital status being positively associated with unobserved, time-invariant productivity characteristics of the worker. However, we want to know how much variation there is in marital status for the workers in our sample. If it is just a few workers who are changing marital status over time, then our estimate may not be representative of the effect of marriage in the population.
Stata provides two commands that can be helpful in this regard. The command xttab reports, among other things, a measure of variable stability across time periods. In the table below, among workers who ever reported being unmarried, they were unmarried for an average of 64.8% of the years in the sample.
Among workers who ever reported being married, they were married for 62.5% of the years in the sample. In this case, changes in marital status are somewhat common. Note that a time-invariant variable would have a “Within Percent” value of 100%.
TRN2(20200425)
Stata provides another command, xttrans, that gives detail about year-to-year variable transitions.
TRN3(20200425)
The rows represent the values in year t, with the columns representing the values in the following year. In this case, 86% of observations that were unmarried at time t, were also unmarried at time t+1. 14% of observations that were unmarried at time t changed status to “married” at time t+1.
Among other things, the xttrans command provides a reminder that the fixed effects estimate of the marriage premium includes the effect of transitioning from married to unmarried: 5% of observations that were married at time t were unmarried at time t+1. The implied assumption is that the effect of marriage on wages is symmetric, something that could be further explored in the data.
While these analyses are useful, they are based on observations, not workers. If there is concern about sample selection biasing the fixed effects estimates (so that “movers” are different from “stayers”), it would be useful to know how many of the 545 workers had experienced a marital status change, since it is the changes that support the fixed effects estimate.
The following set of commands calculate the min and the max values of the explanatory variables for all the workers in the sample. It then creates a dummy variable with the prefix “change” that takes the value 1 anytime the max and min values differ. Finally, it collapses the dataset so that there is one observation per worker, and then takes averages of the change variables.
TRN4(20200425)
The results below indicate that 56.9% of the workers changed their marital status during the sample period. Whether this is a sufficient number of “changers” to represent population “changers” is an open question. However, if the number were only 5 or 10% of workers, the argument for representativeness would be much weaker.
TRN5(20200425)
What does this have to do with replication? Oftentimes treatments are administered over time in panel datasets (say microcredit loans). Fixed effects estimates may be used to identify causal estimates of the treatment. Sample statistics, when they are reported, typically only report the percent of observations receiving treatment. Consider the two samples below.

TRN6(20200425)

In both samples, 30% of the observations are treatment observations. Thus a table of sample statistics would show identical means for the treatment variable in the two samples.
However, in the first sample, 100% of the workers received treatment, and 75% of year-to-year transitions involved a change in treatment status. In the second sample, only 50% of the workers experienced treatment, and 25% of year-to-year transitions involved a change in treatment status.
These are the kinds of differences that the procedures described above can be used to identify.
Bob Reed is a professor of economics at the University of Canterbury in New Zealand. He is also co-organizer of the blogsite The Replication Network. He can be contacted at bob.reed@canterbury.ac.nz. 

NEWS FLASH: Economics Journal Seeking Replication Studies to Publish. Now.

The International Journal for Re-views of Empirical Economics (IREE) is the only economics journal soley dedicated to replication research. They are looking to publish quality replication studies. If you have students conducting replications in your applied economics/econometrics classes, this is a great opportunity for them to convert their classroom research into a journal publication. The journal is online, open access, with quick review and publication times. The editor-in-chief is Joachim Wagner, and the advisory board includes Richard Easterlin and Jeffrey Wooldridge.
To learn more about the journal and how to submit your research, click here.

Looking Back on Metascience 2019

[Excerpts taken from the article, “Metascience: The Science of Doing Science” by Jonathan Schooler, published in Observer magazine, a publication of the Association of Psychological Science. The article appeared in November 2019. TRN apologizes for the belated posting.]
“In September, a symposium on metascience (metascience2019.org), funded by the Fetzer Franklin Fund and held at Stanford University, brought together nearly 500 attendees to help consolidate the field. The symposium included over 50 speakers from a remarkable variety of scientific disciplines, including psychology, philosophy, biology, sociology, network science, economics informatics, quantitative methodology, history, statistics, political science, medicine, business, and chemical and biological engineering.”
“I organized the event with APS Fellows Brian Nosek (University of Virginia) and Jon Krosnick at Stanford, psychological scientist Leif D. Nelson of University of California, Berkeley, and Fetzer Franklin Fund director Jan Walleczek.”
“The meeting addressed pressing questions surrounding the issue of scientific reproducibility including: “What is replication and its impact and value?” and “How are statistics, methods, and measurement practices affecting our capacity to identify robust findings?” However, it broadened the discussion to address a host of other aspects of the scientific process, such as “How do scientists generate ideas?” ‘How do scientists interpret and treat evidence?” and “What are the cultures and norms of science?’”
“The Stanford metascience meeting demonstrated the fundamentally interdisciplinary nature of the field…the meeting illustrated how scientists across domains, united by shared interests, can converse about the common elements underpinning the scientific process.”
“One criticism of the metascience meeting involved its subtitle: “the emerging field of research on the scientific process.” Some viewed this characterization as overlooking the many lines of work on this general topic that have been carried out for decades…”
“Whereas specialized scientists such as Ioannidis have been discussing problems with scientific reproducibility for some time, the mainstream research community has only recently thas taken note of this challenge only recently.”
“Furthermore, while independent lines of work have been carried out across disciplines, the consolidation of these areas into an overarching field has been limited. Thus, although it might be misleading to characterize the field of metascience as “emerging,” it certainly is consolidating and gaining momentum as never before.”
“The increasing role of metascience in science holds both great promise and some risk…On the one hand, exposure to a mirror is known to enhance conscientiousness, and indeed it seems likely that the emergence of metascientific concerns may be encouraging scientists to be more disciplined in the way they conduct their research.”
“However, mirrors can also make people self-conscious, and it seems plausible that scrutiny of the scientific process could (at least sometimes) stifle scientific creativity and risk-taking…when metascience is used as a platform for making attacks on the credibility of researchers whose work has failed to be replicated, both science in general and metascience in particular are bound to suffer indignities.”
“For better or worse, the metascience genie is out of the bottle. The zeitgeist is shifting. As metascience takes on an increasingly central role in science, it remains to be seen what discoveries it will make and what impact it will have.”
To read the article, click here.

GAARDER & JIMENEZ: Replication Research on Financial Inclusion in Developing Countries

Capture
In support of recent efforts by social scientists to address the ‘reproducibility crisis’, the Journal of Development Effectiveness (JDEff) recently devoted a special issue on replication research studies in its last issue of 2019.  Most journals continue to favor new research rather than replication work and to publish research whose data and codes have not been tested. As editors of the journal, and as the past and current executive director of the International Initiative for Impact Evaluation (3ie) which hosts it, we felt that such an issue could increase awareness among funders and researchers of how replication strengthens the reliability, rigor and relevance of their investment.  It would also ensure that the replication research will be acknowledged and appreciated by the larger development community.
The special issue was devoted to the topic of enhancing financial inclusion in developing countries.  3ie, which has championed replication research for many years, had worked closely with the Bill and Melinda Gates Foundation’s Financial Services for the Poor (FSP) program to identify the studies, screen the applicants, and quality assure the replication research in this important area.  FSP invests millions of dollars to broaden the reach of low-cost digital financial services for the poor by supporting the most catalytic approaches to financial inclusion, such as the development of digital payment systems, advancing gender equality, and supporting national and regional strategies.  In doing so, it relies heavily on research and evidence studies, many of which, although cited and referenced heavily, have not been replicated.
We hope that these replications can be used appropriately by FSP and other stakeholders to inform future investments in an important part of the development toolkit – expanding financial services to the poor.  About 1.7 billion people worldwide are excluded from formal financial services, such as savings, payments, insurance, and credit. In developing economies, nearly one billion are still left out of the formal financial system, and there is a 9 percent persistent gender gap in financial inclusion in developing economies. 
Most poor households instead, operate almost entirely through a cash economy. They are cut off from potentially stabilizing and uplifting opportunities like building credit or getting a loan to start a business. And it’s harder to weather common financial setbacks, such as serious illness, a poor harvest, or an economic downturn, such as the one the world experiencing now due to the coronavirus epidemic.  In fact, just last week, one of us, who had just moved from Delhi to Manila, was able to help his former Indian housekeeper with a cash transfer, only because she had a bank account.  It took all of 15 minutes to send her much needed financial support, which would have been much more difficult otherwise.  Millions of others have no access to such mechanisms.
The issue in JDEff replicates 6 important financial inclusion studies.  These include a study on providing banking access to farmers; three studies that evaluated interventions to introduce innovative alternatives to traditional banking, such as using mobile phones or biometric smartcards as payment mechanisms for transfers; and two studies that studied the effects of different kinds of transfers (cash versus kind; conditional versus unconditional) that are distributed through financial institutions.  Importantly, the replications were able to reproduce the principal results of all of the studies.  It is as important to highlight this finding, as much as it is to get notoriety through “gotcha” replications that appear to overturn results, such as in the “worm wars” of a few years ago.
The replications also make several useful findings about nuances that the original research may have missed, such as about heterogeneous effects.
Research must meet the higher bar of being good enough for decision-making that affects human lives (not merely good enough for publication. Organizations like 3ie which consider replication as an important tool in making research more rigorous takes the following lessons from this JDEff issue on doing replications in the future.  One is to ensure that, beyond taking the original research at face value, enough attention is dedicated to results not reported (to avoid reporting bias), to the policy significance of the reported results, to reflections about possible rival explanations for the results, and to how the main variables were constructed. Another replication deficiency in current research relates to the replication of qualitative research. While there is increasing acceptance that replication of quantitative research is part of best practice by funders and journals, replication in the qualitative research field is nascent. In a new initiative, 3ie is partnering with the Qualitative Data Repository (Syracuse University) to get their assistance in archiving and sharing select qualitative data, learn from the experience and thereby contribute to lessons and guidance on how to do this in the future.  Finally, within the evidence architecture, it is worthwhile promoting systematic reviews as a set of replications of studies in differentiated real life settings.
Marie Gaarder is the current Executive Director of the International Initiative for Impact Evaluation (3ie).  Emmanuel Jimenez is a Senior Fellow at and former Executive Director of 3ie, and Editor-in-Chief of the Journal of Development Effectiveness.  They can be contacted, respectively, at mgaarder@3ieimpact.org and ejimenez@3ieimpact.org.

Bad Science = Good Citations

[Excerpts are taken from two articles, “The Unfortunately Long Life of Some Retracted Biomedical Research Publications”, by James M. Hagberg, published in the Journal of Applied Physiology; and “Inflated citations and metrics of journals discontinued from Scopus for publication concerns: the GhoS(t)copus Project”, by Andrea Cortegiani et al., posted at BioArXiv]
The Unfortunately Long Life of Some Retracted Biomedical Research Publications
“In 2005 the scientific misconduct case of a noted researcher concluded with, among other things, the retraction of 10 papers. However, these articles continue to be cited at relatively high rates.”
“While it initially appears there was a relative “cleansing”, as citation rates for these articles did decrease after retraction, the reductions in citation rates for these articles (-28%) were the same as those for matched non-retracted publications both by the same author (-28%) and by another investigator (-29%) over the same time frame.”
To read the article, click here. (NOTE: Article is behind a paywall.)
Inflated citations and metrics of journals discontinued from Scopus for publication concerns: the GhoS(t)copus Project
“The journals included in Scopus are periodically re-evaluated to ensure they meet indexing criteria. Afterwards, some journals might be discontinued for publication concerns. Despite their discontinuation, previously published articles remain indexed and continue to be cited.”
“This study aimed (1) to evaluate the main features and citation metrics of journals discontinued from Scopus for publication concerns, before and after discontinuation, and (2) to determine the extent of predatory journals among the discontinued journals.”
“A total of 317 journals were evaluated. The mean number of citations per year after discontinuation was significantly higher than before (median of difference 64 citations, p<0.0001), and so was the number of citations per document (median of difference 0.4 citations, p<0.0001).”
“Twenty-two percent (72/317) of the journals were included in the Cabell blacklist.”
To read the article, click here.

Old Boys Network: 1; Open Science: 0.

[Excerpts taken from the article “The Stewart Retractions: A Quantitative and Qualitative Analysis”, by Justin Pickett, published in Econ Journal Watch]
“This study analyzes the recent retraction of five articles from three sociology journals—Social Problems, Criminology, and Law & Society Review.”
“The only coauthor on all five retracted articles was Dr. Eric Stewart. He was the data holder and analyst for each article. I coauthored one of the retracted articles…”
“I organize my analysis of the quantitative and qualitative data into three sections: (1) what happened in the articles, (2) what happened among the coauthors, and (3) what happened at the journals. Everything—data, code, emails, text messages, Excel files, drafts, and university documents—needed to verify my claims is provided online.”
What happened in the five articles
— “Consistently incorrect means and standard deviations”
— “Non-uniform terminal-digit distributions”
— “Unverifiable surveys”
— “Identical statistics after changes in…everything else”
— “Inexplicable sample sizes and statistics”
— “Unreported, implausible county clusters”
What happened among the coauthors
“The coauthors of the five retracted articles include two past editors of Criminology, the flagship journal of the American Society of Criminology (ASC), as well as three ASC Fellows and two ASC vice presidents.”
“Two coauthors, Brian Johnson and Eric Baumer, have written articles about the importance of research ethics (e.g., “What Scholars Should Know about ‘Self-Plagiarism’” (Lauritsen et al. 2019); “Salami-Slicing, Peek-a-Boo, and LPUS: Addressing the Problem of Piecemeal Publication” (Gartner et al. 2012)).”
“To my knowledge, none of the coauthors have spoken publicly about what happened in the retracted articles, except to insist in the retraction notices that the irregularities resulted from “coding mistakes” and “transcription errors” (Law & Society Review 2020; Criminology 2020a; b), and to defend the accuracy of the retracted findings (Law & Society Review 2020).”
“Scientific fraud occurs all too frequently—approximately 1 in 50 scientists admit to fabricating or falsifying data (Fanelli 2009)—and I believe it is the most likely explanation for the data irregularities in the five retracted articles…The retraction notices say honest error, not fraud, is the explanation. Fortunately, if that is true, Dr. Stewart could easily prove it: recreate the original sample (N = 1,184) that produces the findings in Johnson et al. (2011) and then publicly explain how he did it.”
“…many authors are reluctant to share data publicly, and sometimes there are legitimate privacy concerns or externally imposed restrictions. Sharing data with coauthors, however, should be uncontroversial and feasible.”
“Yet without institutional support, coauthors may feel uncomfortable requesting data. For example, once irregularities were identified in their articles, Dr. Stewart’s coauthors were reluctant to press him for the data, probably because of concerns related to friendship and loyalty.”
What happened at the journals
“None of the editors followed COPE’s guidelines when alerted to the irregularities in Dr. Stewart’s articles. One editor seemingly tried to coordinate a collective response of ignoring the allegations, even though she recognized their potential seriousness…At Criminology, what seems to have driven how the co-editors responded was sympathy for some of the authors and a low opinion of critics.”
“Connections between the co-editors and authors are likely to blame; Dr. Johnson, the lead author of one article, was a co-editor, and Dr. Stewart, the lead author of the other, was to become a co-editor.”
Conclusion and recommendations
“The Stewart scandal took place over five months, and it required considerable time and effort from the editors involved. The editors corresponded extensively with each other and with other parties…the Criminology co-editors wrote multiple public statements about the steps they were taking to address the problems. Why did they not simply ask Dr. Stewart for his data?”
To read the article, click here.
NOTE FROM TRN: The excerpts above do not do justice to the article. The article should be read!

No Respect for Replication Even on the Big Bang Theory

[Excerpts taken from the transcript to Series 2, Episode 15 of the Big Bang Theory – “The Maternal Capacitance”]
LEONARD’S MOTHER: Leonard, it’s one o’clock, weren’t you going to show me your laboratory at one o’clock?
LEONARD: There’s no hurry, Mother…
LEONARD’S MOTHER: But it’s one o’clock, you were going to show me your laboratory at one o’clock.
SHELDON: Her reasoning is unassailable. It is one o’clock.
LEONARD: Fine. Let’s go. I think you’ll find my work pretty interesting. I’m attempting to replicate the dark matter signal found in sodium iodide crystals by the Italians.
LEONARD’S MOTHER: So, no original research?
LEONARD: No.
LEONARD’S MOTHER: Well, what’s the point of my seeing it? I could just read the paper the Italians wrote.
To read the full transcript, click here.

Is It a Replication? A Reproduction? A Robustness Check? You’re Asking the Wrong Question

[Excerpts taken from the article, “What is replication?” by Brian Nosek and Tim Errington, published in PloS Biology]
“Credibility of scientific claims is established with evidence for their replicability using new data. This is distinct from retesting a claim using the same analyses and same data (usually referred to as reproducibility or computational reproducibility) and using the same data with different analyses (usually referred to as robustness).”
“Prior commentators have drawn distinctions between types of replication such as “direct” versus “conceptual” replication and argue in favor of valuing one over the other…By contrast, we argue that distinctions between “direct” and “conceptual” are at least irrelevant and possibly counterproductive for understanding replication and its role in advancing knowledge.”
“We propose an alternative definition for replication that is more inclusive of all research and more relevant for the role of replication in advancing knowledge. Replication is a study for which any outcome would be considered diagnostic evidence about a claim from prior research. This definition reduces emphasis on operational characteristics of the study and increases emphasis on the interpretation of possible outcomes.”
“To be a replication, 2 things must be true: outcomes consistent with a prior claim would increase confidence in the claim, and outcomes inconsistent with a prior claim would decrease confidence in the claim.”
“Because replication is defined based on theoretical expectations, not everyone will agree that one study is a replication of another.”
“Because there is no exact replication, every replication test assesses generalizability to the new study’s unique conditions. However, every generalizability test is not a replication…there are many conditions in which the claim might be supported, but failures would not discredit the original claim.”
“This exposes an inevitable ambiguity in failures-to-replicate. Was the original evidence a false positive or the replication a false negative, or does the replication identify a boundary condition of the claim?”
“We can never know for certain that earlier evidence was a false positive. It is always possible that it was “real,” and we cannot identify or recreate the conditions necessary to replicate successfully…Accumulating failures-to-replicate could result in a much narrower but more precise set of circumstances in which evidence for the claim is replicable, or it may result in failure to ever establish conditions for replicability and relegate the claim to irrelevance.”
“The term “conceptual replication” has been applied to studies that use different methods to test the same question as a prior study. This is a useful research activity for advancing understanding, but many studies with this label are not replications by our definition.”
“Recall that “to be a replication, 2 things must be true: outcomes consistent with a prior claim would increase confidence in the claim, and outcomes inconsistent with a prior claim would decrease confidence in the claim.”
“Many “conceptual replications” meet the first criterion and fail the second…“conceptual replications” are often generalizability tests. Failures are interpreted, at most, as identifying boundary conditions. A self-assessment of whether one is testing replicability or generalizability is answering—would an outcome inconsistent with prior findings cause me to lose confidence in the theoretical claims? If no, then it is a generalizability test.”
To read the article, click here.

Raise Your Hand If You’ve Messed Up

[Excerpts taken from the article “When We’re Wrong, It’s Our Responsibility as Scientists to Say So” by  Ariella Kristal et al., published in Scientific American.]
“What simple, costless interventions can we use to try to reduce tax fraud? As behavioral scientists, we tried to answer this question using what we already know from psychology: People want to see themselves as good.”
“…we thought that by reminding people of being truthful before reporting their income, they would be more honest. Building on this idea, in 2012, we came up with a seemingly costless simple intervention: Get people to sign a tax or insurance audit form before they reported critical information (versus after, the common business practice).”
“While our original set of studies found that this intervention worked in the lab and in one field experiment, we no longer believe that signing before versus after is a simple costless fix….Seven years and hundreds of citations and media mentions later, we want to update the record.”
“Based on research we recently conducted—with a larger number of people—we found abundant evidence that signing a veracity statement at the beginning of a form does not increase honesty compared to signing at the end.”
“Why are we updating the record? In an attempt to replicate and extend our original findings, three people on our team (Kristal, Whillans and Bazerman) found no evidence for the observed effects across five studies with 4,559 participants.”
”We brought the original team together and reran an identical lab experiment from the original paper (Experiment 1). The only thing we changed was the sample size: we had 20 times more participants per condition. And we found no difference in the amount of cheating between signing at the top of the form and signing at the bottom.”
“This matters because governments worldwide have spent considerable money and time trying to put this intervention into practice with limited success.”
“We also hope that this collaboration serves as a positive example, whereby upon learning that something they had been promoting for nearly a decade may not be true, the original authors confronted the issue directly by running new and more rigorous studies, and the original journal was open to publishing a new peer-reviewed article documenting the correction.”
“We believe that incentives need to continue to change in research, such that researchers are able to publish what they find and that the rigor and usefulness of their results, not their sensationalism, is what is rewarded.”
To read the article, click here.