Incentives to Replicate: Show Me the Money!

[From the article “Go Forth and Replicate: On Creating Incentives for Repeat Studies” by Michael Schulson at the web magazine Undark]
“Suggested reasons for the [replication] crisis are many. Some researchers have blamed the scientific publishing culture itself …. But some stakeholders also point to a more concrete and obvious root cause: a lack of direct incentives to replicate other researchers’ work, including very little funding to do replications. This raises some key questions, including whether the federal government — the largest single funder of basic science in the world — ought to be doing more to encourage and underwrite basic replication research. At the moment, only token reforms have been undertaken in the U.S., but there are signs that novel programs are gestating elsewhere, including a European pilot program that could serve as a model for replication funding in other countries.”
To read more, click here.

All Researchers Want Open Data, Right?

[From the blog by Justin Gallagher entitled “Public Data that Isn’t (or Wasn’t) Public” posted at BITSS]
“I recently completed a coauthored working paper, together with Paul J. Fisher, that examines whether electronic monitoring via red light traffic cameras leads to fewer vehicle accidents and injuries.  As part of the project, we encountered the bizarre situation where researchers at The University of Texas at Austin, a flagship public state university, fought tooth and nail against releasing vehicle accident data originally collected by the Texas Department of Transportation (TxDOT).”
To read more, click here.

CHU, HENDERSON, AND WANG: US Food Aid — Good Intentions, Bad Outcomes

[NOTE: This post is based on the paper, “The Robust Relationship between US Food Aid and Civil Conflict”, Journal of Applied Econometrics, 2017]
Replication can often be thought of as a useful tool to train graduate students or as a starting point for a new line of research, but sometimes replication is necessary as a means to check the robustness of results that can directly influence policy. Recently, Nunn and Qian (US Food Aid and Civil Conflict, American Economic Review 2014; 104: 1630–1666) found that United States (US) food aid increases the incidence and duration of civil conflict in recipient countries. This paper has received significant attention and has even been noticed by the United States Agency for International Development (USAID).
If the results of their study are robust, policymakers can attempt to minimize the predicted negative impacts. In our paper, we first were able to successfully replicate the results of Nunn and Qian (2014) using their data, but alternative software (R instead of Stata). We then attempted to further scrutinize one of their conclusions. Specifically, the authors claim that the adverse effect of US food aid on conflict does not vary across pre-determined characteristics of aid recipient countries (a seemingly strong assumption across sometimes vastly different nations). Nunn and Qian (2014) made attempts to allow for heterogeneity in their regression models by interacting US food aid with these pre-determined characteristics, but this simply amounted to group averages which may miss the underlying heterogeneity.
In order to check for more sophisticated forms of heterogeneity, we used a semiparametric estimation procedure. While the results visually suggested the presence of some amount of heterogeneity, this could not be determined statistically as we were unable to formally reject any of the parametric specifications in Nunn and Qian (2014). The conclusion of such a replication is that their models cannot be rejected using their data and we argue that the results of their paper are robust.
While we rightly criticize studies that cannot be replicated, we should also make note of those that can be replicated. It is typically a non-trivial task and while we were successful, we suggest that this study be further tested with samples from a different set of countries (both recipients and donors) and/or time periods.
Chi-Yang Chu is an assistant professor of economics at National Taipei University. Daniel J. Henderson is a professor of economics and the J. Weldon and Delores Cole Faculty Fellow at the University of Alabama. Le Wang is an associate professor of economics and the Chong K. Liew Chair in Economics at the University of Oklahoma. Correspondence about this blog should be directed to Daniel Henderson at djhender@culverhouse.ua.edu.
References
[1]  N. Nunn and N. Qian, US Food Aid and Civil Conflict. American Economic Review. 104, 1630–1666 (2014)
[2]  C.-Y. Chu, D. J. Henderson and L. Wang. The Robust Relationship between US Food Aid and Civil Conflict. Journal of Applied Econometrics. 32, 1027-1032 (2017)

What?! You Don’t Believe in Badges?!!

[The following is taken from a blog by Hilda Bastian at the blogsite “Absolutely Maybe” at PLOS Blogs]
“As I’ve spent time with the badges “magic bullet” – simple! cheap! no side effects! dramatic benefits! – supported by a single uncontrolled study by an influential opinion leader, with a biased design in a narrow unrepresentative context, very small number of events, and short timeframe….I’ve come to think its biggest lesson may be that even many open science advocates have yet to fully absorb the implications of science’s reliability problems.”
To read more, click here.

Hmmm. Maybe the R-Factor is NOT the answer.

In a recent post, TRN highlighted a recent working paper touting the benefits of something called an “R-Factor.” The R-Factor is a metric that would report — for each published empirical study — the reproducibility rate of that study in subsequent research.  In a recent blog at Discover magazine’s website, Neuroskeptic blogs:
“A new tool called the R-factor could help ensure that science is reproducible and valid, according to a preprint posted on biorxiv: Science with no fiction. The authors, led by Peter Grabitz, are so confident in their idea that they’ve created a company called Verum Analytics to promote it. But how useful is this new metric going to be?”
Neuroskeptic’s answer: “Not very.” To read more, click here.

 

WEICHENRIEDER: FinanzArchiv/Public Finance Analysis Wants Your Insignificant Results!

There is considerable concern among scholars that empirical papers face a drastically smaller chance of being published if the results looking to confirm an established theory turn out to be statistically insignificant. Such a publication bias can provide a wrong picture of economic magnitudes and mechanisms.
Against this background, the journal FinanzArchiv/Public Finance Analysis recently posted a call for papers for a special issue on “Insignificant Results in Public Finance”. The editors are inviting the submission of carefully executed empirical papers that – despite using state of the art empirical methods – fail to find significant estimates for important economic effects that have widespread acceptance.
It has been estimated that studies in economic behavioral research and psychology are ten times more likely to be published if they present statistically significant effects. Because a significant result may happen by chance, too much weight is attributed to them in the scientific literature. The associated publication bias can produce overestimates of the effectiveness of economic policy measures, psychological impacts or even medical medications.
While several ways to address this issue exist, a correction that is most directly related to the problem concerns the attitude of the scientific editors. Publication bias and the negative incentives for researchers are tackled at the root when studies are assessed on the basis of the methodology used and the quality of the data — and not on the results obtained. It requires a certain self-commitment of the journals to the increased publication of so-called “non-significant” results. Such a self-commitment was recently submitted by the editors of FinanzArchiv/Public Finance Analysis and is reflected in its call for papers.
The deadline for submissions to the special issue is 15 September 2017.  Papers can be uploaded here:  Submitting authors should indicate that their paper is being submitted to the special issue “Insignificant Results in Public Finance”. The editors would like to note that if any insignificant results transform into statistically significant results as an outcome of the refereeing process, this will not be held as an argument against publication. In this case, the paper may be shifted into a different issue of the Journal.
FinanzArchiv was first published in 1884, which makes it one of the world’s oldest professional journals in economics and the oldest journal of public finance. The current editors are Katherine Cuff, Ronnie Schöb and Alfons Weichenrieder. Within public economics, a strong focus is on topics as taxation, public debt, public goods, public choice, federalism, market failure, social policy and the welfare state.
Alfons Weichenrieder is Professor of Economics and Public Finance at Goethe-University Frankfurt and a guest research professor at the Institute of International Taxation of Vienna University of Economics and Business.  He can be contacted via email at a.weichenrieder@em.uni-frankfurt.de.

Pre-register. Make a $1000. Really?

[From the Center for Open Science webpage.]
“If you have a project that is entering the planning or data collection phase, we’d like you to try out a preregistration. Through our $1 Million Preregistration Challenge, we’re giving away $1,000 to 1,000 researchers who preregister their projects before they publish them. It’s straightforward to complete and will really enhance your research output.”
To read more, click here.

FYI: ScienceOpen Has a Collection of Papers on How to Fix the Replicability Crisis

ScienceOpen has a collection entitled: “Remedies to the Reproducibility Crisis”.  The collection is introduced thusly:
“Psychology, Medicine, Neuroscience and many other research fields, are facing a serious reproducibility crisis, that is, most of the findings published in peer-review journals, independently from their prestige, are not replicable. This collection aims at offering all remedies suggested to fix this problem.”
The collection currently consists of 24 articles on topics such as:
– “A manifesto for reproducibile science” (Munafo et al.)
– Scientific Standards: Promoting an open research culture (Nosek et al.)
– “The New Statistics: why and how” (Cumming)
– “Badges to acknowledge open practices: A simple, low-cost, effective method for increasing transparency” (Kidwell, et al.)
– “Calculating and reporting effect sizes to facilitate cumulative science” (Lakens)
– “The Peer Reviewers’ Openess Initiative: incentivizing open research practices through peer review” (Morey, et al.)
– “The influence of journal submission guidelines on authors’ reporting of statistics and use of open research practices” (Giofre et al.)
– “On the reproducibility of meta-analyses: six practical recommendations” (Lakens, Hilgard, and Staaks)
– “Equivalence Tests” (Lakens)
– “An agenda for purely confirmatory research” (Wagenmakers et al.)
– And more.
To see the collection and read more, click here.

A Pop Quiz on Significant Effects with Small Sample Sizes

QUICK: Does finding a significant effect when the sample size is small make it more likely that the effects are real and important?  Or less?
James Heckman, Nobel Prize winning economist, says more:
“Also holding back progress are those who claim that Perry and ABC are experiments with samples too small to accurately predict widespread impact and return on investment. This is a nonsensical argument. Their relatively small sample sizes actually speak for — not against — the strength of their findings. Dramatic differences between treatment and control-group outcomes are usually not found in small sample experiments, yet the differences in Perry and ABC are big and consistent in rigorous analyses of these data.” (click here for source).
Andrew Gelman, Professor Statistics and Political Science at Columbia University, and blogger extraordinaire, says less:
“I agree with Stuart Buck that Heckman is wrong here. Actually, the smaller sample sizes (and also the high variation in these studies) speaks against—not for—the strength of the published claims.”
Who do YOU think is right?
To read more from Gelman, click here. His argument is elaborated in this working paper.

Is the R-Factor the Answer?

In a recent working paper (“Science with no fiction: measuring the veracity of  scientific reports by citation analysis”), Peter Grabitz, Yuri Lazebnik,  Josh Nicholson, and Sean Rife suggest that one solution to the “crisis” in scientific credibility is publication of an article’s “R-Factor”.   To calculate the R-Factor for a given study, one would comb through all the papers that cite a given study, then count up the number of attempts to confirm the findings from the original study.  The R-Factor is simply the ratio of confirming studies over total attempts.  R-Factors close to 1 indicate a study is likely to be true.  R-Factors close to 0, not so much.  The authors give an example from three studies in biomedical research.  And how would this be done for thousands and thousands of studies, with results being continuously updated?  The authors suggest this could be done through machine learning technology.  
To read more, click here.