ETIENNE LEBEL, in a blog for BITSS, gives a brief but wide-ranging summary of the status of “open science” in psychology. Topics include: (i) the use of “badges” to encourage provision of research materials, (ii) pre-registration, (iii) reproducibility, (iv) replications, (v) peer review, and (vi) meta-analysis, among other topics. To read more, click here.
[From the article “Come Again”]: “The GRIM test, short for granularity-related inconsistency of means, is a simple way of checking whether the results of small studies of the sort beloved of psychologists (those with fewer than 100 participants) could be correct, even in principle. … “
To understand the GRIM test, consider an experiment in which participants were asked to assess something (someone else’s friendliness, say) on an integer scale of one to seven. The resulting paper says there were 49 participants and the mean of their assessments was 5.93. It might appear that multiplying these numbers should give an integer product—ie, a whole number—since the mean is the result of dividing one integer by another. If the product is not an integer (as in this case, where the answer is 290.57), something looks wrong.”
When the authors of the GRIM test took their simple test to analyse 71 papers in three leading psychology journals, they found that over half the papers failed the test. To read more, click here.
There has been a huge amount of attention focused on “open data.” A casual reading of the blogosphere is that Open Data is good, Secret Data is bad.
Remarkably, there has been very little discussion given to the property right issues associated with open data. The Open Data Movement wants to turn a private good (datasets) into a public good. Economists know something about public goods. They tend to get under-produced. This introduces a trade-off between the propagation of data for use by multiple researchers, a social good (though see here for a discussion where this is not necessarily so), versus the disincentive this causes for producing data, a social bad. How best to make this trade-off is unclear.
In a recent blog entitled “Open data, authorship, and the early career scientist”, MARGARET KOSMALA, a postdoctoral fellow at Harvard University, argues that making one’s data available to others hurts the data-producing scholar, particularly younger scholars. The argument is not so much that the data-producing scholar will be scooped by other scientists on the associated research. Rather, it is that subsequent research projects that could have resulted in publications for the data-producing scholar will end up being undertaken by other scientists. And while Kosmala does not make this point explicitly, this serves as a disincentive for scientists to produce data, if only because younger scholars may not be able to produce sufficient publications to get the funding and tenure they need to continue their careers.
What is really interesting about this blog is that it led to a discussion between a reader and the author about the ethics of “requiring co-authorship” when authors use data produced by another scientist. Missing from the discussion was the recognition that “requiring co-authorship” provides a potential solution to the problem of open data. It is a way for the data-producing scientist to reap the rewards of data production, while still allowing other authors to use it.
Of course, there are issues associated with implementing a policy like this. Once data are released, how will the data-producing author be able to ensure that others who use the data will extend co-authorship to him/her? And suppose the data-producing author does not wish to have their data used in a certain way. Should they have the right to restrict its use? While the answers are debatable, the questions are illuminating, because they make us realise that the debate over open data is just another application of the larger subject of intellectual property rights.
So another study finds that X affects Y, and you are a sufficiently cynical TRN reader that you wonder if the authors have p-hacked their way to get their result. Don’t have time (or the incentive) to do a replication? You might consider using a “p-curve” analysis to determine whether the effect has “evidentiary value.” How does one do that? Let’s take as a given that most journals will not publish a result unless it is statisically significant. Even if the journals only report significant results, one can examine the distribution of p-values to determine whether or not the effect is true. Want to learn more about “p-curves?” The original article by Simonsohn, Nelson, & Simmons (2011), “P-Curve: A Key to the File Drawer” can be found here. A straightforward explanation of the technique by Will Gervais can be found here. A critique by Bruns & Ioannidis can be found here. And an excellent response by the original authors can be found here. P-curves are definitely worth a look!
[From the article “Muddled meanings hamper efforts to fix reproducibility crisis” in Nature] “A semantic confusion is clouding one of the most talked-about issues in research. Scientists agree that there is a crisis in reproducibility, but they can’t agree on what ‘reproducibility’ means.” To read more, click here.
[From the podcast “When Great Minds Think Unlike: Inside Science’s ‘Replication Crisis” from NPR’s Hidden Brain series] This podcast is distinguished by its discussion of what it means – and what it doesn’t mean – when a replication “fails.” It is about 30 minutes long.
It is self-described as follows: “This week, Hidden Brain looks at the “replication crisis” through zooming in on one seminal paper that was the focus of two replication efforts: one succeeded in replicating the original finding, the other failed.
The original study, authored by Margaret Shih, Todd Pittinsky, and Nalini Ambady in 1999, found that Asian women performed worse on a math test when primed to think about their female identity, but better when they were primed to think about their Asian identity.
Nearly two decades later, Nosek and the Reproducibility Project noticed that this study, which by then had been widely disseminated in textbooks and psychology education, had never itself been replicated. So he assigned two teams to run it again—one in Georgia and the other in California. They came back with different results. And this gets at one of the biggest questions explored in this episode: when scientific studies come to different conclusions, what should we think of as true?” To listen to the podcast, click here.
From the article “1500 Scientists lift the lid on reproducibility” published in Nature: “More than 70% of researchers have tried and failed to reproduce another scientist’s experiments, and more than half have failed to reproduce their own experiments. Those are some of the telling figures that emerged from Nature‘s survey of 1,576 researchers who took a brief online questionnaire on reproducibility in research.”
Other questions explored in the survey included: (i) Is there a reproducibility crisis in science?”, (ii) “How much published work in your field is reproducible?”, (iii) “Have you failed to reproduce an experiment”, and (iv) “Have you ever tried to publish a reproduction attempt?”. Respondents came from a wide variety of subject areas including chemistry, biology, physics and engineering, medicine, and earth and environmental sciences. To read more, click here.
Earlier this month, the Psychonomic Society meetings held a session on Open Science. The session was recorded and is available on YouTube (click here). It consisted of four presentations.
— “The Peer Reviewers’ Openness Initiative” by RICHARD MOREY of Cardiff University (1:11)
— “The Availability of Psychological Research Data” by WOLF VANPAEMEL of KU Leuven (23:20)
— “On Knowing How the Sausage is Made” by ROLF ZWAAN of Erasmus University Rotterdam (43:50)
— “The Dark Side of Open Science: Weaponizing Transparency” by STEPHAN LEWANDOWSKY (1:01:23)
The times at which the talks appear in the video are given in parentheses above. The last talk makes a number of arguments that caution that data sharing may not be an unambiguously good thing. All of the talks are interesting and recommended. But, then again, TRN may be a little biased.
In an article entitled “Why Do So Many Studies Fail to Replicate,” Jay Van Bavel, an associate professor of psychology at NYU, writes: “In a paper published on Monday in the Proceedings of the National Academy of Sciences, my collaborators and I shed new light on this issue. Our results suggest that many of the studies failed to replicate because it was difficult to recreate, in another time and place, the exact same conditions as those of the original study.” The example the author gives is a study he did in Canada in 2006 on people’s emotional responses to famous people. The replication was to be done in 2016, in the US. Likely the “famous people” used in the Canadian study (e.g., Jean Chretien, Don Cherry and Karla Homolka) would not carry over the geographical and time divide to the replication. To read more, click here.
In a recent interview on Retraction Watch, Andrew Gelman reveals that what keeps him up at night isn’t scientific fraud, it’s “the sheer number of unreliable studies — uncorrected, unretracted — that have littered the literature.” He then goes on to argue that retractions cannot be the answer. His argument is simple. The scales don’t match. “Millions of scientific papers are published each year. If 1% are fatally flawed, that’s thousands of corrections to be made. And that’s not gonna happen.”
Actually, if 1% of studies are fatally flawed, the problem is probably manageable. Assuming a typical journal publishes 10 articles an issue, 4 issues a year, that means one retraction every two and a half years, which is certainly feasible for a journal. Problems arise only when the percent substantially rises. Gelman goes on to say that he personally thinks the error rate to be large as 50% in some journals, where “half the papers claim evidence that they don’t really have.” At that point retractions are not the solution.
If revealed preference is any indication, hopes for a solution appear centered on “data transparency.” Data transparency means different things to different people, but a common core is that researchers make their data and programming code publicly available.
The Center for Open Science, Dataverse, and EUDAT are but a few examples of the high-profile explosion in efforts to make research data more “open,” transparent and shareable. In a recent guest blog at The Replication Network (reblogged from BITSS), Stephanie Wykstra promotes the related topic of data re-use.
In an encouraging sign, these efforts appear to have had an impact. A recent survey article by Duvendack et al. report that, of 333 journals categorized as “economics journals” by Thompson Reuter’s Journal Citation Reports, 27, or a little more than 8 percent, regularly published data and code to accompany empirical research studies. As some of these journals are exclusively theory journals, the effective rate is somewhat higher.
Noteworthy is that many of these journals only recently instituted a policy of publishing data and code. So while one can argue whether the glass is, say, 20 percent full or 80 percent empty, the fact is that the glass used to contain virtually nothing. That is progress.
But making data more “open” does not, by itself, address the problem of scientific unreliability. Researchers have to be motivated to go through these data, examine them carefully, and determine if they are sufficient to support the claims of the original study. Further, they need to have an avenue to publicize their findings in a way that informs the literature.
This is what replications are supposed to do. Replications provide a way to confirm/disconfirm the results of other studies. They are scalable to fit the size of the problem. With so many studies potentially unreliable, researchers would prioritize the most important findings that are worthy of further analysis. The self-selection mechanism of researchers’ time and interests would insure that the most important, most influential studies are appropriately vetted.
But after obtaining their results, researchers need a place to publicize their findings.
Unfortunately, on this dimension, the Duvendack et al. study is less encouraging. They report that only 3 percent of “economics” journals explicitly state that that they publish replications. Most of these are specialty/field journals, so that an author of a replication study only has a very few outlets, maybe as few as one or two, in which they can hope to publish their research.
And just because a journal states that it publishes replication studies, doesn’t mean that it does it very often. Duvendack et al. report that 6 journals account for 60 percent of all replication studies ever published in Web of Science “economics” journals. Further, only 10 journals have ever published more than 3 replication studies. In their entire history.
Without an outlet to publish their findings, researchers will be unmotivated to spend substantial effort re-analysing other researchers’ data. Or to put it differently, the open science/data sharing movement only addresses the supply side of the scientific market. Unless the demand side is addressed, these efforts are unlikely to be successful in providing a solution to the problem of scientific unreliability.
The irony is this: The problem has been identified. There is a solution. The pieces are all there. But in the end, the gatekeepers of scientific findings, the journals, need to open up space to allow science to be self-correcting. Until that happens, there’s not much hope of Professor Gelman getting any more sleep.
Bob Reed is Professor of Economics at the University of Canterbury in New Zealand, and co-organizer of The Replication Network.
You must be logged in to post a comment.