STEPHANIE WYKSTRA: On Data Re-use

[THIS BLOG ORIGINALLY APPEARED ON THE BITSS WEBSITE]  As advocates for open data, my colleagues and I often point to re-use of data for further research as a major benefit of data-sharing. In fact there are many cases in which shared data was clearly very useful for further research. Take the Sloan Digital Sky Survey (SDSS) data, which researchers have used for nearly 6,000 papers. Or take Genbank, within bioinformatics, which is a widely used database of nucleotide and protein sequence data. Within social science, large-scale surveys such as the Demographic and Health Survey (DHS) are used by many, many researchers as well as policy-makers.
Research data re-use: where are the cases?
In spite of the obviousness of the value of data-sharing in general, I realized that we didn’t have many cases of re-use of research data. By “research data” here, I have in mind data which were collected by an individual researcher or research team for their own project (e.g. from a field experiment), and then shared along with the publication. This differs from the databases like SDSS, Genbank and DHS in a few different ways:
— The data are often much smaller scale than DHS or SDSS; they are often studies of a few hundred to a few thousand subjects.
— They are not part of a unified data-gathering effort using common measures (as are SDSS and DHS), but rather use their own instruments, often with their own non-standardized measures.
— While it’s fairly clear that researchers can use SDSS data for their own research, and bioinformaticists can use Genbank data, it’s less clear how social scientists would re-use data that other researchers collected for the purpose of their own study. In general, they could use data for secondary analysis or meta-analysis; however, we haven’t seen numerous examples.
A call for cases studies of data re-use
After a brainstorming session with Stephanie Wright, a colleague at Mozilla Science Lab, we decided to put out a call for cases of data re-use. For this project, we were particularly interested in cases of re-use within economics or political science. Since we support data-sharing among researchers and research staff, we want to be able to point to cases of real world re-use, and to delve into what made the data particularly useful. We wrote a post on our project, along with a survey on data re-use, and shared in venues such as IASSIST, Polmeth, Berkeley Initiative for Transparency in the Social Sciences’ blog, Open Science Collaboration’s discussion board, Mozilla Science Lab’s blog and various data librarian email lists.
What did we find through our call?
We received 14 responses to our call, including 10 responses to our survey and 4 emailed responses. While the number and quality of responses isn’t sufficient for us to learn a great deal, we want to share what we found in any case, for two reasons: (1) This call and response could be informative to those who are considering putting out a similar survey and (2) we think our findings do provide some evidence which confirms our initial feeling, which is that this is an area which warrants further work and research.
Our 10 survey respondents are in a variety of fields: one in political science, two in psychology, one in education and most of the rest in biochemistry. While all respondents did mention some data that were re-used, only three gave examples of the kind that we had requested e.g. data that had been collected by other researchers for their study, and then re-used for further research. The three cases included:
— Re-use of data from a collaboration of Psychology instructors, which collected data on emerging adulthood and politics. The data were not initially used for a publication, as intended, but were archived and were used for nine published articles later on.
— A researcher in political science gave several of his own research re-use cases in which publicly available data were used for a) a replication to “illustrate the usefulness of a new fit assessment technique for binary DV [dependent variable] models,” b) for pedagogical purposes in a book on causal inference and c) to test a new theory.
— Researchers in psychology used data from two large-scale studies on the benefits and transfer effects of a cognitive training for older adults. The data were used to test whether a subset of one test (the Useful Field of View test) were able to predict scores on another test (the Instrumental Activities of Daily Living test).
Beyond the cases above, we heard about re-use of protein sequence data and genomics data from databases such as ArrayExpress and Protein Data Bank, as well as government data from Open Data Toronto and Statistics Canada. See our spreadsheet for further details (we asked for permission to share responses).
In addition to the cases gathered through our survey, we received four emails with tips about where to find additional cases. One of the suggestions mentioned the Global Biodiversity Information Facility (GBIF), a database on global biodiversity, as well as International Polar Year (IPY), a coordination of research on the Polar regions. A second suggestion from a political scientist pointed to several sites, Uppsala Conflict Data Program and the Correlates of War Program. Both sites offer data which are widely used by scholars within international relations, and include variables which are constructed by scholars for their own research, and then submitted to the databases for others to re-use.
Finally, we received several suggestions from fellow open data advocates, of places to look for cases of re-use. The first source, Dissemination Information Packages for Information Re-use (DIPIR) is a study of data re-use in three communities (quantitative social scientists, archaeologists, and zoologists). The second is ICPSR’s bibliography of data-related literature, which is a searchable database of “over 70,000 citations of published and unpublished works resulting from analyses of data held in the ICPSR archive.” The third is UK Data Archive’s list of case studies of data re-use.
Data re-use: key for rewarding data-sharing
The data-sharing movement is gaining steam. From funders requiring data-sharing to new guidelines for journals (TOP guidelines) and journal requirements, to the rise of many data repositories, there is plenty of effort going into requiring and supporting data-sharing. Yet there are huge issues to confront, as we move forward. One of the biggest is how to change from a culture in which data-sharing is not a norm among researchers (as is still the case in many scientific fields) to one in which it is.
Researchers are rewarded for publishing, not for sharing data, and many researchers cite barriers to sharing data such as lack of time and lack of support (Tenopir et al. 2011). How will we shift to rewarding researchers for sharing their data, so that they have professional incentives to take the time to prepare and share data? One of the most-discussed ways is to develop good data-citation norms, and then to reward researchers (via tenure committee decisions) when others re-use and cite their data.
So, the question of how to promote and encourage data re-use is of clear importance. Yet, as practitioners in the open science movement, we have many questions. When it comes to re-using data from colleagues’ studies, particularly in the social sciences, what factors make datasets particularly helpful to researchers? What challenges arise in re-using data? As data curators and open data advocates, what could we do better to facilitate re-use?  Is there something we can do to encourage others to look at and reuse existing data when they are considering new research projects? How can we increase opportunities for re-using data and decrease barriers?
Next steps
Particularly when it comes to data shared by researchers in the social sciences, we still need more examples of re-use. We also need much more investigation into what would make researchers more likely to re-use data from colleagues for their own research. We can think of a couple of interesting projects that we could undertake:
— Delving into the archives from ICPSR and UK Data Archive, as well as others mentioned above, and attempting to glean lessons from specific cases of re-use found there.
— Contacting researchers that have downloaded data from archives such as IPA’s data repository (we track data users and ask them permission to contact them, when they download data). We could gather more detailed information about what was or wasn’t helpful for re-use about the data and other materials as presented in the repository. We could also try to gather more information on whether data were re-used for further research (and if so, what made them particularly attractive for re-use).
We’re certainly open to further suggestions, so please get in touch if you have ideas!
Stephanie Wykstra directs the Research Transparency Initiative at Innovations for Poverty Action, and also works as an independent research consultant. She may be contacted at stephanie.wykstra@gmail.com.

OkCupid: Where Love and Data Transparency Don’t Match

Welcome to the tale of Emil Kirkegaard, a Danish postgraduate student, who has achieved worldwide notoriety for publishing data from the dating site, OkCupid.  The story is well-told in a Vox article by Brian Resnick (click here).  In addition to a committing a number of ethical sins, both mortal and venal, the case of Mr. Kirkegaard raises important issues for the data transparency movement.  Mr. Kirkegaard stated that he needed to make his data — which was collated from the OkCupid website and “publicly” accessible to OkCupid users — available because that was the condition for submission to an open access journal where he was hoping to publish his research (hopes that were no doubt encouraged by the fact that he was the editor of the journal).  
Perhaps more interesting are the ethical issues this case raises.  Some of these issues are discussed in a blog by Oliver Keyes (click here).  OkCupid will likely file a legal complaint which may involve Open Science Framework (OSF), the online host of Kirkegaard’s data.  Given the open hosting nature of OSF, it seems unlikely that OSF will be at much legal risk. But how about journals (such as Economics Letters) that encourage submitters to post their data on open access data sites like OSF and Dataverse?  Are they legally liable for data improprieties?  Or how about journals that post article supplementary material such as data and code on the journal website?  Are they legally responsible if the data violates copyright or other legal requirements?  And if this is a legal grey area, will this have a chilling effect on the data transparency movement?  Stay tuned.

 

Workshop: The Quest for Reproducible Science

The Association for Library Collections & Technical Services, a division of the American Library Association, is hosting a one-day workshop on issues of reproducibility.  The workshop will feature “scholars, librarians, and technologists” discussing “tools and techniques to manage data, enable research transparency, and promote reproducible science. Attendees will learn strategies for fostering and supporting transparent research practices at their institutions”.  Participants will (i) “learn tools and techniques that can be immediately employed at their institution to foster and support a culture of transparent research practices”; (ii) “know how to develop a research project on an open platform to manage data and other digital objects throughout the research lifecycle; and (iii) “have access to techniques for organizing empirical research projects in such a way that they can be easily and exactly reproduced.”  The conference will be held in Orlando, Florida, on Friday, June 24th.  To learn more, click here.

ROBERT GELFOND and RYAN MURPHY: Out-of-Sample Tests and Macroeconomics

The replication crisis has elicited a number of recommendations, from betting on beliefs, to open data, to improved norms in academic journals regarding replication studies. In our recent working paper, “A Call for Out-of-Sample Testing in Macroeconomics” (available at SSRN), we argue that a renewed focus on out-of-sample tests will significantly mitigate the issues with replication, and we document the fact that out-of-sample tests are absent from entire literatures in economics.
Our starting point is the observation that the literature regarding the government spending multiplier lacked almost any result supported by an out-of-sample test. For this we use a fairly forgiving definition of “out-of-sample test;” so long as a model is parameterized in one period and then applied to new data, we count it. We review 87 empirical papers estimating the multiplier. Out-of-sample tests do not make an appearance, with only a few exceptions. Given that this question is perhaps the most important in macroeconomics, with quite literally trillions of dollars on the line, this result is jarring.
It was in 1953 that Milton Friedman published The Methodology of Positive Economics, urging economists to use prediction as their criterion for comparing the worthiness of competing theories. Clearly, philosophy of science and practical econometrics have moved beyond this simplistic dictum, but does it make sense to cast aside out-of-sample predictions altogether when comparing theories and models? Are we all that confident that the results found using the methods which claim the throne of the “credibility revolution in empirical economics” will withstand the scrutiny of truly out-of-sample tests?
The primary exception to our result is a 2007 paper by Frank Smets and Rafael Wouters, who ably perform an out-of-sample test against a series of baseline models. However, in the absence of other papers performing such tests, it is difficult to say how strong of a result it is. Even more laudable is the lengthy attempt in 2012 by Volker Wieland and his colleagues in comparing the performance of a number of macroeconomic models, although it is difficult to parse this study to answer the narrower question regarding the government spending multiplier. Another example is a 2016 paper by Jorda and Taylor, which creates a “counterfactual forecast,” which is similar to, but not quite, an out-of-sample test.
The other two examples we were able to identify were published in 1964 and 1967.
There are clearly additional criteria that economists can and should use for evaluating theories. Nonetheless, the paucity of examples of these tests points to p-hacking, specification searches, and the whole slew of problems associated with the replication crisis of social science. And perhaps macroeconomics is “hard” and things like recessions cannot be reasonably forecasted. Fine. Meteorologists cannot forecast more than a week or so ahead, but they still forecast what they can forecast. What models work the best is still an extremely pertinent question, even if all models fail miserably when a recession hits.
Rather, doing away with out-of-sample tests and other similar tests does away with the scientific ideal of Conjectures and Refutations, with scientific knowledge evolving as bold ideas starkly stated compete for the title of least wrong.
Bob Gelfond is the CEO of MQS Management LLC and the chairman and founder of MagiQ Technologies. Ryan Murphy is a research assistant professor at the O’Neil Center for Global Markets and Freedom at SMU Cox School of Business.

Journal Finds that Digital Badges Increase Data Sharing

[From the Retraction Watch website] “In January 2014, Psychological Science began rewarding digital badges to authors who committed to open science practices such as sharing data and materials. A study published today in PLOS Biology looks at whether publicizing such behavior helps encourage others to follow their leads. The authors summarize their main findings in the paper: `Before badges, less than 3% of Psychological Science articles reported open data.  After badges, 23% reported open data, with an accelerating trend; 39% reported open data in the first half of 2015″.  To read more, click here.

John Oliver and Last Week Tonight on Replications and Scientific Reliability

How does one know when replication has hit the big time?  When JOHN OLIVER and LAST WEEK TONIGHT do an entire episode on it.  For readers of TRN, much of what he talks about will be familiar.  Just a lot funnier.  Check it out here.

BOB REED: Replications Can Make Things Worse? Really?

In a recent article in Slate entitled “The Unintended Consequences of Trying to Replicate Research,” IVAN ORANSKY and ADAM MARCUS from Retraction Watch argue that replications can exacerbate research unreliability.  The argument assumes that publication bias is more likely to favour confirming replication studies over disconfirming studies. To read more, click here. This is the same argument that Michele Nuijten makes in her guest blog for TRN, which you can read here.  
Whether this is a real concern depends on the replication policies at journals.  At least two economics journals have publication policies that explicitly state they are neutral towards the conclusion of replication studies.  In their “Call for Replication Studies”, Burmann et al. state: “Public Finance Review will publish all … kinds of replication studies, those that validate and those that invalidate previous research” (see here).  And the journal Economics: The Open-Access, Open-Assessment E-Journal states: “The journal will publish both confirmations and disconfirmations of original studies. The only consideration will be quality of the replicating study” (see here).
Further, in their recent study, “Replications in Economics: A Progress Report” (see here), Duvendack et al. find that most published replication studies in economics disconfirm the original research.  So while it is possible that replications could make things worse, perhaps this is more a worry in theory than in practice.  At least in economics.

 

A Replication Crisis in Cancer Research?

[From the article “Cancer Research is Broken” in Slate]  “The deeper problem is that much of cancer research in the lab—maybe even most of it—simply can’t be trusted. The data are corrupt. The findings are unstable. The science doesn’t work.  In other words, we face a replication crisis in the field of biomedicine, not unlike the one we’ve seen in psychology but with far more dire implications.”  To read more, click here.

Publication Bias in Action: The Case of Oxytocin and Trust

[From the article, “How scientists fell in and out of love with the hormone oxytocin” in Vox:Science & Health]  This article recounts how initial laboratory research showing the hormone oxytocin induced trust between people eventually was demonstrated to be mostly Type I error.  In this case, it was the original research team (lead by psychologist ANTHONY LANE) who came to realize that that their own lab had focused on statistically significant findings while ignoring insignificant experiments.  Of particular interest is when Lane and colleagues tried to publish their revised, null results, they found that journals were unreceptive to the new, less sensational evidence.  To read more, click here.  To read a previous, related post from The Replication Network, click here.

IN THE NEWS: Slate (April 15, 2016)

(FROM THE ARTICLE “The Reproducibility Crisis Is Good for Science”) The author, an editor at Nature, reports on ways the reproducibility crisis is promoting change in science. An excerpt: “For what it’s worth, articles about confirmation bias and the misuse of p-values are consistently among Nature’s most-read stories. Opportunities to get credit for careful work that does not yield a flashy, new result are also expanding. In the past year, journals as diverse as the American Journal of Gastroenterology and Scientific Data actively solicited replication studies or negative results. Information Systems, a data science journal, has introduced a new type of article in which independent experts are explicitly invited to verify work from a previous publication. Last November, the U.K.’s Royal Society introduced a system known as registered reports: The decision to publish is made before results are obtained based on a pre-specified plan to address an experimental question. The F1000 Preclinical and Reproducibility Channel, launched in February, aims to give drug companies an easy way to show which scientific papers promising new paths to drugs might not deliver.”  To read more, click here.