Open research enables all aspects of the research cycle to be shared freely for others to reuse. Ben Catt talks about the rise of open research practice during the Covid-19 pandemic, and recent initiatives for open research at York.
Showing posts with label Open data. Show all posts
Showing posts with label Open data. Show all posts
Tuesday, 5 January 2021
Friday, 30 November 2018
Open Data in Practice: success stories and cautionary tales
Open Data in Practice is a series of events that provide researchers and those who work with them an opportunity to share their experiences of data management and open data, including the opportunities it creates and the challenges it presents.
We held the first, of what we hope will be many more, Open Data in Practice event on Thursday 15 November 2018. On the day, staff from different departments shared their data stories including their success with open data, insights into project managing research data and the development of open research initiatives.
Aidan Horner (Psychology): 'Psychology's open science working group'
The event was opened by Aidan Horner, one of our lecturers in the Department of Psychology, who spoke eloquently about what open science is and why we should care about it. Aidan came along to talk about Psychology's Open Science Interest Group, a group that discusses and shares best practice in open science and provides support for those outside of the group who wish to engage in open science. Aidan went on to give a valuable insight into his own open research practice of sharing data, sharing code, using the Open Science Framework to share project information and sharing preprints.Fleur Hughes (Social Policy and Social Work): 'Data Management in the Welfare Conditionality Research Project'
Fleur Hughes, project manager for the Welfare Conditionality research project, gave those who attended an appreciation of what is it like to manage and also prepare data for archiving for a large and complex project. This collaborative project involving researchers and PhD students from six universities, required planning to achieve its goal to share and archive the valuable longitudinal research data it generated. Fleur spoke about the sensitivities of the data collected and the decision to archive the data with the Timescapes Archive, a specialist resource of qualitative longitudinal research data “which serves as a safe place for primary researchers to store large volumes of data for ongoing use”.Cylcia Bolibaugh (Education): 'Reproducibility, open data, & GDPR'
Cylcia Bolibaugh of the Centre for Research in Language Learning and Use in the Department of Education was next to take the floor. Cylcia spoke briefly about Education Researchers for Open Science, an open science working group within the department, and then went on to talk about her concerns and the difficulties encountered in defining personal data, with anonymisation and sharing.Kevin Cowtan (Chemistry): 'Open data and the scientific gift culture'
Last but by no means least, Kevin Cowan gave what one attendee described as a “really inspirational” talk on the significant benefits he has gained from openly sharing his research data. Kevin is an interdisciplinary data scientist working in the fields of X-ray crystallography and climate science. If Kevin’s slides whet your appetite why not read his blog post on the value of open data for scientific research.Questions
In addition to questions about data management (e.g. recording datasets in PURE, restricting access to data), a number of questions were asked about preprints, for example: when can you or can’t you post a preprint; how are DOIs for preprints reconciled with DOIs then assigned to published versions; what are the benefits?For more information see:
- Crossref: Posted content (includes preprints) with information on associating posted content with published content (AM / VOR)
- To check if a journal/publisher allows preprints, the best source of information is always the journal/publisher website (e.g. Wiley’s Preprints Policy). You can also search SHERPA/RoMEO or a crowd-sourced list of journal policies on preprints on Wikipedia.
- A search in a web browser on the benefits of preprints returns many results, e.g. FOSTER Sharing Preprints, PLOS Why choose preprints?, Centre for Open Science Preprints: The What, The Why, The How. Perform your own search for information relevant to you and your research discipline.
Want to join in future conversations?
You can attend future Open Data in Practice events and benefit from your colleagues’ experiences, or come and present your own experiences. We welcome talks and input from early career researchers as well as from more experienced academics or research support staff; research students are welcome to attend. Speaker slots are available for our next Open Data in Practice event so please get in touch. Your talk should not be longer than 20 minutes.
If you have any questions about Open Data in Practice, contact the Library’s Research Support Team. See our web pages for guidance on: Research Data Management and Open Access.
Tuesday, 23 October 2018
Open data and the scientific gift culture
Continuing our theme for International Open Access Week, Professor Kevin Cowtan, Department of Chemistry, writes about the value of open data for scientific research.
If you've applied for a research council grant recently, you'll know that research councils have become rather keen on 'open data' in recent years. Funders would like us, not just to produce new results, but also to provide all the data used in deriving those results. Many journals are introducing similar requirements.
At first glance this might appear as research funders imposing more bureaucracy on grant holders. However I would like to suggest that open data is fundamental to how science works, and in addition that releasing research data can provide significant benefits to the researcher themself.
All science involves building on the work of others, or 'standing on the shoulders of giants'. This makes science a gift culture - we take the gift of the work of others and in turn gift our own work to others for them to build on. Making our results available sooner increases the opportunities for others to build on them, or if necessary to point out our errors, both of which increase human knowledge. Releasing our data often increases the value of our work, because other researchers can test our hypotheses and others against the data. In open source software, these benefits are characterized by the slogans 'release early, release often', and 'given enough eyeballs, all bugs are shallow'.
Or that is what is supposed to happen. But does it work in practice? I would like to highlight three experiences from my own career which suggest that it does.
Example 1: In the 1990s Dr Paul Emsley and myself developed a new piece of software for X-ray crystallography, called 'Coot'. University culture at the time was heavily focussed on the commercialisation of software outputs, however we (not without difficulty) made our work 'open source', meaning anyone else could build on our work, and we in turn could incorporate the work of others. This turned out to be a very good decision: Coot quickly surpassed and largely replaced all competing tools, and for the past few years the software has typically been cited in around 10 new peer-reviewed papers every day. The use of the software in industry as well as in academia produces an economic impact.
Example 2: Around 2013 I became interested in climate science, and identified a problem with how a major historical temperature dataset was being used. Users assumed that the data were global in coverage, when in fact they were not. I published a paper on estimating an unbiased global mean from the incomplete data, but also released the data and monthly updates from then on. The dataset has attracted over 200 citations and been used in official reports from government organizations. The name recognition this has generated has made it easy for me to build collaborations with climate scientists - which is not always easy when starting in a new field.
Example 3: In 2015 I identified a problem in how climate model simulations are compared with observations - the most commonly used method did not provide an 'apples to apples' comparison because of complexities of the historical data. A correct comparison involved some dull but careful data analysis. Again, I released the software as well as the data. Several subsequent comparisons have made use of this code, leading to both citations and co-authorships, at least one of which will be REF returnable.
| Image courtesy of XKCD, https://xkcd.com/1827 under a CC BY-NC 2.5 licence |
Now, this may all have been luck. After all, had I not had success in releasing data and computer code, I would not have been asked to write this blog post. There could be hundreds of people releasing data and not seeing any benefits. I could be the beneficiary of 'survivorship bias', explained by Randall Munroe in the comic XKCD.
However there are objective reasons to believe that releasing data does benefit the researcher. In 2013, Piwowar and Vision found that after controlling for a range of other factors, papers with open data received more citations than papers without open data. Open data also provides economic impact, estimated for example by Houghton and Gruen in 2014, which when measurable may be useful to the department and the researcher for REF "impact" studies.
In summary, open data is a natural extension of the principles of good scientific research: science is and has always been a social activity, and the gifting of information is fundamental to that activity. Studies of open data publications show benefits both to the researcher and to the wider economy. My own research career has been built on giving away data and computer code: not every case has led to benefits, but the net benefit over the course of my career has far outweighed the time cost of releasing the data.
Professor Cowtan is an interdisciplinary data scientist working in the fields of X-ray crystallography and climate science. While most of his career has been at the University of York, he has also spent sabbaticals at San Diego Supercomputer Centre. He is the chair of the university Research Data Management committee.
Subscribe to:
Posts (Atom)

