Publication and archiving of research data

Sharing research data enhances knowledge transfer and increases the societal benefits of research. Publishing data is also advantageous for researchers, as it supports the reproducibility, transparency and scientific validity of their projects. Published research data encourages others to analyse it from new perspectives, providing researchers with valuable visibility, citations, and opportunities for collaboration.

It is often assumed that data collected during a project is too specific and therefore unlikely to be reused. In reality, precisely such data may later prove highly valuable, serving as comparative data or even forming the basis for future studies, or inspiring someone who identifies potential applications and opportunities for innovation within the dataset. If there are no restrictions preventing the publication of the data, it should be made available.

University of Tartu's research data repository

Learn more

Free online platform that supports the management of research data throughout its entire lifecycle, from data collection to publication

Learn more

Which data should be published?

For publication and long-term archiving, it is advisable to select data that is necessary for replicating the research project and that will retain its value after the completion of the study. Particularly valuable is data that would be costly, difficult or impossible to collect again, such as data from long-term cohort studies or interviews with participants who may no longer be accessible to researchers in the future.

If the data contains personal information, this does not automatically mean that it cannot be published in any form. In many cases, data can be anonymised or pseudonymised so that it can be shared under certain conditions. It is worth considering whether this is feasible and whether the value of the data would be preserved in such a form. Properly anonymised data can be published with open access, which is preferable to leaving the data unpublished.

Anonymisation does not always mean simply removing names or identifiers. Data may contain several types of identifying information.

  • Direct and indirect identifiers. Direct identifiers (name, personal identification code, email address, phone number) must be removed. With indirect identifiers, identification may be possible through a combination of characteristics, for example, a rare occupation combined with advanced age. In such cases, the relevant characteristics should be grouped into broader categories: age can be replaced with an age group, and occupation with a field of activity.
  • Responses to open questions. Free-text responses often contain unintentionally identifiable information. Respondents may mention their workplace, place of residence, or a life event that makes them recognisable. Such responses should be reviewed and, where necessary, edited or removed.

Good to know

When anonymising data tables, the Amnesia tool developed by OpenAIRE can be helpful. Available for Windows and Linux, it helps to remove or generalise identifying characteristics within datasets using the k-anonymity method.

Where to publish data?

For researchers of the University of Tartu Faculty of Social Sciences, DataDOI is the recommended repository. It is a dedicated research data repository managed by the University of Tartu Library and is available free of charge to researchers in Estonia. DataDOI is based on the widely recognised open-source Dataverse platform, assigns datasets a persistent identifier (DOI) through the DataCite Estonia consortium, and makes them discoverable through a range of international data catalogues.

DataDOI supports the DDI (Data Documentation Initiative) metadata standard for the social sciences. After uploading a dataset to DataDOI, researchers should additionally complete the social sciences and humanities metadata section. These fields comply with the international DDI standard for describing and documenting research data and help to improve the findability and reusability of datasets. Detailed guidance is available in the DataDOI User Guide.

Image
Screenshot. Adding social sciences metadata in DataDOI.
Screenshot. Adding social sciences metadata in DataDOI. Author: University of Tartu

International subject-specific and general-purpose data repositories are suitable for storing data from international research projects, or where a funder or publisher requires the use of a preferred repository. Useful resources for selecting an appropriate repository include the Re3data.org registry of research data repositories and the General Repository Comparison platform.

Researchers should also explore the list of partner archives within CESSDA (Consortium of European Social Science Data Archives). These repositories provide support for social science metadata standards and ensure that datasets are findable through international data catalogues such as the CESSDA Data Catalogue and DataCite Commons.

  • Zenodo – funded by the European Union, up to 50 GB free. The default recommended repository for projects funded under Horizon Europe
  • Figshare
  • Open Science Framework (OSF) – suitable for managing research data throughout the entire project lifecycle, from data collection to publication

Please note

Many academic journals, particularly high-level ones, have their own data policies that may specify requirements regarding the repository used, licensing arrangements, or the accessibility of underlying data. It is therefore advisable to review a journal's requirements before submitting a manuscript. Increasingly, journals also require a data availability statement that briefly explains where the data underlying the article can be found and under what conditions it may be accessed. Such a statement will normally include a reference to the data repository and the dataset’s persistent identifier (DOI).