Collecting and organising data

In the social sciences, researchers frequently work with data that vary widely in both type and source. To keep datasets manageable, a number of good practices have emerged concerning the selection of appropriate file formats, the creation of a logical folder structure, and the use of consistent naming conventions for files and variables.

Survey data collection platform that the students and staff can access using their University of Tartu user account

Learn more

Provides practical guidance on how to organise and document research data properly to ensure that data remains of high quality and usable both during the research project and after its completion

Learn more

Type of dataClosed formatRecommended open format
Text documentsWord (.docx), Rich Text Format (.rtf)PDF/A (.pdf), ODT (.odt), Unicode text (.txt), Markdown (.md)
Tabular dataExcel (.xlsx)CSV (.csv)
DatabasesAccess (.mdb, .accdb)SQL (.sql)
Statistical analysisSPSS (.sav), STATA (.dta)CSV (.csv), R (.R, .Rdata), Python (.py)
Qualitative analysisNVivo (.nvp), ATLAS.ti (.atlproj)REFI-QDA (.qdpx)
ImagesAdobe Photoshop (.psd)PNG (.png), JPEG (.jpg), TIFF (.tiff)
Video recordingsAVI (.avi), QuickTime (.mov)MPEG-4 (.mp4)
Audio recordingsWindows Media Audio (.wma)FLAC (.flac), WAV (.wav), MP3 (.mp3)

Metadata

Metadata means information that describes a dataset: what the data contains, how it was collected, under what conditions it can be accessed, and to which research project it relates. Metadata are recorded using widely adopted machine-readable standards that enable datasets to be indexed in catalogues and search engines, making them discoverable and visible to potential users. In the social sciences, the most widely used international standard is the Data Documentation Initiative (DDI).

Researchers do not need to create complex machine-readable metadata records manually. Research data repositories typically collect metadata via structured online forms and ensure the information is formatted according to relevant standards. If researchers already know which repository they intend to use for archiving their data, it is advisable to familiarise themselves with that repository’s requirements and capabilities at an early stage. Doing so makes it possible to describe data consistently throughout the research project in a format that aligns with the repository’s requirements, so the final archiving process becomes much more straightforward, as there is no need to reconstruct the precise origin or characteristics of the data retrospectively.

Documentation

In addition to metadata, it is advisable to prepare a simple README file (for example, with a .txt or .md extension) to provide context for the data. This file should include:

  • the project title, authors, funder and contact person;
  • the file structure and the purpose of each file;
  • methodological information, incl. the data collection method and period, sampling strategy, and any data cleaning or analytical procedures that were carried out;
  • conditions governing access to, citation of and use of the data, if these are not already covered by the metadata.

Besides documenting the project itself, it is also important to provide a detailed description of the nature and structure of the data. Survey datasets should be accompanied by a codebook that explains each variable, data types, units of measurement, and response categories. For qualitative data, a coding scheme serves a similar purpose by documenting the structure used to analyse texts or interviews. Both codebooks and coding schemes may be prepared as text documents; however, a tabular format (.csv) is generally preferable, as it facilitates data reuse.