Skip to content
PlagiatScanner.de
Research data · procedure · DAT-09

Minimum metadata for research data

In brief: A useful minimum record identifies creators, title, plain-language description, temporal and spatial coverage, data types and formats, collection and processing methods, version, publication date, persistent identifier, repository, rights, access, related objects and a durable contact. Complete each field specifically or mark it as not applicable with a reason. A long abstract does not repair missing structured fields.

Field list with completed dataset exampleSources reviewed 4 September 2026

Metadata must identify the object without the paper

A filename such as results_final.csv cannot tell a future reader who created the records, what population they represent or which release was analysed. Metadata creates that context in structured form. It supports discovery and citation, but also the earlier decision about whether the dataset is relevant, intelligible and accessible.

This minimum is deliberately cross-disciplinary. A language corpus, genomic sequence collection, survey and satellite product each need richer specialist fields. Use the core as a floor, then add the repository schema and accepted disciplinary vocabulary. Minimal should mean sufficient for an honest first assessment, not the fewest boxes a deposit form will accept.

Authority asset: minimum field list

Core fields and a test for each entry
FieldRecordQuality test
CreatorsPeople or organisations with consistent names and identifiersReflect responsibility for the data object
TitleSpecific name for this datasetIncludes subject, place and period when material
DescriptionPurpose, contents, unit and major limitsUnderstandable without reading an article
CoverageCollection dates, represented period and geographyUses unambiguous dates and place terms
MethodsCollection, selection, instruments and processingLinks to a protocol or methods object
Files and formatsInventory, variables, units, open or proprietary formatsMatches the README and codebook
Version and dateReleased state and publication dateNo silent replacement of the cited object
Identifier and repositoryDOI or other persistent identifier and publisherResolves to a landing page
Rights and accessRights holder, licence, embargo or controlled routeDoes not promise access that is unavailable
RelationshipsArticle, code, prior version or source datasetNames the relationship direction
ContactDurable role or institutional addressNot solely a temporary personal account

Completed fictional dataset record

The following record is invented to demonstrate precision. Creators: Learning Routes Research Group, Example University; individual ORCID identifiers would appear in structured creator fields. Title: “Anonymised weekly records of self-directed study, North Campus pilot, spring 2026”.

Description: “The package contains 84 synthetic weekly logs created for teaching, with time spent, work-setting category and self-rating. It contains no observations of real participants and must not be interpreted as an empirical study.” Temporal coverage: simulated weeks from 2 March to 24 May 2026. Spatial coverage: fictional North Campus, not a real location.

Method: rule-based generation following method-v1.2.pdf; seed and code preserved in a linked software release. Files: learning_logs.csv, UTF-8, 84 rows; codebook.csv; README.txt. Version: 1.0.0, issued 4 September 2026. Identifier: no real DOI; 10.0000/fictional.learning.v1 is deliberately non-resolving example text.

Rights and access: public fictional teaching data; an actual deposit would require a separate licence decision. Relationships: “IsDocumentedBy” the methods protocol, “IsSupplementTo” the teaching article and “IsVersionOf” the concept dataset. Contact: an institutional role address for the fictional data team. Nothing in this example represents a registered study, repository record or real university.

Give titles, descriptions and keywords different jobs

The title identifies the object succinctly. The description explains purpose, construction and limitations in full sentences. Keywords connect the record with disciplinary search terms and controlled vocabularies. Repeating one sentence in all three fields wastes structure and still leaves the dataset difficult to interpret.

Use a recognised subject vocabulary when the repository supports one, adding plain-language synonyms where researchers may search differently. Expand acronyms at least once. Connect people, organisations and places to persistent identifiers where the schema, evidence and privacy conditions permit.

Put structured values in structured fields. Enter a date as a date, a licence through its standard identifier and a relationship through a controlled relation type. A PDF README remains valuable, but it should not be the only location for information that the catalogue can expose consistently.

Audit the metadata against the deposited files

Open the release in a clean folder. Compare filenames, formats, row counts, variables, units and date ranges with the record and codebook. Confirm that every released object appears in the inventory and that excluded or restricted material is not accidentally described as public. A complete-looking form can still refer to the wrong package.

  1. Match creator names and title to the preferred citation.
  2. Compare version and issue date with the released files.
  3. Resolve the identifier and inspect the public landing page.
  4. Test typed links to software, publications and earlier versions.
  5. Read rights and access statements from an external user's perspective.
  6. Replace unexplained abbreviations and private working filenames.
  7. Ask a disciplinary reader to describe the dataset using only the record.

Maintain metadata as a versioned research output

A material data correction needs a new version and a concise change note. State whether files were added, values corrected, labels changed or only descriptions improved. Keep a route to earlier cited releases and express their relation to the new version according to the repository's policy.

Review the record whenever a version is issued, a contact role changes or access conditions are updated. Broken contacts and stale relationships reduce reuse even when files remain online. Export the released metadata and store it beside the deposit package so the project retains evidence of what was published.

The repository selection guide tests whether a service can represent the required fields. The data citation guide draws a reference from the verified core. The research data hub connects this record to versioning and preservation, while the knowledge guide supplies the single wider foundation link.

Release checklist

  • Every mandatory field is specific or explicitly not applicable.
  • The dataset is broadly understandable without its companion article.
  • Metadata, version and files describe the same release.
  • Methods, codebook and README are linked unambiguously.
  • Licence and access status are treated as separate statements.
  • Relations identify both object and direction.
  • Identifier and contact work for an external reader.

Sources and metadata standards

  1. DataCite Metadata Schema – properties for identity, responsibility, versions, rights and relationships.
  2. DCMI Metadata Terms – general vocabulary for describing digital resources.
  3. GO FAIR, FAIR Principles – machine-readable findability, accessibility, interoperability and reuse.
  4. German Research Foundation, Handling of Research Data – responsible documentation and provision.

Sources reviewed 4 September 2026. Add every discipline-specific and repository-mandated field to this baseline.