Metadata must identify the object without the paper
A filename such as results_final.csv cannot tell a future reader who created the records, what population they represent or which release was analysed. Metadata creates that context in structured form. It supports discovery and citation, but also the earlier decision about whether the dataset is relevant, intelligible and accessible.
This minimum is deliberately cross-disciplinary. A language corpus, genomic sequence collection, survey and satellite product each need richer specialist fields. Use the core as a floor, then add the repository schema and accepted disciplinary vocabulary. Minimal should mean sufficient for an honest first assessment, not the fewest boxes a deposit form will accept.
Completed fictional dataset record
The following record is invented to demonstrate precision. Creators: Learning Routes Research Group, Example University; individual ORCID identifiers would appear in structured creator fields. Title: “Anonymised weekly records of self-directed study, North Campus pilot, spring 2026”.
Description: “The package contains 84 synthetic weekly logs created for teaching, with time spent, work-setting category and self-rating. It contains no observations of real participants and must not be interpreted as an empirical study.” Temporal coverage: simulated weeks from 2 March to 24 May 2026. Spatial coverage: fictional North Campus, not a real location.
Method: rule-based generation following method-v1.2.pdf; seed and code preserved in a linked software release. Files: learning_logs.csv, UTF-8, 84 rows; codebook.csv; README.txt. Version: 1.0.0, issued 4 September 2026. Identifier: no real DOI; 10.0000/fictional.learning.v1 is deliberately non-resolving example text.
Rights and access: public fictional teaching data; an actual deposit would require a separate licence decision. Relationships: “IsDocumentedBy” the methods protocol, “IsSupplementTo” the teaching article and “IsVersionOf” the concept dataset. Contact: an institutional role address for the fictional data team. Nothing in this example represents a registered study, repository record or real university.
Give titles, descriptions and keywords different jobs
The title identifies the object succinctly. The description explains purpose, construction and limitations in full sentences. Keywords connect the record with disciplinary search terms and controlled vocabularies. Repeating one sentence in all three fields wastes structure and still leaves the dataset difficult to interpret.
Use a recognised subject vocabulary when the repository supports one, adding plain-language synonyms where researchers may search differently. Expand acronyms at least once. Connect people, organisations and places to persistent identifiers where the schema, evidence and privacy conditions permit.
Put structured values in structured fields. Enter a date as a date, a licence through its standard identifier and a relationship through a controlled relation type. A PDF README remains valuable, but it should not be the only location for information that the catalogue can expose consistently.
Audit the metadata against the deposited files
Open the release in a clean folder. Compare filenames, formats, row counts, variables, units and date ranges with the record and codebook. Confirm that every released object appears in the inventory and that excluded or restricted material is not accidentally described as public. A complete-looking form can still refer to the wrong package.
- Match creator names and title to the preferred citation.
- Compare version and issue date with the released files.
- Resolve the identifier and inspect the public landing page.
- Test typed links to software, publications and earlier versions.
- Read rights and access statements from an external user's perspective.
- Replace unexplained abbreviations and private working filenames.
- Ask a disciplinary reader to describe the dataset using only the record.
Maintain metadata as a versioned research output
A material data correction needs a new version and a concise change note. State whether files were added, values corrected, labels changed or only descriptions improved. Keep a route to earlier cited releases and express their relation to the new version according to the repository's policy.
Review the record whenever a version is issued, a contact role changes or access conditions are updated. Broken contacts and stale relationships reduce reuse even when files remain online. Export the released metadata and store it beside the deposit package so the project retains evidence of what was published.
The repository selection guide tests whether a service can represent the required fields. The data citation guide draws a reference from the verified core. The research data hub connects this record to versioning and preservation, while the knowledge guide supplies the single wider foundation link.
Release checklist
- Every mandatory field is specific or explicitly not applicable.
- The dataset is broadly understandable without its companion article.
- Metadata, version and files describe the same release.
- Methods, codebook and README are linked unambiguously.
- Licence and access status are treated as separate statements.
- Relations identify both object and direction.
- Identifier and contact work for an external reader.
