Definition
The practice of publishing structured data on the web using resolvable HTTP identifiers (URIs), semantic data models (commonly RDF), and interoperable vocabularies, together with explicit interlinks to other datasets and open licensing or clear access terms, to enable machine‑readable integration, discovery and reuse across domains.
Principle
Principle
Publishing resources as URI‑identified, semantically modeled triples and linking them to external URIs enables heterogeneous datasets to be connected into an interoperable graph that supports cross‑dataset queries and semantic integration, provided identifiers, vocabularies and provenance are managed coherently.
Demonstration
Demonstration
Illustrative scenario: Situation — A cultural heritage institution publishes item metadata as RDF with resolvable URIs and links persons to external authority URIs (e.g., ORCID/VIAF‑style identifiers) and to related works in other datasets. Recognition — Consumers discover the URIs, follow links and combine triples from multiple datasets via SPARQL or federated queries. Action — A researcher issues a query that joins person URIs with related works across datasets to assemble an enriched profile. Consequence — Data from different providers integrates automatically, enabling richer discovery and analysis without prior centralized harmonization.
Misapplication
Misapplication
Equating ‘linked’ with mere hyperlinking or assuming that publishing datasets on the web without stable URIs, semantic modelling or clear reuse licenses constitutes LOD; the error is conflating superficial connectivity or availability with the technical and governance practices that produce interoperable linked graphs.
Consequence
Consequence
LOD can materially increase interoperability, discoverability and automated reuse of datasets and enable new cross‑domain applications; it also imposes obligations for identifier persistence, vocabulary management, provenance and licensing, and can surface privacy, licensing or quality issues when linking is performed without governance.
Reversal
Reversal
When data contain sensitive personal information, proprietary commercial restrictions, or strict jurisdictional privacy rules, publishing as open, linked data may be inappropriate or legally restricted; in some uses, simpler structured formats (CSV, JSON without external linking) with controlled access are preferable.
Boundary
Boundary
Clearly within — data published with resolvable HTTP URIs, a semantic model (e.g., RDF or equivalent), explicit interlinks to external URIs, and clear terms permitting reuse. Boundary case — datasets using JSON‑LD or other serializations that map to RDF but lacking persistent URI practices or clear licensing. Clearly outside — proprietary APIs or datasets behind access restrictions with no resolvable identifiers or interlinks, or simple downloadable CSVs with no semantic identifiers.
Semantic Tension
Semantic Tension
Openness/interoperability versus governance/privacy: maximizing linkability and reuse increases integration possibilities but raises risks for privacy, data quality and control, requiring governance mechanisms that can constrain openness.
Synthesis
Synthesis
Linked Open Data is both a technical stack and a publishing practice: its power derives from combining resolvable identifiers, semantic modelling and open reuse terms to assemble interoperable graphs, but realizing that power requires sustained identifier persistence, vocabulary alignment and governance of provenance and privacy.