Chapter 4 — Data Stewardship, Privacy and Ethics
The chapter with the deepest data-standard obligation of the whole Guidebook. Provenance, lineage, purpose limitation, retention, sensitivity classification and machine-readable metadata are exactly what an ontology is for. It is also the chapter that raises one of SPDTF's most genuinely new modelling questions: AI-decision disclosure.
Its through-line: data stewardship is economic infrastructure. Governance choices about access, interoperability and standards shape market contestability, not just compliance.
Source · derived · insight · profiling
The chapter requires schemes to explicitly distinguish four categories of data:
| Category | What it is |
|---|---|
| Source data | Provided by data holders. |
| Derived data | Transformed or enriched. |
| Insights | Analytical outputs or recommendations. |
| Profiling | Inferences about individuals or behaviours. |
Because decisions increasingly rest on derived or inferred data that is "less visible to users and more difficult to interpret", the chapter demands: users are informed when outcomes are based on derived insights; responsibility for derived outputs is clearly assigned; and high-impact decisions from derived data remain explainable and contestable.
This is the chapter's most consequential implication for OPDA. A property pack assembled to
the PDTF schema is an aggregation of claims from many sources. If the pack is wrong but
every source claim was right, who is accountable? The separate schema-derived ontology models
source claims well — prov:wasDerivedFrom is SHACL-mandatory on every
opda:Claim — and derived claims not at all.
Gap SD4.
Provenance and lineage
"Schemes should ensure the origin and lineage of data can be verified, particularly in high-value or legally sensitive use cases."
Accountability must remain "traceable across the full data journey" and must not fragment as data moves between actors. Responsibility transfers with the data (Clause I) — at each stage the receiving party assumes responsibility for lawful, secure and ethical use, and scheme rules continue to apply downstream (Clause H).
One honest caveat the chapter itself supplies: "Full traceability may not always be feasible beyond initial sharing points." That is a get-out clause which lowers the bar — but it also invites a weaker downstream standard. The schema-derived ontology demonstrates that end-to-end lineage is feasible, giving OPDA evidence for a stronger position.
The three-level transparency model
| Level | Users should understand |
|---|---|
| Participant | Who is accessing their data, and the role of each participant. |
| Data | What data is being used, and whether it is original or derived. |
| Decision | How decisions are generated, and whether automation or AI has been used. |
Plus a user-visible "data receipt" — metadata and logs confirming transactions, with visibility of the actors involved.
AI governance — the genuinely new obligation
"AI governance is becoming inseparable from data governance" (one of the chapter's five findings). Concretely it requires: proportionate risk-based AI governance with human oversight; limits on fully automated decision-making in high-impact contexts; enhanced audit for critical systems; and — the sharp one — that users understand whether automation or AI has been used to generate a decision affecting them.
Neither the PDTF schema nor the schema-derived ontology represents this; SPDTF has not yet planned for it. It needs a decision/outcome artefact, an AI-involvement flag, the model that produced it, a human-oversight level, and the inputs it consumed tagged source-vs-derived. DPV-AI is a near-exact upstream fit and the schema-derived ontology has adopted none of it.
It is hard not because the vocabulary is missing but because the ontology today models property facts and nothing else — it has no concept of an actor's output. That is a scope extension, and a council question. Gap SD6.
Machine-readable and human-readable
"Machine-readable to support interoperability, automation and reuse… Human-readable to support transparency, user understanding and trust."
Applied to four artefact classes: consent records · portability outputs · trust signals and audit information · decision explanations. The human-readable half is not UI copy — it must be carried alongside the machine encoding. In RDF terms, labels and definitions are a conformance requirement, not a documentation nicety.
The schema-derived ontology supports this structurally through SKOS labels and OWL annotations. But "we can label everything" is not "everything is labelled" — ship a SHACL shape requiring it and the soft claim becomes auditable (gap SD11, a day's work).
What this could mean for SPDTF
Full table on the overlap page. The honest scorecard:
What the existing evidence already supports
- Machine-readability. Clause D and Clause F closely match the reason the existing PDTF schema was created.
- Claim-level provenance. Real in the separate schema-derived ontology and SHACL-enforced there.
- Data quality. The chapter asks for "responsibility for data quality"; the separate framework produced alongside the PDTF schema already has six dimensions.
- Temporal state and long-lived validity — the chapter's property paragraph specifically calls out "version control" and reuse over extended periods.
What SPDTF must not claim
Model provenance is not instance provenance. The schema-derived ontology's dct:source chain
from every ontology term back to its JSON Schema leaf is genuine provenance of the model.
Citing it as evidence of the instance-data lineage this chapter asks for would be a
material overclaim that would not survive anyone actually reading the ontology. (The real evidence
is the mandatory prov:wasDerivedFrom on opda:Claim — cite that instead.)
The tension to expect
§3.1.7 says derived data "may represent proprietary analysis or intellectual property" and "not all derived outputs can or should be treated identically to source data" — yet lineage must be verifiable and users must know when outcomes rest on derived data. A standard that makes derivation edges mandatory and public will meet resistance from members whose derived products are their margin. The resolution is probably lineage is asserted, but the derivation logic is not disclosed — and the chapter's own words support exactly that compromise: "explainability should prioritise clarity of outcomes rather than full technical transparency". Quote it.
Caveats
- It frames property as the laggard. "Low interoperability and fragmented systems, compared to more mature Smart Data sectors such as Open Banking." Flattering elsewhere ("focus on provenance and data governance") but OPDA should decide whether to accept that framing or contest it with evidence — the ontology, the schema and the quality framework are all real, and more than most sectors have.
- The risk model is use-case-shaped; the existing evidence is dataset-shaped. The chapter insists risk is assessed "at the use case and ecosystem level". Neither the PDTF schema nor its derived ontology has a
UseCaseconcept. SPDTF must either add the dimension or scope out of use-case risk explicitly — drifting between them is not an option (gap SD7). - Watch footnote 9. The only concrete machine-readable data-documentation standard the chapter cites is MLCommons Croissant — an ML dataset-documentation format, not a semantic-web stack. If DBT converges on Croissant as the reference, SPDTF's proposed DCAT/DPV/PROV approach becomes an alternative that must be argued for rather than an obvious fit. Worth an early intervention.
- "Trust signals" is undefined. The chapter demands "standardised metadata and trust signals" without saying what a trust signal is or who defines the vocabulary. That is a standards-shaped hole, and OPDA is well placed to fill it — if someone claims it before a scheme operator invents a bespoke one.
Related
- Gap register — SD3, SD4, SD6, SD7 all originate here
- Data quality framework · Ontology — provenance
- Ch.3 — User lifecycle — the consent half of the same fabric
Comments
Loading comments…
Sign in to post a comment