Gap register

    The asks in the overlap analysis that the existing PDTF schema and its separate schema-derived ontology cannot currently encode, and that SPDTF may need to model. Ontology claims were verified against the committed TTL corpus, not its layer plan. Ranked by effort to close — noting that the highest-value gap (SD4) is mid-difficulty rather than hard, and the highest-leverage one (SD3) is genuinely easy.

    Scope note: these are capability gaps, not compliance failures

    Smart Data and SPDTF are two different initiatives. DBT's Smart Data is a cross-sector programme under the Data (Use and Access) Act. The PDTF schema is the existing Digital Property Pack schema. The ontology extracted from it is separate evidence. SPDTF is the first scheme draft being authored collaboratively across industry and stakeholders. Neither initiative is bound by the Guidebook, and the existing artefacts cannot "conform" to it or "fail" it.

    The relationship is conditional: if and when property becomes a designated Smart Data scheme, SPDTF participants would need to decide how to model the things listed below, using both existing evidence sets. Everything here is therefore a statement about capability, never about non-compliance. The single exception is SD8, which is a defect in OPDA's own published claim and would be wrong even if the Guidebook did not exist.

    This is analysis, not the deferred-work register

    OPDA has one canonical register of committed work: ADR-0005, mirrored at /governance/deferred-work. This page is not a second one, and its SDn items are deliberately numbered to avoid colliding with ADR-0005's A–G series.

    These are engineering judgements from reading external draft guidance against the corpus. They become deferred work only when a council or WG adopts them — at which point the item moves to ADR-0005 and this row should link to it. Anything that changes the ontology's shape — particularly SD7 (a UseCase class), SD6 (modelling decisions and outcomes) and SD13 (reinstating an assurance ladder a council removed) — is a council question, not a unilateral one.

    The register

    # Gap Effort Why it matters
    SD1 Retention & disposition terms. No way to say how long data is kept, or whether it is deleted or anonymised when its purpose ends. 🟢 Very easy Closes Ch.4 Clause C outright. DPV already has the terms (StorageDuration, Erase vs Anonymise); binding them is a shape and a handful of triples.
    SD2 Lifecycle stages with a responsible party. No vocabulary for collection → processing → sharing → retention → deletion, and no hasResponsibleParty per stage. 🟢 Easy Closes Ch.4 §3.1.4 and Clause I ("responsibility transfers with the data"). DPV's processing terms map 1:1 onto the five stages; PROV supplies the actor attribution.
    SD3 The origin facet: source / derived / insight / profiling. Neither the PDTF schema nor the schema-derived ontology can say whether a datum was asserted or computed; SPDTF participants must decide the facet. 🟢 Easy — do this first Cheap and unlocks three other gaps. Per house doctrine it ships as a facet (sh:in over a four-term SKOS scheme), not a subclass tree. It is the precondition for SD4 and SD6, and it delivers Ch.3's user-facing "verified / self-reported / derived" signal and Ch.4's transparency data level. Add the individual-linkable vs anonymised flag alongside it — that flag drives Ch.4's entire proportionality carve-out.
    SD4 Derived-value provenance. A property pack is an aggregation of claims from many sources. If the pack is wrong but every source claim was right — who is accountable? The schema-derived ontology models source claims well and derived claims not at all. 🟡 Moderate — highest value OPDA's largest modelling gap and its largest liability gap are the same gap. The Ch.2 scoping paper says liability is hardest "where value is created through aggregation, transformation, or derived outputs" — and names property packs as its first test case. PROV is already adopted and SHACL-enforced on source claims; the work is choosing the granularity of derivation, requiring it on anything flagged "derived" (SD3), and not confusing instance lineage with the existing model provenance chain.
    SD5 Consent, delegation and the trust-signal cluster. No consent record, no delegated authority, no accreditation status, no audit event. Four missing artefacts that the Guidebook treats as one fabric. 🟡 Moderate — largest hole Chapter 1's data-holder role is defined by relying on a consent signal it must not re-interpret; neither existing evidence set can emit that signal. Chapter 3's dashboard requires active/expired/withdrawn permissions with history. Property is the delegation-heavy sector — LPA, probate, executors, divorce, co-ownership — and none of it is modelled. Needs ODRL (currently zero occurrences in the ontology corpus). Full structure: the consent-record model Chapter 3 implies for SPDTF.
    SD6 AI-decision disclosure. "Whether automation or AI has been used" — and what it did to the outcome. 🔴 Moderate-to-hard — genuinely new Hard not because the vocabulary is missing (DPV-AI is a near-exact fit and is completely unadopted) but because it would require SPDTF to model decisions and outcomes — a new domain concept. The schema-derived ontology has no notion of an actor's output at all, only of property facts. This is a scope extension and a council question.
    SD7 Use-case as a first-class object. The Guidebook insists risk is assessed "at the use case and ecosystem level" — data combination, re-identification, inference, cumulative exposure. Neither the PDTF schema nor its derived ontology has a UseCase concept or risk-tier vocabulary. 🔴 Hardest — structural The deepest mismatch between the Guidebook's mental model and the existing evidence. Both existing layers are property/dataset-centric; risk attaches to data, not to a use of data. SPDTF must either add the dimension through a council decision or explicitly leave use-case risk to scheme operators. Both are defensible. Drifting between them is not.
    SD8 Selective disclosure — and a false claim to retract. The proof suite is Ed25519Signature2020 Linked Data Proofs (MUST, per the access specification), which does not support selective disclosure. 🟡 Moderate — retraction is urgent The only item here that is a defect in its own right, independent of anything DBT says. trust-framework/docs/governance.md asserts "Data minimization and selective disclosure enforced" — a claim OPDA publishes and cannot honour with an Ed25519-only proof suite. That is OPDA failing OPDA's own stated capability, and it would be wrong whether or not the Smart Data Guidebook existed. Two actions, and the first is not optional: (1) withdraw that sentence; (2) decide whether to adopt a BBS or SD-JWT proof suite — a trust-framework decision, not just an ontology one. (The Guidebook's Principle 2 and Example D would also want selective disclosure, but that is a *separate*, conditional matter — see the scope note above.)
    SD9 Policy propagation on onward sharing. "Obligations travel with the data" (Ch.4 Clauses H & I). 🟡 Moderate Distinct from SD5: recording a share and binding the downstream policy are different asks. ODRL's nextPolicy/Duty is exactly the primitive — and is the same reason the "confidentiality survives withdrawal" requirement (Ch.3 §2.3) cannot currently be expressed. A permission-only model cannot say "you may no longer receive this, and you must still keep secret what you already received."
    SD10 Data-quality metadata. A quality framework exists as policy; there is no DQV/ISO-25012 vocabulary attaching accuracy/completeness/currency to a value. 🟢 Easy Converts an existing OPDA strength into a machine-checkable claim. We already over-deliver on the substance (six dimensions vs the Guidebook's vague ask); we just cannot yet say it in RDF.
    SD11 Label coverage as a conformance property. Ch.4 requires outputs be human-readable and machine-readable. 🟢 Trivial "We can label everything" is not the same as "everything is labelled". Ship a SHACL shape requiring skos:prefLabel on every concept and rdfs:label on every class/property, then assert coverage. A day's work that converts a soft claim into an auditable one.
    SD12 Dynamic trust signals. Participant suspension status, dynamic risk score, security incident notification, accreditation status change, credential validity — as machine-readable signals. 🟡 Moderate — live deadline This is Chapter 5's one and only data-standard ask (Position 3, "Trust should be dynamic") — and Chapter 5 is the chapter with the 24 July 2026 review deadline. Neither existing evidence set can emit these signals. Overlaps heavily with SD5's accreditation-status half, but is listed separately because it is the one thing OPDA could commit to in the Chapter 5 response.
    SD13 Assurance-level vocabulary. opda:assuranceLevel was removed on 2026-07-05 (ODR-0009 — exact source wording: "zero PDTF schema basis"). There is no assurance ladder in the schema-derived ontology corpus today. 🟡 Moderate — needs a council Chapter 2 asks for a provenance schema carrying "issuer, inputs, timestamps, assurance level, revocation checks"; Chapter 1's CBOM needs an assurance type per credential. Do not re-add it by reflex — it was removed for a good reason (no schema basis), and re-adding it to satisfy DBT would repeat the original mistake. The right move is to ask whether an evidential-provenance scale (how authoritative is this fact about this house?) is warranted on its own merits, distinct from GPG 44/45 identity assurance. If it is, the schema-derived ontology's mandatory claim provenance is the natural substrate — and no other sector in the Guidebook has one.

    Things that look like gaps and are not

    Equally important to record, so effort is not wasted defending against them or — worse — building them:

    • Liability allocation. SPDTF does not necessarily need a liableParty property. Chapter 2 allocates liability by rule, not by data field. What the scheme model must supply is the evidence the rule operates on: who issued a signal, when, at what assurance level, and whether a status check was performed at the point of reliance. Encode the evidence, not the verdict.
    • Security controls. FAPI, ISO 27001, encryption, incident response. These are scheme-operator obligations. SPDTF's modelling role is at most to encode that a participant holds a given certification — which is part of SD5, not a separate security gap.
    • Model provenance ≠ instance provenance. The schema-derived ontology's dct:source chain from every ontology term back to the JSON Schema leaf it was minted from is genuine, verifiable provenance — of the model. Claiming it satisfies Chapter 4's instance-data lineage ask would be a material overclaim that would not survive scrutiny by anyone who reads the ontology. (The good news, established while writing this register: the schema-derived ontology does have real instance-level claim provenance — prov:wasDerivedFrom is SHACL-mandatory on every opda:Claim. The strength is real; just don't cite the wrong evidence for it.)

    Suggested sequencing

    1. SD3 first (origin facet) — cheap, and SD4 and SD6 both depend on it.
    2. SD11, SD1, SD2, SD10 — a batch of easy wins that between them convert several soft conformance claims into auditable ones.
    3. SD4 (derived-value provenance) — the highest-value gap, and the one property is being watched on. Chapter 2's scoping paper has property packs as its named test case, and its stated output is template clauses. There is an unusually clear invitation here to write the property accountability map rather than receive one.
    4. SD5 (consent/delegation/trust signals) — the largest hole, and the one where the window may close. Every other sector already has an answer (Open Banking's consent dashboard, energy's RECCo Consumer Consent Solution). Property has nothing named. Either SPDTF scopes consent into its model, or another body will own it for property — and consent is where the scheme power sits.
    5. SD8's retraction is immediate and unconditional — the false selective-disclosure claim in governance.md should be withdrawn regardless of when the proof-suite work happens. The proof-suite decision itself needs the trust framework, so start that conversation early.
    6. SD12 — the only gap with a dated deadline (Chapter 5, 24 July). Even if it cannot be closed by then, OPDA's Chapter 5 response should name it rather than let the review pass silently.
    7. SD6, SD7, SD13 — council questions. Do not start unilaterally, and in SD13's case do not simply reinstate what a council previously removed.
    One of our own artefacts was wrong, independently of DBT — now fixed

    trust-framework/docs/governance.md claimed "Data minimization and selective disclosure enforced". OPDA cannot enforce selective disclosure with an Ed25519Signature2020-only proof suite. The sentence has been withdrawn and replaced with an accurate statement of what the trust framework does do (data minimisation) and what it does not yet do (selective disclosure). See SD8.

    Retraction: a defect this register wrongly reported

    An earlier draft claimed the legal-basis SHACL shape "can never pass" because it validates for dpv:hasLegalBasis while the corpus asserts opda:lawfulBasis. That was wrong — they are different layers, not a mismatch. opda:lawfulBasis is the class-level co-annotation; dpv:hasLegalBasis is the instance-level predicate, which ODR-0012 makes Phase-2 (a lawful basis is an assertion about a processing act). The shape is correct and simply not yet exercised. dpv:hasLegalBasis stands.

    Comments

    Loading comments…