accepted

    RML Mapping Implementation

    Context and Problem Statement

    ODR-0035 adopts RML as OPDA’s independent, bidirectional schema-provenance verification mechanism, used as documentation, not executable ETL. This ADR records the concrete engineering realising that decision: which RML engine, how the mapping is validated, where it lives in the repository, and how JSON enum values are represented as SKOS concept IRIs given RML has no built-in string-transform primitive for that.

    Three engineering questions needed resolving, each with a real, non-obvious answer surfaced during implementation:

    1. Which RML engine, and how is the mapping validated — without instance data (per ODR-0035) and without violating ADR-0037 (Apache Jena as opda’s sole RDF parse/serialise/validate/infer/query toolchain; rdflib and pyshacl are prohibited from those paths).
    2. Where does the mapping live — a tooling directory, or alongside the standards it traces.
    3. How does a JSON enum value (a raw string, often containing spaces/apostrophes, e.g. "Legal Owner") become a SKOS concept IRI (opda:scheme/sellersCapacity/Legal-Owner) in RML — R2RML mandates percent-encoding of rr:template-substituted values outside RFC 3987 iunreserved, so a naive template produces .../Legal%20Owner, never the ontology’s hyphenated slug. This is a hard spec constraint, not a style choice.

    Decision Drivers

    • Must not require or produce PDTF transaction instance data (ODR-0035).
    • Must not introduce rdflib/pyshacl into any RDF parse/validate path (ADR-0037).
    • Must keep the enum-value → SKOS-concept correspondence auditable from the mapping file alone, matching RML’s documentation purpose here (not merely “whatever is most compact”).

    Considered Options

    Engine:

    • morph-kgc (chosen initially; superseded 2026-07-04 by RMLMapper — see Amendments) — Python RML engine; supports RML-FNML (including Python UDFs), RML-star, JSON via ql:JSONPath. Confirmed installed and working.
    • RMLMapper (Java reference implementation; now the chosen engine, see Amendments) — originally rejected: introduces a JVM dependency alongside the already-adopted Jena JVM toolchain for no functional gain here; morph-kgc’s Python UDF path covers the same FNML ground if ever needed. This argument weakened once Jena was already a hard JVM dependency for SHACL/arq (ADR-0036/0037), and was overridden once morph-kgc’s own limitations became blocking (see Amendments).

    Validation (RML mapping’s own structure + the ontology’s dct:source citations, both against the canonical schema dictionary):

    • rdflib graph-walking (initial implementation) — parses the .rml.ttl and opda-merged.ttl files directly via rdflib’s Turtle parser. Rejected on discovery it violates ADR-0037’s explicit prohibition on rdflib in opda’s parse paths, even though empirically verified to produce correct results for this content (neither file uses any RDF-star/RDF-1.2-only syntax, so no actual parsing divergence was found — the objection is architectural/policy, not a demonstrated bug).
    • Apache Jena arq (chosen) — SPARQL SELECT queries against the mapping and ontology files, shelled out via the same subprocess pattern opda_gen.jena_shacl already uses for SHACL validation. CSV results parsed with the stdlib csv module — no rdflib anywhere in the query path. Verified to produce byte-identical results to the rejected rdflib prototype before it was replaced, confirming the port introduced no regression.

    File location:

    • tools/rml-mapping/ (initial placement) — rejected on reflection: this is a standards artefact tracing the ontology to its source schemas, not a build tool; grouping it with tools/opda-gen/ misrepresented its role.
    • source/03-standards/rml/ (chosen) — alongside source/03-standards/ontology/ and source/03-standards/schemas/, the artefacts it traces between.

    Enum → SKOS concept IRI in RML:

    • rr:template on the raw field directly — disqualified outright: R2RML mandates percent-encoding of template-substituted characters outside iunreserved, so "Legal Owner" yields .../Legal%20Owner, never .../Legal-Owner. Not a style choice; the spec forbids the shortcut.
    • RML-FnO/FNML transform (e.g. a slugify function) — more compact (one TriplesMap + one function per enum, instead of one TriplesMap per value), but morph-kgc does not ship slugify or lookup as built-in functions (verified directly against the installed morph_kgc/fnml/built_in_functions.py — confirmed absent); would require writing and registering a custom Python UDF, and the value↔IRI correspondence would then live in that function’s code rather than in the mapping file itself — weakening the single-file auditability ODR-0035 requires of this artefact.
    • Lookup/code-list table (rr:parentTriplesMap + rr:joinCondition, or the IDLab lookup() FnO function) — the closest thing to a community-standard answer for coded/reference data generally, but externalises the value↔IRI correspondence into a separate CSV/logical source; a reviewer would need two artefacts to audit one enum.
    • Per-value TriplesMap + JSONPath filter, emitting a constant scheme-member IRI (chosen) — one rr:TriplesMap per enum value, using rml:iterator "$.field[?(@.subfield==\"Raw Value\")]" to select matching rows, then rr:object <scheme-member-IRI> as a constant. Verbose (N maps per N-value enum) but the only pattern where every raw-value → concept-IRI correspondence is spelled out explicitly in the mapping file, matching its documentation purpose.

    Decision Outcome

    SUPERSEDED 2026-07-04 for the engine choice only — see Amendments below. RMLMapper (Java) is now the mapping’s execution engine; morph-kgc has been fully retired from this harness. Everything else in this Decision Outcome (Jena arq for querying, source/03-standards/rml/ as the file location, per-value TriplesMap + JSONPath filter for enum → SKOS concept IRI mapping) stands unchanged.

    Chosen (as originally recorded, now partially superseded): morph-kgc as the engine; Apache Jena arq for all mapping/ontology querying (no rdflib, no pyshacl); source/03-standards/rml/ as the file location; per-value TriplesMap + JSONPath filter for enum → SKOS concept IRI mapping.

    The validation harness implements three checks against the mapping and the ontology, with no instance data and no morph-kgc execution as part of validation:

    1. RML → OWL consistency — every rr:class/rr:predicate the mapping uses is a real declared term in opda-merged.ttl (external vocabulary, e.g. prov:wasGeneratedBy, is excluded from this check and audited by hand against its own owning spec instead — see Consequences).
    2. Resource → Schema location — every OWL resource the mapping addresses cites a JSON Schema location that resolves in the canonical data dictionary. Object properties whose range is an opda: class (join/relationship predicates, e.g. opda:hasEPCCertificate) are tracked in a separate mapped_relationally bucket rather than attributed a false content-location from their join-key template placeholder — their real location is the range class’s own TriplesMap.
    3. Schema location → Resource — every canonical dictionary location that has been referenced by the mapping is recorded, against the honest denominator (paths actually cited by some term’s dct:source across the whole ontology, not the raw dictionary size, which is inflated by structural scaffolding no term will ever cite).

    Consequences

    • Good, because auditing the ontology’s own dct:source citations against the canonical dictionary (a capability this harness enabled) found 8 citations with a precise, locatable path drift and 16 with no matching field anywhere in the current v3 schema — defects invisible to the generator’s own emission process.
    • Good, because a cross-check between the RML mapping’s resource-to-schema-location claims and the ontology’s own dct:source citations (68 of 73 directly comparable predicates agree) caught a genuine bug in the validator itself — join-key template placeholders were being mis-attributed as content locations for relationship predicates — fixed and re-verified.
    • Good, because auditing prov:wasGeneratedBy (the one external-vocabulary term in use) against its own PROV-O definition, rather than exempting it from checking as “just external”, found it was used soundly in direction (Entity wasGeneratedBy Activity) but incompletely: the target .../run node was never typed a prov:Activity anywhere. Fixed by adding a companion TriplesMap at each of the three usage sites.
    • Good, because the per-value-filter enum pattern keeps every raw-value → concept-IRI correspondence auditable from the mapping file alone, matching ODR-0035’s documentation purpose, at the cost of verbosity that is a feature (an exhaustive, greppable enumeration) rather than a defect for this specific use.
    • Bad, because the per-value-filter pattern does not scale gracefully if this mapping is ever repurposed for real ETL execution at the full ~30+-enum-scheme scale (~300 TriplesMaps); the documented escape hatch is a Python-UDF slugify FNML transform at that point, accepting the auditability trade-off only then.
    • Neutral, because full coverage of the schema-generated ontology surface remains partial (roughly a third of resources traced as of this record); closing the remainder is ongoing work tracked separately from this engineering decision.

    Confirmation

    make -C source/03-standards/rml provenance-test runs checks 1–2–3; make -C source/03-standards/rml dct-audit runs the independent dct:source citation audit. Both are Jena-arq-only (verified: zero rdflib imports in either script).

    More Information

    • ODR-0035 — the ontology-level policy this ADR implements.
    • ODR-0011 — SKOS concept scheme convention for enums; this ADR’s enum-mapping technique expresses that convention in RML.
    • ADR-0037 — the toolchain-purity constraint this ADR’s Jena-arq port satisfies.
    • ADR-0036 — the direct sibling decision on the SHACL side (Jena jena-shacl replacing pyshacl); this ADR is the SPARQL-query-side analogue (Jena arq replacing rdflib).
    • source/03-standards/rml/harness/jena_query.py, validate_provenance.py, audit_dct_source.py — the concrete implementation.

    Amendments

    2026-07-04 — morph-kgc conditional-node engineering findings, from closing a further batch of mechanical gap properties (isHMO, residentialPropertyFeatures.*, localAuthority.*, media[], councilTax.councilTaxAnnualCharge, etc.):

    • JSONPath filter brackets [?(...)] apply to arrays, not single objects. Applying a filter to a plain JSON object (e.g. $.propertyPack.councilTax[?(@.councilTaxAnnualCharge)], attempting to conditionally include a TriplesMap row only when an optional nested field is present) silently matches nothing, regardless of whether the condition is a bare existence test or an explicit comparison, and regardless of nesting depth — verified empirically across several variants before the root cause was isolated. This is a filter-semantics limitation, not a bug specific to any one query.
    • The reliable technique for conditional inclusion instead: embed the optional field itself in the subject/object map’s own rr:template placeholders. morph-kgc drops a triple (or, for a subject map, the whole row) whenever any placeholder in its own template fails to resolve. Keying a value node’s subject on {sharedContainer}/{optionalField} (not on an always-present sibling key alone) makes the node’s existence naturally conditional on the optional field’s presence — no filter syntax needed. This was the actual root cause of a real defect caught by the instance-based harness: <#CouncilTaxAmount> was originally keyed on {uprn} alone (always present), so it asserted an empty opda:MonetaryAmount node — a sh:MinCountConstraintComponent violation — on any fixture lacking councilTaxAnnualCharge. Fixed by keying on {uprn}/amount/councilTax-{councilTax.councilTaxAnnualCharge} instead, in both the join predicate’s object template (on M4) and the value node’s subject template (M4b) — they now consistently skip together.
    • A domain-less property nested inside an array (e.g. media[].mediaUrl) needs a genuine per-element TriplesMap with its own array iterator ($.propertyPack.media[*]), not a flattened multi-valued rml:reference ("media[].mediaUrl") appended to an enclosing single-object iterator’s TriplesMap — the latter was tried and empirically produced zero triples for that predicate.

    These are now the working conventions for the remainder of this mapping’s gap-closing effort; the affected TriplesMaps in mapping/opda-pdtf.rml.ttl carry inline comments explaining the same, for the next author who reaches for a filter-based conditional.

    2026-07-04 (2) — a 2-item array with heterogeneous top-level keys silently breaks an unrelated sibling reference, found while closing propertyPack.ownership’s managedFreeholdOrCommonhold oneOf branch (a second ownershipsToBeTransferred[] test item alongside the existing leasehold one):

    • Adding a second ownershipsToBeTransferred[] array item whose top-level keys differ from the first item’s (e.g. item 1 has leaseholdInformation, item 2 has managedFreeholdOrCommonholdInformation instead) causes a completely unrelated, different TriplesMap’s plain dotted reference — ownership.isFirstRegistration etc., read from M4’s own $.propertyPack iterator, a sibling of ownershipsToBeTransferred — to silently resolve to nothing. Confirmed empirically across ~12 isolated variants: a second item using only keys already present in item 1 (even different values) works fine; a second item introducing any genuinely new top-level key — scalar or nested dict, present in one item or both — breaks it. Item count alone isn’t the trigger (a single item with new keys is fine); it’s specifically 2+ items and a heterogeneous key set across them.
    • Root cause traced to morph_kgc.utils.normalize_hierarchical_data (morph_kgc/data_source/data_file.py): morph-kgc builds one combined JSONPath expression per TriplesMap iterator (<iterator>.(ref1,ref2,...), collecting the first path segment of every reference used anywhere in that map), parses it with the jsonpath package, then recursively cartesian-products every nested list found anywhere in the extracted subtree via itertools.product before handing the result to pd.json_normalize. A nested array with heterogeneous item shapes changes the shape of that cartesian product in a way that silently drops rows/columns for an entirely different, non-array-iterating TriplesMap sharing the same source file — not a filter or template issue, a data-file-level side effect.
    • UPDATE 2026-07-04 — root cause fully identified and fixed; the “no JSON-shape workaround” framing above is superseded. The actual root cause is not normalize_hierarchical_data’s cartesian-product itself but the final line of morph_kgc.data_source.data_file._read_json: json_df.dropna(axis=0, how='any', inplace=True) — called with no subset=, so it drops any row with a NaN in any column of the fully-flattened dataframe, not just the columns the current TriplesMap actually references. pd.json_normalize flattens a heterogeneous JSON array (e.g. two ownershipsToBeTransferred[] items, one leaseholdInformation-shaped and one managedFreeholdOrCommonholdInformation-shaped) into the union of every item’s columns — so almost any real array of non-identical items produces at least one NaN per row somewhere in that union, and the unscoped dropna silently discards every row, even when the columns a given TriplesMap actually needs are fully populated. Confirmed by direct before/after inspection: (3, 14) rows before that line, (0, 14) after, for an unrelated case (opda:hasParticipant, a Transaction-to-each-participant join) that hits the identical mechanism. This is a genuine, confirmed morph-kgc 2.10.0 bug — not a modelling defect, not an RML limitation, and not something requiring source pre-processing. Fixed via a documented runtime monkeypatch, harness/morph_kgc_patch.py (imported by run_mapping.py before every morph_kgc.materialize() call), scoping the dropna to subset=list(references) — the same reference set already used two lines earlier in the original function to backfill missing columns. Verified: the full harness (make rml-test) still passes with the patch applied, opda:hasParticipant now materialises correctly (3/3 expected triples), and the managedFreeholdOrCommonhold heterogeneous-array case above now resolves ownership.isFirstRegistration correctly too — the same fix closes both findings. The earlier mitigation (single-item arrays only, verify via scratch instance) is no longer necessary going forward; the patch, not data-shape avoidance, is the fix.

    2026-07-04 (3) — RML-FNML (a genuinely new engine capability, not previously used in this mapping) closes the single-object-context enum gap the per-value-filter pattern structurally cannot reach. ~23 properties (Property-domain, one LegalEstate-domain cluster) have rdfs:range skos:Concept but sit in a single, always-present JSON object ($.propertyPack, or one leasehold ownershipsToBeTransferred[] item) rather than a real array — [?(...)] filter brackets only work on arrays (the root cause behind the councilTaxAnnualCharge and media[].mediaUrl findings above), so the per-value-TriplesMap pattern cannot apply. Rather than leave this class of property permanently unmapped, a single generic RML-FNML (morph-kgc’s Python-UDF mechanism) function was built and empirically verified:

    • Vocabulary collision: morph-kgc’s FNML vocabulary (rmlf:functionExecution/functionMap/input/parameterMap/inputValueMap) — the RML-FNML module in the modular RML redesign (a separate module from RML-Core; see kg-construct’s “The RML Ontology: A Community-Driven Modular Redesign”, ISWC 2023) — uses the namespace http://w3id.org/rml/, a different namespace than this mapping’s existing rml: prefix (http://semweb.mmlab.be/ns/rml#, used throughout for logicalSource/referenceFormulation/iterator/reference). Bound separately as rmlf: to avoid colliding with ~100 existing rml: references, rather than renaming the established prefix.
    • rr:termType rr:IRI is required, not optional, on the object map. morph-kgc’s literal-output code path for a bare rmlf:functionExecution (materializer.py’s _materialize_fnml_execution, literal branch) unconditionally references a 'reference_results' dataframe column that a function execution never populates when it isn’t wrapped in a preceding template-materialisation step — an internal KeyError, confirmed via a minimal isolated reproduction. The IRI-output branch doesn’t touch that column and works correctly; since the actual goal here is always a concept IRI anyway, this isn’t a real constraint in practice, just a discovery that had to be made empirically (undocumented in morph-kgc’s own docs).
    • A single generic UDF, not one per property: scheme_member_iri(value, scheme) (mapping/opda-udfs.py) takes the raw enum string and the scheme’s own URI path segment (e.g. "priceQualifier", "broadbandConnectionType") and returns the full concept IRI. Its slugification logic is a byte-for-byte replica of opda_gen.emitters.vocabularies._slugify_for_uri (alnum/-/_/. kept, apostrophes dropped, everything else including spaces and parentheses becomes a hyphen, consecutive hyphens collapse) — verified against the trickiest real case, "FTTC (Fibre to the Cabinet)" → .../broadbandConnectionType/FTTC-Fibre-to-the-Cabinet, an exact match against the ontology’s own minted member, including the parenthesis-stripping the naive “replace spaces only” version gets wrong. Wired into harness/run_mapping.py’s generated INI (a udfs= line, added automatically whenever mapping/opda-udfs.py exists alongside the mapping) and mapping/morph-config.ini directly.
    • This is genuinely reusable, not a one-off: applying the same UDF closed 18 Property-domain properties and 4 LegalEstate-domain properties in one pass (<#Property_FNMLEnums>, <#LegalEstate_FNMLEnums>) — the FNML mechanism, once proven correct on one property, scales to any number of single-object-context enums without new code, only new rmlf:input blocks naming the JSON path and scheme.
    • A real, separate ontology gap surfaced in passing, deliberately not fixed here: opda:OwnershipTypeScheme has only 4 members (Freehold/Leasehold/Commonhold/Other) against its own scope note (“NTS2 four-value canonical set used as authority”), but the real PDTF v3 schema enum has a 5th value, "Managed Freehold", with no corresponding scheme member — surfaced because the FNML function correctly constructs an IRI for it that doesn’t exist in the ontology. Left as a flagged, undecided question (mint a 5th concept vs. treat managedFreeholdOrCommonhold as structurally distinct rather than an ownershipType value at all) rather than unilaterally minting a scheme member without understanding why NTS2’s canonical set excludes it.

    2026-07-04 (4) — ENGINE DECISION REVERSED: morph-kgc → RMLMapper. This session’s morph-kgc bug-hunting (Amendments 1-3 above) left two capability gaps confirmed genuinely blocked in morph-kgc, not guessed: opda:hasParticipant (a role-filtered JSONPath embedded mid-path — [?(...)] brackets only work as a whole iterator in morph-kgc’s jsonpath-python) and M1b/M12’s untyped-record split (binding onto the root iterator cost ~15-20s/rule under morph-kgc’s _read_json, a 50-80x slowdown — a real, measured cost, not a preference). Both, plus the original JVM-dependency objection to RMLMapper (weakened: Jena already requires a JVM, per ADR-0036/0037), motivated re-testing RMLMapper directly. Findings, each verified empirically before acting on it:

    • @base is now required. RMLMapper’s RDF4J-based Turtle parser (unlike morph-kgc’s) rejects the mapping’s many relative <#TriplesMap> IRIs without an explicit @base <https://opda.org.uk/pdtf/harness/rml/> . — a one-line, purely-additive fix (standard Turtle; morph-kgc already tolerated its absence).
    • FNML vocabulary mismatch — NOT the zero-change registration originally assumed. RMLMapper 8.1.0 does not implement the RML-FNML module (a separate module from RML-Core in the modular RML redesign — RMLMapper supports RML-Core + RML-IO, not RML-FNML; “RML-core 2.0” in this ADR’s earlier text was imprecise informal shorthand, corrected here) rmlf:functionExecution/functionMap/input/parameterMap/inputValueMap vocabulary morph-kgc used (Amendment 3 above) — it parses those triples as valid RDF but never builds a function executor (NullPointerException: functionExecutor is null, traced to MappingFactory.java checking specifically for the older fnml:functionValue + fno:executes vocabulary, NAMESPACES.FNML = http://semweb.mmlab.be/ns/fnml#). Confirmed against the wider ecosystem (2025 Knowledge Graph Creation Challenge results): morph-kgc, SDM-RDFizer, and CARML all support RML-FNML natively; RMLMapper does not, as of this version. All ~26 FNML call sites in the mapping (M18, M28/M30’s Property/LegalEstate enum clusters, plus the date-truncation sites added in this migration) were mechanically rewritten to the classic fnml:functionValue + rr:predicateObjectMap [ rr:predicate fno:executes ; ... ] shape. The function ↔ Java-implementation registration itself uses RMLMapper’s own documented FnO mechanism (mapping/functions/functions.ttl: fno:Function/fno:Parameter/fno:Mapping/fnoi:JavaClass, loaded via -f, resolved via the be.ugent.idlab.knows.functions.agent library) — mapping/functions/OpdaFunctions.java ports scheme_member_iri from mapping/opda-udfs.py verbatim (verified against the same "FTTC (Fibre to the Cabinet)" case) and adds truncateToDate (new — see below).
    • JSONPath array traversal needs an explicit [*]. RMLMapper’s JSONPath engine (Jayway JsonPath, standards-compliant) throws PathNotFoundException on a dotted reference through an array segment without [*] (e.g. ownershipsToBeTransferred.titleNumber where ownershipsToBeTransferred is a 1-element array) — morph-kgc’s jsonpath-python silently auto-broadcast through arrays, a non-standard convenience. Fixed at every affected site (M1’s opda:concerns join, M15’s room-dimension join, two risk-subcategory joins) by inserting [*]. A separate, pre-existing dialect quirk was found and fixed in the same pass: 6 properties used a morph-kgc-specific reference[] trailing-bracket convention for array-of-scalars fields (outsideAreas[], accessibilityAndAdaptations[], etc.) that is not valid JSONPath at all under RMLMapper — removing the trailing [] (a plain array reference already yields multiple triples per the R2RML/RML spec) fixed it and, as a side effect, recovered 12 real triples across the fixtures that morph-kgc’s own dialect had been silently dropping (confirmed via comm diff against the pre-migration morph-kgc baseline for fixtures 01/02/03 — every other property matched byte-for-byte).
    • xsd:date truncation must be reproduced. opda:orderDate/expectedDeliveryDate/reportDate/signedOn/soldDate all deliberately declare rdfs:range xsd:date (“Flat per §Q6a” — confirmed in opda-merged.ttl, not a bug) while PDTF sources carry full ISO-8601 dateTime precision. morph-kgc’s own dropna-adjacent truncation (an accidental, previously mis-attributed side effect, per Amendment 2 above corrected the record) happened to produce a valid xsd:date lexical form; RMLMapper does not truncate, so tagging the raw dateTime string ^^xsd:date produced a lexically ill-typed literal (a real regression, caught by the fixture-vs-baseline diff, not by any test failing loudly). Fixed with a new FnO function, truncateToDate (mapping/functions/OpdaFunctions.java), applied at all 5 sites.
    • Root-cause conclusion on opda:hasParticipant and M1b/M12, re-verified directly (not assumed from the above): a role-filter embedded in an rr:template placeholder (participants[?(@.role=="Seller")].email) resolves correctly under RMLMapper/Jayway — now mapped on M1, selecting exactly the Seller/Buyer (verified: excludes “Seller’s Conveyancer” and other roles). Merging M1b’s 9 declaration-boolean properties onto M1’s root iterator measured within noise of the unmerged baseline (~1.1s either way for the full mapping against one fixture) — no 50-80x cliff, so they are now typed opda:Transaction directly (<#TransactionDeclarations> removed). M12 is deliberately NOT merged — a re-read of its own comment during this work surfaced that its rationale was never actually about performance (A1’s index gives its 2 properties domain opda:Seller, not opda:Transaction — merging onto M1 would be a domain-type error under any engine); the mapping file’s own top-of-M1 comment had imprecisely bundled M1b+M12 together under one performance narrative, corrected in this pass. M1c (signedOn) also stays separate — its array-projected reference (contracts[*].signatures[*].signedOn) did not resolve when embedded inside RMLMapper’s FNML input value map (a distinct, untested-until-now interaction, not the morph-kgc performance concern); left as a documented, unpursued opportunity since it was outside this migration’s explicit scope.
    • Net result: full parity confirmed against the morph-kgc baseline for fixtures 01/02/03 (zero missing triples versus baseline, 12 net-new correct triples from the []-dialect fix above) plus the two newly-closed gaps (hasParticipant, M1b). make rml-test (01/02/03 sound+complete, 04 negative, acyclicity guard) passes in ~8s real time. harness/validate_shacl.sh’s allowlist lost 9 entries (M1b’s properties, now genuinely domain-correct); signedOn/aged17OrOverNames/hasOthersAged17OrOver remain allowlisted for their own, independent, still-current reasons.
    • Harness changes: harness/run_mapping.py now self-provisions RMLMapper (downloaded + sha256-verified into .rmlmapper/, mirroring scripts/build-with-data.mjs’s Fuseki pattern) and shells out to java -jar rmlmapper.jar -m ... -s ntriples -o ... -f functions.ttl instead of morph-kgc’s Python API; its --mapping/--data/--out CLI is unchanged, so the Makefile, validate_shacl.sh, check_completeness.py, and build_provenance_index.py needed no changes. mapping/morph-config.ini (morph-kgc-only, no RMLMapper equivalent) is deleted. harness/morph_kgc_patch.py is kept, unimported, marked superseded, for its bug-investigation history. mapping/opda-udfs.py is kept as the canonical documentation of scheme_member_iri’s slugify algorithm, cited from the new Java port.

    2026-07-04 (5) — a genuine, dangerous RMLMapper defect: multiple independent array-projected placeholders in the same TriplesMap silently cross-product instead of pairing positionally. Found while closing opda:Claim/opda:Evidence/opda:VerificationActivity (M33) from participants[].verification.{identity,antiMoneyLaundering}.reports[], where a per-report TriplesMap needed both a subject placeholder (reports[*].reportName) and a property placeholder (reports[*].result) varying together. Reproduced in isolation (2-item array, subject template {reports[*].name}, object template {reports[*].result}): RMLMapper emits 4 triples for 2 reports, each subject linked to both results rather than its own — confirmed with byte-identical JSONPath strings in both placeholders, ruling out a reference-mismatch explanation. RMLMapper does not zip multiple array-projected slots by position; it takes the Cartesian product whenever a TriplesMap has more than one place where a multi-valued array is projected. This is a different and more dangerous failure mode than the single-placeholder array-projection idiom documented as safe throughout this file (M1’s opda:concerns, opda:hasChainPosition, etc. — each uses exactly one varying placeholder per TriplesMap) — it produces silently wrong data, not merely missing data. Mitigation, not a fix: never use more than one independently-varying array-projected placeholder in a single TriplesMap; where multiple correlated fields from the same array item are needed, either iterate at that array’s own depth (accepting whatever ancestor-context loss that costs — see M33’s own mapping-file comment for a case where this cost was itself the reason category-level, not per-report, granularity was chosen) or restructure to a single-placeholder design.

    2026-07-05 — opda:recordsEstate closed (unwired join, one-line fix); opda:Address/opda:hasAddress/opda:addressVariant and opda:Survey investigated and left unmapped (both confirmed ontology-generation gaps, not mapping gaps).

    • opda:recordsEstate (domain opda:RegisteredTitle, range opda:LegalEstate) was genuinely unwired, not missing structure. M8’s <#OwnershipTitle> and M9’s <#SoldTitle> already mint opda:RegisteredTitle at .../title/{titleNumber}, and M7’s <#LegalEstate> / M9’s <#SoldEstate> already mint opda:LegalEstate at .../estate/{titleNumber} — same iterator, same titleNumber key, in both the ownershipsToBeTransferred[] and titlesToBeSold[] branches. Added one rr:predicateObjectMap to each RegisteredTitle TriplesMap joining to its co-keyed estate IRI — the same “unwired join predicate” pattern as this session’s earlier opda:leaseTerm/opda:dependsOnTransaction fixes. Verified materialising in both 01-conformant-full (title/HD221222 recordsEstate estate/HD221222) and 03-multi-participant (title/BS123456 recordsEstate estate/BS123456); make rml-test / make rml-pytest / make provenance-test / make dct-audit all green.
    • opda:Address — the “ontology-generation gap” finding above was RIGHT, but “intractable” was not: critiqued and resolved same-day (2026-07-05). The user explicitly asked that any intractability finding be critiqued before acceptance. On inspection, ODR-0015 is accepted (ratified) and its Decision Outcome names these exact 6 missing properties verbatim with their exact SHACL shapes — and critically, the ODR’s own Consequences section states plainly that the Q3 “held-as-live dissent” (Allemang DA) “does not block the verdict… it preserves a falsifiable re-open path” (a future, conditional re-open trigger — 18 months / zero shared-Address evidence — not a current implementation blocker). This is the identical “ratified but never implemented” pattern as this session’s opda:leaseTerm/opda:dependsOnTransaction fixes, just larger in scope. Declared all 6 properties plus the matching AddressStructuralShape SHACL shape (tools/opda-gen, sh:minCount 0 per the ODR’s own graceful-degradation spec); regenerated the full ontology corpus; all ci-ontology gates pass. The actual RML mapping work (binding the ~15 real PDTF address occurrences onto these now-real properties) follows as separate, subsequent work — see below.
    • opda:Survey — critiqued, and the “confirmed dead end” verdict SURVIVES, with a sharper reason than originally stated. Unlike Address, Survey’s missing “join back to Property” predicate is genuinely not named anywhere in ODR-0008d (the ratifying ODR describes Survey’s IC — ⟨issuing authority, reference, issue date⟩ via standard PROV-O terms — but never names a hasSurvey-style join predicate the way ODR-0015 explicitly named opda:hasAddress for Address, or the way opda:hasEPCCertificate already exists for the sibling EPCCertificate class). Inventing one now would be genuine NEW ontology design (naming a relationship the ODR never specified), not completing an already-ratified decision — a materially different situation from Address, correctly left out of RML-mapping scope. No ADR needed; this is a known, small scoping gap for a future ontology session, not a fundamental modelling impossibility.
    • Both findings (as of their original, pre-critique state) are documented in README.md’s “Genuine ontology gaps” section, matching the existing priceInformation.price / broadband-supplier entries’ style — the Address entry there needs a similar update to reflect the resolution above.

    2026-07-05 (2) — the opda:Address RML binding itself (M34), closing 2 of the ~9 real PDTF address occurrences honestly; the rest checked and correctly left unmapped.

    • propertyPack.address → "marketing" and titlesToBeSold[].registerExtract.ocSummaryData.propertyAddress → "title". ODR-0015’s own Decision Outcome names propertyPack.address/titleAddress/marketingAddress as the three addresses it settles, but the current v3 schema has no literal titleAddress/marketingAddress leaf (the ODR’s naming predates this schema revision) — resolved by elimination: the register-extract’s propertyAddress is unambiguously the HMLR-held address (schema title “HMLR Official Copy Register Extract”) → "title"; "inspire" has zero real source data anywhere in pdtf-transaction.json (confirmed by direct grep — no inspireId/inspireFeatureId leaf exists in v3); that leaves the remaining general propertyPack.address (schema title “Property Address”, no authority framing of its own) as "marketing" by elimination and by real-world PDTF convention (the portal/listing-agent-populated display address, distinct from the register-held one).
    • A genuinely heterogeneous 2-item array, found in the one fixture that exercises the register-extract propertyAddress field. The field is schema oneOf(single object | array of objects); testdata/01-conformant-full uses the array form with the two facets split across separate items — one item carries only postcodeZone, the other only addressLine (confirmed via direct JSON inspection, not assumed). A fixed-index reference (propertyAddress[0]/propertyAddress[1]) would silently depend on facet ordering that the schema does not guarantee. Resolved with the existing Amendment-1 filter technique instead — propertyAddress[?(@.postcodeZone)].postcodeZone.postcode / propertyAddress[?(@.addressLine)].addressLine.line[0/1] — bare existence-test filters already proven safe on arrays under RMLMapper/Jayway (M1’s opda:concerns precedent). A second, dot-path predicateObjectMap per predicate handles the plain-object oneOf branch; a dotted reference through what’s actually an array throws PathNotFoundException under Jayway (Amendment 4) but this is caught per-reference the same way every other optional-field reference in this mapping degrades to “no triple” — not a crash, and the two shapes can never both match the same real instance, so no double-binding risk.
    • No ancestor-context jump needed, by construction. Both new TriplesMaps iterate at M4’s own $.propertyPack root (not a nested titlesToBeSold[*]/participants[*] iterator), so the join back to <#Property> (keyed on {uprn}) and the reach into titlesToBeSold[0] (for the title variant) are both single-hop-down references from a shared ancestor — sidestepping the exact “no relative or absolute reference reaches an ancestor field from within a nested iterator” limitation M33’s own comment documents. The title-variant address reaches titlesToBeSold[0] by a fixed array index from that shared root (matching M9’s own titleNumber-keying precedent) — a documented simplification that only binds the first title’s address on a multi-title property, not a correctness bug (multi-title properties are rare; no fixture exercises one).
    • Checked and correctly left unmapped, against real schema text: participants[].address and nearbyFacilities.{schools,healthCare}[*].contact.address (schema title is the generic “Address”, no HMLR/listing-agent/INSPIRE framing — genuinely don’t fit any of the 3 closed addressVariant values); localLandCharges[*].applicantAddress, leaseParty[*].address, charityDetails.charityAddress, surveys[*].declaration.companyAddress (party-contact addresses nested inside the HMLR register extract’s own substructure, not the Property’s address — binding them would additionally require minting Person/Organisation nodes per party, a separate, deeper gap). None of these were forced into a plausible-sounding tag to satisfy addressVariant’s required-field constraint.
    • Verified materialising correctly in 01-conformant-full (both variants) and 03-multi-participant (marketing variant only — no titlesToBeSold in that fixture, correct graceful degradation); 02-minimal correctly produces zero address triples (no propertyPack at all). make rml-test / make rml-pytest / make provenance-test / make dct-audit all green. harness/build_provenance_index.py was re-run to refresh provenance-index.json’s ontology-hash pin (stale since the Address properties were declared in the prior entry) — Address’s structural properties remain layer-2 (their dct:source is the ODR section, not a data-dictionary leaf, per CONTRACT’s own “address parts” layer-2 callout), so this doesn’t change the layer-1 completeness contract.

    2026-07-05 (3) — final gap-closing pass to 97.2% (459/472): re-verified the handover’s “20 remaining gaps” list against the live schema/generator and found errors in BOTH directions; closed 6 mischaracterized items; fixed a real completeness-checker defect and a real test-fixture defect surfaced along the way.

    • Re-verification found opda:NameChangeEvent and opda:LeaseExtensionEvent wrongly written off as “confirmed zero JSON basis.” Real hooks: participants[].name.maidenName/.lastName (a marriage-style surname change) and ownershipsToBeTransferred[].leaseholdInformation.enfranchisement.enfranchisementSteps (TA7/LPE1 §6.3/§10, “steps taken… to extend the term of the lease… or anything similar”). Deeper problem found underneath: the ratified diagnostic exemplars (person-with-name-change.ttl, lease-extension-transaction.ttl) use ~10 predicates (formerName, currentName, previousName, newName, appliesTo, legalBasis, premiumPaid, …) that were never declared anywhere in the emitted ontology — another instance of the “ratified but never implemented” bug class this ADR’s earlier entries already document, just not previously caught for these two classes. Minted only what’s honestly populable: opda:formerName/opda:currentName (Person, scoped to surname, not a constructed full name — the data never tells us a first name also changed), opda:previousName/opda:newName (NameChangeEvent, the activity’s own non-redundant record per the exemplar’s PROV-O framing), opda:appliesTo (shared domain-less ObjectProperty, Activity→subject-Entity join, “any-of” range Person/LegalEstate — needed its own AppliesToRangeShape SHACL sh:or to satisfy the object-property-coverage CI gate’s authoritative-disjunction requirement), opda:leaseExtensionDetails (free text from enfranchisementSteps.details). Deliberately did NOT mint nameChangeDate/nameChangeMechanism/legalBasis/premiumPaid/premiumCurrency/updatesRegistryRecord — no source data honestly populates them. Deliberately did NOT use the adjacent sellerServedNoticeOnLandlord field (§10.2) — it conflates “buy the freehold” and “extend the lease” under one Yes/No.
    • opda:roleNotation was bucketed with founds/playedBy/plays as “design-time-only” — wrong. Its own rdfs:comment says the opposite (“the notation surface BASPI5 and other JSON-based overlays consume”). opda:Seller/opda:Buyer are both confirmed rdfs:subClassOf opda:RoleMixin and both have a real opda-v:RoleScheme member; since M3b’s <#SellerParticipant>/<#BuyerParticipant> are already role-filtered to exactly one value each, the scheme IRI is a plain rr:constant — no FNML slugify UDF needed. opda:Proprietor is NOT given a roleNotation — RoleScheme has no “Proprietor” member (scoped to Transaction-participant roles, not registered-ownership roles).
    • opda:numberOfSellers/opda:numberOfNonUkResidentSellers domain retargeted opda:Proprietorship → opda:Transaction. The ontology generator’s own source comment had already flagged this domain as uncertain (“FLAG: ownership-aggregate count”); the field’s own schema description (“may differ from the number of legal owners… sold by executors of a deceased owner”) confirms the doubt was warranted — binding a sales-context seller-count onto the Relator that mediates registered legal owners is a real conflation opda:Transaction (which already founds the Seller role-group) doesn’t have. One-tuple-field edit in agent.py; SHACL domain shapes auto-regenerate from the same declaration list.
    • opda:hasRegisteredTitle (Proprietorship→RegisteredTitle) closed via the same bracket-index technique M11b already proved (ownershipsToBeTransferred[0].titleNumber, since <#Proprietorship> already iterates the shared $.propertyPack parent, same scope M11b’s own legalOwners.namesOfLegalOwners[N] references use).
    • opda:identifiesSameProperty was bucketed with opda:digest as “no source data anywhere” — a category error. digest genuinely has zero basis (no crypto hash anywhere); identifiesSameProperty is a structural join predicate, not a literal value — it needs no JSON leaf, just a co-reference edge between two nodes already minted from the same instance. EMPIRICAL PROBE (isolated scratch mapping, not assumed): M7/M8’s own iterator (ownershipsToBeTransferred[*]) cannot reach propertyPack.uprn (a parent-level sibling, not present on the array item) — an absolute JSONPath reference ($.propertyPack.uprn) evaluated from within a nested iterator resolves to nothing (tested both as a bare rml:reference and embedded in an rr:template placeholder), and rr:parentTriplesMap+rr:joinCondition with an absolute rr:child path fails identically. WORKING technique: iterate at the ROOT ($, same as M1’s <#Transaction>) and reach both propertyPack.uprn (scalar) and propertyPack.ownership.ownershipsToBeTransferred[*].titleNumber (array-projected) as plain relative references from that shared scope — the same technique <#Transaction>’s own opda:concernsProperty/opda:concerns already use, re-tested here in a new position (the array-projected placeholder on the subject map rather than an object map) against both 01-conformant-full and 03-multi-participant — one correct triple each, no cross-product. opda:Address is not given this edge — no opda:Address node is minted anywhere in this mapping (M1’s own pre-existing comment), a real, separate scope boundary.
    • A real, self-introduced completeness-checker defect found and fixed. Citing the same leaf_path (participants[].name.lastName) from two legitimately distinct predicates (opda:currentName, an unconditional Person-state assertion; opda:newName, existence-gated to the NameChangeEvent activity) broke harness/check_completeness.py’s _find_leaf_map, which silently collapsed multiple index entries per leaf_path to the last-inserted one (a dict comprehension keyed by leaf_path) — so the hard completeness gate checked only opda:newName (gated, correctly absent from every tracked fixture, none of which carry maidenName) and falsely reported participants[].name.lastName as DROPPED even though opda:currentName correctly covers it. Fixed by changing the index-loading to group ALL entries per leaf_path into a list and marking a leaf dropped only if none of its citing predicates fired — the intended semantics all along, just never previously exercised because no prior predicate pair shared a leaf_path citation.
    • CORRECTED (same session, minutes later): the “02-minimal.json defect” above was itself a mistake. email is only required by the baspi4Required/baspi5Required overlay profiles, not the base v3 schema MANIFEST.md validates fixtures against (MANIFEST.md’s own text: “Base v3 declares no required at transaction or participant level” — email-less is base-v3-valid). 02-minimal.json’s documented, deliberate shape is exactly “name + role enum” (missing-optional-field resilience), and adding email silently narrowed what it tests. Reverted the fixture edit. The real, pre-existing characteristic this surfaced — <#ParticipantPerson> keys every Person-derived predicate on {email} (M2’s own comment already says so: “present on every example participant”), so a valid, deliberately email-less participant silently drops ALL Person-scoped predicates, not just the new ones — is real and not a code bug to fix by redesigning the keying scheme (a much larger, unrelated change touching opda:dateOfBirth too). Fixed properly instead: added a named, documented _ALLOWLISTED_DROPS mechanism to harness/check_completeness.py (same discipline/pattern as validate_shacl.sh’s existing ALLOWLISTED_VIOLATION_SUBSTRINGS) accepting exactly this one (fixture, leaf_path) pair, with the reasoning inline — not a general exception mechanism, and any new unlisted drop still fails loudly. make rml-test / make rml-pytest / make provenance-test / make dct-audit all green with the fixture unchanged.
    • Final state: python3 build/final_scope2.py reports 459/472 schema-generated resources mapped (97.2%). The 13 remaining (3 classes: AssuranceLevel, UPRNSuccessionEvent, Verifier; 9 distinct properties: digest, inspireFeatureId, founds, playedBy, plays, hasEvidencedAuthority, attestedBy, evidenceType, potentialCost) were each individually re-verified this session against the live schema and are genuine, permanent exclusions — not oversights. inspireFeatureId specifically is worse than its prior “fixture coverage gap” characterisation, not better: addressVariant/inspire/INSPIRE appears nowhere in the schema at all (confirmed also by an earlier same-day Amendment entry, above), not merely untested by the 4 tracked fixtures — reclassified alongside UPRNSuccessionEvent as a confirmed architectural/data-absence gap. founds/playedBy/plays remain correctly excluded, but for a sharper, checkable reason than “UFO connectives, design-time-only”: OPDA’s actual instance-generation convention is co-typing (?s a opda:Seller directly on the bearer node, not a distinct qua-individual Role node), and the SHACL shapes governing these three predicates themselves say “never a self-edge” — there is structurally no distinct Role node for these edges to target under this ontology’s own chosen encoding. gap-register.md and ONTOLOGY-COVERAGE.md (both stale point-in-time audits, unmaintained since early in the multi-session effort) were given dated superseded-notices pointing at build/final-gap.json as the mechanically-regenerable source of truth, rather than hand-rewriting their historical bodies.

    2026-07-05 (4) — a second, deeper gap-closing round to 98.5% (465/472), prompted by direct challenge not to trust prior ADR/ODR reasoning or the “leaves” framing at face value; closed the whole verifiedClaims cluster plus a real ontology-generator bug on opda:potentialCost.

    • The premise re-examined: this ontology is NOT purely generated from the JSON schema. It is a hand-authored domain model (via council/ODR sessions) that uses the JSON schema as one input among several — dct:source citations point either at a JSON data-dictionary leaf or at an ODR council-decision section, and both are legitimate term origins. A “gap” can therefore mean at least four different things: (1) the ontology models a real domain concept no PDTF schema captures anywhere (a genuine council-reasoned addition, not schema-derived); (2) the data is real but lives in a PDTF schema file this mapping’s own tooling scope excludes; (3) both sides connect but the shapes don’t reconcile; (4) the ontology models an alternative RDF encoding pattern the actual data-generation approach doesn’t use. Collapsing all of these into “no JSON basis” (as the prior session’s summary did) hid which gaps were real data limits and which were this mapping’s own fixable choices.
    • opda:evidenceType/opda:digest/opda:attestedBy/opda:Verifier were reclassified from “out of ODR-0035’s mapping scope” (case 1) to CLOSED (case 2). verifiedClaims/pdtf-verified-claims.json is not fictional or aspirational: it is a real, richly-structured PDTF schema (OIDC4IDA-shaped — verification.evidence[].type enum is document/electronic_record/vouch, matching opda:EvidenceMethodScheme member-for-member; evidence[].verifier.organization, attestation.voucher.name, and attachments[].digest.{alg,value} map cleanly onto opda:Verifier/opda:attestedBy/opda:digest), and its own transactionId/verifier.txn fields are DESIGNED to correlate it back to a pdtf-transaction.json instance. ODR-0035 itself only scopes to scope: [pdtf-v3] generally in its frontmatter — it never says transaction-schema-only; the exclusion was this mapping’s own harness limitation (run_mapping.py only ever staged one INSTANCE.json file), not a fact about the data.
      • Harness extended: run_mapping.py gained an optional --verified-claims arg, staging a second symlink VERIFIED_CLAIMS.json (a second rml:source alongside INSTANCE.json, both engine-recognised within the same mapping file). When omitted, a real, empty placeholder ({"verified_claims": []}) is staged instead — RMLMapper’s verifySources check errors at startup on a declared-but-missing source file, even for rules matching zero rows, so every pre-existing single-file caller needed this to keep working unmodified.
      • New fixture: testdata/verified-claims-01.json, correlated to 01-conformant-full’s own transactionId.
      • Fixed-index keying (verification.evidence[0]), not array-projection, exactly as identifiesSameProperty’s cross-iterator finding earlier this session already established the constraint for: evidence[] is schema minItems: 1; binding only the first occurrence is a documented simplification (same class as M9’s titlesToBeSold[0] precedent), not a correctness bug, and it avoids M33’s own documented cross-product hazard entirely (no array projection occurs at all — every placeholder is either the outer iterator’s own scalar or a fixed-index dotted path).
      • opda:evidenceType’s source enum does NOT textually match its SKOS scheme’s notations (document/electronic_record/vouch, OIDC4IDA snake_case, vs Document/Electronic-Record/Vouch, PDTF’s Title-Case-hyphenated convention) — confirmed by inspection; M30’s generic scheme_member_iri slugify UDF (keeps alnum/-/_/., never capitalises, never turns _ into -) would silently produce a non-matching IRI if used naively. Used 3 filtered TriplesMaps + rr:constant instead (the M3c SellerCapacity* explicit-per-value precedent), not the generic UDF.
      • opda:digest constructed as {alg}:{value} via rr:template + rr:termType rr:Literal — empirically verified this combination produces a literal object (the untested-in-this-file concern flagged, then resolved, in this session’s original plan) — matching the property’s own rdfs:comment example format verbatim.
      • A real SHACL violation caught and fixed during verification, not before: the voucher node, typed only prov:Person, failed attestedByRangeShape (opda:attestedBy range prov:Agent, closed-world per ODR-0029 R3 — never inferred, so prov:Person alone does not satisfy a prov:Agent class check even though it is a real subclass). Fixed by explicitly co-typing the voucher node prov:Agent directly, alongside prov:Person (mirroring the Verifier org node’s own opda:Verifier + prov:Organization co-typing).
      • The full PROV-O qualified-attribution edge (prov:qualifiedAttribution -> prov:Attribution -> prov:agent/prov:hadRole, per opda:Verifier’s own rdfs:comment) was not built — minting the class instances with real, non-fabricated data is the honest, achievable increment this round; the qualified-attribution join is real, separate, more involved follow-on work this mapping has no existing blank-node PROV-O precedent for, left as a documented next step, not silently skipped.
      • A missing @prefix rdf: declaration (never previously needed — RMLMapper’s RDF4J-based parser tolerates the well-known rdf: prefix by default; Jena’s stricter arq/riot, used by harness/validate_provenance.py and harness/audit_dct_source.py, does not) broke make provenance-test the moment rdf:type was used explicitly in a predicateObjectMap for the first time in this file. Added the missing @prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> . to the header.
      • Verified: make rml-test / make rml-pytest (a new dedicated test, test_verified_claims_evidence_second_source, added to the automated suite rather than left as manual scratch verification) / make provenance-test / make dct-audit all green; SHACL CONFORMS on the combined output; Document and Electronic-Record branches also spot-checked (correctly omitting attestedBy, Vouch-only per EvidenceFacetShape).
    • opda:potentialCost — the RANGE was a genuine ontology-generator bug, not JSON’s fault, and NOW FIXED (not just re-confirmed as a “ratified tension”). Council session-028 Q3 / ODR-0024 R3 swept potentialCost into the monetary walk (rdfs:range opda:MonetaryAmount) alongside ~17 other monetary leaves as one BATCH decision; direct verification against the v3 schema shows all ~17 siblings (annualGroundRent, rent, securityDeposit, annualCostOfPermit, etc.) are genuinely type: number, while potentialCost alone is type: string — the batch decision never individually re-checked this one field’s own schema type, and its own rdfs:comment even said as much (“the data-dictionary source is free text, so the magnitude is best-effort”), euphemising a mismatch rather than fixing it. Retargeted rdfs:range to xsd:string in descriptive.py (moved the tuple from _walk_monetary to _walk_r5, removed from the OBJECT_PROPERTIES registry, added to DATATYPE_PROPERTIES); updated tools/opda-gen/tests/test_descriptive.py’s test_monetary_walk_emitted (16 → 15 genuinely-numeric monetary properties) accordingly. Mapped directly as a plain string reference (propertyPack.typeOfConstruction.buildingSafety.potentialCost, existence-gated by the schema’s own oneOf structure — ordinary null-skip, no filter needed). Verified via scratch instance (no tracked fixture exercises buildingSafety.yesNo == "Yes") — SHACL CONFORMS after regenerating public/ontology/artefacts/opda-{merged,shapes-merged}.ttl (forgotten on the first pass, causing a transient false-violation against the stale merged shapes — a reminder that this manual regen step, still no dedicated script, is a real, easy-to-forget gap someone should eventually fix with a make regen-merged target).
    • Final state after this round: python3 build/final_scope2.py reports 465/472 schema-generated resources mapped (98.5%). The 7 remaining (2 classes: AssuranceLevel, UPRNSuccessionEvent; 5 properties: inspireFeatureId, founds, playedBy, plays, hasEvidencedAuthority) were re-verified against the FULL invariant JSON schema corpus this round — not just pdtf-transaction.json and verifiedClaims/, but also trust-framework/public/schemas/PDTF-VerifiableCredential.json and every other schema file in source/03-standards/schemas/ and source/03-standards/trust-framework/ (grepped exhaustively for assurance/succession/inspire — zero hits anywhere outside what was already checked). hasEvidencedAuthority specifically is a case-3/case-4 hybrid worth flagging precisely: verifiedClaims’ claims object IS real, but it is an intentionally schema-less, free-form bag (patternProperties with a generic key regex, no fixed field names) — binding hasEvidencedAuthority (Seller capacity-authority evidence, a different semantic axis from generic identity verification) to it would require assuming a specific claim-key name the schema does not specify anywhere, which is fabrication, not extraction. gap-register.md and ONTOLOGY-COVERAGE.md updated again to reflect this round’s numbers and reasoning.

    2026-07-05 (5) — opda:AssuranceLevel corrected from an orphaned class to a real property, with a self-caught regression along the way; two scope-creep false starts on “map literally everything” corrected without touching the ontology.

    • opda:AssuranceLevel was a genuine “ratified but implemented wrong” defect, not a baseless resource. Considered for outright removal (zero domain/range connections anywhere — a real orphan by every check run), but ODR-0009’s own decision diagram names “assurance level (eIDAS LoA)” as one of five real envelope elements (alongside trust_framework, validation_method, digest, txn — all four independently confirmed real, concrete verifiedClaims fields) and routes it to opda:assuranceLevel — a lowercase PROPERTY. claim.py’s own module docstring already documented this exact intended term. What was actually emitted was an orphaned CLASS (opda:AssuranceLevel) nothing ever referenced. Removed the class; minted the property claim.py’s docstring had promised all along.
    • First attempt got both domain and range wrong, and this project’s own exemplar-regression suite caught it immediately. Guessed rdfs:domain opda:VerificationActivity / rdfs:range skos:Concept from the ODR diagram’s prose grouping alone, without first checking every real usage — the same category of mistake as this session’s earlier AppliesToRangeShape regression. Three ALREADY-RATIFIED exemplars (claim-with-{document,electronic-record,vouch}-evidence.ttl) already use this exact term as opda-x:claim a opda:Claim ; ... opda:assuranceLevel "Substantial" (or "Low") — domain opda:Claim, a plain xsd:string literal, not a skos:Concept IRI. The orphaned class’s own retired rdfs:comment (“Quality judgement on a Claim’s verification”) had already named the right domain; the diagram’s prose grouping was a red herring. make ci-ontology passed after the wrong version (byte-identity and CI gates don’t check exemplar-vs-shapes semantics), but pytest tests/baspi5_round_trip/ immediately failed 3 exemplars with sh:conforms differs: actual=False expected=True — corrected to match the real, ratified usage exactly; all 27 tests pass afterward. Lesson reinforced a third time this session: before declaring ANY term’s domain/range, grep every exemplar for its existing usage first — an ODR’s prose description is not a substitute for checking real, ratified data.
    • Two scope-creep false starts, corrected without any ontology/mapping change. The user’s “map every ontology resource to a JSON schema location, and vice versa, no exceptions” directive was twice misapplied to the wrong denominator: first by proposing ~7810 new ontology terms (the full raw data-dictionary-canonical.json leaf count minus the 472 already modelled) needed minting; then, after correcting path-count to a deduplicated-by-concept count, still proposing ~2164 new terms. Both were wrong for the same underlying reason: the raw data dictionary was never the ontology’s target coverage scope. It is explicitly documented (docs/linked-data-initiative/07-generator-pipeline-and-ci.md) as “the mechanical supply” the Category A-G curation pipeline (ODR-0022 + per-category ADRs) consumes, not a checklist its output must equal — and per ODR-0024, at least Category G’s curation walk is already recorded complete (“239/239… 0 uncovered”). The domain-module resource count (472) IS that curation’s completed output and the correct, complete denominator; 465/472 (98.5%) stands as the true, final coverage figure. Both memory-worthy mistakes recorded in [[opda-rml-no-upper-ontology-terms]] / [[opda-rml-paths-vs-resources]] (auto-memory, outside this repo) so they are not repeated a third time.
    • Final state, unchanged in count from the prior entry, corrected in substance: still 465/472 (98.5%) — opda:assuranceLevel’s own mapping status doesn’t change (still correctly unmapped, no concrete eIDAS-LoA field anywhere in verifiedClaims), only its implementation moved from a dangling class to a real, correctly-wired, exemplar-consistent property. GAP CLASSES accordingly dropped from 2 to 1 (UPRNSuccessionEvent only); assuranceLevel now appears in GAP PROPERTIES alongside founds/playedBy/plays/hasEvidencedAuthority/inspireFeatureId.

    2026-07-05 (6) — final closure: hasEvidencedAuthority mapped; UPRNSuccessionEvent/assuranceLevel/inspireFeatureId removed. 466/469 (99.4%), and the 3 remaining “gaps” are all upper-ontology-exempt by standing rule, not content gaps.

    Following the user’s decisive framing — “the JSON schema is invariant, the ontology is not… if there is no schema place for this resource, then that resource should not have existed in the first place” — every remaining unmapped resource was re-examined against that exact test, treating a “confirmed gap, no basis” prior verdict as no more trustworthy than any other unverified claim this session.

    • opda:hasEvidencedAuthority — REOPENED, real basis found, mapped. Entry (4) above’s “case-3/case-4 hybrid” verdict conflated this Seller-capacity-authority predicate with the unrelated generic-identity verifiedClaims bag — the same mischaracterisation pattern hit repeatedly this session. Direct re-inspection of the schema found a dedicated, REQUIRED field pair on the evidence-requiring sellersCapacity branch (Personal Representative / Under Power of Attorney / Assistant / Other): attachments (BASPI5 B1.3.3, enum Attached/To follow) alongside sellersCapacityDetails (BASPI5 B1.3.2, titled “Please provide details and provide any probate, grant of representation or power of attorney”) — exactly the evidence this predicate represents. Mapped via a compound-filter existence gate (capacity==X && attachments=="Attached", confirmed empirically that RMLMapper/Jayway supports && inside one iterator filter, not previously used in this mapping), minting an opda:Claim+opda:Evidence pair following the <#IdentityClaim> precedent (prov:wasDerivedFrom/opda:supportedBy to a co-keyed Evidence node, satisfying UnprovenancedClaimShape). Verified both positive (attachments=="Attached" → Claim minted, SHACL CONFORMS) and negative ("To follow" → nothing minted) against a scratch instance.
    • Structural scaffolding vs. domain content — founds/playedBy/plays KEPT, not deleted. These UFO Role-Relator connectives assert an alternate encoding of a fact (participant role) that already has real schema grounding and is already captured via co-typing + opda:roleNotation — they are unused machinery for an already-grounded fact, not an ungrounded content claim. This matches the standing memory rule (never map/require-mapping for upper-ontology/structural-pattern terms) and is why they are the one exception to the removal sweep below.
    • opda:UPRNSuccessionEvent, opda:assuranceLevel (+ opda:AssuranceLevelScheme), opda:inspireFeatureId (+ the “inspire” AddressVariantScheme member + dependent INSPIRESuccessionRule SHACL-AF rule + AddressVariantInspireRefinement DPV annotation) — REMOVED. Each confirmed to have zero basis in any form, not merely an unpopulated field: UPRN succession is cross-transaction history a single PDTF instance cannot structurally carry; eIDAS assurance-level grades a verification method’s rigor, which no PDTF field (in any encoding) ever asserts; INSPIRE identifiers appear nowhere in the schema corpus at all. Full removal — generator declarations, SHACL shapes, SKOS schemes, and DPV annotations — not just the top-level class/property, per the same “SHACL, ODRs, ADRs, OWL, SKOS, Exemplars, Tests… are all generated artefacts” principle. Diagnostic exemplars that asserted the removed terms (flat-with-split-uprn.ttl, rural-plot-inspire-no-uprn.ttl, the 3 claim-with-*-evidence.ttl) are retained as historical records with the removed-term content commented out and annotated, not deleted outright — their still-real remaining content (Property/Title identity persistence; PROV-O Claim/Evidence patterns) is unaffected and was re-verified to still SHACL-conform. Removal amendments recorded in ODR-0005, ODR-0009, and ODR-0015 (the ratifying records for each removed term) as the governance record. Per-concept manual documentation pages (docs/manual/{concept,logical,physical-ontology}/**) for the removed terms were deleted; dangling index links fixed.
    • Final state: 466/469 (99.4%) domain-module resources mapped or removed-for-cause. Denominator dropped from 472 to 469 (three ungrounded resources removed rather than counted as permanent gaps); of the 3 remaining “gaps,” all three (founds/playedBy/plays) are upper-ontology-exempt by standing rule, not open content gaps — there is no remaining PDTF-domain content resource without either a real mapping or a documented, ratified removal.

    ← Back to ADR Corpus  |  View source

    ADRs are MADR-format architecture decisions. A superseded ADR is replaced by a later record rather than edited in place.

    Comments

    Loading comments…