How it works

    The mapping is one Turtle file, one validation harness, and a handful of deliberate engineering choices forced by RML/R2RML's own constraints. This page summarises them for a reader; the full engineering log — including the dead ends — is ADR-0057.

    Engine & toolchain

    RMLMapper (the Java reference implementation) materialises the mapping against a PDTF instance. It replaced an earlier engine, morph-kgc (Python) — the switch is documented in ADR-0057's amendments, prompted by two morph-kgc limitations that turned out to be genuinely blocking (a role-filtered JSONPath embedded mid-path, and a ~50–80× slowdown binding onto a root iterator), not merely inconvenient.

    Validation runs entirely on Apache Jena — arq for SPARQL querying, shacl for conformance — never rdflib or pyshacl, per the toolchain-purity rule in ADR-0037. An early prototype used rdflib for graph-walking; it was replaced once that violated the rule, even though it had already been verified to produce byte-identical results.

    The two validation gates

    source/03-standards/rml/ ships two independent checks, run via make from that directory:

    GateChecksNeeds instance data?
    make provenance-test
    primary
    Three checks, no instance data: (1) every rr:class/rr:predicate the mapping uses is a real declared term in the merged ontology; (2) every OWL resource the mapping addresses cites a schema location that resolves in the canonical data dictionary; (3) every dictionary location the mapping has referenced is recorded, against the honest denominator of paths some term actually cites. No
    make dct-audit Independently audits the ontology's own dct:source citations against the canonical schema dictionary — a different direction of check than provenance-test, which checks the mapping's own references. No

    A secondary harness (make rml-test / rml-map / rml-shacl / rml-complete / rml-pytest) materialises the mapping against four small fixtures (three conformant, one deliberately negative) and checks the output is SHACL-sound and layer-1-complete. It exists to keep the mapping file itself executable and regression-tested — see Running & validating for the full list.

    The leaf→predicate index the mapping's own dct:source inversion is checked against (provenance-index.json) currently indexes 41 classes and 338 schema leaves (293 datatype / 45 object) — read live from that committed file.

    Enum values → SKOS concept IRIs

    A JSON enum value like "Legal Owner" needs to become a SKOS concept IRI — .../scheme/sellersCapacity/Legal-Owner — not a literal. The obvious approach, rr:template on the raw field, is disqualified outright: R2RML mandates percent-encoding of template-substituted characters, so a raw value with a space produces .../Legal%20Owner, never the hyphenated slug the ontology actually mints.

    The mapping uses two techniques, chosen for auditability over compactness:

    • Per-value TriplesMap + JSONPath filter (the default, for values inside an array) — one rr:TriplesMap per enum value, using a JSONPath filter to select matching rows and a rr:constant IRI as the object. Verbose (N maps for an N-value enum) but every raw-value → concept-IRI correspondence is spelled out explicitly in the mapping file itself — no external lookup table or function to also audit.
    • An FnO function (for enums sitting in a single, always-present JSON object, where filter brackets don't apply) — a JSONPath filter only works against an array; roughly two dozen properties needed a generic scheme_member_iri(value, scheme) function instead, ported byte-for-byte from the ontology generator's own slugification rule so the two never diverge.

    A real example of the first pattern — one of five TriplesMaps for the Ofsted rating enum:

    <#OfstedRatingOutstanding> a rr:TriplesMap ;
      rml:logicalSource [ rml:source "INSTANCE.json" ; rml:referenceFormulation ql:JSONPath ;
                          rml:iterator "$.propertyPack.nearbyFacilities.schools[?(@.ofstedRating==\"Outstanding\")]" ] ;
      rr:subjectMap [ rr:template "https://opda.org.uk/pdtf/harness/data/facility/school/{name}" ] ;
      rr:predicateObjectMap [ rr:predicate opda:ofstedRating ;
          rr:objectMap [ rr:constant <https://opda.org.uk/pdtf/scheme/ofstedRating/Outstanding> ] ] .

    Field notes

    Two sharp, non-obvious findings worth knowing before touching the mapping file — both fully written up in ADR-0057's amendments:

    • RMLMapper cross-products independently-varying array placeholders, silently. If a single TriplesMap has more than one place where a multi-valued array is projected, RMLMapper does not zip them by position — it takes the Cartesian product, producing wrong (not merely missing) data. The mitigation is structural: never use more than one independently-varying array-projected placeholder in one TriplesMap.
    • JSONPath array traversal needs an explicit [*]. RMLMapper's JSONPath engine (Jayway, standards-compliant) throws on a dotted reference through an array segment without it — morph-kgc silently auto-broadcast through arrays as a non-standard convenience, so this surfaced only during the engine migration.

    Comments

    Loading comments…