Organise OPDA evidence into purpose-specific NotebookLM portfolios
Context and Problem Statement
OPDA wants to use NotebookLM to produce grounded presentations, reports, infographics, audio and video about the programme, its governance, semantic modelling, working-group participation and the data model. The available corpus contains primary evidence, public website explanations, governance decisions, ontology decisions, generated references, transcripts, presentations, source schemas and several representations of the same facts.
One notebook containing the whole repository would mix incompatible audiences, authority levels and maturity states. It would also encourage generated outputs to conflate:
- OPDA’s programme with one technical input;
- an accepted decision with a proposed operating model;
- source evidence with an endorsed model;
- the independent Property Pack ontology with a complete SPDTF ontology that does not yet exist; and
- the PDTF schema, its OPDA-derived ontology and the future SPDTF standard.
The current data model is the Property Pack ontology. It is an independent delivery with its own bounded scope and governance, intended to become a component of the future Smart Property Data Trust Framework. OPDA has not yet gathered the evidence or developed the additional contextual-boundary models needed to claim a broader SPDTF ontology.
The operator’s NotebookLM plan permits 500 sources per notebook. That capacity removes any need to merge documents pre-emptively. Individual files, public URLs, meeting transcripts and rendered website routes are easier to cite, replace and audit when they remain separate. Exact duplicates and out-of-scope projections may still be excluded, but source count alone is not a reason to combine material while a notebook remains within its 500-source limit.
This ADR decides the notebook information architecture, authority boundaries, source-selection and source-unit rules, and records the resulting private implementation. It does not approve generated preparation notes or authorise sharing or publication. Explicit owner instructions may authorise private draft generation; homepage use and publication remain separately gated.
Decision Drivers
- Give each notebook one coherent subject, audience and production purpose.
- Keep programme narrative, formal governance, modelling method and participant operations distinct while allowing controlled shared context.
- Describe the Property Pack ontology accurately as an independent delivery and future SPDTF component.
- Prevent historic PDTF evidence from appearing to be an OPDA-endorsed predecessor standard or current SPDTF authority.
- Preserve decision status, source authority, provenance, dates and dissent for every independently ingested source.
- Prefer canonical glossaries, dictionaries, ontologies and registers over hundreds of generated navigational projections.
- Use the 500-source allowance deliberately without inventing a smaller target or merging sources before a demonstrated capacity or compatibility need exists.
- Support both non-technical communication artefacts and detailed technical review without forcing either audience through the other’s corpus.
- Make every notebook independently intelligible because notebooks do not inherit the complete context of their peers.
Considered Options
- One repository-wide notebook as the only production workspace. Rejected because it mixes authority and audience and makes duplicate or historic material disproportionately influential. A separate private aggregate was retained for cross-portfolio discovery, not governed artefact production.
- Four notebooks matching the initial subjects. Rejected because programme and standards governance require different narratives, while the proposed working- group notebook combines participant operations, technical modelling and formal decision authority.
- One notebook for every website section or bounded context. Rejected for the initial portfolio because the present corpus would fragment shared meaning and create many thin notebooks before additional domain evidence exists.
- Six purpose-specific core notebooks with optional archive and production notebooks (chosen). This separates audience and authority while retaining a small controlled facts pack across related notebooks.
Decision Outcome
OPDA will create six core notebooks grouped into two related portfolios. Optional notebooks may be added only when their distinct production purpose or corpus size justifies the boundary.
Programme and participation portfolio
| Notebook | Purpose | Primary audience |
|---|---|---|
| Programme, policy and history | Why the work exists; the Government’s Smart Data programme; evidenced government milestones and their status; OPDA’s role; roadmap; and development history | Executives, policymakers and new participants |
| Standards governance | Authority, decision rights, maturity stages, working-group ownership, interoperability decisions, consultation, consensus, ratification and maintenance | Chairs, reviewers and programme leaders |
| Semantic modelling method | Evidence-led modelling, AI-assisted extraction, contextual boundaries, common elements, SKOS mappings, the current SSSOM decision and non-adoption boundary, validation, provenance and modelling rules | Modellers, technical reviewers and interested participants |
| Working-group participant guide | Why to participate; what members contribute; meetings; Teams; SharePoint; source submission; model review; feedback; iteration; and how decisions affect the model | Current and prospective participants |
Models and evidence portfolio
| Notebook | Purpose | Primary audience |
|---|---|---|
| Property Pack ontology | The current independent data-model delivery: ontology, business glossary, data dictionary, contextual boundaries, diagrams, mappings, validation and coverage | Domain experts, implementers and reviewers |
| PDTF lineage and historical evidence | The PDTF schema’s historical role, the OPDA-derived ontology, lessons learned, traceability, migration evidence and what was retained, challenged or superseded | Historians, migration teams and auditors |
The canonical relationship statement for the model notebook and every shared facts pack is:
The Property Pack ontology is an independent delivery being developed as a future component of the Smart Property Data Trust Framework.
The portfolio must not call the Property Pack ontology “the SPDTF ontology” or imply that the wider model already exists. A separate SPDTF data-model notebook may be introduced only after additional contextual-boundary evidence and models have been developed through the governed process.
Optional notebooks
The following notebooks are permitted but are not part of the initial six:
- Evidence and external landscape — a research library of government publications, external standards, sector evidence, surveys and recordings used for fact-checking without overwhelming narrative notebooks.
- Decision and provenance archive — the complete ADR and ODR corpus organised by subject, owner, status and date. Narrative notebooks receive selected decisions rather than the complete archive.
- Communications studio — a deliberately small collection of approved summaries, terminology, key diagrams and scripts used to create outward-facing artefacts. It receives curated outputs from authoritative notebooks and never becomes an alternative source of truth.
- Complete research corpus — a private, deduplicated union of all sources approved for the six core notebooks. It supports cross-portfolio discovery but does not replace the core notebooks’ authority, audience or prompt boundaries. Its current manifest contains 398 unique sources from 443 scoped placements.
Authority and maturity rules
Notebook sources and generated artefacts must preserve the status of the underlying decision record:
- ADR-0063’s domain-led working groups and ADR-0066/ADR-0067’s Property Pack modelling approach are accepted.
- ADR-0075’s treatment of the Property Pack ontology as a future SPDTF component is accepted.
- ADR-0070’s uniform Microsoft 365 working-group workspace pattern is accepted.
- ADR-0065’s AI-assisted evidence-to-model workflow remains proposed.
- ADR-0068’s standards lifecycle, consensus and ratification model remains proposed.
- ADR-0077 keeps the PDTF schema as attributed third-party input rather than an endorsed predecessor scheme or authority for SPDTF meaning.
- ADR-0039, ADR-0057, ADR-0062, ADR-0069, ADR-0071, ADR-0072 and ADR-0078 are accepted; ADR-0051 remains proposed. Resource groups that include them must preserve those individual statuses rather than acquire a collective status from the resource-group label.
Notebook instructions must prevent proposed policy from being paraphrased as current operative policy. Historic facts, current implementation, accepted direction and proposed governance must remain distinguishable in both sources and prompts.
Shared facts and terminology pack
Each core notebook receives a small, versioned shared pack containing only canonical cross-portfolio facts:
- full names and roles of OPDA, SPDTF, the Property Pack ontology and the PDTF schema;
- the relationship between the independent Property Pack delivery and future SPDTF;
- programme purpose, government context, separately sourced milestone dates and their legal or policy status;
- current maturity and authority statements;
- contextual-boundary and working-group names;
- approved definitions for recurring programme terms; and
- the pack version, production date and links to its authoritative sources.
The pack is defined and built once by the portfolio shared-facts configuration. Every notebook configuration references that one output; no notebook may redefine it from a different source set. This controlled duplication makes each notebook self-contained. Narrative summaries or generated claims must not be copied between notebooks without their authority and status metadata.
Property Pack ontology source policy
The Property Pack ontology notebook includes:
- one canonical ontology representation suitable for ingestion;
- one canonical business glossary;
- one canonical data dictionary, split by contextual boundary only if source size or use makes that necessary;
- contextual-boundary summaries and model diagrams;
- incoming and outgoing relationship summaries;
- the current mapping state, governed SKOS mappings when they exist, and the documented SSSOM decision and non-adoption boundary;
- coverage and traceability summaries;
- validation and conformance reports;
- selected ODRs explaining material modelling choices; and
- a short status document explaining the ontology’s independent delivery and future relationship to SPDTF.
It excludes:
- Property Pack source schemas;
- PDTF JSON Schemas and overlays;
- generated page-per-class documentation;
- generated page-per-property documentation;
- duplicate serialisations of the same ontology unless a representation carries unique evidence;
- duplicate glossary and dictionary projections;
- raw build output and obsolete website copies;
- both SRT and VTT versions of the same transcript; and
- both an original source and a generated copy when they carry no distinct content.
The generated class and property website pages are navigation and presentation projections. The ontology, glossary and data dictionary represent their substantive information more compactly and are the canonical notebook sources.
PDTF lineage source policy
The lineage notebook excludes the PDTF and Property Pack source schemas themselves. It may include:
- the canonical OPDA-derived PDTF ontology;
- the business glossary and data dictionary needed to understand that derivation;
- provenance, mapping, traceability and migration summaries;
- selected model diagrams and historical explanations; and
- decisions that establish the source’s third-party, non-normative status.
The notebook must distinguish the third-party PDTF schema from OPDA’s technical derivation and from SPDTF development.
Source-unit, conversion and manifest rules
Each core notebook has a 500-source limit. The default source unit is one repository file, one public URL, one meeting transcript or one rendered website route. Resource groups in the configuration organise prompt scope; they are not instructions to concatenate their members. There is no lower target range and no reserve percentage that justifies premature merging.
Keep sources discrete even when several records concern the same topic. In particular:
- ingest each ADR and ODR as its own source;
- ingest one transcript per meeting or presentation;
- render each selected website route to its own source;
- ingest each external URL or downloaded publication separately;
- convert unsupported JSON, YAML, TOML, Turtle, SPARQL, Astro or other formats one-to-one into a supported textual source while retaining the original path and checksum; and
- keep glossaries, dictionaries, ontologies, mappings and validation reports separate unless a source is already a single maintained document.
Combining sources is permitted only when the manifest demonstrates a concrete need: the selected discrete sources would exceed 500, NotebookLM cannot ingest a source even after one-to-one conversion, or the records are inseparable parts of one maintained work. The shared facts and terminology pack is the one initial intentional synthesis because its purpose is controlled cross-notebook consistency, not source- count reduction.
Documents from different authorities must never be flattened into unattributed prose. Every converted source, and any exceptional combined source, carries a machine-readable or plainly formatted manifest with:
- document title;
- originating person or organisation;
- original date and conversion or combination date;
- status and maturity;
- original repository path or public URL;
- the reason for inclusion;
- clear boundaries between component documents; and
- a checksum or version identifier when the input is maintained in the repository.
Each notebook will have a maintained source manifest recording inclusion, exclusion, deduplication, one-to-one conversion and any exceptional combination. Generated summaries are outputs, not silent replacements for primary evidence.
Notebook configurations and preparation-prompt pipeline
Each core notebook has a version-controlled configuration that records its NotebookLM identifier, purpose, audiences, intended artefacts, authority guardrails, candidate resources, exclusions, missing dependencies, prepared-source outputs, ordered preparation prompts, execution fields and artefact-generation gate:
| Notebook | Configuration | Preparation prompts |
|---|---|---|
| Programme, policy and history | programme-policy-history.yaml | PPH-01 to PPH-08 |
| Standards governance | standards-governance.yaml | SG-01 to SG-08 |
| Semantic modelling method | semantic-modelling-method.yaml | SMM-01 to SMM-09 |
| Working-group participant guide | working-group-participant-guide.yaml | WG-01 to WG-10 |
| Property Pack ontology | property-pack-ontology.yaml | PP-01 to PP-09 |
| PDTF lineage and historical evidence | pdtf-lineage-historical-evidence.yaml | PDTF-01 to PDTF-09 |
The prompt pipeline prepares grounded data for later production; it does not generate the final artefacts. Its execution sequence is:
- subject the completed ADR and configuration set to an adversarial Opus review and record or apply the findings;
- review and approve the candidate source manifest for each notebook;
- implement and approve the missing preparation builder, then build the shared facts pack and convert every unsupported source one-to-one with manifests and checksums; deterministic extraction and rendering require integrity checks, while any judgement-based summarisation or exceptional source combination requires named human sign-off;
- record the repository owner’s public-source authorisation, ingest the selected sources privately and verify NotebookLM processing;
- execute each notebook’s preparation prompts in configured order and source scope; save each output as a named note and a clearly labelled, non-authoritative derived source selected by dependent prompts; independent notebook chains may run in parallel;
- append event-sourced JSONL receipts containing prompt and dependency versions, stable and NotebookLM source identifiers, note and conversation identifiers, output checksums, citation counts and review state;
- run the deterministic receipt and dependency verifier, then have a human review every final preparation pack for source authority, maturity leakage, unresolved contradictions and notebook-specific risks;
- require a human-reviewed pack for homepage or published artefacts; private review drafts require separate, explicit owner authorisation; and
- begin artefact briefs and generation as a separate downstream step.
Prompt outputs are derived working data. They must retain their source citations, may not silently replace primary evidence and may not be re-ingested as authoritative sources. Explicit owner authorisation may permit private draft generation from human-review-pending packs, but public or homepage use still requires the configuration gate. Every generated draft must preserve caveats and prohibited claims.
The Opus adversarial review completed on 2026-08-31. It identified four controls: canonical treatment of unsupported milestone claims, dependency-note carry-forward, implementation of the preparation builder, and rights/data classification for the personal NotebookLM workspace. The repository owner subsequently confirmed that all repository material is public, duplicates material already available on public SharePoint sites and is authorised for NotebookLM upload. That decision closes the rights/data-classification control. The milestone guardrails, preparation builder and dependency transport are implemented; human review of the final packs remains in force.
Implemented state
The private preparation run completed on 2026-08-31:
| Notebook | Primary sources | Preparation prompts |
|---|---|---|
| Programme, policy and history | 33 | 8 of 8 |
| Standards governance | 51 | 8 of 8 |
| Semantic modelling method | 52 | 9 of 9 |
| Working-group participant guide | 62 | 10 of 10 |
| Property Pack ontology | 82 | 9 of 9 |
| PDTF lineage and historical evidence | 163 | 9 of 9 |
The six core notebooks contain 443 scoped source placements. The private Complete
Research Corpus contains 398 unique sources, leaving 102 sources below the configured
500-source limit. The shared facts pack was rebuilt from the controlled source list
with SHA-256 a438578be73bf2a4e949549ae9ab0f650a40670c99b90a43c745b589d5681d78 and
uploaded to every core notebook and the aggregate.
All 53 version-2 preparation prompts completed. Dependency transport uses selected derived preparation sources, identified by prompt ID and output SHA-256, because the CLI does not preserve an in-memory conversation cache between processes. These sources are explicitly labelled non-authoritative working data. The strict verifier confirmed current prompt fingerprints, source and dependency identifiers, hashes, NotebookLM note IDs, citations and derived-source IDs for all 53 runs.
Automated verification does not approve the semantic content. Each final note remains labelled human review pending. The repository owner authorised the requested cinematic, explainer and short videos as private review drafts only; homepage use and publication remain blocked. All notebooks remain private.
Consequences
- Good, because artefacts can address a defined audience without importing the entire technical and historical corpus.
- Good, because the Property Pack ontology cannot be mistaken for a complete SPDTF ontology or for a direct translation of the source schema.
- Good, because proposed governance and AI-assisted methods remain visibly proposed.
- Good, because independently ingested sources preserve NotebookLM’s source-level citations, replacement boundaries and provenance.
- Good, because a controlled facts pack keeps recurring programme statements consistent across otherwise independent notebooks.
- Good, because a complete decision archive can remain available without allowing its volume to dominate narrative notebooks.
- Bad, because selected material and shared facts must be updated deliberately when authority, maturity or programme facts change.
- Bad, because unsupported formats and rendered routes require maintained one-to-one conversion, manifests and integrity checks whenever their inputs change.
- Bad, because prompt execution is non-deterministic and requires a retained run record, citation review and human approval before reuse.
- Bad, because cross-portfolio context must be repeated explicitly through the controlled shared facts pack.
- Neutral, because the source corpus remains authoritative in the repository; the notebooks are derived research and production workspaces.
- Neutral, because this ADR authorises no public sharing, publication or deployment.
Confirmation
- The six core notebook names, purposes and audiences are accepted by the operator.
- On 2026-08-31 the six core notebooks were populated privately through the dedicated
personalCLI profile with 443 scoped source placements; none was shared. - The Property Pack ontology is described as an independent delivery and future SPDTF component, never as the complete SPDTF ontology.
- The source-selection rules exclude source schemas and generated page-per-term documentation while retaining canonical ontology, glossary, dictionary, mapping, coverage and decision evidence.
- Six notebook configurations, one centrally owned shared-facts configuration and one aggregate-corpus configuration record the ingested resources, prompt sequences, receipts and remaining human approval gates.
- The maturity statements above match the status and update dates of ADR-0063, ADR-0065, ADR-0066, ADR-0067, ADR-0068, ADR-0070, ADR-0075 and ADR-0077 on 2026-08-31.
- The Opus adversarial review has been completed and its accepted findings are encoded in this ADR and the notebook configurations.
- On 2026-08-31 the repository owner confirmed that every repository source is public, duplicates material already available on public SharePoint sites and is authorised for NotebookLM upload. No configured source may be deferred on privacy, confidentiality, personal-data or rights grounds.
- The operative subscription limit is 500 sources per notebook. Configured resource groups expand to discrete sources by default; documents are combined only after a manifest records a concrete capacity, format or inseparability reason.
- The preparation builder, discrete source manifests, private ingestion and all 53 configured prompts are complete and machine-verified. Human authority, contradiction and fitness review remains outstanding. Private draft generation was separately authorised; homepage use and publication remain gated.
- Notebook creation and source upload are private workspace actions; this ADR does not authorise sharing notebooks or publishing generated artefacts.
Amendment History
- 2026-08-31 — Configured source and preparation-prompt plans. Added the centrally owned shared facts pack, six notebook configuration manifests, ordered prompt execution and review, and an explicit gate separating preparation from later artefact generation; corrected the SSSOM description to its current non-adoption boundary.
- 2026-08-31 — Applied Opus adversarial review. Removed the unsupported 2030 canonical fact, added rights and builder gates, made dependency-note transport and append-only run evidence explicit, corrected decision maturity and repository- tracking assumptions, and required human sign-off for judgement-based transformations.
- 2026-08-31 — Recorded public-source authorisation and discrete-source policy. Closed the rights/data-classification gate on the repository owner’s authority, recorded the 500-source-per-notebook limit and replaced pre-emptive bundling with one-source-per-file, URL, transcript or rendered route by default.
- 2026-08-31 — Implemented private source ingestion. Built the controlled shared facts pack and one-to-one source projections, populated all six core notebooks with 433 placements, and populated the private aggregate with 395 unique sources.
- 2026-08-31 — Completed preparation prompts. Executed and machine-verified all 50 prompts, recorded dependency-grounded outputs and receipts, and retained the human review gate before homepage use or publication.
- 2026-08-31 — Completed targeted private video drafts. Added primary legislation,
PDTF-era modelling pages and the post-release participant journey; executed three
additional prompts and generated nine private review videos: programme
be4cf704,dc19d3fe,59790bf3; semantic modellingb493abed,7931f00a,e313f77b; and working-group participation7700adfe,4d6a9c4f,08c49f5b. Homepage use, sharing and publication remain subject to human review.
More Information
- ADR-0063 — Domain-led bounded-context working groups
- ADR-0065 — AI-assisted evidence-to-model workflow
- ADR-0066 — Property Pack closed seed scope
- ADR-0067 — First-principles Property Pack ontology
- ADR-0068 — OPDA standards lifecycle
- ADR-0070 — Uniform Microsoft 365 working-group workspaces
- ADR-0075 — Property Pack ontology as an accelerated SPDTF component
- ADR-0077 — PDTF schema as a third-party input
- NotebookLM shared facts and terminology configuration
- Programme, policy and history notebook configuration
- Standards governance notebook configuration
- Semantic modelling method notebook configuration
- Working-group participant guide notebook configuration
- Property Pack ontology notebook configuration
- PDTF lineage and historical evidence notebook configuration
- Complete research corpus notebook configuration
- NotebookLM preparation runner
- NotebookLM execution verifier
- NotebookLM source and usage limits
Comments
Loading comments…
Sign in to post a comment