accepted

    AWS hosting CI/CD pipeline

    Cache-safe static releases, 2026-09-11. Normal and explicitly authorised break-glass releases use the shared scripts/site-cache-release.mjs publisher and the same opda-site-${{ github.ref }} production lock. Their existing validation and activation policies remain distinct. This amendment supersedes the blanket S3 synchronisation/deletion and /* invalidation described below. UI CSS, JavaScript, fonts and referenced assets use dependency-aware filename hashes. Publish _astro/, _images/ and _ui/ before HTML and retain historical hashed objects for open pages and rollback. These objects receive public,max-age=31536000,immutable. HTML and mutable original URLs both receive public,max-age=0,s-maxage=86400,must-revalidate: browsers revalidate while CloudFront can retain them for one day between targeted release invalidations.

    A SHA-256 cache-manifest.json identifies changed uploads and explicitly removed previously managed objects. Invalidation covers changed/deleted paths and their extensionless, trailing-slash and index.html aliases. A complete changed subtree may be compacted only without flushing known unchanged objects; the root and immutable prefixes are never blanket-purged. First adoption preserves unknown existing keys: an inventory is not proof of deletion ownership. Its cache-bootstrap-baseline.json is conditionally persisted before mutable publication and reused after interruption rather than overwritten.

    CloudFront caller references include the deployment run/attempt, both manifests and the exact invalidation batch. Identical retries are idempotent; changed batches and later roll-forward executions cannot accidentally reuse an earlier purge. The successful manifest advances only after uploads, scoped deletions and invalidation completion. Both deployment paths retain protected historical resource prefixes and use the existing short-lived GitHub OIDC role.

    Public delivery restored, 2026-09-11. The owner explicitly authorizes public access to all pages, illustrations, downloads and comment reading. This supersedes the development barrier below. CloudFront uses a network-free path-rewrite function for static content, with no Lambda/session/database checks and no forced browser no-store policy. Private S3 origins remain accessible through CloudFront. Regional APIs retain their own validation. Only comment posting requires current approved membership; anonymous, withdrawn or unavailable sessions do not block public reading. Comment loading remains deferred until after page load. CI detaches the legacy gate without deleting its replicated versions or identity records. Existing workspace and onboarding permissions are unchanged.

    Gate restored, 2026-09-09 (ADR-0038). Infra again packages a numbered Lambda@Edge version in us-east-1 before updating the regional site. The package copies the regional session reader byte-for-byte; it introduces no SSM member list or secret. Site deployment replaces all six viewer-request associations with that version while preserving private OAC origins and browser no-store. Bootstrap permits the SAM transform in both required regions. The existing main/GitHub-OIDC release process remains authoritative; no unguarded step exists. The restored function and role are named opda-session-gate: the retired opda-gate function remained after its replicated version could not be deleted. Do not delete or reuse that orphan to force the deployment through.

    Amended 2026-09-07 by ADR-0083. .github/workflows/deploy-aws.yml is now a thin trigger for the reusable site-release.yml workflow. A deterministic path classifier selects editorial, application or ontology evidence; every release keeps a small common safety envelope, while ontology changes retain the complete model toolchain. The release candidate is built once, stamped with its source commit and validation lane, checked against a 500 MB budget, and deployed unchanged. Whole-site, broad accessibility/responsive and visual checks run in the non-deploying site-assurance.yml schedule instead of blocking every release.

    Amended 2026-09-07. A deliberately manual, audited break-glass workflow (deploy-aws-break-glass.yml) may build and deploy a named main commit without the normal validation jobs when the operator explicitly confirms an unvalidated production release. It retains GitHub OIDC, the production S3 exclusions and CloudFront invalidation. It is an incident escape hatch, not the routine release path; use of it is visible in the Actions audit trail.

    Amended 2026-09-03. Site validation and deployment now form one real release gate in .github/workflows/deploy-aws.yml. Independent contracts run in parallel with creation and browser validation of a single release candidate. Deployment downloads that exact candidate only after both jobs pass; AWS credentials are available only to the deploy job. Ordinary content changes use the committed ontology projection and a pure Astro build. Changes to ontology inputs run the complete generator, BASPI5, exemplar, schema and model drift gates before deployment. The former independent ontology workflows are retired because they could fail after the site had already published. Main-branch deploys are serialised; superseded runs no longer race to write the same S3 bucket. The completed IA migration checker remains an explicit audit tool rather than an evergreen release gate.

    Amended 2026-08-27 by ADR-0079. GitHub OIDC, CI-only CloudFormation, the regional site packaging bucket and Artalk single-writer deployment remain current. The Lambda@Edge packaging, gate configuration and gate-version hand-off are retired. During removal, CI deploys the ungated site distribution before reconciling opda-edge to its certificate-only template; later runs retain that certificate/site ordering.

    Amended 2026-09-01 by ADR-0079. The regional site package now also carries a nested, path-scoped Auth0 session Lambda and HTTP API. CI supplies the Auth0 domain, public PKCE client ID and authorised-commenter list to that nested application when it deploys the site stack. No Lambda@Edge packaging or function-version hand-off has returned, and ordinary site deploys remain independent of authentication.

    Context and Problem Statement

    ADR-0038 moves the site’s hosting to AWS: S3 + CloudFront static delivery in eu-west-2, a Lambda@Edge OAuth/PKCE gate against Auth0 (the gate and ACM certificate being the us-east-1 control-plane exceptions), Artalk comments on ECS Fargate with Litestream→S3 persistence, and all infrastructure declared as CloudFormation committed to this repository. That decision needs a deployment pipeline.

    The incumbent pipeline (.github/workflows/deploy.yml) deploys to Cloudflare Pages: path-filtered on the production bundle, it runs the full data build (npm run build:data — self-provisioned Fuseki + build-time GRLC API per ADR-0021/0036/0037, only dist/ ships), then wrangler pages deploy. Deploys are CI-only — push to main is the only deployment path; manual deploys are an escape hatch to be avoided.

    The AWS architecture has three deployable artefacts where Cloudflare had one, each with a different cadence and blast radius:

    1. The static site (dist/) — changes often; must reach S3 and be visible through CloudFront promptly.
    2. The infrastructure (CloudFormation stacks in two regions) — changes rarely; a bad change can take the site down or violate the single-writer SQLite invariant.
    3. The Artalk container image — built from the sparkling/Artalk fork (a separate repository); changes rarely; a deploy must never run two writer tasks concurrently.

    How do these three artefacts deploy, with what credentials, triggered by what, in what order?

    Decision Drivers

    • CI-only deploys — preserve the existing discipline: push to main deploys; no routine manual aws CLI or console deploys.
    • No long-lived AWS credentials in GitHub repository secrets — standing keys are the dominant CI compromise vector.
    • Keep releases deterministic and proportionate — ordinary site changes build from committed inputs; ontology changes additionally refresh and verify the committed model through the Jena toolchain.
    • Everything-as-code — the pipeline that deploys the CloudFormation stacks must itself be reviewable in this repository (workflows are code), and the IAM role it assumes must be defined in CloudFormation, not hand-built.
    • Single-writer invariant (ADR-0038): no deployment path may ever run two Artalk tasks concurrently.
    • 3-user scale — no staging environment, no blue/green, no approval gates beyond main being green; the cheapest pipeline that is safe.

    Considered Options

    • A — GitHub Actions + IAM OIDC role + CloudFormation deploys — extend the existing GitHub Actions setup; jobs assume a short-lived role via GitHub’s OIDC provider; aws cloudformation deploy for infra, aws s3 sync + invalidation for the site; the Artalk fork repo pushes to ECR and bumps the service.
    • B — AWS CodePipeline / CodeBuild — AWS-native pipeline triggered from GitHub via CodeStar connection.
    • C — GitHub Actions + long-lived IAM user access keys — same workflows as A but authenticated with static keys in repository secrets.

    Decision Outcome

    Chosen option: A — GitHub Actions + IAM OIDC role + CloudFormation deploys, because it keeps CI where the repository, the existing gates (byte-identity, BASPI5 round-trip), and the path-filter logic already live; eliminates standing credentials entirely (the role is assumable only by this repository’s main branch, for minutes at a time); and adds no new vendor surface — option B introduces a second CI system for zero gain at this scale, and option C reintroduces exactly the credential class the drivers exclude.

    Pipeline components

    1. Authentication — GitHub OIDC → IAM role.

    • A bootstrap CloudFormation stack (config/aws/bootstrap-stack.yaml, eu-west-2) creates the token.actions.githubusercontent.com OIDC identity provider and a single deploy role. This stack is the one deliberate exception to CI-only: it is applied once, manually, because the role it creates is what CI authenticates as (chicken-and-egg). Everything after it deploys through CI.
    • The role’s trust policy restricts sub to repo:sparkling/opda:ref:refs/heads/main (plus the Artalk fork’s claim, below) — a workflow run from any other repository, branch, or PR cannot assume it.
    • The role’s permissions policy covers exactly the pipeline’s verbs: cloudformation:* on the project’s stacks, s3:PutObject/DeleteObject/ListBucket on the site bucket, cloudfront:CreateInvalidation on the distribution, ecr:* push on the Artalk repository, ecs:UpdateService/DescribeServices on the Artalk service, plus the iam:PassRole/regional service permissions CloudFormation needs to manage the declared resources.
    • GitHub repository configuration holds only the role ARN and account ID — neither is a secret.

    2. Site deploys — deploy-aws.yml and reusable site-release.yml.

    • Trigger: push to main filtered to actual site inputs, including the docs/ content collections; NotebookLM-only scripts do not trigger a site release.
    • Classification: a deterministic, contract-tested path classifier selects the editorial, application and ontology lanes. Unknown inputs fail safe into the application lane. Infrastructure remains independently owned by infra.yml.
    • Build: pnpm run build for editorial and application changes. Ontology-input changes run pnpm run build:data, then fail if the refreshed committed model or graph differs from the repository.
    • Validation: every site release checks the design and ADR registries, the test inventory, stable release contracts, changed routes and exactly five critical browser journeys. Application changes add focused component contracts and navigation/runtime journeys. Ontology changes add Jena, generator, BASPI5, schema, byte-identity, generated-model and resource-receipt evidence.
    • Scheduled assurance: the complete Node and Playwright suites, whole-site crawl, broad accessibility/responsive checks and curated visual baselines run without deployment. External links run weekly. A failed schedule creates or updates a tracked issue and cannot retroactively invalidate a deployed artefact.
    • Handoff: the build job stamps dist/release.json with the source commit, lane and run identity, enforces the 500 MB ceiling, and uploads the publishable tree under an immutable commit/run-attempt name. The deploy job starts only after validation passes, verifies that identity, downloads the same artefact, and never rebuilds it. Historical generated-tool prefixes already excluded from S3 synchronisation are also excluded from the handoff artefact.
    • Deploy: replace the wrangler step with aws s3 sync dist/ s3://<site-bucket> --delete followed by aws cloudfront create-invalidation --paths '/*' (at ~20 views/day, a full invalidation is simpler than hashed-path bookkeeping and within the 1,000 free invalidation paths/month).
    • Permissions: validation jobs receive only contents: read; the final deploy job alone receives id-token: write for OIDC.

    2a. Break-glass site deploys — explicit manual exception.

    • Trigger: workflow_dispatch with an explicit boolean confirmation, or a main commit that changes only this workflow’s path and carries the explicit [break-glass-release] marker. The push trigger exists for operators who can push main but cannot dispatch Actions; other workflow edits do not deploy.
    • Build and deployment: install the locked dependency graph, build from the selected main commit, assume the same short-lived OIDC role, apply the same protected-prefix exclusions, synchronise S3 and invalidate CloudFront.
    • Auditability: the GitHub Actions run records the initiating operator and the exact commit. No standing credential or local AWS deployment path is added.
    • Boundary: this path exists for recovery from a disproportionate or defective gate. It does not redefine a failed gate as passing and must not become the normal publication mechanism.

    3. Infrastructure deploys — infra.yml (new).

    • Trigger: push to main filtered to config/aws/** and the workflow itself, plus workflow_dispatch.
    • Steps: validate (aws cloudformation validate-template + cfn-lint), then deploy in dependency order — a small us-east-1 artifacts stack (one private S3 bucket: Lambda@Edge code exceeds CloudFormation’s 4 KB inline limit, so the gate is staged via aws cloudformation package, and the packaging bucket must live in us-east-1 with the function), then the us-east-1 edge stack (ACM certificate + Lambda@Edge function, publishing a new numbered function version when the code changed) before the eu-west-2 site stack (which consumes the certificate and function-version ARNs as parameters).
    • aws cloudformation deploy (change-set based) with --no-fail-on-empty-changeset, so a run that touches only one stack is a no-op for the other.
    • Concurrency: the workflow declares a concurrency group so two infra runs never interleave change sets.

    4. Artalk image deploys — in the sparkling/Artalk fork repository.

    • The image is built where its source lives: a workflow in the fork builds the image, pushes it to the ECR repository in this account (authenticating via the same OIDC provider — the deploy role’s trust policy carries a second sub claim for repo:sparkling/Artalk:ref:refs/heads/main), then runs aws ecs update-service --force-new-deployment.
    • The service’s deployment configuration (minimumHealthyPercent=0, maximumPercent=100, desired count 1 — declared in the site stack per ADR-0038) is what enforces stop-old-then-start-new, so no deployment path can ever run two writer tasks; the pipeline relies on the declared configuration rather than scripting the invariant.
    • This replaces the current flyctl deploy flow from the fork.

    5. Ordering and coupling.

    • deploy.yml and infra.yml are independent — a site deploy never touches CloudFormation and vice versa. The one coupled case (an infra change that renames/replaces the site bucket or distribution) requires an infra run followed by a site run; both expose workflow_dispatch precisely so the operator can sequence them. This is accepted as a manual-coordination case rather than automated, because it is rare and automation would couple the workflows for every run.
    • Ontology byte identity and BASPI5 remain mandatory model-change gates, now invoked by the deployment workflow rather than independent workflows that could not block that deployment.

    Consequences

    • Good, because no standing AWS credentials exist anywhere — a leaked GitHub secret yields a non-secret role ARN; assuming the role requires a workflow run on sparkling/opda@main (or the fork’s main for ECR/ECS verbs only).
    • Good, because the pipeline itself is code in this repository: the bootstrap stack, both workflows, and the deploy role’s permissions are all reviewable and diffable.
    • Good, because a normal site release has no Java service dependency and every proportionate required failure prevents deployment in the same workflow.
    • Good, because independent contracts overlap the build and browser path while the site tree is built, browser-tested and uploaded only once.
    • Good, because slow or presentation-sensitive breadth remains automated on a schedule without making routine publication depend on unrelated pages.
    • Good, because the deployed tree is the validated release candidate rather than the result of a second build on the deployment runner.
    • Good, because ontology changes retain the Jena, generator, round-trip, exemplar and byte-identity evidence without launching 17 duplicate runners.
    • Good, because the single-writer invariant is enforced declaratively (ECS deployment configuration in the stack) rather than by pipeline scripting that could be bypassed.
    • Bad, because three artefacts mean three deploy paths (site, infra, image) where Cloudflare had one — more workflows to understand, though each is individually simpler.
    • Bad, because the large, gitignored ontology tool renderings remain protected from aws s3 sync --delete; until they gain a versioned release source, that retained prefix is intentionally history-dependent rather than reproducible.
    • Bad, because the bootstrap stack is a manual one-time step that must be documented (in the stack’s own header comment) or it becomes invisible tribal knowledge.
    • Bad, because Lambda@Edge function updates are slow to converge (publish in us-east-1, global replication — ADR-0038) so an infra run that changes the gate takes materially longer than a site deploy, and rollback of a bad gate version has the same latency.
    • Neutral, because cross-repository coupling (the fork deploys into this account’s ECR/ECS) mirrors the current Fly arrangement — the fork already owns the comments deploy today.
    • Neutral, because there is no staging environment: main deploys straight to production. At 3 users this is the existing, accepted posture.

    Confirmation

    • GitHub repository secrets contain no AWS access keys — only the role ARN / account ID variables.
    • A push to main touching src/** deploys the site: the change is visible through CloudFront within minutes, and the run log shows OIDC role assumption (no static credentials).
    • A push to main touching config/aws/** updates the stacks via change sets, edge stack before site stack, and a no-op template produces an empty-change-set success, not a failure.
    • A workflow run from a branch other than main, or from a PR, fails to assume the deploy role (trust-policy rejection) — verified once deliberately.
    • An Artalk image push from the fork results in exactly one running task throughout the rollout (aws ecs describe-services shows runningCount ≤ 1 at all times).
    • The bootstrap stack is the only resource set not deployed by CI, and its template lives at config/aws/bootstrap-stack.yaml with the manual-application instruction in its header.

    Pros and Cons of the Options

    A — GitHub Actions + IAM OIDC role + CloudFormation deploys

    • Good, because zero standing credentials; trust scoped to repo + branch.
    • Good, because CI stays in one place (GitHub) with the existing gates and path filters.
    • Good, because aws cloudformation deploy + change sets give idempotent, diffable infra updates with no extra tooling.
    • Bad, because the OIDC provider + role need a manually-applied bootstrap stack before the first CI deploy.
    • Bad, because IAM trust/permission policies are easy to over-scope; the role policy needs review discipline.

    B — AWS CodePipeline / CodeBuild

    • Good, because deploy credentials never leave AWS at all (no OIDC bridge needed).
    • Bad, because it splits CI across two systems — the quality gates stay in GitHub Actions, deploys move to CodePipeline — doubling the operational surface for a 3-user site.
    • Bad, because pipeline definition, GitHub connection, and build images are all additional AWS resources to declare and maintain.

    C — GitHub Actions + long-lived IAM access keys

    • Good, because simplest to set up (no OIDC provider, no bootstrap ordering).
    • Bad, because static keys in GitHub secrets are standing credentials with no expiry — the exact risk class OIDC was designed to remove; rotation becomes a recurring manual chore that will be skipped.

    More Information

    • Implements ADR-0038 (hosting, auth, and comments architecture — AWS): this ADR realises its “Infrastructure as code — AWS CloudFormation” and CI-deploy commitments; the stack split (us-east-1 edge / eu-west-2 site) and the single-writer ECS configuration are decided there.
    • Depends on ADR-0021 as amended (entity pages read the committed model; model changes refresh it through Fuseki) and ADR-0037 (Apache Jena remains the ontology validation toolchain, conditionally provisioned for model work).
    • Cutover sequence: bootstrap stack (manual, once) → infra.yml deploys edge + site stacks → rewritten deploy.yml ships the site to S3 → fork’s workflow ships Artalk to ECR/Fargate → ADR-0038’s confirmation checklist → decommission Cloudflare Pages deploy (the wrangler step and its CLOUDFLARE_* secrets), Fly.io, and the Auth0 tenant.
    • The current Cloudflare deploy workflow is not recorded in any ADR (ADR-0038 is the first infrastructure ADR), so this ADR supersedes no record; it replaces .github/workflows/deploy.yml’s deploy step in implementation.
    • GitHub OIDC federation reference: Configuring OpenID Connect in Amazon Web Services.

    ← Back to ADR Corpus  |  View source

    ADRs are MADR-format architecture decisions. A superseded ADR is replaced by a later record rather than edited in place.

    Comments

    Loading comments…