AWS hosting CI/CD pipeline
Cache-safe static releases, 2026-09-11. Normal and explicitly authorised break-glass releases use the shared
scripts/site-cache-release.mjspublisher and the sameopda-site-${{ github.ref }}production lock. Their existing validation and activation policies remain distinct. This amendment supersedes the blanket S3 synchronisation/deletion and/*invalidation described below. UI CSS, JavaScript, fonts and referenced assets use dependency-aware filename hashes. Publish_astro/,_images/and_ui/before HTML and retain historical hashed objects for open pages and rollback. These objects receivepublic,max-age=31536000,immutable. HTML and mutable original URLs both receivepublic,max-age=0,s-maxage=86400,must-revalidate: browsers revalidate while CloudFront can retain them for one day between targeted release invalidations.A SHA-256
cache-manifest.jsonidentifies changed uploads and explicitly removed previously managed objects. Invalidation covers changed/deleted paths and their extensionless, trailing-slash andindex.htmlaliases. A complete changed subtree may be compacted only without flushing known unchanged objects; the root and immutable prefixes are never blanket-purged. First adoption preserves unknown existing keys: an inventory is not proof of deletion ownership. Itscache-bootstrap-baseline.jsonis conditionally persisted before mutable publication and reused after interruption rather than overwritten.CloudFront caller references include the deployment run/attempt, both manifests and the exact invalidation batch. Identical retries are idempotent; changed batches and later roll-forward executions cannot accidentally reuse an earlier purge. The successful manifest advances only after uploads, scoped deletions and invalidation completion. Both deployment paths retain protected historical resource prefixes and use the existing short-lived GitHub OIDC role.
Public delivery restored, 2026-09-11. The owner explicitly authorizes public access to all pages, illustrations, downloads and comment reading. This supersedes the development barrier below. CloudFront uses a network-free path-rewrite function for static content, with no Lambda/session/database checks and no forced browser no-store policy. Private S3 origins remain accessible through CloudFront. Regional APIs retain their own validation. Only comment posting requires current approved membership; anonymous, withdrawn or unavailable sessions do not block public reading. Comment loading remains deferred until after page load. CI detaches the legacy gate without deleting its replicated versions or identity records. Existing workspace and onboarding permissions are unchanged.
Gate restored, 2026-09-09 (ADR-0038). Infra again packages a numbered Lambda@Edge version in
us-east-1before updating the regional site. The package copies the regional session reader byte-for-byte; it introduces no SSM member list or secret. Site deployment replaces all six viewer-request associations with that version while preserving private OAC origins and browser no-store. Bootstrap permits the SAM transform in both required regions. The existingmain/GitHub-OIDC release process remains authoritative; no unguarded step exists. The restored function and role are namedopda-session-gate: the retiredopda-gatefunction remained after its replicated version could not be deleted. Do not delete or reuse that orphan to force the deployment through.
Amended 2026-09-07 by ADR-0083.
.github/workflows/deploy-aws.ymlis now a thin trigger for the reusablesite-release.ymlworkflow. A deterministic path classifier selects editorial, application or ontology evidence; every release keeps a small common safety envelope, while ontology changes retain the complete model toolchain. The release candidate is built once, stamped with its source commit and validation lane, checked against a 500 MB budget, and deployed unchanged. Whole-site, broad accessibility/responsive and visual checks run in the non-deployingsite-assurance.ymlschedule instead of blocking every release.
Amended 2026-09-07. A deliberately manual, audited break-glass workflow (
deploy-aws-break-glass.yml) may build and deploy a namedmaincommit without the normal validation jobs when the operator explicitly confirms an unvalidated production release. It retains GitHub OIDC, the production S3 exclusions and CloudFront invalidation. It is an incident escape hatch, not the routine release path; use of it is visible in the Actions audit trail.
Amended 2026-09-03. Site validation and deployment now form one real release gate in
.github/workflows/deploy-aws.yml. Independent contracts run in parallel with creation and browser validation of a single release candidate. Deployment downloads that exact candidate only after both jobs pass; AWS credentials are available only to the deploy job. Ordinary content changes use the committed ontology projection and a pure Astro build. Changes to ontology inputs run the complete generator, BASPI5, exemplar, schema and model drift gates before deployment. The former independent ontology workflows are retired because they could fail after the site had already published. Main-branch deploys are serialised; superseded runs no longer race to write the same S3 bucket. The completed IA migration checker remains an explicit audit tool rather than an evergreen release gate.
Amended 2026-08-27 by ADR-0079. GitHub OIDC, CI-only CloudFormation, the regional site packaging bucket and Artalk single-writer deployment remain current. The Lambda@Edge packaging, gate configuration and gate-version hand-off are retired. During removal, CI deploys the ungated site distribution before reconciling
opda-edgeto its certificate-only template; later runs retain that certificate/site ordering.Amended 2026-09-01 by ADR-0079. The regional site package now also carries a nested, path-scoped Auth0 session Lambda and HTTP API. CI supplies the Auth0 domain, public PKCE client ID and authorised-commenter list to that nested application when it deploys the site stack. No Lambda@Edge packaging or function-version hand-off has returned, and ordinary site deploys remain independent of authentication.
Context and Problem Statement
ADR-0038 moves the site’s hosting to AWS: S3 + CloudFront static delivery in eu-west-2, a Lambda@Edge OAuth/PKCE gate against Auth0 (the gate and ACM certificate being the us-east-1 control-plane exceptions), Artalk comments on ECS Fargate with Litestream→S3 persistence, and all infrastructure declared as CloudFormation committed to this repository. That decision needs a deployment pipeline.
The incumbent pipeline (.github/workflows/deploy.yml) deploys to Cloudflare Pages: path-filtered on the production bundle, it runs the full data build (npm run build:data — self-provisioned Fuseki + build-time GRLC API per ADR-0021/0036/0037, only dist/ ships), then wrangler pages deploy. Deploys are CI-only — push to main is the only deployment path; manual deploys are an escape hatch to be avoided.
The AWS architecture has three deployable artefacts where Cloudflare had one, each with a different cadence and blast radius:
- The static site (
dist/) — changes often; must reach S3 and be visible through CloudFront promptly. - The infrastructure (CloudFormation stacks in two regions) — changes rarely; a bad change can take the site down or violate the single-writer SQLite invariant.
- The Artalk container image — built from the
sparkling/Artalkfork (a separate repository); changes rarely; a deploy must never run two writer tasks concurrently.
How do these three artefacts deploy, with what credentials, triggered by what, in what order?
Decision Drivers
- CI-only deploys — preserve the existing discipline: push to
maindeploys; no routine manualawsCLI or console deploys. - No long-lived AWS credentials in GitHub repository secrets — standing keys are the dominant CI compromise vector.
- Keep releases deterministic and proportionate — ordinary site changes build from committed inputs; ontology changes additionally refresh and verify the committed model through the Jena toolchain.
- Everything-as-code — the pipeline that deploys the CloudFormation stacks must itself be reviewable in this repository (workflows are code), and the IAM role it assumes must be defined in CloudFormation, not hand-built.
- Single-writer invariant (ADR-0038): no deployment path may ever run two Artalk tasks concurrently.
- 3-user scale — no staging environment, no blue/green, no approval gates beyond
mainbeing green; the cheapest pipeline that is safe.
Considered Options
- A — GitHub Actions + IAM OIDC role + CloudFormation deploys — extend the existing GitHub Actions setup; jobs assume a short-lived role via GitHub’s OIDC provider;
aws cloudformation deployfor infra,aws s3 sync+ invalidation for the site; the Artalk fork repo pushes to ECR and bumps the service. - B — AWS CodePipeline / CodeBuild — AWS-native pipeline triggered from GitHub via CodeStar connection.
- C — GitHub Actions + long-lived IAM user access keys — same workflows as A but authenticated with static keys in repository secrets.
Decision Outcome
Chosen option: A — GitHub Actions + IAM OIDC role + CloudFormation deploys, because it keeps CI where the repository, the existing gates (byte-identity, BASPI5 round-trip), and the path-filter logic already live; eliminates standing credentials entirely (the role is assumable only by this repository’s main branch, for minutes at a time); and adds no new vendor surface — option B introduces a second CI system for zero gain at this scale, and option C reintroduces exactly the credential class the drivers exclude.
Pipeline components
1. Authentication — GitHub OIDC → IAM role.
- A bootstrap CloudFormation stack (
config/aws/bootstrap-stack.yaml,eu-west-2) creates thetoken.actions.githubusercontent.comOIDC identity provider and a single deploy role. This stack is the one deliberate exception to CI-only: it is applied once, manually, because the role it creates is what CI authenticates as (chicken-and-egg). Everything after it deploys through CI. - The role’s trust policy restricts
subtorepo:sparkling/opda:ref:refs/heads/main(plus the Artalk fork’s claim, below) — a workflow run from any other repository, branch, or PR cannot assume it. - The role’s permissions policy covers exactly the pipeline’s verbs:
cloudformation:*on the project’s stacks,s3:PutObject/DeleteObject/ListBucketon the site bucket,cloudfront:CreateInvalidationon the distribution,ecr:*push on the Artalk repository,ecs:UpdateService/DescribeServiceson the Artalk service, plus theiam:PassRole/regional service permissions CloudFormation needs to manage the declared resources. - GitHub repository configuration holds only the role ARN and account ID — neither is a secret.
2. Site deploys — deploy-aws.yml and reusable site-release.yml.
- Trigger: push to
mainfiltered to actual site inputs, including thedocs/content collections; NotebookLM-only scripts do not trigger a site release. - Classification: a deterministic, contract-tested path classifier selects the
editorial, application and ontology lanes. Unknown inputs fail safe into the
application lane. Infrastructure remains independently owned by
infra.yml. - Build:
pnpm run buildfor editorial and application changes. Ontology-input changes runpnpm run build:data, then fail if the refreshed committed model or graph differs from the repository. - Validation: every site release checks the design and ADR registries, the test inventory, stable release contracts, changed routes and exactly five critical browser journeys. Application changes add focused component contracts and navigation/runtime journeys. Ontology changes add Jena, generator, BASPI5, schema, byte-identity, generated-model and resource-receipt evidence.
- Scheduled assurance: the complete Node and Playwright suites, whole-site crawl, broad accessibility/responsive checks and curated visual baselines run without deployment. External links run weekly. A failed schedule creates or updates a tracked issue and cannot retroactively invalidate a deployed artefact.
- Handoff: the build job stamps
dist/release.jsonwith the source commit, lane and run identity, enforces the 500 MB ceiling, and uploads the publishable tree under an immutable commit/run-attempt name. The deploy job starts only after validation passes, verifies that identity, downloads the same artefact, and never rebuilds it. Historical generated-tool prefixes already excluded from S3 synchronisation are also excluded from the handoff artefact. - Deploy: replace the wrangler step with
aws s3 sync dist/ s3://<site-bucket> --deletefollowed byaws cloudfront create-invalidation --paths '/*'(at ~20 views/day, a full invalidation is simpler than hashed-path bookkeeping and within the 1,000 free invalidation paths/month). - Permissions: validation jobs receive only
contents: read; the final deploy job alone receivesid-token: writefor OIDC.
2a. Break-glass site deploys — explicit manual exception.
- Trigger:
workflow_dispatchwith an explicit boolean confirmation, or amaincommit that changes only this workflow’s path and carries the explicit[break-glass-release]marker. The push trigger exists for operators who can pushmainbut cannot dispatch Actions; other workflow edits do not deploy. - Build and deployment: install the locked dependency graph, build from the
selected
maincommit, assume the same short-lived OIDC role, apply the same protected-prefix exclusions, synchronise S3 and invalidate CloudFront. - Auditability: the GitHub Actions run records the initiating operator and the exact commit. No standing credential or local AWS deployment path is added.
- Boundary: this path exists for recovery from a disproportionate or defective gate. It does not redefine a failed gate as passing and must not become the normal publication mechanism.
3. Infrastructure deploys — infra.yml (new).
- Trigger: push to
mainfiltered toconfig/aws/**and the workflow itself, plusworkflow_dispatch. - Steps: validate (
aws cloudformation validate-template+cfn-lint), then deploy in dependency order — a smallus-east-1artifacts stack (one private S3 bucket: Lambda@Edge code exceeds CloudFormation’s 4 KB inline limit, so the gate is staged viaaws cloudformation package, and the packaging bucket must live inus-east-1with the function), then theus-east-1edge stack (ACM certificate + Lambda@Edge function, publishing a new numbered function version when the code changed) before theeu-west-2site stack (which consumes the certificate and function-version ARNs as parameters). aws cloudformation deploy(change-set based) with--no-fail-on-empty-changeset, so a run that touches only one stack is a no-op for the other.- Concurrency: the workflow declares a
concurrencygroup so two infra runs never interleave change sets.
4. Artalk image deploys — in the sparkling/Artalk fork repository.
- The image is built where its source lives: a workflow in the fork builds the image, pushes it to the ECR repository in this account (authenticating via the same OIDC provider — the deploy role’s trust policy carries a second
subclaim forrepo:sparkling/Artalk:ref:refs/heads/main), then runsaws ecs update-service --force-new-deployment. - The service’s deployment configuration (
minimumHealthyPercent=0,maximumPercent=100, desired count 1 — declared in the site stack per ADR-0038) is what enforces stop-old-then-start-new, so no deployment path can ever run two writer tasks; the pipeline relies on the declared configuration rather than scripting the invariant. - This replaces the current
flyctl deployflow from the fork.
5. Ordering and coupling.
deploy.ymlandinfra.ymlare independent — a site deploy never touches CloudFormation and vice versa. The one coupled case (an infra change that renames/replaces the site bucket or distribution) requires an infra run followed by a site run; both exposeworkflow_dispatchprecisely so the operator can sequence them. This is accepted as a manual-coordination case rather than automated, because it is rare and automation would couple the workflows for every run.- Ontology byte identity and BASPI5 remain mandatory model-change gates, now invoked by the deployment workflow rather than independent workflows that could not block that deployment.
Consequences
- Good, because no standing AWS credentials exist anywhere — a leaked GitHub secret yields a non-secret role ARN; assuming the role requires a workflow run on
sparkling/opda@main(or the fork’smainfor ECR/ECS verbs only). - Good, because the pipeline itself is code in this repository: the bootstrap stack, both workflows, and the deploy role’s permissions are all reviewable and diffable.
- Good, because a normal site release has no Java service dependency and every proportionate required failure prevents deployment in the same workflow.
- Good, because independent contracts overlap the build and browser path while the site tree is built, browser-tested and uploaded only once.
- Good, because slow or presentation-sensitive breadth remains automated on a schedule without making routine publication depend on unrelated pages.
- Good, because the deployed tree is the validated release candidate rather than the result of a second build on the deployment runner.
- Good, because ontology changes retain the Jena, generator, round-trip, exemplar and byte-identity evidence without launching 17 duplicate runners.
- Good, because the single-writer invariant is enforced declaratively (ECS deployment configuration in the stack) rather than by pipeline scripting that could be bypassed.
- Bad, because three artefacts mean three deploy paths (site, infra, image) where Cloudflare had one — more workflows to understand, though each is individually simpler.
- Bad, because the large, gitignored ontology tool renderings remain protected
from
aws s3 sync --delete; until they gain a versioned release source, that retained prefix is intentionally history-dependent rather than reproducible. - Bad, because the bootstrap stack is a manual one-time step that must be documented (in the stack’s own header comment) or it becomes invisible tribal knowledge.
- Bad, because Lambda@Edge function updates are slow to converge (publish in
us-east-1, global replication — ADR-0038) so an infra run that changes the gate takes materially longer than a site deploy, and rollback of a bad gate version has the same latency. - Neutral, because cross-repository coupling (the fork deploys into this account’s ECR/ECS) mirrors the current Fly arrangement — the fork already owns the comments deploy today.
- Neutral, because there is no staging environment:
maindeploys straight to production. At 3 users this is the existing, accepted posture.
Confirmation
- GitHub repository secrets contain no AWS access keys — only the role ARN / account ID variables.
- A push to
maintouchingsrc/**deploys the site: the change is visible through CloudFront within minutes, and the run log shows OIDC role assumption (no static credentials). - A push to
maintouchingconfig/aws/**updates the stacks via change sets, edge stack before site stack, and a no-op template produces an empty-change-set success, not a failure. - A workflow run from a branch other than
main, or from a PR, fails to assume the deploy role (trust-policy rejection) — verified once deliberately. - An Artalk image push from the fork results in exactly one running task throughout the rollout (
aws ecs describe-servicesshowsrunningCount≤ 1 at all times). - The bootstrap stack is the only resource set not deployed by CI, and its template lives at
config/aws/bootstrap-stack.yamlwith the manual-application instruction in its header.
Pros and Cons of the Options
A — GitHub Actions + IAM OIDC role + CloudFormation deploys
- Good, because zero standing credentials; trust scoped to repo + branch.
- Good, because CI stays in one place (GitHub) with the existing gates and path filters.
- Good, because
aws cloudformation deploy+ change sets give idempotent, diffable infra updates with no extra tooling. - Bad, because the OIDC provider + role need a manually-applied bootstrap stack before the first CI deploy.
- Bad, because IAM trust/permission policies are easy to over-scope; the role policy needs review discipline.
B — AWS CodePipeline / CodeBuild
- Good, because deploy credentials never leave AWS at all (no OIDC bridge needed).
- Bad, because it splits CI across two systems — the quality gates stay in GitHub Actions, deploys move to CodePipeline — doubling the operational surface for a 3-user site.
- Bad, because pipeline definition, GitHub connection, and build images are all additional AWS resources to declare and maintain.
C — GitHub Actions + long-lived IAM access keys
- Good, because simplest to set up (no OIDC provider, no bootstrap ordering).
- Bad, because static keys in GitHub secrets are standing credentials with no expiry — the exact risk class OIDC was designed to remove; rotation becomes a recurring manual chore that will be skipped.
More Information
- Implements ADR-0038 (hosting, auth, and comments architecture — AWS): this ADR realises its “Infrastructure as code — AWS CloudFormation” and CI-deploy commitments; the stack split (
us-east-1edge /eu-west-2site) and the single-writer ECS configuration are decided there. - Depends on ADR-0021 as amended (entity pages read the committed model; model changes refresh it through Fuseki) and ADR-0037 (Apache Jena remains the ontology validation toolchain, conditionally provisioned for model work).
- Cutover sequence: bootstrap stack (manual, once) →
infra.ymldeploys edge + site stacks → rewrittendeploy.ymlships the site to S3 → fork’s workflow ships Artalk to ECR/Fargate → ADR-0038’s confirmation checklist → decommission Cloudflare Pages deploy (the wrangler step and itsCLOUDFLARE_*secrets), Fly.io, and the Auth0 tenant. - The current Cloudflare deploy workflow is not recorded in any ADR (ADR-0038 is the first infrastructure ADR), so this ADR supersedes no record; it replaces
.github/workflows/deploy.yml’s deploy step in implementation. - GitHub OIDC federation reference: Configuring OpenID Connect in Amazon Web Services.
Comments
Loading comments…
Sign in to post a comment