Move nectar-charges DB to project account + VPC (peered to shared)

TLDR: The Aurora cluster was created in the shared account and shared VPC. Recreate it in the nectar-charges project account inside a dedicated private VPC peered to shared, with dynamic (random_pet) instance names — and fix the infra/database skills so this cannot happen again.

Status: proposed Created: 2026-08-04 Owner: @oporpino


Context

modules/backend/.infra/terraform provisions the Aurora Serverless v2 cluster against the shared account and shared VPC:

  • data.tf has a single provider "aws" that assumes TerraformRole in var.aws_shared_account_id → every resource lands in the shared account.
  • database.tf uses var.aws_shared_vpc_id for the security group and data.aws_subnets.shared for the DB subnet group → the cluster lives in the shared VPC.
  • The RDS instance is named nectar-charges-cluster-writer-${env}-1 — a static -writer name, which lies after any failover.

This violates the org standard (see wehive:infra → “AWS account model” and “Per-account VPC + peering to shared”): each project has its own AWS account and its own VPC; app resources live there and reach shared only via VPC peering. The cluster is already applied in the wrong place, and there is no data to preserve, so it will be recreated (not snapshot-migrated) in the correct topology.

The root cause is partly the skills themselves: the wehive:database skill’s code template still shows -writer instance names and says the cluster is “created in the shared VPC”, contradicting the correct rules elsewhere. Those are fixed here so the next project is born correct.

Objectives

  • Provision the nectar-charges AWS account with its TerraformRole bootstrap (prerequisite — the account does not exist yet).
  • Add a project VPC per environment, private-only (private subnets, NAT egress, no inbound from internet), with unique CIDRs: staging 10.1.0.0/16, production 10.2.0.0/16.
  • Add a VPC peering connection between each project VPC and the shared VPC, with routes on both sides, so the shared EKS pod reaches the DB on private IPs.
  • Convert data.tf to the two-provider pattern: default provider on var.aws_account_id (project) for all app resources; aws.shared alias on var.aws_shared_account_id used ONLY by the EKS data sources.
  • Recreate the Aurora cluster in the project VPC’s private subnet group / SG, with dynamic instance names via random_pet (no -writer/-reader, no ordinal).
  • Delete the old cluster, instances, subnet group, and security group left in the shared account/VPC — no leftovers.
  • Fix the wehive:database skill so its template and prose match the rules (project VPC, random_pet names). The wehive:infra skill golden rule is already added.

Non-goals

  • No public reader / public VPC — everything is internal (EKS reaches the DB over peering). A public reader would need a separate public VPC and is out of scope.
  • No data migration — the current cluster has no data worth preserving; it is recreated empty, not snapshot→restored.
  • No changes to the frontend module (it has no AWS app resources; its single shared provider for EKS is correct).
  • No changes to k8s manifests beyond what the new DATABASE_URL outputs require (URLs are already derived from RDS outputs in secrets.tf).

Changes

Prerequisite (operator, in the infrastructure repo): - Create the nectar-charges AWS account and bootstrap its TerraformRole via ACCOUNT_NAME=nectar-charges ACCOUNT_EMAIL=aws+nectar-charges@wehive.tech make aws.account.create (one account for the whole project; staging + production share it in separate VPCs — 10.1/10.2. Idempotent — imports if the account already exists). - Copy the resulting account_id into the project ward vault as TF_VAR_aws_account_id (infra level). The project never reads it from the infrastructure repo.

modules/backend/.infra/terraform/variables.tf - Add aws_account_id (this project’s account id). - Keep aws_shared_account_id, aws_shared_vpc_id, aws_shared_vpc_cidr (aws_shared_vpc_id now used ONLY to look up shared for wiring the peering). - Add project VPC CIDR var(s) if not derived from a locals map (staging 10.1, production 10.2).

modules/backend/.infra/terraform/data.tf - Split into two providers: default aws → ${var.aws_account_id} role; aliased aws.shared → ${var.aws_shared_account_id} role. - Add provider = aws.shared to aws_eks_cluster.shared and aws_eks_cluster_auth.shared. - Remove data.aws_subnets.shared (app resources no longer use shared subnets).

modules/backend/.infra/terraform/network.tf (new) - aws_vpc (project), private subnets across AZs, route table, NAT gateway + EIP, per env, CIDR from the staging/production allocation. - aws_vpc_peering_connection to shared + aws_route entries on both route tables (project→shared CIDR via peering; shared→project CIDR via peering). Follow the skill rule: no inline route {} mixed with standalone aws_route.

modules/backend/.infra/terraform/database.tf - aws_security_group.aurora.vpc_id → project VPC id. - aws_db_subnet_group.aurora.subnet_ids → project private subnets. - Ingress cidr_blocks → [var.aws_shared_vpc_cidr] (allow the peered shared VPC in for EKS reachability — this is the only sanctioned use of the shared CIDR). - Add random_pet and set the instance identifier to nectar-charges-${local.environment}-${random_pet.db.id} — drop the -writer-...-1 name.

Cleanup (same work): - After the new cluster is validated, destroy the old cluster/instance/subnet group/SG in the shared account/VPC and reconcile terraform state so state == reality. Verify zero residue.

How to verify

  • make terraform.staging.plan shows the cluster’s db_subnet_group_name and vpc_security_group_ids resolving to project-VPC resources; the default provider assumes the project account role.
  • After apply: the cluster and instance exist in the nectar-charges account/VPC; the instance name is nectar-charges-staging-<pet> (no -writer).
  • From a pod in the shared EKS cluster, psql/nc to the writer endpoint succeeds over the peering (private IP); the DB has no public IP.
  • aws rds describe-db-clusters shows nothing named nectar-charges-* remaining in the shared account; terraform plan is clean in both workspaces (no leftovers).
  • Run through the wehive:infra “Pre-apply account/VPC checklist” — all boxes pass.

Documentation

  • wehive:infra — golden rule section already added (done in this change, v3.1.0).
  • wehive:database — fix the instance template to use random_pet (align with the existing rule at the bottom of the skill), and correct the TLDR/flow text that says the cluster is “created in the shared VPC” → project VPC peered to shared.
  • wehive:infra — document the make aws.account.create operator target (run in the infrastructure repo) and fix the naming table that still showed -writer/-reader ordinal instance names (both done in this change, v3.1.0).
  • .project/docs/learnings/ — add a learning: “app DB belongs in the project account/VPC, shared reached only via peering; instance names are dynamic (random_pet).”