Skip to content

Fable

Verdict

REVISE — the direction (adopt RDWS first, rebuild second, standalone read-only) is right and most mechanics check out against AWS/Terraform docs, but three gates do not enforce what they claim (scaling target vs. desired-count-0, check --retire's Requests test, ACM validation-record collision at retire) and must be fixed before v0.2b is approved.

Blocking issues (P0/P1)

  1. P1 — §4.2 "Scaling" vs §4.6 "Desired count 0 at create". aws_appautoscaling_target with min_capacity = MinSize (≥1) scales the service up as soon as it is applied, bypassing the background-work gate; ignore_changes = [desired_count] makes Terraform blind to it. Edit: in the rebuild manifest emit the scalable target with min_capacity = 0 (or omit target+policy from the rebuild phase); raise min to App Runner's MinSize only in runbook step 4, as a tfvars change gated by a check --phase cutover-style rule that allows only min_capacity on the one expected address. State this in ADR-0019.

  2. P1 — §4.8 check --retire cannot "catch clients still on the default URL". AWS/AppRunner Requests is service-wide with no per-hostname dimension (https://docs.aws.amazon.com/apprunner/latest/dg/monitor-cw.html); *.awsapprunner.com is a public, PSL-registered namespace that is scanned continuously, so Requests > 0 will be true forever and the "operator threshold" escape defeats the claim. Edit: gate on 2xxStatusResponses (scanners mostly produce 4xx) below an operator threshold and require an operator attestation after sampling the App Runner application log group for Host: values during the window; reword the row to "measures residual traffic; cannot attribute it to the default URL". Whether App Runner's own HTTP health checks count toward Requests is UNCONFIRMED — do not rely on zero.

  3. P1 — §3.2, §4.2 "Certificate", §4.8 check --phase retire: ACM validation-CNAME collision. ACM's DNS-validation CNAME is "created specifically for your domain and your account" and is reused by every certificate for that FQDN in the account (https://docs.aws.amazon.com/acm/latest/userguide/dns-validation.html). App Runner's managed certificate is "stored in AWS Certificate Manager" (https://docs.aws.amazon.com/apprunner/latest/dg/manage-custom-domains.html); if it lives in the customer's account, the validation CNAME imported in v0.2a is the same record aws_acm_certificate_validation needs in v0.2b (create conflict), and the retire step's "delete its validation CNAMEs" removes the record the new certificate renews with. Edit: verify-cutover --stage 0 compares domain_validation_options names/values against DescribeCustomDomains.CertificateValidationRecords; on a match, the rebuild emits no new record (reference the imported one) and the retire manifest excludes it. Cheap either way; the account question stays UNCONFIRMED until e2e.

  4. P1 — §4.8 check --phase cutover is under-specified. Route 53 routes to all records with equal probability when every weight in the set is 0 (https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/resource-record-sets-values-weighted.html), so a plan that sets apprunner to 0 before ecs is non-zero silently sends ~50% to ECS. Edit: the gate must reject plans where both weights are 0, where TTL/records/set_identifier change, and where the step moves weight in the wrong direction unless the runbook step is marked rollback.

  5. P1 — §3.2 custom-domain row: enable_www_subdomain. Copilot's handler calls AssociateCustomDomain with only DomainName/ServiceArn (https://raw.githubusercontent.com/aws/copilot-cli/mainline/cf-custom-resources/lib/custom-domain-app-runner.js); the API default is EnableWWWSubdomain: true (https://docs.aws.amazon.com/apprunner/latest/api/API_AssociateCustomDomain.html). So every Copilot RDWS association also covers www.<alias> with no DNS record for it. The Terraform default (true) matches live, so the import stays import-only (https://raw.githubusercontent.com/hashicorp/terraform-provider-aws/main/website/docs/r/apprunner_custom_domain_association.html.markdown), but ecsodus must read EnableWWWSubdomain from DescribeCustomDomains rather than assume, and the v0.2b report must say www.<alias> is not carried to ECS. Edit the row and ADR-0016; drop the UNCONFIRMED tag (import ID is domain_name,service_arn, confirmed).

Other findings

  • P2 — §3.1/§3.2/§12 ADR-0016: AWS::AppRunner::ObservabilityConfiguration is not a Copilot stack resource. Copilot's rd-web/cf.yml sets ObservabilityConfigurationArn: …observabilityconfiguration/DefaultConfiguration/1/00000000000000000000000000000001 (AWS-managed) on the Service; nothing is created (https://raw.githubusercontent.com/aws/copilot-cli/mainline/internal/pkg/template/templates/workloads/services/rd-web/cf.yml). Fate: external-reference, not import; the §3.1 "(UNCONFIRMED: render details)" note is now resolvable for tracing and private ingress (AppRunnerVpcIngressConnection with VpcEndpointId from EnvControllerAction.AppRunnerVpcEndpointId or a literal).
  • P2 — §3.2 AutoScalingConfigurationArn. Copilot renders autoscalingconfiguration/{{.Count}} where count is "the name of an existing autoscaling configuration" (high-availability/3), i.e. user-created and shared, never stack-owned; the template ARN lacks the ID suffix, DescribeService returns the full ARN. Use the live ARN; see Q6.
  • P2 — §4.5 pre-flight "a zone in this account". The handler writes to ListHostedZonesByName(AppDNSName) through the app DNS role (AppDNSRole), i.e. the app/tools account in multi-account setups; v0.1 already models this for the NS record. Reword to "the zone the handler used, via the app-account provider alias"; failing closed on not-found is fine.
  • P2 — §4.5 alias alternative is dismissed too fast. App Runner supports Route 53 alias A records for services created after 2022-08-01, with per-region hosted-zone IDs (https://docs.aws.amazon.com/apprunner/latest/dg/manage-custom-domains-route53.html, https://docs.aws.amazon.com/general/latest/gr/apprunner.html). A weighted alias-A pair (App Runner alias + ALB alias) is valid, has no TTL to lower, costs nothing per query, and is what AWS's own guide uses for the ECS side (https://docs.aws.amazon.com/apprunner/latest/dg/apprunner-availability-change.html). Keep the CNAME pair as default (less to verify) but offer alias as an option and fix the sentence "an ALB alias A record is not used" to say why.
  • P2 — §4.2 security groups / §4.8 retire set. ServiceSecurityGroup and its env-SG ingress are reused by ECS; list them explicitly as "never in the retire set", together with the env AppRunnerVpcEndpoint (shared).
  • P2 — §3.2 prevent_destroy. Put it on aws_apprunner_custom_domain_association and the domain CNAME too, not only the service; a Terraform destroy of the association is the same outage as the Copilot Delete handler.
  • P2 — §4.3 X86_64. No official statement; App Runner has no architecture setting and runs images as amd64 (https://github.com/aws/apprunner-roadmap/issues/79, https://repost.aws/questions/QUFkj5Ih-QTwqWOrCyH3m8LQ/does-apprunner-support-arm64-images). Keep X86_64, and when ecr:DescribeImages shows a manifest list, pin the digest of the linux/amd64 child so both sides run identical bits.
  • P2 — §8 eligibility probe. create-auto-scaling-configuration is not a service and may succeed for an ineligible account (false positive). Probe with create-service on public.ecr.aws/aws-containers/hello-app-runner:latest, 0.25 vCPU, then delete-service; which call the closure blocks is UNCONFIRMED (no docs).
  • P2 — §4.5 state gap. Between the atomic batch and terraform state rm + import, Terraform state still holds the simple record; any apply would attempt to recreate it and fail on InvalidChangeBatch (safe) — say "no applies between these two steps" in the runbook.
  • P2 — §4.2 idle_timeout = 120. Confirmed: App Runner's limit is 120 s total request time (https://docs.aws.amazon.com/apprunner/latest/dg/develop.html); ALB default idle timeout 60 s (https://docs.aws.amazon.com/elasticloadbalancing/latest/application/application-load-balancers.html). They are not the same quantity (idle vs. total); the report should say so.
  • P3 — §3.3 "verified 2026-09-30". That run verified Retain-suppression for the custom resources present (env-controller etc.), not CustomDomainAction; the mechanism is CloudFormation-generic, so say "verified for Custom::EnvControllerFunction; same mechanism".
  • P3 — §4.3 "App Runner terminates TLS and redirects HTTP": confirmed (301 redirect since 2023-02-22, https://docs.aws.amazon.com/apprunner/latest/relnotes/release-2023-02-22-http-https-support.html).
  • P3 — §4.4 Express facts confirmed: HTTPS 443 host-header listener only, internal ALB for private subnets, public IPs on tasks in public subnets, up to 25 services per ALB, ALBs deprovisioned as services go (https://docs.aws.amazon.com/AmazonECS/latest/developerguide/express-service-work.html); aws_ecs_express_gateway_service exists with cpu 256–4096 / memory 512–8192 (https://raw.githubusercontent.com/hashicorp/terraform-provider-aws/main/website/docs/r/ecs_express_gateway_service.html.markdown). Note Express also supports ARM64 and taskDefinitionArn, which the plan need not use.
  • P3 — §4.5 weighted-record import IDs need the set identifier: ZONEID_NAME_TYPE_SETID (https://raw.githubusercontent.com/hashicorp/terraform-provider-aws/main/website/docs/r/route53_record.html.markdown).
  • P3 — §1 dates confirmed: nodejs20.x deprecated 2026-04-30, block create 2027-07-29, block update 2027-08-31 (https://docs.aws.amazon.com/lambda/latest/dg/lambda-runtimes.html).
  • P3 — §4.3 health-check ranges confirmed and the UNCONFIRMED tag can go: ALB interval 5–300, timeout 2–120, healthy/unhealthy 2–10 (https://docs.aws.amazon.com/elasticloadbalancing/latest/application/target-group-health-checks.html); App Runner interval/timeout/thresholds 1–20, defaults TCP/5/2/1/5 (https://docs.aws.amazon.com/apprunner/latest/api/API_HealthCheckConfiguration.html).
  • P3 — §4.3 CPU/memory confirmed: App Runner offers 0.25/{0.5,1}, 0.5/1, 1/{2,3,4}, 2/{4,6}, 4/{8,10,12} GB (https://docs.aws.amazon.com/apprunner/latest/dg/architecture.html); all are valid Fargate pairs. Drop the tag.
  • P3 — Scope/sequencing. v0.2a → v0.2b is right; the one-path rule is what keeps the env-controller out of the cutover. Standalone read-only is right. The §8 testability answer is honest; add that check --phase retire and DisassociateCustomDomain behaviour will have been exercised only on synthetic plans and recorded responses, and that the UNVERIFIED flag applies to the whole retire section, not just the apply.

Answers to the open questions

  1. 0.3.0 and 0.4.0. Stop writing "v0.2" for this scope in docs; call the milestones M5a (RDWS adopt) and M5b (rebuild) to avoid the package-version collision.
  2. Keep (b) read-only. Rebuild also needs the owner at the freeze step, for SG reuse and for retire; a report measures demand at near-zero cost. Revisit when a design partner asks.
  3. Require v0.2a first. No exception for a still-Copilot RDWS: the env-controller and CustomDomainAction must be retained and out of the loop before any DNS change.
  4. Agree: plain ECS + ALB default, Express opt-in. ALB: operator decision in decisions.yml, default new ALB per service; allow a host rule on the env's Terraform-owned ALB only when it exists, has a 443 listener, the hostname is covered by its certificate (or an aws_lb_listener_certificate is emitted), and a rule priority is chosen with the v0.1 priority knowledge. Coupling is acceptable when the operator chooses it.
  5. Flat resources. One emitter, one validate harness, no module version in the matrix.
  6. external-reference by default: it is a user-created, shared resource Copilot never owned; AWS-managed DefaultConfiguration always external. Offer --adopt-shared later to import as aws_apprunner_auto_scaling_configuration_version (immutable, clean import, ARN confirmed: https://raw.githubusercontent.com/hashicorp/terraform-provider-aws/main/website/docs/r/apprunner_auto_scaling_configuration_version.html.markdown).
  7. Block in v0.2b. App Runner does not support private hosted zones for custom domains (https://docs.aws.amazon.com/apprunner/latest/dg/manage-custom-domains.html), so a private RDWS has no shared hostname and no gradual cutover; clients change URL regardless. Document the manual path.
  8. Allow with an explicit tracing: accept-loss decision key and a report banner. Copilot tracing is only the AWS-managed default configuration; blocking forces users to stay on a platform that gets no new features.
  9. Only the maintainer can answer; check each account's App Runner console for existing services or a usable Create button before M5.0. If none, go straight to the design-partner path.
  10. Register one cheap domain dedicated to ecsodus e2e (Route 53 Domains, kept year to year) and use it for issue #6 as well; never a personal or production domain.
  11. Keep retirement in the runbook, but as a separate opt-in section (generate --retire) that is omitted by default so the default runbook stops at 100%. Operators will delete anyway; a gated delete is safer than none. Needs the fixes in P1 #2 and #3.
  12. Yes, read-only, with the IAM actions listed in the runbook; treat alarms and target health as hard gates, metrics as advisory, and keep the operator curl --resolve mandatory.
  13. Always require the operator's list. Pre-fill ["sh","-c","<StartCommand>"] in decisions.template.yml with a # confirm marker; never default it, since ecsodus cannot see whether the image has /bin/sh.

Unconfirmed claims

  • App Runner managed certificate location (customer account vs. AWS-owned) and hence whether its validation CNAME names equal the new ACM certificate's — §3.2/§4.8. Confirm: in an eligible account, compare DescribeCustomDomains.CertificateValidationRecords with aws acm request-certificate DomainValidationOptions for the same FQDN.
  • Definition of "existing customer" and whether deleting the last service ends it — §1, §4.7, §8. No doc states it (https://docs.aws.amazon.com/apprunner/latest/dg/apprunner-availability-change.html). Confirm: AWS Support case, or probe in an eligible account after deleting all services.
  • Which API call is blocked for ineligible accounts — §8. Confirm: run the probe; record the error code.
  • App Runner runs only x86_64 images — §4.3. Community-confirmed, no official statement. Confirm: deploy an arm64-only image in an eligible account and observe failure.
  • Whether App Runner's HTTP health checks count in Requests — §4.8. Confirm: eligible account, HTTP health check, zero external traffic, read the metric.
  • aws_apprunner_vpc_connector replacement on any change — §3.2. Provider docs mark no ForceNew (https://raw.githubusercontent.com/hashicorp/terraform-provider-aws/main/website/docs/r/apprunner_vpc_connector.html.markdown); the API has no Update operation so replacement is implied. Confirm: terraform plan with a changed subnet against a stub.
  • Import-only plan for a service with the default TCP health check and no health_check_configuration block — §3.2. Confirm: real import in an eligible or design-partner account.
  • Cost table — §8. Not checked against current prices.
  • Terraform's aws_apprunner_service honouring ignore_changes on source_configuration[0].image_repository[0].image_identifier without drift on auto_deployments_enabled/image_configuration — §3.4. Confirm: import then start-deployment, then plan must be empty.