Fable
Verdict¶
REVISE — the direction (adopt RDWS first, rebuild second, standalone read-only) is right and most mechanics check out against AWS/Terraform docs, but three gates do not enforce what they claim (scaling target vs. desired-count-0, check --retire's Requests test, ACM validation-record collision at retire) and must be fixed before v0.2b is approved.
Blocking issues (P0/P1)¶
-
P1 — §4.2 "Scaling" vs §4.6 "Desired count 0 at create".
aws_appautoscaling_targetwithmin_capacity = MinSize(≥1) scales the service up as soon as it is applied, bypassing the background-work gate;ignore_changes = [desired_count]makes Terraform blind to it. Edit: in the rebuild manifest emit the scalable target withmin_capacity = 0(or omit target+policy from the rebuild phase); raise min to App Runner'sMinSizeonly in runbook step 4, as a tfvars change gated by acheck --phase cutover-style rule that allows onlymin_capacityon the one expected address. State this in ADR-0019. -
P1 — §4.8
check --retirecannot "catch clients still on the default URL".AWS/AppRunnerRequestsis service-wide with no per-hostname dimension (https://docs.aws.amazon.com/apprunner/latest/dg/monitor-cw.html);*.awsapprunner.comis a public, PSL-registered namespace that is scanned continuously, soRequests > 0will be true forever and the "operator threshold" escape defeats the claim. Edit: gate on2xxStatusResponses(scanners mostly produce 4xx) below an operator threshold and require an operator attestation after sampling the App Runner application log group forHost:values during the window; reword the row to "measures residual traffic; cannot attribute it to the default URL". Whether App Runner's own HTTP health checks count towardRequestsis UNCONFIRMED — do not rely on zero. -
P1 — §3.2, §4.2 "Certificate", §4.8
check --phase retire: ACM validation-CNAME collision. ACM's DNS-validation CNAME is "created specifically for your domain and your account" and is reused by every certificate for that FQDN in the account (https://docs.aws.amazon.com/acm/latest/userguide/dns-validation.html). App Runner's managed certificate is "stored in AWS Certificate Manager" (https://docs.aws.amazon.com/apprunner/latest/dg/manage-custom-domains.html); if it lives in the customer's account, the validation CNAME imported in v0.2a is the same recordaws_acm_certificate_validationneeds in v0.2b (create conflict), and the retire step's "delete its validation CNAMEs" removes the record the new certificate renews with. Edit:verify-cutover --stage 0comparesdomain_validation_optionsnames/values againstDescribeCustomDomains.CertificateValidationRecords; on a match, the rebuild emits no new record (reference the imported one) and the retire manifest excludes it. Cheap either way; the account question stays UNCONFIRMED until e2e. -
P1 — §4.8
check --phase cutoveris under-specified. Route 53 routes to all records with equal probability when every weight in the set is 0 (https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/resource-record-sets-values-weighted.html), so a plan that setsapprunnerto 0 beforeecsis non-zero silently sends ~50% to ECS. Edit: the gate must reject plans where both weights are 0, where TTL/records/set_identifierchange, and where the step moves weight in the wrong direction unless the runbook step is marked rollback. -
P1 — §3.2 custom-domain row:
enable_www_subdomain. Copilot's handler callsAssociateCustomDomainwith onlyDomainName/ServiceArn(https://raw.githubusercontent.com/aws/copilot-cli/mainline/cf-custom-resources/lib/custom-domain-app-runner.js); the API default isEnableWWWSubdomain: true(https://docs.aws.amazon.com/apprunner/latest/api/API_AssociateCustomDomain.html). So every Copilot RDWS association also coverswww.<alias>with no DNS record for it. The Terraform default (true) matches live, so the import stays import-only (https://raw.githubusercontent.com/hashicorp/terraform-provider-aws/main/website/docs/r/apprunner_custom_domain_association.html.markdown), but ecsodus must readEnableWWWSubdomainfromDescribeCustomDomainsrather than assume, and the v0.2b report must saywww.<alias>is not carried to ECS. Edit the row and ADR-0016; drop the UNCONFIRMED tag (import ID isdomain_name,service_arn, confirmed).
Other findings¶
- P2 — §3.1/§3.2/§12 ADR-0016:
AWS::AppRunner::ObservabilityConfigurationis not a Copilot stack resource. Copilot'srd-web/cf.ymlsetsObservabilityConfigurationArn: …observabilityconfiguration/DefaultConfiguration/1/00000000000000000000000000000001(AWS-managed) on the Service; nothing is created (https://raw.githubusercontent.com/aws/copilot-cli/mainline/internal/pkg/template/templates/workloads/services/rd-web/cf.yml). Fate:external-reference, notimport; the §3.1 "(UNCONFIRMED: render details)" note is now resolvable for tracing and private ingress (AppRunnerVpcIngressConnectionwithVpcEndpointIdfromEnvControllerAction.AppRunnerVpcEndpointIdor a literal). - P2 — §3.2
AutoScalingConfigurationArn. Copilot rendersautoscalingconfiguration/{{.Count}}wherecountis "the name of an existing autoscaling configuration" (high-availability/3), i.e. user-created and shared, never stack-owned; the template ARN lacks the ID suffix,DescribeServicereturns the full ARN. Use the live ARN; see Q6. - P2 — §4.5 pre-flight "a zone in this account". The handler writes to
ListHostedZonesByName(AppDNSName)through the app DNS role (AppDNSRole), i.e. the app/tools account in multi-account setups; v0.1 already models this for the NS record. Reword to "the zone the handler used, via the app-account provider alias"; failing closed on not-found is fine. - P2 — §4.5 alias alternative is dismissed too fast. App Runner supports Route 53 alias A records for services created after 2022-08-01, with per-region hosted-zone IDs (https://docs.aws.amazon.com/apprunner/latest/dg/manage-custom-domains-route53.html, https://docs.aws.amazon.com/general/latest/gr/apprunner.html). A weighted alias-A pair (App Runner alias + ALB alias) is valid, has no TTL to lower, costs nothing per query, and is what AWS's own guide uses for the ECS side (https://docs.aws.amazon.com/apprunner/latest/dg/apprunner-availability-change.html). Keep the CNAME pair as default (less to verify) but offer alias as an option and fix the sentence "an ALB alias A record is not used" to say why.
- P2 — §4.2 security groups / §4.8 retire set.
ServiceSecurityGroupand its env-SG ingress are reused by ECS; list them explicitly as "never in the retire set", together with the envAppRunnerVpcEndpoint(shared). - P2 — §3.2
prevent_destroy. Put it onaws_apprunner_custom_domain_associationand the domain CNAME too, not only the service; a Terraform destroy of the association is the same outage as the Copilot Delete handler. - P2 — §4.3 X86_64. No official statement; App Runner has no architecture setting and runs images as amd64 (https://github.com/aws/apprunner-roadmap/issues/79, https://repost.aws/questions/QUFkj5Ih-QTwqWOrCyH3m8LQ/does-apprunner-support-arm64-images). Keep
X86_64, and whenecr:DescribeImagesshows a manifest list, pin the digest of thelinux/amd64child so both sides run identical bits. - P2 — §8 eligibility probe.
create-auto-scaling-configurationis not a service and may succeed for an ineligible account (false positive). Probe withcreate-serviceonpublic.ecr.aws/aws-containers/hello-app-runner:latest, 0.25 vCPU, thendelete-service; which call the closure blocks is UNCONFIRMED (no docs). - P2 — §4.5 state gap. Between the atomic batch and
terraform state rm+ import, Terraform state still holds the simple record; any apply would attempt to recreate it and fail onInvalidChangeBatch(safe) — say "no applies between these two steps" in the runbook. - P2 — §4.2
idle_timeout = 120. Confirmed: App Runner's limit is 120 s total request time (https://docs.aws.amazon.com/apprunner/latest/dg/develop.html); ALB default idle timeout 60 s (https://docs.aws.amazon.com/elasticloadbalancing/latest/application/application-load-balancers.html). They are not the same quantity (idle vs. total); the report should say so. - P3 — §3.3 "verified 2026-09-30". That run verified Retain-suppression for the custom resources present (env-controller etc.), not
CustomDomainAction; the mechanism is CloudFormation-generic, so say "verified for Custom::EnvControllerFunction; same mechanism". - P3 — §4.3 "App Runner terminates TLS and redirects HTTP": confirmed (301 redirect since 2023-02-22, https://docs.aws.amazon.com/apprunner/latest/relnotes/release-2023-02-22-http-https-support.html).
- P3 — §4.4 Express facts confirmed: HTTPS 443 host-header listener only, internal ALB for private subnets, public IPs on tasks in public subnets, up to 25 services per ALB, ALBs deprovisioned as services go (https://docs.aws.amazon.com/AmazonECS/latest/developerguide/express-service-work.html);
aws_ecs_express_gateway_serviceexists with cpu 256–4096 / memory 512–8192 (https://raw.githubusercontent.com/hashicorp/terraform-provider-aws/main/website/docs/r/ecs_express_gateway_service.html.markdown). Note Express also supportsARM64andtaskDefinitionArn, which the plan need not use. - P3 — §4.5 weighted-record import IDs need the set identifier:
ZONEID_NAME_TYPE_SETID(https://raw.githubusercontent.com/hashicorp/terraform-provider-aws/main/website/docs/r/route53_record.html.markdown). - P3 — §1 dates confirmed:
nodejs20.xdeprecated 2026-04-30, block create 2027-07-29, block update 2027-08-31 (https://docs.aws.amazon.com/lambda/latest/dg/lambda-runtimes.html). - P3 — §4.3 health-check ranges confirmed and the UNCONFIRMED tag can go: ALB interval 5–300, timeout 2–120, healthy/unhealthy 2–10 (https://docs.aws.amazon.com/elasticloadbalancing/latest/application/target-group-health-checks.html); App Runner interval/timeout/thresholds 1–20, defaults TCP/5/2/1/5 (https://docs.aws.amazon.com/apprunner/latest/api/API_HealthCheckConfiguration.html).
- P3 — §4.3 CPU/memory confirmed: App Runner offers 0.25/{0.5,1}, 0.5/1, 1/{2,3,4}, 2/{4,6}, 4/{8,10,12} GB (https://docs.aws.amazon.com/apprunner/latest/dg/architecture.html); all are valid Fargate pairs. Drop the tag.
- P3 — Scope/sequencing. v0.2a → v0.2b is right; the one-path rule is what keeps the env-controller out of the cutover. Standalone read-only is right. The §8 testability answer is honest; add that
check --phase retireandDisassociateCustomDomainbehaviour will have been exercised only on synthetic plans and recorded responses, and that the UNVERIFIED flag applies to the whole retire section, not just the apply.
Answers to the open questions¶
- 0.3.0 and 0.4.0. Stop writing "v0.2" for this scope in docs; call the milestones M5a (RDWS adopt) and M5b (rebuild) to avoid the package-version collision.
- Keep (b) read-only. Rebuild also needs the owner at the freeze step, for SG reuse and for retire; a report measures demand at near-zero cost. Revisit when a design partner asks.
- Require v0.2a first. No exception for a still-Copilot RDWS: the env-controller and
CustomDomainActionmust be retained and out of the loop before any DNS change. - Agree: plain ECS + ALB default, Express opt-in. ALB: operator decision in
decisions.yml, default new ALB per service; allow a host rule on the env's Terraform-owned ALB only when it exists, has a 443 listener, the hostname is covered by its certificate (or anaws_lb_listener_certificateis emitted), and a rule priority is chosen with the v0.1 priority knowledge. Coupling is acceptable when the operator chooses it. - Flat resources. One emitter, one validate harness, no module version in the matrix.
external-referenceby default: it is a user-created, shared resource Copilot never owned; AWS-managedDefaultConfigurationalways external. Offer--adopt-sharedlater to import asaws_apprunner_auto_scaling_configuration_version(immutable, clean import, ARN confirmed: https://raw.githubusercontent.com/hashicorp/terraform-provider-aws/main/website/docs/r/apprunner_auto_scaling_configuration_version.html.markdown).- Block in v0.2b. App Runner does not support private hosted zones for custom domains (https://docs.aws.amazon.com/apprunner/latest/dg/manage-custom-domains.html), so a private RDWS has no shared hostname and no gradual cutover; clients change URL regardless. Document the manual path.
- Allow with an explicit
tracing: accept-lossdecision key and a report banner. Copilot tracing is only the AWS-managed default configuration; blocking forces users to stay on a platform that gets no new features. - Only the maintainer can answer; check each account's App Runner console for existing services or a usable Create button before M5.0. If none, go straight to the design-partner path.
- Register one cheap domain dedicated to ecsodus e2e (Route 53 Domains, kept year to year) and use it for issue #6 as well; never a personal or production domain.
- Keep retirement in the runbook, but as a separate opt-in section (
generate --retire) that is omitted by default so the default runbook stops at 100%. Operators will delete anyway; a gated delete is safer than none. Needs the fixes in P1 #2 and #3. - Yes, read-only, with the IAM actions listed in the runbook; treat alarms and target health as hard gates, metrics as advisory, and keep the operator
curl --resolvemandatory. - Always require the operator's list. Pre-fill
["sh","-c","<StartCommand>"]indecisions.template.ymlwith a# confirmmarker; never default it, since ecsodus cannot see whether the image has/bin/sh.
Unconfirmed claims¶
- App Runner managed certificate location (customer account vs. AWS-owned) and hence whether its validation CNAME names equal the new ACM certificate's — §3.2/§4.8. Confirm: in an eligible account, compare
DescribeCustomDomains.CertificateValidationRecordswithaws acm request-certificateDomainValidationOptionsfor the same FQDN. - Definition of "existing customer" and whether deleting the last service ends it — §1, §4.7, §8. No doc states it (https://docs.aws.amazon.com/apprunner/latest/dg/apprunner-availability-change.html). Confirm: AWS Support case, or probe in an eligible account after deleting all services.
- Which API call is blocked for ineligible accounts — §8. Confirm: run the probe; record the error code.
- App Runner runs only x86_64 images — §4.3. Community-confirmed, no official statement. Confirm: deploy an arm64-only image in an eligible account and observe failure.
- Whether App Runner's HTTP health checks count in
Requests— §4.8. Confirm: eligible account, HTTP health check, zero external traffic, read the metric. aws_apprunner_vpc_connectorreplacement on any change — §3.2. Provider docs mark no ForceNew (https://raw.githubusercontent.com/hashicorp/terraform-provider-aws/main/website/docs/r/apprunner_vpc_connector.html.markdown); the API has no Update operation so replacement is implied. Confirm:terraform planwith a changed subnet against a stub.- Import-only plan for a service with the default TCP health check and no
health_check_configurationblock — §3.2. Confirm: real import in an eligible or design-partner account. - Cost table — §8. Not checked against current prices.
- Terraform's
aws_apprunner_servicehonouringignore_changesonsource_configuration[0].image_repository[0].image_identifierwithout drift onauto_deployments_enabled/image_configuration— §3.4. Confirm: import thenstart-deployment, thenplanmust be empty.