Council review: Claude Fable 5.1¶
(Verbatim final report of the Fable 5.1 reviewer, read-only, web access; 2026-09-29.)
Verdict¶
Revise before building. The idea is sound and the timing facts check out (Copilot EOS 2026-06-12, repo archived 2026-06-22 [1]; App Runner closed to new customers with Express Mode + Route 53 weighted DNS as the official path [2]; Proton EOS 2026-10-07 [3]). But the plan's safety spine is wrong in two places that would cause real outages if shipped as written: (a) delete-stack --retain-resources only works on stacks already in DELETE_FAILED [4], so step 4 of the product cannot work as described; and (b) Copilot's Lambda custom resources delete out-of-band resources on stack delete (the ACM certificate and Route 53 alias records) and the env-controller removes the ALB/NAT/EFS from the environment stack when the last workload stack is deleted [5][6][7]. The single most important change: replace "teardown with --retain-resources" with a Retain-patch step (ecsodus emits the deployed templates with DeletionPolicy: Retain/UpdateReplacePolicy: Retain on every imported resource plus the exact update-stack commands, applied before any delete), and treat every custom resource as a resource lifecycle to neutralize, not just a template line to replace. Then cut v1 to Copilot LBWS/Backend only; App Runner image-based is a thin second milestone, everything else is later.
Findings¶
-
[P0]
--retain-resourcescannot be used on a healthy stack — §2 step 4, §4 "Stack ownership order" — The CLI doc:--retain-resourcesis "for stacks in theDELETE_FAILEDstate" [4]. A firstdelete-stackon a healthy Copilot service stack deletes everything (including imported RDS/DDB/S3/secrets) before you ever get to pass the flag. — Following the runbook destroys the very resources Terraform just imported. — Fix: runbook step "Retain-patch": for each stack (workload, nested addons, env, app),get-template→ addDeletionPolicy: Retain+UpdateReplacePolicy: Retainto every resource with fate=import (and to theAWS::CloudFormation::Stackaddons resource) →update-stackwith the same parameters → verifydescribe-stack-resources→ thendelete-stack. Retain also covers resources removed during updates [8], which matters for finding 3. Nested addons stacks must be patched separately (retaining the nested stack resource leaves an orphan stack you delete afterwards, itself Retain-patched). -
[P0] Custom resources delete ACM certs and Route 53 records on stack delete — §4 Copilot custom resources —
dns-cert-validator.js(env stack) andwkld-cert-validator.js(service stack) request the ACM certificate and, on Delete, wait for it to be unused, thenDeleteCertificateand remove validation records [5];wkld-custom-domain.jsUPSERTs the alias A-records in the env/app/root hosted zones and deletes them on Delete [6]. These are notAWS::resources;DescribeStackResourcesshows only theCustom::handle, so an inventory built from stack resources will miss them, and an "import the cert into Terraform" plan silently loses it at teardown. — TLS/DNS outage at the moment the runbook says "safe teardown". — Fix: (a) inventory must discover them via ACM/Route 53 APIs and record provenance = "custom resource X"; (b) recommended fate for the Copilot cert is recreate (newaws_acm_certificate+ validation records on the new listener) and for alias records cut over to new records before teardown; (c) Retain-patch theCustom::logical IDs too (UNSURE whetherDeletionPolicy: Retainon a custom resource skips the Delete invocation — verify in e2e; regardless, ordering matters: delete stacks only after DNS points at the new ALB and the new cert is in place). -
[P0] Env-controller tears down shared env resources when the last workload leaves — §4 "Stack ownership order" — The service stack's
EnvControllerActioncallsUpdateStackon the env stack on Create/Update/Delete; on Delete it removes the workload fromALBWorkloads/NATWorkloads/EFSWorkloads/InternalALBWorkloadsandAliases[7]. The env template creates the public ALB, NAT gateways and EFS only when those lists are non-empty (CreateALB: !Not [!Equals [!Ref ALBWorkloads, ""]], same for NAT/EFS) [9]. So deleting the last LB service stack deletes the env ALB, NAT gateways and any EFS file system (stateful!) — even though the plan says "env last". — Data loss for EFS, outage if the ALB was imported into Terraform. — Fix: Retain-patch the env stack first (finding 1); flag EFS as stateful with fate=import; document that env teardown effectively starts on the first service delete. -
[P1] Express Mode fit criteria are wrong/incomplete — §2, §4 App Runner "Networking" — Verified Express Mode facts [10][11]: (a) ALB scheme is decided by the subnets: private subnets → internal ALB, public subnets → internet-facing ALB with public IPs on tasks; the first Express service in a VPC fixes the ALB's AZs/scheme for that VPC. So "public App Runner endpoint + VPC connector → tasks in private subnets behind an ALB" is not possible with Express Mode; it needs either public subnets (no NAT, public task IPs) or the plain
servicesubmodule with its own ALB. (b) Listener is HTTPS 443 only, host-header routed, shared by up to 25 services, auto-deprovisioned. Copilot LBWS without a domain is reached over plain HTTP on the ALB DNS name — no shared hostname, no gradual shift (AWS says the same for default App Runner URLs [2]). (c) Canary is the only deployment strategy; load-balancer config cannot be updated; single container unless you bringtaskDefinitionArn[10]. (d) Terraform provider limits:cpu256–4096,memory512–8192, notask_definition_arn, nocpu_architecturein the provider schema [12][13]; App Runner instances go to 4 vCPU/12 GB, so some don't fit. — Fix: write the "fits Express" predicate explicitly in the report: has custom domain or accepts a new hostname; single container; no service discovery/Service Connect consumers; no NLB/gRPC/HTTP-only needs; cpu/memory in provider range; subnet scheme matches desired exposure. Otherwiseservicesubmodule. -
[P1] ADOT sidecar plan collides with Express Mode via Terraform — §4 App Runner "Tracing" — Sidecars require
taskDefinitionArn(API) [12]; the Terraform resource has no such argument [13], and the module README says task definitions aren't managed [14]. Editing the task definition out-of-band is allowed and Express won't overwrite unless conflicting [15], but Terraform then fights the drift. — Fix: services with X-Ray tracing get fate "plainservicesubmodule + ADOT sidecar" in v1; document Express+sidecar as a manual path. -
[P1] Express Mode custom domains are outside Terraform's Express resource — §4 "Custom domains" — Custom domains are done by editing the Express-created ALB listener rule (host-header OR condition) and adding a cert to the listener [15]. Terraform must
data-lookup an ALB/listener it does not own and addaws_lb_listener_rule+aws_lb_listener_certificate; the module explicitly says custom domains need external config [14]. Removing the Express service deprovisions shared ALBs [10], destroying those rules. — Fix: generate the rule/cert resources with a clear "depends on Express-owned ALB" header and a rule-priority choice; document the dependency in the runbook. -
[P1] Copilot has an app-level stack (StackSet) the plan ignores — §3/§4 —
copilot app initcreates per-region app infrastructure via a StackSet: ECR repositories per service, the S3 artifact bucket, KMS key, and (with--domain) theapp.domain/env.app.domainhosted zones with NS delegation viadns-delegation.js. ECR repos hold the only copies of the images; hosted zones hold production DNS. — Fix: inventory the app stack; fates: ECR repos import, hosted zones import, artifact bucket drop-after; StackSet instance deletion uses--retain-stacksthen Retain-patched stack delete. -
[P1] Addon deletion defaults are worse than "retain" — §3 addons — Copilot addon templates carry no
DeletionPolicy: AuroraAWS::RDS::DBClusterdefaults toSnapshot(cluster deleted, snapshot kept → outage),AWS::SecretsManager::Secretis deleted without recovery, DynamoDB tables and S3 buckets default toDelete[16][8]. The fetched Aurora template showsEngineMode: serverless(v1) while docs say v2 is now the default; Aurora Serverless v1 was end-of-lifed by AWS, so real clusters may have been converted — read the live RDS state, not the template (UNSURE on converted state per account). — Fix: findings 1 + emitdeletion_protection = true,lifecycle { prevent_destroy = true }, and final-snapshot handling on every imported stateful resource; treat the Aurora secret as import. -
[P1] Imported resources should not live inside modules — §5 emit/terraform, §10 Q3 — Import blocks may target
module.x.resourcefrom root since 1.5 (in-module import blocks only since 1.16), but-generate-config-outcannot generate for module resources orcount/for_eachinstances and remains experimental [17][18]; importing a Copilot ALB/target group/cluster into terraform-aws-modules internals means matching that module's exact arguments orplanshows replacements (ALBname/subnetsare ForceNew). — Fix: emit imported resources as flat root resources whose arguments are derived from the deployed template; useexpress-service/servicesubmodules only for new compute. Don't rely on config generation. -
[P1] Scope does not fit ~60–90 evening hours — §7 — Two sources, Dockerfile generator, ADOT, CI, cost model,
verify, runbook generator and two paid e2e runs. The Copilot half alone has the app/env/workload/addons stack model plus 14 custom resources [19]. — Fix: v1 = Copilot LBWS + Backend (env + addons + app stack) → Terraform + Retain-patch runbook, one e2e. Dropverify, cost estimation, Dockerfile generation, ADOT, Proton from v1. App Runner image-based as a 1-week M2 if time remains. -
[P2] "None map to a Terraform resource" is overstated; some custom resources are for later-scope workloads — §4 — Actual
cf-custom-resources/lib: alb-rule-priority-generator, backlog-per-task-calculator (Worker), bucket-cleaner, cert-replicator (Static Site/CloudFront), custom-domain-app-runner (RDWS), custom-domain, desired-count-delegation, dns-cert-validator, dns-delegation, env-controller, trigger-state-machine (Job), unique-json-values, wkld-cert-validator, wkld-custom-domain [19]. Most outputs map cleanly: priority → literalpriority; desired count →desired_count+ignore_changes; cert validator →aws_acm_certificate+aws_acm_certificate_validation; custom domain →aws_route53_recordalias; dns-delegation →aws_route53_zone+ NS records; bucket-cleaner →force_destroy; env-controller → nothing (its parameters become explicit config). — Fix: mapping table with columns: source file, what it creates out-of-band, Delete behaviour, TF equivalent, fate. -
[P2] Deploy-path conflict for the Express resource — §2 output 3 — If CI deploys images with
aws-actions/amazon-ecs-deploy-express-service[2] while Terraform declaresprimary_container.image, everyterraform applyrolls the image back. — Fix:lifecycle { ignore_changes = [primary_container[0].image] }when a CI workflow is generated, or make Terraform the deployer. -
[P2] Health-check default mismatch — §4 — API default
healthCheckPathis/ping[12]; the ECS guide says/; the module defaults to/ping[14]; Copilot's default is/. — Fix: always emit the path explicitly from the deployed target group. -
[P2] Shared/stateful resources inside the service stack — §3 — LBWS
publish:SNS topics, service log groups, task/execution roles with addon managed policies, and Service Connect/Cloud Map registrations live in the workload stack; consumers (workers) break when it's deleted. — Fix: fates for topics = import; roles = recreate; runbook note on subscribers. -
[P2] App Runner feature list is incomplete — §3 — Missing: private endpoint (VPC Ingress Connection → internal ALB in private subnets) [20], health check config (TCP vs HTTP, path, intervals), instance size, encryption KMS key, runtime secrets, instance role → task role, access role → execution role, source-code connection and
apprunner.yaml(the API only returns build/run commands when config source isAPI; withREPOSITORYyou must read the repo), IP address type, WAF association. — Fix: table ofServicefields → fates [21]. -
[P2] Double-run hazard even for LBWS — §2 step 4 — A web service that also runs in-process consumers/cron processes twice during the parallel run. — Fix: report line "parallel run is safe only for stateless request handling; found: [sqs/sns/eventbridge refs]".
-
[P3] Plan tests against LocalStack — §6 — LocalStack does not list Express Mode APIs (UNSURE) and
terraform planwith import blocks is read-only against the real account anyway; use the sandbox for plan tests (only the deployed Copilot sample costs money; one evening of ALB + 2 NAT ≈ a few dollars). -
[P3] Kiro claim — §1/§9 — The Kiro + MCP guide AWS published is EC2 → Express Mode [22]; there is no Copilot-specific Kiro path yet. Say "generic MCP tooling".
Open questions (§10)¶
- Copilot first (LBWS + Backend), App Runner image-based as a thin M2. Copilot is where the custom-resource/CFN-teardown risk (the differentiator) lives; App Runner → Express is one Terraform resource and AWS already documents it end to end.
- Python. Nothing is CPU-bound; boto3 + CFN-aware YAML loader + Jinja +
terraform fmtin CI suffices;uvx ecsodusis acceptable distribution. - Hybrid, Terraform only. Flat root resources for everything imported;
express-service/servicesubmodules for new compute. No CDK/CFN emit in v1 — but document a "stay on CloudFormation" option (strip custom resources, keep stacks, deploy withaws cloudformation deploy+ ecspresso), the lowest-risk path AWS itself recommends [1]. - No. Eight days is too short and after 2026-10-07 the Proton API is gone [3], leaving a generic CFN-adoption problem. At most one docs page.
- Fine. Distinct, searchable, no "Copilot" collision; confirm PyPI/npm availability today, title docs pages exactly as in §2.
Missing from the plan¶
- The Copilot app stack / StackSet layer (ECR repos, artifact bucket, KMS, hosted zones + NS delegation) and its retention procedure.
- A Retain-patch step and the ordering rule: DNS cutover → new cert live → Retain-patch every stack → delete workloads → delete addons orphans → delete env → delete app.
- Custom-resource Delete behaviour as a first-class inventory field (what each Lambda deletes out-of-band).
- Env-level resources conditional on workload lists (ALB, NAT, EFS, internal ALB) and the env-controller side effect on first service delete.
- Copilot SSM metadata (
/copilot/applications/...) andcopilot secret initSecureString parameters: use for inventory, decide fates. - Terraform safety rails in generated code:
prevent_destroy,deletion_protection,ignore_changeson desired_count/image; a "plan must show 0 destroy / 0 replace on fate=import" check the runbook can grep for. - Precise Express Mode fit predicate (finding 4) and the fallback to
service. - Imported-ALB vs new-ALB decision per path (Express path: always new ALB;
servicepath: option to import Copilot's ALB and listener). - Terraform/provider minimums from the module: Terraform ≥ 1.5.7, AWS provider ≥ 6.41 [14].
- Manifest features that change architecture:
network.connect(Service Connect),nlb,http.additional_rules,sidecars,storage.volumes(EFS),observability.tracing,taskdef_overrides, CDK/YAML overrides (read the deployed template, but also diff manifest→template so the report can say "override detected"). - Multi-account/multi-region environments.
- Time budget for the mapping-table research: the 14 Lambdas plus partials are the bulk of M0 and should be scoped as their own deliverable.
Sources: [1] https://aws.amazon.com/blogs/containers/announcing-the-end-of-support-for-the-aws-copilot-cli , https://github.com/aws/copilot-cli · [2] https://docs.aws.amazon.com/apprunner/latest/dg/apprunner-availability-change.html · [3] https://docs.aws.amazon.com/proton/latest/userguide/proton-end-of-support.html · [4] https://docs.aws.amazon.com/cli/latest/reference/cloudformation/delete-stack.html · [5] https://github.com/aws/copilot-cli/blob/mainline/cf-custom-resources/lib/dns-cert-validator.js , .../wkld-cert-validator.js · [6] https://github.com/aws/copilot-cli/blob/mainline/cf-custom-resources/lib/wkld-custom-domain.js · [7] https://github.com/aws/copilot-cli/blob/mainline/cf-custom-resources/lib/env-controller.js · [8] https://docs.aws.amazon.com/AWSCloudFormation/latest/TemplateReference/aws-attribute-deletionpolicy.html · [9] https://github.com/aws/copilot-cli/blob/mainline/internal/pkg/template/templates/environment/cf.yml · [10] https://docs.aws.amazon.com/AmazonECS/latest/developerguide/express-service-work.html · [11] https://docs.aws.amazon.com/AmazonECS/latest/developerguide/express-service-overview.html · [12] https://docs.aws.amazon.com/AmazonECS/latest/APIReference/API_CreateExpressGatewayService.html · [13] https://github.com/hashicorp/terraform-provider-aws/blob/main/internal/service/ecs/express_gateway_service.go , https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/ecs_express_gateway_service · [14] https://github.com/terraform-aws-modules/terraform-aws-ecs/blob/master/modules/express-service/README.md · [15] https://docs.aws.amazon.com/AmazonECS/latest/developerguide/express-service-advanced-customization.html · [16] https://github.com/aws/copilot-cli/tree/mainline/internal/pkg/template/templates/addons · [17] https://developer.hashicorp.com/terraform/language/import/generating-configuration · [18] https://www.naviteq.io/blog/terraform-1-16-import-blocks-finally-work-inside-modules/ · [19] https://github.com/aws/copilot-cli/tree/mainline/cf-custom-resources/lib · [20] https://docs.aws.amazon.com/apprunner/latest/dg/network-pl.html · [21] https://docs.aws.amazon.com/apprunner/latest/api/API_Service.html · [22] https://aws.amazon.com/blogs/containers/migrate-amazon-ec2-to-ecs-express-mode-using-kiro-cli-and-mcp-servers