Astra
Verdict¶
Revise. The useful product is a deterministic migration assessment with a tested ownership handoff—not broad Terraform generation alone. The single most important change is to replace the unsafe stack-teardown recipe with an explicit, verified CloudFormation-to-Terraform handoff. For six weeks of evenings, ship one narrowly supported migration path; the proposed two-source scope, arbitrary addons, source builds, tracing, networking and teardown is too large.
Findings¶
-
[P0]
--retain-resourcesis the wrong protection mechanism — §2, §4 — It applies to stacks already inDELETE_FAILED; it does not protect imported resources during normal stack deletion. Terraform import does not remove CloudFormation ownership. Fix: generate a reviewed retention-only template change usingDeletionPolicy: Retain, plusUpdateReplacePolicy: Retainwhere appropriate; verify it before detaching resources or deleting stacks. Handle nested stacks explicitly. Never deliberately fail deletion to obtain retention. Confidence: high. AWS deletion guidance, DeletionPolicy semantics. -
[P0] Custom-resource deletion can undo retained infrastructure — §4 — Replacing a custom resource’s creation behavior does not neutralize its Delete handler. Retaining a bucket, certificate or DNS resource alone does not establish that another handler cannot empty or remove it. Fix: document Create/Update/Delete side effects separately, identify externally created objects, and retain or safely detach destructive handlers before teardown. Unknown handlers must block teardown. UNSURE: the exact affected Copilot handlers require source-level verification; the proposed example list is not sufficient evidence.
-
[P0] “Service-first, environment-last” is unsafe for partial migrations — §3–4 — Unsupported Workers, Scheduled Jobs and other services can still depend on the same environment. Addons exist at both workload and environment levels, including nested stacks. Fix: inventory every consumer, recursively traverse addons, and prohibit environment teardown while any dependency remains. Add
retain under existing ownerandblocked/unknownfates; neither belongs under “drop.” Confidence: high. Copilot addon ownership. -
[P1] Terraform planning is not ownership transfer — §2, §5–6 —
terraform planpreviews imports;applycommits them to state. Import blocks require matching resource configuration and resource-specific IDs.-generate-config-outproduces a starting point, not a valid dependency graph or module configuration automatically. Fix: specify exact addresses, composite IDs and provider aliases; separate import-only adoption from infrastructure changes; require verified state membership and a subsequent no-change plan before teardown. Freeze other deployment controllers during handoff. Confidence: high. Configuration generation, import semantics. -
[P1] The deployed template is not the complete operational truth — §4–5 — It captures deployed overrides, but not necessarily live drift, autoscaled counts, current task revisions, resolved image digests or objects created indirectly by custom resources.
DescribeStackResourcesis also limited to 100 resources. Fix: combine templates, stack parameters/outputs, paginated resource enumeration and live service reads; record disagreements and unavailable values instead of silently choosing defaults. -
[P1] Express Mode needs a concrete compatibility matrix — §2–4, §9 — Its bundled ownership includes ECS/Fargate, load balancing, security groups and scaling. Crucially, private subnets produce an internal ALB; the proposed public endpoint plus private tasks topology cannot simply be assumed from its basic configuration. Existing Copilot ECS resources also cannot be assumed importable as an Express service. Fix: distinguish new Express deployment from ordinary ECS adoption; define networking, routing, storage and workload eligibility, and keep Express-managed children out of competing Terraform ownership. Confidence: high. Express resource and networking behavior.
-
[P1] AWS API support and Terraform support are different gates — §3, §9 — Express Terraform support exists; treating it as merely speculative is outdated. Conversely, AWS now accepts a custom task definition, including extra containers, while the inspected module does not expose a corresponding task-definition input. ADOT therefore needs a tested implementation path. Fix: pin and test the Terraform CLI, provider and module together; publish supported capabilities for that exact combination. “Terraform ≥1.5” alone is insufficient. Confidence: high on API; UNSURE on parity across every provider/module release. AWS API, provider resource, module implementation.
-
[P1] Concurrency cannot be translated directly into request-count scaling — §4 — Concurrent requests and requests per minute are different quantities; latency changes the relationship. Calling the mapping “approximate” does not supply a defensible target. Fix: require measured latency/load and load testing, or an explicit operator-selected scaling policy. Preserve min/max capacity separately. Confidence: high. Target-tracking units.
-
[P1] App Runner discovery omits security-critical features — §3–5 — A VPC connector describes outbound access, not private ingress. App Runner also supports private endpoints, dual-stack ingress and WAF associations. Missing these can expose an internal application or remove protection. Fix: discover ingress connections, endpoint security groups, WAF, address family, IAM/KMS dependencies and custom domains through their relevant APIs; block unsupported configurations. VPC endpoints do not substitute for NAT when arbitrary internet access is required. Confidence: high. App Runner private networking.
-
[P1] Source-service containerization is not recoverable from runtime commands alone — §4–5 — Configuration can come from repository
apprunner.yaml, with source-directory and build/runtime assumptions absent from the proposed inventory. Generated Dockerfiles can build successfully yet behave differently. Fix: defer source services, or require the source checkout and validated runtime recipe, covering build secrets, native dependencies, commands and image architecture. Confidence: high. Source repository configuration. -
[P1] DNS cutover and rollback lack necessary preconditions — §2, §4 — Weighted DNS requires a hostname the user controls; the AWS-owned App Runner URL cannot move. DNS weights are not exact request percentages, and cached connections outlive changes. Fix: branch the runbook by hostname/record type; issue and validate target TLS first, preserve delegation and validation records, test Host/SNI routing, and retain the source through a defined rollback window. Database changes must remain backward-compatible. Confidence: high. AWS migration guidance.
-
[P1] Parallel execution can duplicate work even for web services — §2, §3, §5 — Web containers may run schedulers, queue consumers or startup migrations. Excluding Worker/Scheduled types does not make parallel running safe. HTTP requests also are not guaranteed side-effect-free. Fix: require an operator declaration about background work; keep replacement consumers disabled until handoff; define job/queue draining and use explicitly safe verification endpoints.
-
[P1] Secret preservation needs an ownership and disclosure policy — §3–5 — Secret references alone do not preserve KMS access, rotation, resource policies or role permissions. Copying plaintext environment values into inventory/HCL can leak credentials; deleting a CloudFormation-managed secret can remove it without recovery. Fix: preserve references and dependencies, avoid fetching secret values, redact exports, and separately map task execution versus application IAM permissions. Confidence: high. CloudFormation secret deletion behavior.
-
[P1] CI can fight Terraform and does not replace every image-push trigger — §4–5 — A GitHub source workflow does not observe images pushed by arbitrary external producers. Independent deployment tooling can also advance task definitions that Terraform later rolls back. Fix: choose one deployment owner, pin image digests, document any intentional ignored attributes, and either support registry events or report the trigger change explicitly.
-
[P1] Validation is too late and too shallow — §6–7 — One happy-path migration per source cannot validate stateful teardown. LocalStack cannot establish real IAM, custom-resource deletion, DNS/TLS or Express lifecycle correctness. The sandbox may be ineligible to create App Runner services after closure to new customers. Fix: confirm eligibility and obtain per-run approval early; front-load one real migration. Test retention, interrupted handoff, partial migration and rollback with sentinel data. Gate on unexpected updates as well as destroys. UNSURE: this account’s App Runner eligibility. Availability rules.
-
[P1] Six-week scope is internally inconsistent and overcommitted — §3, §7 — Backend Service is both v1 and later; npm distribution appears despite a Python-only implementation plan. The milestones leave safety research and real testing until after most generation work. Fix: choose one source and one supported topology; defer source builds, arbitrary addons, Proton, npm and additional targets. Make the support matrix and first real migration the initial milestone.
-
[P2] Tracing requires more than adding a sidecar — §4 — App Runner tracing already depends on application instrumentation; an observability configuration alone does not prove traces work. Fix: preserve instrumentation and OTLP configuration, configure collector/export permissions, and verify an actual trace before declaring parity. Confidence: high. App Runner tracing requirements.
-
[P2] Competitive positioning overstates novelty and urgency — §1, §8–9 — The retirement dates check out, but App Runner remains supported for existing customers, including new services. AWS already supplies the basic migration flow; extract-cf2tf already emits imports. The located Kiro guide concerns EC2 migration, and fortem.dev currently describes Kubernetes diagnostics. Fix: remove “what does not exist”; position around verified discovery, explicit blockers and safe ownership transfer. Against Encore, emphasize retaining existing application code and delivering maintainable Terraform. Validate demand with three design partners before broadening. AWS status, extract-cf2tf, Kiro guide, Fortem, Encore approach.
-
[P2] Inventory alone cannot produce a credible monthly cost delta — §4 — Traffic, active versus idle time, scaling history, NAT throughput and shared-resource allocation are missing. Fix: provide a dated, region-specific baseline plus low/expected/high scenarios with explicit assumptions and temporary parallel-run costs.
Open questions (§10)¶
- App Runner image-only first, conditional on test-account eligibility. One public HTTP topology is more achievable in six weeks; defer source builds and Copilot teardown. If eligibility is unavailable, choose a narrow Copilot adoption path instead.
- Python. It matches the maintainer’s skills and boto3 ecosystem; Go adds delivery friction without solving the hard problems.
- Terraform only; thin pinned module wrapper for new compute, explicit resources for adoption. Module addresses and defaults complicate imports; CloudFormation/CDK outputs multiply the validation surface.
- No Proton implementation in v1. Publish an immediate export checklist if useful: AWS says Proton data disappears after October 7, so later reconstruction cannot promise complete provenance. Proton migration guide.
- Keep ecsodus provisionally. It is memorable enough; add “AWS container migration assessment and Terraform handoff.” UNSURE: package/domain availability and naming conflicts.
Missing from the plan¶
- Versioned support matrix and refusal rules; unsupported must never mean silently dropped.
- Account/region identity checks, cross-account roles, pagination, inventory timestamps and stale-inventory detection.
- App-level Copilot resources, StackSets, ECR artifacts, EFS, imported VPCs and externally owned dependencies.
- Encrypted state, locking, recovery checkpoints and a resource-to-owner ledger.
- Backups and restore verification before any stateful ownership change.
- Machine-readable plan checks and post-teardown resource/data verification.
- Health-check, timeout, session, routing and application-behavior parity criteria.
- Explicit licensing/attribution review for reused converters and Copilot code.
- Success metrics based on completed supported migrations and operator effort—not stars alone.