ecsodus: plan (r2, final after council)¶
Safe exit from AWS Copilot CLI (and later App Runner) to Terraform-managed ECS. 2026-09-29.
The r1 draft and the three council reviews are in council/.
1. Problem and timing¶
| Tool | Status (2026-09-29) | Deadline | AWS path |
|---|---|---|---|
| Copilot CLI | End of support 2026-06-12, repo archived 2026-06-22 (3.7k stars, 446 forks) | None. Stacks keep running, the CLI goes stale | ECS Express Mode, CDK L3, or keep the CloudFormation |
| App Runner | Closed to new customers 2026-04-30. Existing services run, no new features | None | ECS Express Mode + Route 53 weighted DNS |
| Proton | End of support 2026-10-07 | Yes, 8 days away | CloudFormation Git Sync, CodePipeline, GitHub Actions |
The gap. Leaving Copilot for Terraform means importing what Copilot built and then deleting Copilot's CloudFormation stacks without deleting the infrastructure. Doing that naively destroys data and takes production down:
-
Custom resources. Copilot's Lambda custom resources delete ACM certificates, DNS records and NS delegations, and empty buckets, when their stack is deleted.
-
Env controller. Deleting the last service stack makes the env controller delete the shared ALB, the NAT gateways and the EFS.
-
Addons. Addon databases sit in nested stacks with no
DeletionPolicy. - App stack. The app-level StackSet owns the ECR repositories and the hosted zone.
Generic tools cover none of this: extract-cf2tf, -generate-config-out and AWS's prose guides.
ecsodus is the tool that knows where Copilot's landmines are. That is the product.
2. Product¶
ecsodus is a read-only Python CLI. It never applies Terraform, never modifies or deletes AWS
resources, never shifts traffic and never reads secret values.
ecsodus inventory --app myapp --env prod [--copilot-dir ./copilot] -> inventory.json
ecsodus report inventory.json -> REPORT.md
ecsodus generate inventory.json --out ./infra -> Terraform + retain patches + RUNBOOK.md
ecsodus check plan.json --inventory inventory.json -> exit 1 if any import-fated address would be destroyed/replaced
-
inventory. Walks the app stack and StackSet instances, the env stack, the workload stacks and their nested addon stacks, recursively and with pagination. Reads the live service state and SSM metadata (
/copilot/applications/...). Finds out-of-band resources created by custom resources (ACM certificates, Route 53 records) through the ACM and Route 53 APIs. Records provenance, account, region and a timestamp for every field. -
report. Assigns every resource a fate:
import,recreate,retain-under-existing-owner,drop-after-cutoverorblocked.blockedmeans unsupported, and ecsodus refuses to generate anything for it, so nothing is ever dropped silently. The report also lists:- per stack, what deleting it would destroy as a side effect,
- which services could run Express Mode, and why the others can't,
- a double-work question for each service,
- "keep CloudFormation" as the zero-risk baseline.
- generate.
-
Flat root-module
aws_*resources withimportblocks for everything fatedimport. Their arguments come from the deployed template and the live reads, not from module internals. -
prevent_destroyanddeletion_protectionon stateful resources. ignore_changesondesired_count, and on the image when CI deploys.-
Retain-patch templates for every stack, nested stacks included: the deployed template with
DeletionPolicy: RetainandUpdateReplacePolicy: Retainon every imported resource and everyCustom::*resource, includingEnvControllerAction. -
RUNBOOK.md. - check. Parses
terraform show -jsonand fails if any imported address would be destroyed or replaced. The runbook requirescheckto pass before every apply.
Two modes¶
-
Adopt in place (default for Copilot). Import the existing cluster, service, task definition family, ALB, listeners, rules, target groups, security groups, VPC, subnets, Cloud Map, log groups, addons, ECR repos and hosted zone. Traffic does not move.
-
Rebuild in parallel (opt-in). For changing shape, for example to Express Mode, and later for App Runner. Adds new compute, a new ALB where needed, an ACM certificate issued ahead of time, atomic Route 53 change batches, and weighted cutover with the source kept alive for TTL + 24h.
Runbook order (adopt in place)¶
-
Freeze all Copilot use: no
copilot deploy, pipelines paused. Take backups and snapshots, and verify that they restore. -
terraform plan→ecsodus check→ import-onlyapply, with no attribute changes. Then a second plan must show zero changes. -
Apply the retain patches with
update-stackand the same parameters, in this order: app, env, then each workload with its nested addons. Verify withget-templatethat the policies are present. -
Delete the workload stacks, then any orphaned addon stacks, then the env stack. For the app stack, delete StackSet instances with
--retain-stacks, then the retain-patched stack. -
After the teardown, confirm that every imported resource still exists. Plan again: zero changes. List the leftover Lambdas, roles and log groups for manual cleanup. Note that retained resources keep their
aws:cloudformation:*tags. -
Never run
copilot app deleteorcopilot env deleteafter migrating.
3. Scope¶
v0.1 (in):
-
Copilot Load Balanced Web Service and Backend Service, with their env stack (created or imported VPC), workload and env addons (Aurora, RDS, DynamoDB, S3), secrets by reference, aliases and custom domains, the app stack and StackSet, and adopt-in-place mode.
-
Detected but
blocked, so the report names them: Worker Service, Scheduled Job, Request-Driven Web Service, Static Site, NLB, CloudFront CDN, sidecars, Service Connect consumers, managed EFS (supported if time allows),copilot pipeline, and multi-account environments.
Later:
-
v0.2: App Runner image-based services in rebuild-in-parallel mode. The inventory covers custom domains, VPC connectors, VPC ingress connections (private services), WAF, health check (TCP by default, so the target needs an HTTP path), port, instance size, KMS key and roles. Scaling is operator-selected, because concurrency cannot be converted to a request-count target.
-
v0.3: Worker Service and Scheduled Job, rebuild mode into Express, App Runner source-based (requires the repo and its
apprunner.yaml), ADOT tracing.
Non-goals:
- A deploy CLI (ecspresso exists).
- CloudFormation or CDK output.
- Proton tooling. Instead, one docs page that points to generic CloudFormation adoption.
- Cost estimation. The report gives a static table of line items instead.
- A hosted service, and npm.
4. Knowledge base (M0, its own deliverable)¶
docs/copilot-custom-resources.md is a table of Copilot's 14 custom resources, built from the
archived source (Apache-2.0):
-
Columns: the resource; what it creates outside CloudFormation; what its Delete handler does; the Terraform equivalent; the fate.
-
Resources:
alb-rule-priority-generator,backlog-per-task-calculator,bucket-cleaner,cert-replicator,custom-domain,custom-domain-app-runner,desired-count-delegation,dns-cert-validator,dns-delegation,env-controller,trigger-state-machine,unique-json-values,wkld-cert-validator,wkld-custom-domain.
A second table covers the conditional env resources and addon DeletionPolicy defaults.
These tables double as the SEO docs pages.
5. Architecture¶
-
Stack: Python 3.11+, boto3, a CloudFormation-aware YAML loader, Jinja2. Distributed with
uvx ecsodusorpipx. -
Pins: Terraform ≥ 1.5.7 and AWS provider ≥ 6.41 (tested together); terraform-aws-modules/ecs v7.x, used only for newly created compute.
-
Modules:
sources/copilot: stack graph walker, manifest-vs-template override detection, live reads.model/: target-neutral, with provenance on every field.mappers/: fate and gap rules, one per knowledge-base row.emit/terraform: resource blocks plus import blocks.emit/retain_patch: rewrites each template, keeping it byte-stable otherwise.emit/runbook.check/: the plan guard.
- Rules: never fetch secret values, and redact the environment from
inventory.json.
6. Validation¶
-
Golden tests: fixtures from real Copilot-generated templates (the archived repo's integration templates) and recorded API responses; snapshot the generated Terraform and retain patches; run
terraform validateandtflintin CI. -
checktests: syntheticterraform show -jsonplans with destroy and replace actions. -
One real end-to-end run, early (M1), not at the end:
-
Deploy a Copilot sample app (LBWS + Aurora addon + custom domain) with the archived CLI, and write sentinel data to it.
-
Run ecsodus, apply, retain-patch and tear down.
-
Assert that the sentinel data, the ALB, the certificate, the DNS and the ECR images all survive.
-
Settle the three unverified questions listed in council/README.md.
- Test an interrupted handoff and a rollback.
- Needs a Route 53 domain and costs a few dollars (ALB, NAT, Aurora): propose the run, get approval, and tear down the same day.
-
7. Milestones (≈6 weeks of evenings)¶
| # | Deliverable | Est. |
|---|---|---|
| M0 | Repo, CI, fixture corpus, custom-resource knowledge base | 1 wk |
| M1 | inventory + report for Copilot (report usable on its own); first real e2e deploy of the sample |
1.5 wk |
| M2 | generate (Terraform + retain patches + runbook) + check; e2e teardown with sentinel data |
2 wk |
| M3 | v0.1 release: docs pages titled for the searches people make, PyPI, launch posts (r/aws, dev.to, HN "Show") | 0.5 wk |
| M4 | App Runner image-based (v0.2) | 1 wk, if time allows |
Ship the report publicly after M1 even if generate slips. A readiness report that lists the
landmines is useful on its own.
8. Success measures (first 90 days)¶
- 3 completed migrations reported by users. 10 reports run and shared as feedback issues.
- Zero reports of lost data or a production outage caused by following the output.
- Cited in at least one migration guide or thread. Stars are secondary.
9. Risks¶
-
Small audience. Keep v0.1 narrow and ship the report first. Recruit 2–3 design partners from the copilot-cli issues and discussions, and from Reddit threads, before M2.
-
Wrong import or teardown causes an outage. Mitigations: read-only tool, retain patches,
prevent_destroy, thecheckgate, backups in runbook step 1,blockedover guessing, and the sentinel e2e test. -
Provider or module drift. Pin versions and publish a support matrix for each release.
- Competitors: AWS guides, Encore, fortem.dev, generic MCP tooling. They cover rebuilds. The wedge is safe adoption of existing Copilot infrastructure into maintainable Terraform, with no rewrite of the application code.
10. Decisions (from council)¶
- Copilot first, adopt in place. App Runner image-only goes to v0.2.
- Python.
-
Terraform only. Flat root resources for imports, terraform-aws-modules for new compute, and Express only under the fit test.
-
No Proton tooling. One docs page.
- Name: ecsodus. Tagline: "Safely migrate AWS Copilot apps to Terraform." Put the exact search phrases in the repo description, the PyPI summary and the docs titles.