ecsodus: plan M5, App Runner (r4, approved)¶
r4 (approved unanimously in council round 3, 2026-10-08)
App Runner support (M5a) and a rebuild-in-parallel mode (M5b). 2026-10-08.
Status: r4, approved. All three reviewers voted APPROVE WITH NOTES on r3 in round 3; r4 folds
every round-3 note into the sections below, and §18 lists them as numbered build requirements.
History (kept in the repository, not on the docs site): docs/council/v0.2/plan-r1.md,
plan-r2.md, plan-r3.md; the council record is
council/v0.2/README.md. Round 1: brief and
reviews gpt-6-astra, Claude Opus 5.5,
Claude Fable 5.1, all REVISE (11, 4 and 5 P1s). Round 2:
brief and votes gpt-6-astra (REJECT,
2 P1s), Claude Opus 5.5 and Claude Fable 5.1
(both APPROVE WITH NOTES). Round 3: brief and votes
gpt-6-astra, Claude Opus 5.5 and
Claude Fable 5.1, all APPROVE WITH NOTES. §15, §16 and §17 map every
round-1, round-2 and round-3 finding to the section that addresses it; §17 also maps the
maintainer's review of PR #28 (the new_host record is staged after the ALB exists). This plan extends
PLAN.md r4 (approved) and changes no v0.1 decision. Where it relies on behaviour
nobody has verified, it says (UNCONFIRMED).
Naming. PLAN.md §3 and §7 call this scope "v0.2". The package is already at 0.2.0
(Worker Services and Scheduled Jobs, see CHANGELOG.md), so this plan does not use "v0.2" for
the scope. The milestones are M5a (adopt RDWS in place, release 0.3.0) and M5b
(rebuild in parallel, release 0.4.0). Retirement of the App Runner service is M5c, a
separate artifact with no release until it has run on real App Runner (§5). The items PLAN.md §3
lists under "v0.3" (source-based services, ADOT, Express rebuilds of workers) are "after M5" here;
they are not release 0.3.0. The file name PLAN-v0.2.md is historical.
1. Problem and who has it¶
| Fact (2026-10-08) | Consequence |
|---|---|
| App Runner closed to new customers on 2026-04-30. Existing services keep running, with no new features. | No deadline, but no future. Accounts that never used App Runner cannot create a service. |
| Copilot CLI end of support 2026-06-12; repo archived. | Copilot's RDWS stacks keep running. The CLI goes stale. |
Copilot custom-resource Lambdas run nodejs20.x: create blocked from 2027-07-29, update from 2027-08-31. |
Every Copilot stack, including RDWS stacks, should be off Copilot before mid-2027. |
A Copilot Request-Driven Web Service (RDWS) is an AWS::AppRunner::Service. |
RDWS users have both problems at once. |
Who has it.
-
(a) Copilot RDWS users. Today ecsodus detects the RDWS, keeps its stack, and by kept-status propagation (docstring of
src/ecsodus/mappers/fates.py) also keeps its env stack and the app stack and StackSet. One RDWS in an environment stops the whole app from leaving Copilot. M5a closes this gap. -
(b) Teams with standalone App Runner services (console, CloudFormation, CDK, Terraform, Pulumi). They have no Copilot landmines. Their exit is "rebuild on ECS", which AWS documents (Express Mode plus Route 53 weighted DNS).
Decision: M5 generates for (a) only. (b) gets read-only inventory and report in 0.4.0.
-
ecsodus's safety model needs a known owner. Retain patches, the closure rule and the teardown order all read CloudFormation stacks. For (b) the owner can be anything.
-
The wedge is Copilot (PLAN.md §9). For (b), ecsodus would compete with AWS's own guide without an advantage.
-
The rebuild model (§4) reads the App Runner API, not the stack. Adding (b) later is a source adapter. If it is added, it stops at cutover: with no known owner there is no retirement.
2. Two options, evaluated¶
v0.1's adopt-in-place (ADR-0003) imports an ECS service that already exists. An RDWS has no ECS service, task definition, target group or ALB. "Adopt in place to ECS" does not exist for App Runner. There are two real options.
2.1 Option A: adopt the App Runner service in place, into Terraform¶
Import AWS::AppRunner::Service as aws_apprunner_service, the VPC connector and ingress
connection as their aws_apprunner_* resources, the custom-domain association and its CNAMEs as
aws_apprunner_custom_domain_association and aws_route53_record. The rest of an RDWS stack
(roles, security groups, SNS topics, addons) already has a mapper (src/ecsodus/mappers/). Then
the stack hands off like any other, and the env and app stacks can follow.
2.2 Option B: rebuild on ECS Fargate in parallel, cut over DNS, retire App Runner¶
Generate a new ECS service, task definition, target group, ALB and certificate. Run both. Shift the custom domain with Route 53. Retire App Runner later.
2.3 Tradeoffs¶
| A: adopt App Runner | B: rebuild on ECS | |
|---|---|---|
| Traffic moves | No | Yes (DNS cutover) |
| New infrastructure | None | ECS service, task def, execution role, TG, ALB, cert, SG |
| Fits the existing gates | Yes: retain patches, closure, check --phase import, unchanged |
Needs new phases (§4.8) |
| Removes Copilot | Yes, the whole app | Only after A (§4.1) |
| Leaves you on App Runner | Yes. Closed to new customers, no new features, no announced end date | No |
| Cost of a mistake | An import plan that is not import-only fails the gate. A replacement of aws_apprunner_service would give a new default URL; check rejects replacements |
Double-running side effects, DNS outage, TLS mismatch |
Default *.awsapprunner.com URL |
Kept | Cannot move. Clients using it must change URL |
| Operator input needed | None beyond v0.1 | Decisions file (§4.4) |
| Build effort | Small: a few mappers, out-of-band reads, fixtures | Large: new mode, emitter, phases, runbook |
| Testable in our sandbox | Only if the account can create App Runner (§9) | ECS side yes; App Runner side same problem |
A is not a dead end. It takes Copilot out of the loop before anything risky happens. After A, the App Runner service is a plain Terraform resource with no custom-resource Delete handlers and no env-controller. B then starts from that simpler state.
A alone is not enough. App Runner gets no new features. Teams that want to leave it need B.
2.4 Recommendation: M5a then M5b¶
-
M5a (release 0.3.0): adopt RDWS in place. Copilot can be removed from apps that have an RDWS. No traffic moves.
-
M5b (release 0.4.0): rebuild in parallel, through cutover. Input is an App Runner service whose Copilot hand-off is proven complete (§4.1). There is no path for a still-Copilot RDWS. The default runbook stops before retirement.
-
M5c (no release yet): retirement, a separate opt-in artifact, unavailable until it has run on real App Runner (§5).
-
Standalone App Runner (b): read-only inventory and report in 0.4.0.
3. M5a: adopt RDWS in place¶
3.1 What an RDWS stack holds¶
The fixture tests/fixtures/copilot/rendered/workloads/rdws-test.stack.yml (Copilot v1.29.0
render) has 14 resources. Rows marked alias or private come from the archived Copilot
rd-web/cf.yml template (reviewed by Opus and Fable in round 1); M5a.0 adds golden renders for
them.
| Logical ID | Type | Condition | Mapper today | M5a fate |
|---|---|---|---|---|
AccessRole |
AWS::IAM::Role (trust build.apprunner.amazonaws.com, AWSAppRunnerServicePolicyForECRAccess) |
NeedsAccessRole (ECR images) |
yes | import |
InstanceRole |
AWS::IAM::Role (trust tasks.apprunner.amazonaws.com; inline DenyIAM, tag-scoped SSM/Secrets/KMS, Publish2SNS) |
— | yes | import |
Service |
AWS::AppRunner::Service |
— | no | import (new) |
AddonsStack |
AWS::CloudFormation::Stack |
HasAddons |
yes | nested-wrapper |
customersSNSTopic, mytopicfifoSNSTopic |
AWS::SNS::Topic |
— | yes | import |
customersSNSTopicPolicy, mytopicfifoSNSTopicPolicy |
AWS::SNS::TopicPolicy |
— | yes | import |
EnvControllerAction |
Custom::EnvControllerFunction (Parameters: [NATWorkloads,]) |
— | yes | manual-cleanup (retained) |
EnvControllerFunction, EnvControllerRole |
AWS::Lambda::Function (nodejs20.x), AWS::IAM::Role |
— | yes | manual-cleanup (retained) |
ServiceSecurityGroup |
AWS::EC2::SecurityGroup |
— | yes | import |
EnvironmentSecurityGroupIngressFromServiceSecurityGroup |
AWS::EC2::SecurityGroupIngress (env SG trusts ServiceSecurityGroup, all ports) |
— | yes | import |
VpcConnector |
AWS::AppRunner::VpcConnector (SecurityGroups: [ServiceSecurityGroup, imported EnvironmentSecurityGroup]) |
only with network.vpc.placement: private; otherwise EgressType: DEFAULT |
no | import (new) |
alias: CustomDomainAction |
Custom::CustomDomainFunction (custom-domain-app-runner.js) |
alias set |
yes (knowledge base) | manual-cleanup (retained) |
alias: CustomDomainFunction, CustomResourceRole |
AWS::Lambda::Function, AWS::IAM::Role |
alias set |
yes | manual-cleanup (retained) |
| private: ingress connection | AWS::AppRunner::VpcIngressConnection; VpcEndpointId is !GetAtt EnvControllerAction.AppRunnerVpcEndpointId or a literal |
http.private |
no | import (new); endpoint ID from a live read |
Notes:
-
AutoScalingConfigurationArnis a literal ARN built fromcount(autoscalingconfiguration/high-availability/3): a configuration the user created, possibly shared, never owned by the stack. Fate:external-reference, listed in the report, never imported in M5a and never in any delete set. The template ARN has no revision UUID; the value comes fromDescribeService(§3.3). -
Tracing does not create a resource. Copilot sets
ObservabilityConfigurationArnto the AWS-managed literalobservabilityconfiguration/DefaultConfiguration/1/00000000000000000000000000000001. Fate:external-reference. There is no observability mapper. -
App Runner creates its own log groups (
/aws/apprunner/<name>/<id>/applicationand/service). No stack owns them. They survive any stack delete.
3.2 New mappers and out-of-band reads¶
| Source | Terraform | Import ID | Notes |
|---|---|---|---|
AWS::AppRunner::Service |
aws_apprunner_service |
service ARN | every argument from template + DescribeService, provider defaults emitted explicitly (§3.3); prevent_destroy; ignore_changes on the image identifier (deploy owner, §3.5) |
AWS::AppRunner::VpcConnector |
aws_apprunner_vpc_connector |
connector ARN | subnets and security_groups from live DescribeVpcConnector, so the import keeps the actual membership (both SGs for a Copilot render); the provider marks security_groups, subnets and vpc_connector_name ForceNew and updates only tags (provider source, internal/service/apprunner/vpc_connector.go), so any argument mismatch plans a replace, which check rejects |
AWS::AppRunner::VpcIngressConnection |
aws_apprunner_vpc_ingress_connection |
ARN | private RDWS only |
| custom domain (out of band) | aws_apprunner_custom_domain_association |
<domain_name>,<service_arn> |
from DescribeCustomDomains; enable_www_subdomain read from live (§3.3); prevent_destroy |
| domain CNAME (out of band) | aws_route53_record (CNAME, TTL 60) |
<zone-id>_<name>_CNAME |
in the root-domain zone (an RDWS alias is a one-level subdomain of the app's root domain), written by the handler through AppDNSRole; prevent_destroy |
| validation CNAMEs (out of band) | aws_route53_record (CNAME) each |
as above | from CertificateValidationRecords; they keep App Runner's certificate renewing |
env AppRunnerVpcEndpoint + SG |
aws_vpc_endpoint etc. |
— | already mapped; present when AppRunnerPrivateWorkloads is non-empty (ENV_CONDITIONS in src/ecsodus/knowledge.py) |
| auto-scaling and observability configurations | — | — | external-reference (§3.1) |
Zone discovery. The handler takes HostedZones[0] of ListHostedZonesByName. ecsodus does
not copy that. The zone rules come in two parts:
-
Zone rules (every use): exactly one public zone whose name equals the root domain, in the account being migrated, delegated (its NS set matches the parent's delegation, read with a public DNS query).
-
Adoption rule (only when adopting an existing custom domain): the zone also holds a CNAME at the alias whose value is the service's
DNSTarget.
Anything else fails closed. A new hostname for the no-custom-domain branch (new_host, §4.9)
can never meet the adoption rule; it is checked against the zone rules, the CAA check (§4.8 #1)
and the absence of any record of any type at that name. Its alias record is created only after
the ALB exists, in its own gated step (§4.8 #4a). A zone in another account (a
multi-account app, where AppDNSRole sits in the app account) is blocked, consistent with
v0.1 (PLAN.md §3).
New read-only live calls: apprunner:DescribeService, DescribeVpcConnector,
DescribeVpcIngressConnection, DescribeCustomDomains, DescribeAutoScalingConfiguration,
ListTagsForResource; wafv2:GetWebACLForResource; route53:ListHostedZonesByName,
GetHostedZone, ListResourceRecordSets. None returns a secret value. DescribeService returns
RuntimeEnvironmentVariables in plaintext, the same exposure as a task definition's
environment; the v0.1 secrets contract applies (inventory.json 0600, HCL with env values in
.gitignore, encrypted state).
Blocked in M5a (fail closed, named in the report): source-code services (CodeRepository),
a WAF web ACL whose owner is not discoverable, a zone in another account, any DescribeService
field the mapper does not know.
3.3 Provider defaults that differ from what Copilot deploys¶
Each of these would make the import plan non-empty and fail the gate. The knowledge base records them, and the mapper emits the live value explicitly.
| Argument | Provider default | Copilot / live | Rule |
|---|---|---|---|
auto_deployments_enabled |
true |
Copilot sets false |
emit the live value |
enable_www_subdomain (ForceNew) |
true |
Copilot's handler omits it, so the API default true applies |
read EnableWWWSubdomain from DescribeCustomDomains; never assume. A mismatch would plan a disassociate and re-associate |
auto_scaling_configuration_arn |
— | template ARN has no /<uuid>; the provider stores the full ARN |
take the full ARN from DescribeService |
observability_configuration_arn |
— | AWS-managed literal | emit the literal; external-reference |
health_check_configuration |
block defaults TCP, path /, interval 5, timeout 2, healthy 1, unhealthy 5 (provider docs) |
App Runner default is the same, TCP/5/2/1/5 | emit the live block even when it is the default (UNCONFIRMED only that a real import plan shows no diff, §9) |
Every association with EnableWWWSubdomain: true also covers www.<alias>, with no DNS record
for it. The M5a report says so. M5b does not carry www.<alias> to ECS (§4.4).
3.4 Safety: what a plain RDWS stack delete destroys¶
| Resource | Plain delete does | With retain patch |
|---|---|---|
Service |
Deletes the App Runner service. Its *.awsapprunner.com URL is gone for good |
kept |
CustomDomainAction |
Delete handler disassociates the domain and DELETEs the domain CNAME and every validation CNAME: the custom domain stops resolving | handler not invoked. The 2026-09-30 run (e2e) verified this for Custom::EnvControllerFunction; the mechanism is generic CloudFormation, so the same holds here |
EnvControllerAction |
Removes the workload from NATWorkloads and AppRunnerPrivateWorkloads, then updates the env stack. If it was the last user: NAT gateways, their routes and the App Runner VPC endpoint are deleted, breaking every other private workload |
handler not invoked |
ServiceSecurityGroup + ingress |
Deleted; the env SG rule that trusts it goes too | kept |
VpcConnector, roles |
Deleted | kept |
| SNS topics | Deleted; other services' subscriptions break | kept |
AddonsStack |
Nested addons with no DeletionPolicy: Aurora snapshot-and-delete, DynamoDB and secrets deleted |
kept |
The v0.1 mechanics apply unchanged: retain-all patch, change set with check --changeset,
verify-retain, import, check --phase import, check --state, teardown, check --phase steady.
No exception for App Runner.
3.5 Deploy owner after M5a¶
Copilot deployed by changing the ContainerImage parameter. After hand-off the named deploy
owner is CI calling aws apprunner update-service (or start-deployment for a mutable tag).
Terraform gets ignore_changes on
source_configuration[0].image_repository[0].image_identifier, the analogue of
ignore_changes = [task_definition] on ECS (PLAN.md §2.4). If live
AutoDeploymentsEnabled is true, App Runner itself is the deploy owner and the report says so.
(UNCONFIRMED that a start-deployment leaves the plan empty: import, deploy, then plan, §9.)
The report also searches the other workloads' environment variables for the service's
*.awsapprunner.com URL, since those callers cannot follow a DNS cutover later.
4. M5b: rebuild-in-parallel mode¶
4.1 Preconditions¶
Copilot must be provably gone. The M5a manifest records an intended hand-off, not a finished
teardown. A new read-only command, ecsodus verify-handoff --manifest M, proves it:
-
Every stack in the manifest's
handoff_stacksis gone (DescribeStacks: not found orDELETE_COMPLETE). -
No stack tagged with this app and env remains, and the env stack itself was handed off and deleted. If any Copilot workload remains in the env, its env-controller can still rewrite the env stack (NAT gateways, the App Runner endpoint) that the rebuild depends on, so the rebuild is blocked. An app stack kept for other envs is allowed; the report names it.
-
terraform state listholds the App Runner service at its manifest address, and a fresh plan of the M5a root passescheck --phase steady. If the service has a custom domain, state also holds the association and the domain CNAME. With no custom domain (a supported branch, §4.9) these two are not required, andDescribeCustomDomainsmust return none. -
The service ARN in state equals the live service (
DescribeService), in the same account and region, and the state'simage_identifierequals the live one. After a synchronized deployment (below) the HCL still holds the tag;ignore_changesplus refresh carries the live digest into state, and this check proves it, soworker-off(§4.10) sends the digest.
It writes handoff-complete.json (account, region, app, env, service ARN, plan hash, time).
inventory --mode rebuild requires it and re-runs checks 1, 2 and 4 itself; a record older than
24 h fails, as in v0.1's freshness rule.
The deployed image must be known. DescribeService returns the configured identifier, not
proof of what runs. If the identifier is a digest (repo@sha256:…), ecsodus uses it. If it is a
tag, the running digest cannot be established (the tag may have moved since the last
deployment), and generation blocks. The operator then runs a synchronized deployment through
the deploy owner: update-service to the digest the tag resolves to now, wait for the operation
to reach SUCCEEDED (ROLLBACK_SUCCEEDED is a failure: App Runner reverted), then refresh the
M5a state and re-run verify-handoff and inventory. Digests come from ecr:DescribeImages, or
ecr-public:DescribeImages (us-east-1) for ECR_PUBLIC. If the digest is a manifest list,
ecr:BatchGetImage reads it and ecsodus pins the linux/amd64 child on the ECS side.
4.2 Flow and roots¶
ecsodus verify-handoff --manifest m5a/manifest.json -> handoff-complete.json
ecsodus inventory --mode rebuild --handoff handoff-complete.json -> inventory.json
ecsodus report inventory.json -> REPORT.md + decisions.template.yml
ecsodus generate inventory.json --mode rebuild --decisions decisions.yml --out ./rebuild
-> rebuild/infra-rebuild/ (new root), rebuild/infra-patches/ (edits to the
M5a root), rebuild/staged/taskdef-on.tf (8a, outside the root),
rebuild/dns/*.json, rebuild/manifest.json, RUNBOOK.md
ecsodus generate ... --stage rebuild --evidence cert-requested.json (after 2a, before 2b)
ecsodus generate ... --stage new-host --evidence alb-created.json (after step 4, no custom domain only)
ecsodus generate ... --stage worker-on --evidence taskdef-on.json (after step 8a)
ecsodus generate ... --stage taskdef-off --evidence live-revision.json (rollback after hand-over)
ecsodus check plan.json --phase <phase> --manifest rebuild/manifest.json [--step N] [--rollback]
ecsodus check --dns-batch rebuild/dns/convert.json --manifest rebuild/manifest.json
ecsodus verify-cutover --manifest rebuild/manifest.json --step <step> (read-only)
ecsodus stays read-only. It never applies Terraform, never changes DNS, never shifts traffic. Every mutation is a runbook step behind a gate.
Staged generation. Some values the gates need are computed by AWS at apply time: the new
certificate's validation record; with no custom domain, the new ALB's DNS name and canonical
hosted-zone ID, which the new_host alias record targets (§4.9); and the ARN of the
task-definition revision that turns background work on. ecsodus never gates on an unknown value
for them. A gated phase creates the resource (2a, 3, 8a); a read-only verify-cutover step reads
the real value and writes an evidence file bound to the ARN; generate --stage then emits the
next phase's HCL and manifest entry with that value as a literal. The first generate emits only
what does not depend on them.
A rollback after hand-over (§4.11) uses the same pattern once more: verify-cutover --step
live-revision binds the live service revision, and generate --stage taskdef-off emits its off
twin.
The 8a file. The first generate writes aws_ecs_task_definition.on to
rebuild/staged/taskdef-on.tf, outside the infra-rebuild root. The runbook copies it into
rebuild/infra-rebuild/ at §4.12 step 9, after cutover and before the 8a plan, and the manifest
records its hash. If it is in the root at step 3, check --phase rebuild rejects the plan as a
create outside the manifest (§8 has the must-fail test).
Two roots, two states.
-
M5a root (exists): owns the App Runner service, the association, the domain record (later the weighted pair), the instance role and the SGs. M5b only edits it through the
prepare,import,cutover,switchandworker-offphases. -
infra-rebuildroot (new, own state): owns everything M5b creates. Rollback is a destroy of this whole root, never-target. -
With a shared ALB only (§4.7), a third root is edited: the one whose state holds the env ALB.
manifest.json records, per phase, the exact addresses, allowed attributes and expected
before/after values, with the SHA-256 of every generated file.
4.3 What ecsodus generates¶
Default target: a plain ECS Fargate service behind a dedicated new ALB, as flat resources in
the infra-rebuild root. A shared ALB (§4.7) and Express (§4.6) are opt-in.
| Generated | Resource | Notes |
|---|---|---|
| Task definitions | aws_ecs_task_definition off (step 3) and on (step 8a), two addresses of one family |
one container; X86_64; image pinned by digest (§4.1); off carries background work off (§4.10), on differs only in that env var. skip_destroy = true on both, so a replaced or destroyed revision stays ACTIVE and rollback can re-reference it (UpdateService cannot use an INACTIVE revision) |
| Service | aws_ecs_service |
in the env cluster; desired_count = 0; ignore_changes = [desired_count]. Created with task_definition = aws_ecs_task_definition.off.arn; step 8c replaces that with the literal ARN of on. Terraform owns task_definition during the migration (CI deploys stay frozen, §4.12 step 1); the hand-over step adds ignore_changes = [task_definition] and names CI. health_check_grace_period_seconds from a decision (pre-filled 60). deployment_circuit_breaker { enable = true, rollback = false }: a failing deployment stops and shows FAILED, which ready and 8c catch; ECS never reverts task_definition behind Terraform's back, so state and the gates stay true. depends_on the 443 listener (and, shared ALB, the listener rule): the target group must be attached first |
| Scaling | aws_appautoscaling_target + target-tracking policy |
created at min_capacity = 0, max_capacity = 0 (UNCONFIRMED that RegisterScalableTarget accepts MaxCapacity = 0 for ECS; MinCapacity = 0 is documented; tested offline and on the stand-in, §8, §9); the start phase raises them (§4.8). An ALBRequestCountPerTarget policy depends_on the listener and rule too |
| Task role | the App Runner instance role, reused | prepare adds ecs-tasks.amazonaws.com to its trust policy with an aws:SourceAccount condition. Every key, bucket and secret policy that names the role keeps working |
| Execution role | aws_iam_role (new) |
ECR pull, logs; per RuntimeEnvironmentSecrets ARN, secretsmanager:GetSecretValue (Secrets Manager) or ssm:GetParameters (SSM Parameter Store); kms:Decrypt only for a customer-managed key, scoped to that key (App Runner fetched these with the instance role; ECS fetches them with the execution role) |
| Log group | aws_cloudwatch_log_group |
retention copied from the App Runner application log group; kept on rollback (§4.11) |
| Target group | aws_lb_target_group |
ip targets; health check from §4.4 |
| ALB | aws_lb, listener 443 + 80→443 redirect, ALB SG |
idle_timeout from the decision (§4.4); the 443 listener takes aws_acm_certificate_validation.certificate_arn, so it is created only after the certificate is ISSUED |
| Certificate | aws_acm_certificate (step 2a, alone) + validation record if absent + aws_acm_certificate_validation (step 3) |
see "Validation record" below. aws_acm_certificate_validation takes the literal validation_record_fqdns, which gives no implicit edge to the created record, so it depends_on that record when one is created |
new_host record (no custom domain only, §4.9) |
aws_route53_record, alias A to the ALB |
not in step 3. Emitted by generate --stage new-host from alb-created.json (§4.8 #4), with the ALB's DNS name and canonical hosted-zone ID as literals, evaluate_target_health = false and allow_overwrite = false emitted explicitly; created alone in step 4a and gated on every one of those values (§4.8 #4a); deleted by the rollback destroy if it exists (§4.11) |
| Task SG | aws_security_group + aws_vpc_security_group_ingress_rule |
new; ingress only from the ALB SG on the container port |
| WAF | aws_wafv2_web_acl_association |
same regional web ACL, on the new ALB; the ACL is external-reference |
Security groups. Copilot's VPC connector carries both ServiceSecurityGroup and the
imported EnvironmentSecurityGroup (the RDWS fixture, and Copilot's vpc-connector.yml
partial). M5a imports the connector with its actual membership and changes nothing. For the
ECS tasks, ecsodus reads live DescribeVpcConnector.SecurityGroups:
-
Tasks always join
ServiceSecurityGroup(reused, so rules that trust it, including Copilot's addon rules and the env SG's ingress from it, still match) and the new task SG (so the ALB can reach them). -
If the connector holds any other SG (normally
EnvironmentSecurityGroup),decisions.ymlmust answerenv_sg: join | omitper extra SG; a missing answer blocks. The report lists, for each extra SG, the closure rules whose source is that SG (they stop matching the tasks underomit) and the inbound exposure ofjoin(the env SG admits traffic from every env peer, so the tasks become reachable from them on all ports; App Runner ENIs never accepted inbound traffic). -
verify-cutover --step createdchecks the live service's SGs equal this decision.
Validation record. ACM uses one DNS-validation CNAME per FQDN per account, and every
certificate for that FQDN renews through it. App Runner's managed certificate may live in the
customer account (UNCONFIRMED). The new certificate's domain_validation_options exist only
after RequestCertificate, so the comparison runs in two stages:
-
Predictor (read-only,
generateandverify-cutover --step pre-create).acm:ListCertificates(all statuses), thenDescribeCertificateon every certificate whose domain or SANs cover the FQDN. Because ACM reuses the record per FQDN per account, an existing certificate'sResourceRecordis the one the new certificate will get. ecsodus compares it, the importedCertificateValidationRecordsand liveListResourceRecordSetsfor that name. A conflict (same name, different value, or a non-CNAME at the name) blocks early. With no existing certificate the prediction is "unknown", which is allowed. The predictor only reports; it never decides what is created. -
Certificate request (gated, step 2a). A plan that creates only
aws_acm_certificate(check --phase cert-request). Thenverify-cutover --step cert-requestedreads the actualDescribeCertificatevalidation options and compares each with live DNS and both states. It writescert-requested.json(certificate ARN, record name, type, value, outcome), andgenerate --stage rebuildemits from it, record name and value as literals, both the 2b import block and manifest entry (when the outcome needs them) and step 3. It therefore runs before 2b:
| Live record at the name | Outcome |
|---|---|
| absent | infra-rebuild creates it, with allow_overwrite = false emitted explicitly (the provider then sends CREATE, which fails safely if a record appears meanwhile; true would send UPSERT). check --phase rebuild rejects any other value |
present, same value, owned by the M5a root (an imported CertificateValidationRecords entry) |
no record created; the validation's validation_record_fqdns is the literal name |
| present, same value, in neither ecsodus state | ownership unknown (below); with validation_record: import, imported into the M5a root (step 2b, import-only plan, check --phase import, prevent_destroy), then referenced as above; with validation_record: reference, not imported, and the validation references the literal FQDN |
| present, different value or type | blocked; the certificate from 2a is removed by the rollback destroy (§4.11) |
Ownership of an unowned matching record (2b). "In neither ecsodus state" does not mean no
owner: another Terraform state or a CloudFormation stack may hold the record, and matching DNS
values establish compatibility, not ownership. The report says the owner is unknown and that a
deletion by another owner stops renewal for every certificate on that FQDN, the old owner's and
ours. decisions.yml must answer validation_record: import | reference; import also takes a
provenance note in which the operator confirms no other IaC owns the record. A missing answer
blocks.
72-hour validation window. ACM moves a certificate to VALIDATION_TIMED_OUT if it is not
validated within 72 hours. For the "absent" outcome, step 3 must apply within 72 h of 2a.
generate --stage rebuild and check --phase rebuild read DescribeCertificate and fail unless
the status is PENDING_VALIDATION (or ISSUED, for the reference and adopt outcomes) and the
certificate is younger than 72 h; VALIDATION_TIMED_OUT or FAILED fails. Recovery: the rebuild
gate forbids a replace, so the runbook runs the rollback destroy (§4.11) and then 2a again. A
validation record created by an earlier attempt and forgotten by that rollback still matches the
new certificate (ACM reuses it per FQDN), so the second attempt normally takes the 2b row.
Listener activation needs ISSUED: the listener references aws_acm_certificate_validation,
and verify-cutover --step created checks the certificate is ISSUED and on the listener. No
gate may ever delete a record that matches a validation option of any certificate in either
state, including a validation record infra-rebuild created itself: the rollback destroy forgets
it (§4.11), and the retire set never names one (§5.3).
Why both stages: the predictor catches most conflicts before anything is created, but it cannot see a certificate that does not exist yet and rests on documented ACM behaviour. The gated request is the binding check: it compares the real value, and it creates one free, inert resource that the rollback destroy removes.
Why a plain service by default: Copilot's RDWS egresses through a VPC connector in private subnets. Express puts an internal ALB in front of private subnets, so a public RDWS on private subnets cannot be expressed in Express without moving tasks to public subnets with public IPs (council/fable.md finding 4).
Why a dedicated ALB by default: an RDWS-only env usually has no env ALB, and a shared ALB needs updates (timeout, certificate, WAF, rules) that a create-only gate cannot allow.
Flat resources, not terraform-aws-modules/ecs: one emitter, one terraform validate harness,
no module version in the support matrix.
4.4 Config translation¶
Every value comes from DescribeService and related reads, or from decisions.yml. Nothing is
guessed. A missing decision blocks generation, with the key named in the report.
| App Runner | ECS | Rule |
|---|---|---|
ImageRepository.ImageIdentifier, ECR / ECR_PUBLIC |
container image |
digest only (§4.1) |
CodeRepository (source-based) |
— | blocked (after M5; needs the repo and apprunner.yaml) |
ImageConfiguration.Port |
containerPort, TG port, task SG rule port |
literal |
StartCommand set |
command or entryPoint |
decision required, always. App Runner takes one string; ECS takes lists, and whether App Runner overrides ENTRYPOINT or CMD is UNCONFIRMED. The operator gives command (and entry_point if needed). The template pre-fills ["sh","-c","<StartCommand>"] with a # confirm marker; it is never a default |
StartCommand absent |
— | image defaults kept |
RuntimeEnvironmentVariables |
environment |
literal; secrets contract |
RuntimeEnvironmentSecrets |
secrets[].valueFrom |
literal ARNs; execution role gets secretsmanager:GetSecretValue or ssm:GetParameters per ARN, kms:Decrypt only for a customer-managed key (§4.3; §4.5 checks it) |
InstanceRoleArn |
task role | same role (§4.3) |
AuthenticationConfiguration.AccessRoleArn |
execution role | ECR pull |
InstanceConfiguration.Cpu/Memory |
task cpu/memory |
lookup table: 0.25 vCPU → 256 with 512/1024; 0.5 → 512/1024; 1 → 1024 with 2048/3072/4096; 2 → 2048 with 4096/6144; 4 → 4096 with 8192/10240/12288. All are valid Fargate pairs |
egress VPC |
awsvpc in the connector's subnets |
same subnets, so the same NAT EIPs; outbound allowlists keep working |
connector SGs beyond ServiceSecurityGroup |
task SGs | decision required: env_sg is join or omit per extra SG (§4.3) |
egress DEFAULT |
— | decision required: subnets and public-IP choice. The report warns the source IP changes from App Runner's shared ranges to yours |
IsPubliclyAccessible: false + ingress connection |
— | blocked: App Runner custom domains do not support private zones, so clients use the connection's AWS-generated domain, which cannot move. The report gives the manual path |
IpAddressType: DUAL_STACK |
ALB dualstack |
blocked if the VPC has no IPv6 CIDR |
| health check HTTP, in ALB range | TG health check | literal |
| health check TCP (App Runner default) | — | decision required: ALB needs an HTTP path |
| interval < 5 s, timeout < 2 s, timeout ≥ interval, healthy threshold 1, or any threshold > 10 | — | decision required. ALB: interval 5–300, timeout 2–120, thresholds 2–10. App Runner: 1–20 each; default TCP/5/2/1/5 |
MinSize/MaxSize |
scalable target min/max | literal, applied in start (§4.8) |
| — | health_check_grace_period_seconds |
decision required, pre-filled 60 (App Runner has no equivalent) |
MaxConcurrency |
— | decision required: concurrency is not requests per target or CPU; the operator picks metric and target |
ObservabilityConfiguration (AWS-managed X-Ray default) |
— | decision required: tracing: accept-loss. The report states that traces stop on ECS. ADOT is after M5 |
EncryptionConfiguration.KmsKey |
— | reported; no ECS equivalent beyond Fargate ephemeral-storage keys |
| WAF association | aws_wafv2_web_acl_association |
same ACL, dedicated ALB only (§4.7) |
| custom domain | certificate + listener + DNS cutover (§4.9) | required for a gradual cutover |
EnableWWWSubdomain: true |
— | not carried. The report states that www.<alias> stops working once App Runner is gone |
default *.awsapprunner.com URL |
— | cannot move. The report states it |
| background work | task env | decision required (§4.10) |
| request timeout | ALB idle_timeout |
decision required, pre-filled 120. App Runner's 120 s is a total-request limit; the ALB value is an idle timeout. They are different quantities, and the report says so |
Runtime behaviour that differs, stated in the report:
-
App Runner throttles CPU on idle instances. On ECS a container runs at full CPU, so background loops that barely ran on App Runner run for real.
-
App Runner terminates TLS and redirects HTTP to HTTPS (since 2023-02-22). The ALB has both listeners.
-
App Runner runs images as amd64. There is no official statement (UNCONFIRMED), only community reports; pinning the
linux/amd64digest (§4.1) makes both sides run the same bits either way.
4.5 Networking and authorization checks¶
verify-cutover --step created checks, read-only, before any task starts:
-
Routes. Each task subnet passes one of three paths: its route table has
0.0.0.0/0to anavailableNAT gateway; or0.0.0.0/0to an attached internet gateway and the service assigns a public IP (assign_public_ip = true, the §4.4 public-subnet decision for egressDEFAULT); or the VPC has endpoints forecr.api,ecr.dkr, S3 (gateway),logs, andsecretsmanager/ssmas the secrets need. VPC DNS support and hostnames are on. -
Security groups. The live service's SGs equal the task SG,
ServiceSecurityGroup, and each extra connector SG whose decision isjoin(§4.3). -
Ingress. The task SG rule from the ALB SG on the container port exists, and the ALB SG allows 443 and 80 from the declared sources.
-
Authorization.
iam:SimulatePrincipalPolicyon the execution role for each secret ARN and KMS key, and on the reused task role for its existing actions. Simulation cannot evaluate every resource policy; the report lists the key, bucket and secret policies found in the closure that name the instance role, and says the list is closure-limited, not exhaustive.
The first real proof is the task starting: a secret it cannot fetch stops it with
ResourceInitializationError, which the ready step catches.
4.6 Express (opt-in, not available in 0.4.0)¶
The predicate is PLAN.md §2.3, plus:
- egress
DEFAULT, or the operator accepts tasks in public subnets with public IPs; - memory ≤ 8192 MiB and CPU ≤ 4096 (excludes App Runner 4 vCPU / 10 and 12 GB);
tracing: accept-lossif tracing is on; health-check path given.
Three gaps keep --target express failing closed in 0.4.0:
-
No demonstrated zero-task start. The scale-to-zero interlock (§4.8) is not shown for
aws_ecs_express_gateway_serviceand its scaling. -
HTTP clients break. App Runner redirects HTTP to HTTPS; Express creates only a 443 listener. A port-80 redirect would have to live outside Express.
-
Ownership outside the manifest. Express creates and deprovisions the ALB, rules and certificate attachment itself. Terraform would need that ALB's ARN.
checkrejects every data source today, so the ARN would come in as a decision after Express creates it, and the gate could not see those resources.
Express is unlocked only after a real run closes all three. The translation table is shared.
4.7 Shared env ALB (opt-in)¶
alb: shared in decisions.yml uses the env's Terraform-owned ALB instead of a new one. It is
allowed only when the ALB is in the state of a root ecsodus generated, has a 443 listener, and
the host header is not already routed. That root is usually the v0.1 env root (the env stack
handed off by v0.1), not the M5a root; inventory --mode rebuild finds it by the ALB's ARN in
terraform state and the manifest records its path and state lineage. The operator gives the
rule priority (checked against existing rules with the v0.1 priority knowledge).
-
In
infra-rebuild(creates only):aws_lb_listener_certificate,aws_lb_listener_ruleat the given priority, the task SG rule. The listener ARN is a literal from inventory, not a data source. -
In the ALB's root, through its own
prepare-albplan (check --phase prepareagainst that root's manifest entry):idle_timeoutonly, and only if the decision differs from the live value. It applies to every service behind that ALB, and the report says so. -
WAF: blocked if the ALB has a different web ACL, or none while App Runner has one; associating an ACL would change protection for every other service on the ALB.
4.8 Transitions and gates¶
Every phase has a manifest entry: the plan must touch exactly its addresses (a targeted or
incomplete plan fails), with the listed actions and after-values; unknown (computed) values on a
gated attribute fail; data sources and other providers fail, as in v0.1. In a create phase only
the listed after-values are gated; other computed values (ARNs, DNS names, the service's first
task_definition) are allowed because nothing can run at 0/0 and step 4 reads the live values.
The computed values a later gate depends on, the validation record, (no custom domain) the ALB's
DNS name and canonical hosted-zone ID, the on revision's ARN and (rollback after hand-over) the
off twin's ARN, come in as literals through staged generation (§4.2). A resource whose gated
attributes would reference a computed value is never in the phase that creates that value: the
new_host record targets the ALB, so it is not in step 3 but in its own step 4a.
| # | Transition | Root | Gate | Allowed | Fails on |
|---|---|---|---|---|---|
| 0 | Copilot gone | — | verify-handoff |
— | §4.1 |
| 1 | pre-create | — | verify-cutover --step pre-create |
— | image not a digest; validation-record predictor conflict (§4.3). With a custom domain also: zone not exact/public/delegated/in-account; record not a simple CNAME to DNSTarget; weighted set already present; CAA forbids Amazon; App Runner validation records missing. With no custom domain instead: new_host declared; its zone fails the §3.2 zone rules (the adoption rule does not apply); any record of any type already at new_host; CAA forbids Amazon |
| 2 | prepare | M5a (and, shared ALB, the ALB's root as prepare-alb, §4.7) |
check --phase prepare |
update instance role assume_role_policy to the exact expected document; update domain record ttl to 60 if higher; (shared ALB) idle_timeout to the decided value |
anything else, including any other attribute on those addresses |
| 2a | cert-request | infra-rebuild |
check --phase cert-request, then verify-cutover --step cert-requested |
create of aws_acm_certificate only, with the manifest's domain_name, validation_method = "DNS", and subject_alternative_names == [domain_name] (the provider adds domain_name to the SANs at plan time, as ACM does) |
anything else, including any other SAN or an empty list. Verify: the §4.3 outcome table; writes cert-requested.json |
| 2b | adopt validation record (only for "present, same value, in neither ecsodus state" with validation_record: import, §4.3) |
M5a | check --phase import |
import of exactly that record, emitted by generate --stage rebuild |
v0.1 import rule; a missing validation_record answer blocks generation |
| 3 | rebuild (create) | infra-rebuild |
check --phase rebuild |
creates of exactly the manifest addresses (from generate --stage rebuild), plus no-ops of exactly the addresses already in the infra-rebuild state (the 2a certificate); desired_count = 0; scalable target min = max = 0; validation record (if created) with the literal name and value and allow_overwrite = false; task definition off with the off value and skip_destroy = true. Live certificate PENDING_VALIDATION (or ISSUED) and younger than 72 h (§4.3). The new_host record is not in this phase |
any update, delete, replace or import; a no-op of any other address; a create outside the manifest (including aws_ecs_task_definition.on, §4.2, and the new_host record, which belongs to 4a); any other after-value on those attributes |
| 4 | created | — | verify-cutover --step created |
— | live runningCount, pendingCount or desiredCount not 0; scalable target not 0/0; live task definition (DescribeServices → DescribeTaskDefinition, not Terraform's copy) lacks the background-off value or the pinned digest; certificate not ISSUED, not covering the host, or not on the listener; §4.5 checks. With no custom domain also: the ALB ARN in the infra-rebuild state is not a live ALB (DescribeLoadBalancers) in the manifest's account and region; still no record of any type at new_host. It then writes alb-created.json (ALB ARN, live DNSName, CanonicalHostedZoneId, account, region, manifest hash) |
| 4a | new-host (no custom domain only) | infra-rebuild |
check --phase new-host, then verify-cutover --step new-host |
create of exactly the manifest's new_host record (from generate --stage new-host --evidence alb-created.json), plus no-ops of exactly the addresses already in the infra-rebuild state, nothing else. The record: zone_id the manifest's zone, name = new_host, type = "A", one alias block whose name and zone_id equal the DNSName and CanonicalHostedZoneId in alb-created.json (normalised to the form the provider stores; lowercase with no trailing dot is UNCONFIRMED and is checked against the pinned provider when the gate is built), evaluate_target_health = false, allow_overwrite = false. The check re-reads the ALB by its ARN and fails unless its live DNS name and hosted-zone ID still equal the evidence |
an unknown value on any of those attributes; another name, type, zone or alias target; allow_overwrite = true or absent; ttl, records or set_identifier set; a create, update, delete, replace or import of any other address; a no-op of an address not in the infra-rebuild state. Verify: ListResourceRecordSets holds exactly one record at new_host, the alias A to the evidence target; a public DNS query for new_host returns the addresses the ALB's DNS name resolves to |
| 5 | start | infra-rebuild |
check --phase start |
update min_capacity/max_capacity on the one scalable target to App Runner's MinSize/MaxSize |
anything else |
| 6 | ready | — | verify-cutover --step ready |
— | running < MinSize or ≠ desired; primary deployment not COMPLETED (FAILED from the circuit breaker fails); healthy targets < MinSize (never vacuous); listener rule not forwarding the host to the TG; TLS handshake with SNI = host fails or the cert does not match; GET on the declared safe path not the expected status; live task definition as in step 4 |
| 7a | convert to weighted | DNS, then M5a | check --dns-batch, then check --phase import |
one atomic batch (§4.9); then imports of the two weighted records and, for exactly the simple record's manifest address (removed block), either a forget or a no-op with null before and after together with a resource_drift delete entry for that address (§4.9 step 5). After apply, check --state asserts the address is absent |
batch: anything but DELETE of the exact live simple record + CREATE of the two expected weighted records. Import: v0.1 rule, plus a forget, null no-op or drift entry for any other address |
| 7b | shift | M5a | verify-cutover --step shift, then check --phase cutover --step N |
update weight only to the manifest's next pair; a record whose weight the manifest pair leaves unchanged (the waypoint steps) appears as a no-op, and only then |
§4.9 pair rules |
| 7c | simple switch (alternative to 7a–b) | M5a | check --phase switch |
one in-place update of the simple record's records from DNSTarget to the ALB DNS name (reverse with --rollback) |
a replace (it would change the record's identity in state); any change to ttl, name, type; anything else |
| 8a | register on revision (ECS) |
infra-rebuild |
check --phase taskdef-on, then verify-cutover --step taskdef-on |
create of aws_ecs_task_definition.on only, whose container_definitions, compared as parsed JSON (the provider normalizes the string), equal off's except the declared env var's on value; skip_destroy = true. The service does not reference it, so nothing runs it |
anything else, including any change to the service. Verify: revision ACTIVE; live definition differs from the service's live one only in that env var; digest pinned; writes taskdef-on.json (ARN, family, revision) |
| 8b | worker-off (App Runner) | M5a | check --phase worker-off, then verify-cutover --step worker-off |
one in-place update of aws_apprunner_service changing only the declared env var to its off value |
anything else. Verify: latest UPDATE_SERVICE operation SUCCEEDED (ROLLBACK_SUCCEEDED fails) and started after the plan; status RUNNING; live env var off; idle evidence (§4.10) |
| 8c | worker-on (ECS) | infra-rebuild |
check --phase worker-on, then verify-cutover --step worker-on |
update of aws_ecs_service.task_definition only, from the exact live off ARN to the literal on ARN from taskdef-on.json (emitted by generate --stage worker-on) |
anything else, including an unknown task_definition. Verify: deployment COMPLETED; live task definition is the on ARN |
| 9 | window | — | ecsodus retire-evidence (read-only, §5.2) |
— | the default runbook ends here |
| 8a′ | register off twin (rollback after hand-over only, §4.11) | infra-rebuild |
verify-cutover --step live-revision, check --phase taskdef-off, then verify-cutover --step taskdef-off |
create of aws_ecs_task_definition.off_live only, whose parsed container_definitions equal those of the live revision bound in live-revision.json except the declared env var's off value; skip_destroy = true. The service does not reference it |
anything else; a binding to any ARN other than the service's live revision. Verify: revision ACTIVE; differs from the live revision only in that env var; digest pinned; writes taskdef-off.json |
| R | rollback | both | §4.11 | ||
| — | hand-over | infra-rebuild |
check --phase steady after adding ignore_changes = [task_definition] |
none | any change |
8a runs before 8b, so App Runner workers are never turned off while the ECS side still depends on an unapplied or unknown revision. If 8a or its verification fails, App Runner is untouched.
4.9 DNS cutover¶
Pre-flight facts. Copilot's handler writes the domain CNAME with TTL 60, so prepare is
normally a no-op for TTL. If it was raised, prepare lowers it and the runbook waits the old TTL.
Simple record → weighted pair: one atomic batch (7a). A simple CNAME and weighted CNAMEs of
the same name cannot coexist, and Terraform would delete and create in separate calls, leaving a
window where the name does not resolve. generate writes dns/convert.json: DELETE the simple
record (exact live value and TTL), CREATE weighted apprunner (weight 100, DNSTarget) and
ecs (weight 0, ALB DNS name), TTL 60. Batches are transactional.
terraform state pull > pre-convert.tfstate.check --dns-batchcompares the batch with liveListResourceRecordSets.- Apply the batch;
aws route53 wait resource-record-sets-changed --id <change-id>(INSYNC). -
No Terraform apply in the M5a root until step 5 passes. State still holds the simple record; an apply would try to recreate it (Route 53 would refuse, but the runbook forbids it).
-
Swap in
infra-patches/convert.tf: aremoved { from = <simple record>; lifecycle { destroy = false } }block and import blocks for the two weighted records, IDsZONEID_NAME_CNAME_SETID. Plan →check --phase import --forgotten <simple-record address>→ apply. One plan, so state never holds neither or both. Todaysrc/ecsodus/check/plan.pyacceptsforgetonly in the steady phase and only foraws_ecs_task_definition; M5b extends the import phase to accept aforgetof exactly the addresses the manifest names for this step (and for its rollback, §4.11), alongside the imports. Any otherforgetfails. (Splitting it into an import plan and a later steady plan was rejected: between the two, state would hold the simple record and the weighted pair at once.)A refreshed plan shows no
forget. After the batch, the refresh of the simple record looks it up by name, type and set identifier;""does not matchapprunner, so the provider finds nothing and clears the ID, and Terraform plans ano-opwith null before and after instead of aforget(providerinternal/service/route53/record.go; TerraformplanForget). The same holds for the two weighted records in the reverse conversion. The extended import gate therefore accepts, for exactly the manifest-named addresses, either aforget(the-refresh=falsecase, for which theremovedblock stays) or ano-opwith null before and after together with aresource_driftentry showing that address deleted.--forgottencoverage stays exact: every named address must show one of the two forms, and no other address may. After apply,check --stateasserts every named address is absent from state.
Recovery. If the run stops between steps 3 and 5, verify-cutover --step converted sees the
live weighted pair, and the runbook resumes at step 5. From then on generate emits only the
converted config; the old simple-record block is never applied again unless rolling back
(§4.11), which re-adds and imports it.
Both targets are CNAMEs. A weighted pair of alias A records (App Runner supports Route 53 alias targets for services created after 2022-08-01) is also valid and has no TTL. 0.4.0 offers only the CNAME pair, so one path is tested; alias A is a candidate later.
Weight-pair rules (7b). Route 53 treats a set whose weights are all 0 as equal weights, so
(0, 0) sends about half the traffic to each side. Terraform updates the two records in separate
calls, so a plan from (a, b) to (a′, b′) can pass through (a′, b) or (a, b′). check --phase
cutover accepts a plan only if:
-
both weighted addresses are in the plan,
name,type,set_identifier,recordsandttlare unchanged on both, and onlyweightchanges. A record whose weight the manifest pair leaves unchanged (the waypoint steps (0, 100) → (100, 100) and (100, 100) → (100, 0)) appears as a manifest-boundno-op; ano-opon a record whose weight the pair changes fails; -
the before pair equals the live pair (read during the check) and the after pair is the manifest's next pair, or with
--rollbackthe next pair on the rollback path the manifest lists for the before pair; -
none of (a′, b′), (a′, b), (a, b′) is (0, 0).
The default sequence (apprunner, ecs) is (100, 0) → (90, 10) → (50, 50) → (10, 90) →
(0, 100); every step satisfies the last rule. The manifest lists the rollback path from each
pair explicitly: from (90, 10), (50, 50) or (10, 90), one plan to (100, 0); from (0, 100), two
plans through the waypoint (100, 100), then (100, 0). A direct (0, 100) → (100, 0) could pass
through (0, 0), so it is not listed and fails. Before each step, verify-cutover --step shift: live weights equal the expected
pair, targets healthy, the ALB 5xx alarm and the App Runner 5xx alarm OK (ALARM or
INSUFFICIENT_DATA fails), and the runbook waits one TTL after the previous step. DNS weights
are not exact percentages and resolvers cache; the runbook says so.
Provider behaviour (from the provider source, internal/service/route53/record.go). An
update that keeps name, type and set_identifier is sent as one UPSERT; a change of type
or set identifier is sent as DELETE + CREATE in one transactional batch. Create and update
both wait for INSYNC. So each weight step and the 7c switch is atomic per record and
INSYNC-waited. The support matrix pins the provider version, and an offline test asserts this
behaviour at that version.
Simple switch (7c). For low-traffic services: one in-place records update, sent as an
UPSERT. The gate still rejects a replace, because a replace changes the record's identity in
state, not because of a resolution gap. No gradual shift.
No custom domain. No DNS cutover is possible. The custom-domain checks of steps 1, 7 and
§4.11 do not apply. The certificate (2a–2b) needs a hostname: the operator declares a new one in
decisions.yml (new_host), or generation blocks. Its zone must pass the §3.2 zone rules
(exact name, public, delegated, same account) and the CAA check; the adoption rule (a CNAME to
DNSTarget) cannot apply to a new name. pre-create requires that no record of any type exists
at new_host.
-
Traffic record, staged after the ALB. The alias target is the ALB's DNS name and canonical hosted-zone ID, both known only after the ALB is created, so a record in step 3 could not pass the no-unknowns rule. Step 3 therefore creates the ALB without the record.
verify-cutover --step created(step 4) reads the liveDNSNameandCanonicalHostedZoneIdwithDescribeLoadBalancersand writesalb-created.json, bound to the ALB ARN.generate --stage new-hostemits the record with those values as literals,evaluate_target_health = false(a single simple record has nothing to fail over to) andallow_overwrite = false(the provider then sendsCREATE, which fails if a record appears meanwhile). Step 4a creates it alone, beforestart, throughcheck --phase new-host(§4.8 #4a). Its owner is theinfra-rebuildroot;readychecks TLS with SNI =new_hostand the safe-path probe through that name. -
Clients. The runbook is "start ECS, give clients the new URL", and the retirement evidence (§5.2) is the only measure of what still reaches App Runner.
-
Rollback. After 4a, the rollback destroy deletes the
new_hostrecord with the ALB, so clients already given the new URL lose it. Before 4a the record does not exist and is not in state, so the destroy has nothing to delete at that name. In both casesverify-cutover --step rolled-backprints the ALB'sRequestCountsince step 5 (none if step 5 never ran) and requires an explicit operator answer before the destroy proceeds.
4.10 Background work and side effects¶
Both sides run the same image against the same databases, queues and topics.
-
Declaration.
decisions.ymlmust answer, per service, the PLAN.md §2.3 question: "does this container run background work or migrations on boot?" Answers:none;switch: the env var name, its off and on values, and a statement that the switch also stops boot-time migrations. Plus one oflock: <how the app holds a distributed lock>oridle-check: <command that exits 0 only when no job is in flight>. Migrations that run on boot outside the switch are blocked: every App Runner configuration deployment (includingworker-off) reboots instances and would rerun them.
-
Start off. The ECS task definition carries the off value from the first apply. Step 4 verifies it on the live task definition before any task can start.
-
Transfer (steps 8a–8c), App Runner side first. Register and verify the ECS
onrevision (8a, nothing runs it). Then turn work off on App Runner (8b), wait for the operation to reachSUCCEEDED, run the idle check (or confirm the lock), then point the ECS service at the knownonARN (8c). -
Rollback, reverse order. Reverse 8c:
check --phase worker-on --rollbackupdatestask_definitiononly, from the liveonARN back to the literaloffARN (stillACTIVE,skip_destroy); deploymentCOMPLETED, idle check. After hand-over the source is the live revision and the target is its off twin fromtaskdef-off.json(§4.8 #8a′, §4.11). Then reverse 8b:check --phase worker-off --rollbacksets the env var to its on value on App Runner; operationSUCCEEDED(ROLLBACK_SUCCEEDEDfails). 8a needs no reverse: theonrevision stays registered, unused, until the rollback destroy. -
Migrations must be backward compatible for the whole window.
- HTTP is not side-effect free. Verification uses only the declared safe path.
- SNS publish. Both sides may publish during the window. Consumers must be idempotent; the report lists the topics the instance role can publish to.
worker-off is a Terraform change, not an out-of-band update-service, so it leaves no drift.
Its in-place update sends the full source configuration from state; the deploy freeze keeps that
equal to live, and verify-handoff check 4 proved the image identifier in state is the live one.
4.11 Rollback¶
Rollback has its own phase, rollback-rebuild, and its own order. It is allowed at any point
before retirement.
| From | Steps |
|---|---|
| Before 7a | destroy infra-rebuild (below). App Runner untouched |
| During 7b | check --phase cutover --rollback back to (100, 0), as two plans if needed (§4.9) |
| After hand-over (§4.12 step 11) | first restore the migration's controls, before any worker change: re-freeze deploys to App Runner and ECS; refresh both states; remove ignore_changes = [task_definition] and set the service's literal task_definition to the live ARN (the plan must be empty, check --phase steady), so Terraform owns deployment again. Then verify-cutover --step live-revision writes live-revision.json (the service's live revision ARN, its digest, account, region, service ARN, manifest hash), so the fresh rollback manifest and evidence are bound to that revision before any worker mutation. If the live revision is on, the target is the existing off. Otherwise generate --stage taskdef-off --evidence live-revision.json emits aws_ecs_task_definition.off_live, the off twin of the live revision, registered through the taskdef-off gate (§4.8 #8a′) and recorded in taskdef-off.json; it becomes the target of the reverse 8c. Then continue with the rows below |
| After 8a–8c | reverse background work first (§4.10), then the weights |
| Weighted pair at (100, 0) | check --dns-batch --rollback on dns/revert.json (DELETE both weighted records with their live values, CREATE the simple record to DNSTarget), wait INSYNC, then removed blocks for the weighted records and an import of the simple record in one plan (check --phase import --forgotten the two weighted addresses, §4.9) |
| After 7c | check --phase switch --rollback |
| DNS back on App Runner only (or before 7a) | plan the infra-rebuild-rollback/ config → check --phase rollback-rebuild → apply. Then optionally prepare --rollback to remove ecs-tasks.amazonaws.com from the instance role's trust |
The rollback destroy keeps the logs and the validation record. generate writes
infra-rebuild-rollback/: the same backend and state, no resource blocks, and `removed { from =