Skip to content

fix(cli): send the state bucket's own region to the OpenTofu S3 backend - #509

Open
talmo wants to merge 1 commit into
mainfrom
fix/tofu-backend-bucket-region
Open

talmo wants to merge 1 commit into
mainfrom
fix/tofu-backend-bucket-region

Conversation

@talmo

@talmo talmo commented Oct 4, 2026

Copy link
Copy Markdown
Member

Problem

The S3 backend's region identifies the state bucket. The region to provision into is passed separately as -var=region=. _tofu_init sent cfg.app.region for both, so any deployment whose region differs from its state bucket's region fails at init:

Failed to get existing workspaces: operation error S3: ListObjectsV2,
https response error StatusCode: 301 ... api error PermanentRedirect:
The bucket you are attempting to access must be addressed using the
specified endpoint.

OpenTofu's S3 backend does not follow S3's region redirect. The AWS CLI does, which makes this easy to misdiagnose — a head-bucket/ls against the bucket with the "wrong" region succeeds, so the config looks fine right up until tofu init.

This became reachable in 0.3.0, when app.region started actually driving the deployment (before that everything landed in us-west-2 regardless). The bucket name is derived from the account ID and S3 bucket names are global, so a bucket an earlier deployment created in another region is reused by name and silently mismatches.

Fix

Ask S3 where the bucket actually is via get_bucket_location, and send that as the backend region. On any lookup failure it falls back to cfg.app.region — no worse than the assumption it replaces, so a caller without s3:GetBucketLocation is unaffected. _tofu_destroy shares _tofu_init, so destroy is fixed too.

Testing

Four new tests in TestResolveBucketRegion / TestTofuInitBackendRegion: bucket location is used, a null LocationConstraint resolves to us-east-1, lookup failure falls back, and _tofu_init sends the bucket's region rather than app.region.

packages/cli: 835 passed, 1 failed, and that failure (test_docker_isolation.py::test_explicit_subprocess_patch_overrides_the_autouse_guard) is pre-existing and environmental — it fails identically on clean main on this machine because Docker isn't installed. ruff check packages/cli is clean.

Found while deploying to eu-west-1 against a us-west-2 state bucket; the deployment is now live and healthy with the patch applied.

Note: the allocator has the same bug, and it needs a template change too

packages/allocator/src/lablink_allocator_service/main.py (~line 360) builds the client-VM workspace's backend the same way:

f"-backend-config=bucket={bucket_name}",
f"-backend-config=region={cfg.app.region}",

It fails identically, and worse — it crashes allocator startup, so the service never becomes healthy (Flask process exited before becoming ready). I have not fixed it here, because the same fix alone wouldn't work: the allocator calls AWS with its instance role, and the s3_backend_doc policy in lablink-template grants s3:ListBucket and object actions but not s3:GetBucketLocation. The lookup would fail, hit the fallback, and reproduce the bug. Fixing it properly means a coordinated change across both repos, so it seemed better as a separate PR than bundled here.

Workaround in the meantime: point bucket_name at a bucket in app.region, which the allocator honors (unlike the CLI, which derives its own name and ignores cfg.bucket_name — arguably its own inconsistency).

🤖 Generated with Claude Code

https://claude.ai/code/session_014bWRfnmvBt12hWFk5FwAg3

The S3 backend's `region` identifies the bucket; the region to provision
into is passed separately as `-var=region=`. `_tofu_init` sent `app.region`
for both, so any deployment whose region differed from its state bucket's
failed init with:

    Failed to get existing workspaces: operation error S3: ListObjectsV2,
    https response error StatusCode: 301 ... api error PermanentRedirect

OpenTofu's S3 backend does not follow S3's region redirect. The AWS CLI
does, which makes this easy to misdiagnose as working.

This became reachable in 0.3.0, when `app.region` started actually driving
the deployment. The bucket name is derived from the account ID and S3 names
are global, so a bucket an earlier deployment created in another region is
reused by name and silently mismatches.

Ask S3 where the bucket is via `get_bucket_location`, falling back to
`app.region` when the lookup fails — no worse than the previous assumption.
`_tofu_destroy` shares `_tofu_init`, so destroy is fixed too.

Found deploying to eu-west-1 against a us-west-2 state bucket.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014bWRfnmvBt12hWFk5FwAg3

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant