Repository navigation
Finish the public S3Core API extraction from filesystem adapters #1086
Description
Activity
Agreed design
The maintainer agreed to these names and contracts on 2026-10-04, after reviews of a draft by claude-fable-5-1 and Codex.
All operations are public and synchronous, are sent throughS3Core.call(), and preserve the adapters' current requests, unless a point below says otherwise.
Object operations take anS3Path, and bucket operations take a bucket name.1. Object GET and PUT
S3Core.get_object(path, range_=None, **params) -> bytesrange_is(start, end)with an exclusive end, or(start, None)to read to the end.
A negative start with no end reads the last-startbytes.
Nonesends noRangeof its own, so aRangeinparamsstill applies.- An empty range raises
ValueError, because S3 would return the whole object.
A negative start with an end also raisesValueError; no adapter path reaches it. - The core reads the whole body and closes it, also when the read fails.
A read failure of the body propagates as the untranslated botocore exception, and the GET is not retried, as today.
A start past the end raises the translatedOSError, whose__cause__is theClientErrorwith codeInvalidRange. - The adapter keeps the
InvalidRangetob""mapping ofcat_file(), the range resolution based oninfo(), and the parallel range splitting inS3File._fetch_range().
Failures are still observed in completion order, and the parts are joined in range order.
S3FileSystem._get_object(),S3File._format_ranges()andS3File._merge_objects()are removed.
S3Core.put_object(path, body=None, **params) -> S3PutObject- An empty or
Nonebody sends noBodyof its own, so aBodyinparamsstill applies.
The request fields take precedence overparams. ValueErrorif the path has no key or has a version ID.
The adapters keep their own checks and messages first.S3FileSystem._put_object()is removed.
- An empty or
2. Tags and ACLs
S3Core.get_object_tagging(path, **params) -> dict[str, str]andS3Core.put_object_tagging(path, tags, **params) -> None.
Both accept a version ID, includingnull.
put_tags(mode="m")stays in the adapter: two requests, the core get, then the core put of the merged tags.
plan_multipart_copy()reusesget_object_tagging().S3Core.put_object_acl(path, acl, **params) -> NoneandS3Core.put_bucket_acl(bucket, acl, **params) -> Nonetake a canned ACL.
They raiseValueErrorif the ACL is not inS3Core.OBJECT_ACLSorS3Core.BUCKET_ACLS, respectively.
chmod()validates against the same sets before any recursive change.- Breaking (4.0.0 release note):
S3FileSystem.OBJECT_ACLSandS3FileSystem.BUCKET_ACLSare removed, as theMULTIPART_UPLOAD_*limits were.
3. Metadata replacement
S3Core.replace_object_metadata(path, head, metadata, **params) -> Nonesends one CopyObject request.
The request copies the object onto itself withMetadataDirective="REPLACE"andMetadata=metadata.
It retains fromheadthe content headers,Expires,WebsiteRedirectLocationandStorageClass.
Unlessparamsset an encryption parameter, it also retainsServerSideEncryption,SSEKMSKeyIdandBucketKeyEnabled.
paramstake precedence over the retained fields.
A field that the operation itself sets, such asMetadata, still raisesTypeErrorwhen it is given inparams.
headmust be the HeadObject result ofpath.setxattr()keeps its validation before the HEAD, the HEAD throughmetadata(), theNonedeletions, and the cache invalidation.
4. Multipart upload listing
S3ListMultipartUploadsPage(bucket, uploads, is_truncated, next_key_marker, next_upload_id_marker)is a frozen dataclass withfrom_response().
Itsuploadsfield is atuple[S3MultipartUpload, ...].S3Core.list_multipart_uploads_page(bucket, prefix=None, key_marker=None, upload_id_marker=None, **params)lists one page.
TheS3Core.list_multipart_uploads(...)iterator yields the pages.- A
Noneprefix sends noPrefix, as today. - The iterator stops when a page is not truncated or lacks either next marker, as today.
It raisesTypeErrorforKeyMarkerandUploadIdMarkerinparams.
- A
S3FileSystem.list_multipart_uploads()keeps the filter that selects only the key and the keys under it.
5. Presigned URLs
S3Core.generate_presigned_url(path, client_method="get_object", expires_in=3600, **params) -> str.
paramstake precedence over the bucket, key and version ID of the path, as today.
request_kwargsare not added, because the method is not an API operation.
A path without a key is not rejected; botocore validates the parameters.
6. Bucket lifecycle
S3Core.create_bucket(bucket, acl=None, region_name=None, **params) -> None.- The ACL is validated against
BUCKET_ACLS. - The region defaults to the client's region.
- No
LocationConstraintis sent forus-east-1.
- The ACL is validated against
S3Core.delete_bucket(bucket, **params) -> None.- A direct call to either method creates or deletes the bucket.
allow_bucket_creationandallow_bucket_deletionare filesystem options, not core permissions. mkdir()andrmdir()keep the path checks, theexists()checks, the opt-ins, theParamValidationErrortoValueErrortranslation, and the cache eviction.
7. Multipart size check
S3Core.check_multipart_upload_size(size, block_size) -> NoneraisesValueErrorwhensize > block_size * MULTIPART_UPLOAD_MAX_PARTS.
The error condition is unchanged.
The message is generic: it gives the minimum block size, but not the path or the filesystem option.
The sync adapter and the aio transaction helpers call it directly.
S3FileSystem._check_multipart_upload_size()is removed.
Kept as is
S3FileSystem._call()stays a compatibility delegate.- The cache stays in
DirCache. - The adapters keep scheduling, callbacks, transactions, invalidation and fsspec behavior.
Pull requests
- GET and PUT (section 1).
- Tags, ACLs, metadata replacement and presigned URLs (sections 2, 3 and 5).
- Bucket lifecycle, multipart upload listing and the size check (sections 4, 6 and 7).
After the merges, I will audit the remaining SDK call sites and record the disposition of step 5.
I will then close this issue and #1053 if no work in the agreed scope remains.
#1083 stays separate.- added 20 commits that reference this issue
on Oct 4, 2026 Addendum: the requests added by #1084
#1084 (merged after the inventory above was recorded) added two adapter-owned pieces to
S3FileSystem._move_pairs():- a GetBucketVersioning request built in the adapter (
self._call(self._client.get_bucket_versioning, Bucket=...)) - a call to the private
S3Core._is_directory_bucket()
The maintainer agreed on 2026-10-05 to move them in a small fourth PR, with these names:
S3Core.get_bucket_versioning(bucket, **params) -> str | None. It sends one GetBucketVersioning request and returns theStatusof the bucket's versioning ("Enabled"or"Suspended"), orNonefor a bucket whose versioning has never been enabled.S3Path.is_directory_bucket. This public property says whether the path's bucket is a directory bucket (S3 Express One Zone, a name ending with--x-s3). It replaces the privateS3Core._is_directory_bucket(), whichplan_multipart_copy()also uses.
Requests, request counts and errors stay the same.
- a GetBucketVersioning request built in the adapter (
Completed
The agreed scope of this issue is done in four PRs:
PR Scope Merge #1091 Object GET/PUT: S3Core.get_object(),put_object()5a0a5cb2#1090 Tags, ACLs, metadata replacement, presigned URLs: get_object_tagging(),put_object_tagging(),put_object_acl(),put_bucket_acl(),replace_object_metadata(),generate_presigned_url();S3Core.OBJECT_ACLS/BUCKET_ACLSd0e883ae#1089 Bucket lifecycle, multipart upload listing, size check: create_bucket(),delete_bucket(),list_multipart_uploads_page()/list_multipart_uploads()withS3ListMultipartUploadsPage,check_multipart_upload_size()47fb6f31#1098 The request added by #1084 after the inventory: get_bucket_versioning(),S3Path.is_directory_bucket23a54e16Each PR went through both self-review rounds, an independent review, a live AWS run and AWS CI before it was merged.
Release notes for 4.0.0
S3FileSystem.OBJECT_ACLSandS3FileSystem.BUCKET_ACLSare removed. UseS3Core.OBJECT_ACLSandS3Core.BUCKET_ACLSinstead (Move tagging, ACL, metadata and presign requests into S3Core #1090).- The multipart size check raises a generic message. The error condition is unchanged (Move bucket lifecycle, multipart upload listing and the size check into S3Core #1089).
S3Core.create_bucket()keeps the fields of a givenCreateBucketConfigurationand adds the region'sLocationConstraintonly when the configuration has none ofLocationConstraint,LocationandBucket(Move bucket lifecycle, multipart upload listing and the size check into S3Core #1089).
Audit of the remaining SDK calls
These SDK uses remain in
S3FileSystem,AioS3FileSystemandS3File(master after #1098), each for the stated reason:- Client construction in
S3FileSystem.__init__(boto3session/client,Config,UNSIGNED): connection configuration. The filesystem owns it and passes the client toS3Core. self.core.operation_params(...): the adapter selects which of its filesystem-wides3_additional_kwargsand lookup parameters each request receives. This is adapter policy; the core provides the filter.botocore.exceptions.ParamValidationError→ValueErrorinmkdir(): kept in the adapter, as agreed.botocore.exceptions.ClientErrorinspection incat_file()(InvalidRange→b"") andclear_multipart_uploads()(NoSuchUploadfor an upload completed or aborted since listing is ignored): fsspec translation of core errors.S3FileSystem._call(): a compatibility delegate toS3Core.call(). No production code calls it any more.
No S3 request is built in the adapters any more. The aio adapter reaches the core only through the sync filesystem or
asyncio.to_thread.Step 5 (cache keys from
S3Path)Not needed. No extraction step required building
DirCachekeys fromS3Path, soDirCachestorage, options, invalidation and sync/aio sharing are unchanged, as decided on #1053.Sub-issues of #1053
#1049, #1059, #1063, #1077 and this issue are all complete. #1083, the
mv()behavior fix, was handled separately in #1084.The CI of #1098 surfaced an intermittent failure in the #1077 test
test_interrupted_creation, which is unrelated to this extraction. It is tracked separately in #1099.No work remains in the agreed scope, so I am closing this issue and #1053.
Use case
Complete the remaining S3 API extraction in #1053.
Steps 1–4 are merged: paths (#1052), typed listing/HEAD operations (#1060), delete/multipart/copy/pairing (#1064, #1069, #1072, #1074, #1081), and the multipart writer (#1085).
The native sub-issues #1049, #1059, #1063 and #1077 are closed.
On master
855d4a7b652d6dad106c27ee7112264ca45b923e, several operations still build S3 requests and apply S3 rules insideS3FileSystem.They use
S3Core.call()for transport, retries and error translation, but a direct core caller has no typed operation takingS3Pathfor these requests.Moving only GET/PUT, tags and ACLs would leave other request construction in the adapter and would not finish the separation proposed in #1053.
The remaining inventory in
pyathena/filesystem/s3.pyis:_get_object()(2843),_put_object()(2892): request construction, ranges, body handling and PUT result conversionget_tags()(2313),put_tags()(2336),chmod()(2377): request construction, version selection, tag merge and ACL validationsetxattr()(2227): metadata changes, retained system/storage/encryption fields, and the self-copy request; reads already usecore.head_object()list_multipart_uploads()(2429): requests, pagination and conversion to upload objectssign()(2147): S3 request parameters and SDK signingmkdir()(1240),rmdir()(1336): create/delete requests and region/ACL rules_check_multipart_upload_size()(1744): S3 part-count validation also reached by the aio transaction helpersProposed change
Add the remaining public, synchronous operations and pure S3 planning rules to
S3Core,S3MultipartWriter, or an appropriate fsspec-independent type.Object operations accept
S3Path; bucket operations use a clearly documented bucket-only argument.Results expose documented types, including the existing
S3PutObject,S3MetadataandS3MultipartUploadwhere appropriate.Settle the GET result and ownership contract before implementation: whether a result exposes a response body or bytes determines who reads and closes it, and how read failures are handled.
Move request construction, operation-specific parameter rules, version handling, response conversion, and S3 validation out of the adapters for the inventory above.
Extract page-level multipart upload listing separately from iteration, following the existing core listing API.
The core uses S3 prefix semantics; the adapter retains its documented key-or-descendant filter, which excludes sibling keys sharing a prefix.
Make metadata replacement and tag merging use core operations or pure planners while preserving their existing requests and results.
Centralize the multipart size rule for sync and aio callers without changing the existing error conditions.
The adapters keep path parsing and fsspec behavior: directory synthesis/expansion, recursive traversal,
chmod(recursive=True), callbacks, cache invalidation, buffering, transactions, executor scheduling and async bridging.Bucket lifecycle opt-ins (
allow_bucket_creationandallow_bucket_deletion) remain enforced by the filesystem before it invokes a core primitive.The public core bucket methods document that a direct call creates or deletes a bucket; an adapter option is not a core permission mechanism.
Preserve the existing public filesystem signatures, return shapes, errors, request order/count, per-call parameter precedence, checksum behavior and cleanup.
Preserve open-ended and suffix reads, empty-range handling, empty PUT bodies, tag overwrite/merge modes, version-qualified paths, metadata retention/override rules, ACL validation before recursive mutation, bucket region handling and multipart listing markers.
The core stays synchronous and imports no fsspec or aiobotocore; scheduling stays with the adapters as decided on #1053.
New public names and result contracts should be agreed in this issue before implementation.
The implementation can use several independently reviewable PRs, each keeping the adapters functional.
Completion criteria
S3File. Record each remaining site's reason; SDK client compatibility, raw-call compatibility shims, fsspec translation and orchestration remain adapter responsibilities.DirCachestorage, options, invalidation and sync/aio sharing. Record the disposition of step 5: centralize cache keys fromS3Pathonly if this extraction needs it; otherwise explicitly record that the conditional step was not needed, as decided in Separate an S3 core from the fsspec adapter in pyathena.filesystem #1053.pyathena/s3fs/,pyathena/pandas/,pyathena/aio/s3fs/and Polars/fsspec registration.The separate
mv()behavior defect #1083 keeps its own issue and implementation.This extraction preserves existing behavior; new S3 features and behavior changes require separate agreement, as specified by #1053.
Validation plan (if implementing)
Use botocore
Stubberand self-contained tests for the new operations, typed results and planners.Cover inherited/per-call parameter precedence, version IDs including
null, request bodies/ranges, body ownership and read failures, metadata/encryption retention, tag modes, ACL validation, bucket regions, multipart pagination and size limits, plus translated failures.Pin expected requests independently of the moved implementation.
Run
just formatandjust lint, then the affected tests intests/pyathena/filesystem/test_s3_core.py,test_s3_writer.py,test_s3.pyandtest_s3_async.py.Check the existing filesystem request sequences, cache invalidation, transactions, conditional creation, multipart checksums and interruption/cancellation cleanup.
Exercise internal consumers that use these read/write operations and fsspec registration.
Run documentation lint/build for API changes and the applicable full PyAthena AWS CI before declaring an implementation PR ready.
Use the repository's existing AWS test environment and fixtures for live S3 validation.
Serialize live runs with other Test workflows/local tests; do not provision new persistent infrastructure for this extraction.
Bucket creation/deletion coverage must use the explicit opt-ins and disposable test resources.
Record exact tested commits, commands, results and skipped coverage, separating static, offline and live AWS evidence.
This proposal is based on source and GitHub-state inspection; no new tests or AWS operations were run to prepare it.