Skip to Content
SourcesS3-Compatible Storage

S3-Compatible Storage

S3-Compatible Storage

Scan objects from AWS S3, MinIO, Cloudflare R2, Backblaze B2, and other S3-compatible endpoints.

Category
Warehouse & Lakehouse
Source type
S3_COMPATIBLE_STORAGE
Produces
fileimageaudiovideoarchive

Buckets are where data goes to be forgotten: exports, backups, log dumps, “temporary” copies from three years ago. This source scans any S3-compatible endpoint — AWS S3, MinIO, Cloudflare R2, Backblaze B2, Garage, Ceph, Wasabi and the rest — through the same connection.

What you need to connect

A bucket name, plus either an access key pair or nothing at all — leave the credentials empty and the environment’s ambient AWS credential chain (instance role, workload identity, ~/.aws) is used. For non-AWS providers, add the endpoint URL and, where the provider needs it, a region.

Read-only credentials are enough: s3:ListBucket and s3:GetObject on the prefix you want scanned.

What Classifyre reads

Every object under the prefix you choose, filtered by extension allow- and denylists if you want to narrow it further.

Shared behaviour · File and object sources

Every object in the bucket becomes one asset. What it is — a PDF, a spreadsheet, a screenshot, a video — is worked out from the bytes themselves, not from the file name, because storage systems routinely label everything as generic binary data. The full list of readable types is on File Formats.

Files hidden inside other files are pulled out and scanned in their own right: images embedded in a document, and every member of a ZIP, TAR, 7z or RAR archive. Images, audio and video need OCR and transcription switched on before their content can be read.

Large objects are streamed rather than loaded whole, so a multi-gigabyte file costs disk rather than memory, and only the part a sampling window asks for is read.

Objects from this bucket are read with the shared file pipeline: see Supported File Formats for everything it can open, and OCR & Transcription for reading text out of images, audio and video.

Metadata on every asset

Asset kind · file

FieldTypeAlways presentWhat it is
size_bytesintegerYesRaw byte size of the content
mime_typestringYesResolved MIME type
parse_errorstringNoSet when content extraction failed
image_widthintegerNoWidth in pixels
image_heightintegerNoHeight in pixels
page_countintegerNoNumber of pages (pdf)
paragraph_countintegerNoNumber of paragraphs (docx)
table_countintegerNoNumber of tables (docx)
row_countintegerNoNumber of data rows
columnsobject[]NoColumns as {name, type} objects (type may be empty for csv/xlsx)
encodingstringNoDetected character encoding
json_root_typestringNoRoot JSON type: object, array, or scalar
top_level_keysintegerNoNumber of top-level keys when the root is an object
array_lengthintegerNoLength when the root is an array
providerstringYesStorage provider label
object_keystringYesObject key/path
etagstringNoObject entity tag
source_hashstringNoHash of the parent asset (embedded files and archive members only)
locationstringNoLocation within the parent (embedded file location or archive member path)

Asset kind · image

FieldTypeAlways presentWhat it is
size_bytesintegerYesRaw byte size of the content
mime_typestringYesResolved MIME type
parse_errorstringNoSet when content extraction failed
image_widthintegerNoWidth in pixels
image_heightintegerNoHeight in pixels
providerstringNoStorage provider label
object_keystringNoObject key/path
etagstringNoObject entity tag
source_hashstringNoHash of the parent asset (embedded files and archive members only)
locationstringNoLocation within the parent (embedded file location or archive member path)

Asset kind · audio

FieldTypeAlways presentWhat it is
size_bytesintegerYesRaw byte size of the content
mime_typestringYesResolved MIME type
parse_errorstringNoSet when content extraction failed
providerstringYesStorage provider label
object_keystringYesObject key/path
source_hashstringNoHash of the parent asset (embedded files and archive members only)
locationstringNoLocation within the parent (embedded file location or archive member path)

Asset kind · video

FieldTypeAlways presentWhat it is
size_bytesintegerYesRaw byte size of the content
mime_typestringYesResolved MIME type
parse_errorstringNoSet when content extraction failed
providerstringYesStorage provider label
object_keystringYesObject key/path
source_hashstringNoHash of the parent asset (embedded files and archive members only)
locationstringNoLocation within the parent (embedded file location or archive member path)

Asset kind · archive

FieldTypeAlways presentWhat it is
size_bytesintegerYesRaw byte size of the content
mime_typestringYesResolved MIME type
parse_errorstringNoSet when content extraction failed
providerstringYesStorage provider label
object_keystringYesObject key/path
etagstringNoObject entity tag
source_hashstringNoHash of the parent asset (embedded files and archive members only)
locationstringNoLocation within the parent (embedded file location or archive member path)

Lineage

Lineage

This source records no lineage. Nothing in the system it reads describes data moving from one place to another, so no FLOW edges are produced. Related items are still linked — see Lineage & Relationships for what those links mean and how they differ from lineage.

Worth knowing

  • Object type is detected from content. S3-compatible providers routinely label uploads as generic binary data; the format is worked out from the bytes, so those objects are still read properly.
  • Archives are expanded into child assets, with caps on member count, member size, and total uncompressed bytes — the guard against a zip bomb.
  • Big objects don’t blow up memory. Anything past the in-memory threshold is streamed to a temporary file or read by byte range instead.
  • Metadata-only inventories are possible: turn off content preview and you get a catalogue of what’s in the bucket without downloading anything.

Configuration

Beyond the fields below, every source also has the settings shared by all of them: the sampling strategy, the detectors to run, the scan schedule, and the compute limits for its scan jobs.

Required

Without these, the source will not save.

FieldTypeRequiredWhat it doesDefault
requiredobjectYesno extra properties
bucketstringYesBucket name for AWS S3, MinIO, Cloudflare R2, Backblaze B2, Garage, and other S3-compatible endpoints

Secrets

Stored encrypted and never shown again after you save them. See Configuration & Fields.

FieldTypeRequiredWhat it doesDefault
maskedobjectNoOptional static credentials. Leave empty to use ambient AWS credentials chain.no extra properties
aws_access_key_idstringNoS3-compatible access key ID
aws_secret_access_keystringNoS3-compatible secret access key
aws_session_tokenstringNoOptional session token for temporary credentials

Optional

Everything you can tune. Sensible defaults apply when you leave them alone.

FieldTypeRequiredWhat it doesDefault
optionalobjectNono extra properties
connectionobjectNono extra properties
connection.endpoint_urlstringNoCustom endpoint URL for MinIO/R2/B2/Garage and other S3-compatible providersformat uri
connection.max_archive_member_bytesintegerNoMaximum uncompressed bytes read from a single archive membermin 102410485760
connection.max_archive_membersintegerNoMaximum member files expanded from one archive into child assets. Bounds fan-out, and is the guard against a zip bomb.min 1, max 10000200
connection.max_archive_total_bytesintegerNoMaximum uncompressed bytes read across all members of one archive. The decompression-ratio ceiling: a zip bomb hits this before it hits memory.min 1024104857600
connection.max_embedded_filesintegerNoMaximum embedded files (parquet image/audio columns, office media) expanded from one container into child assets per runmin 1, max 10000200
connection.max_file_bytesintegerNoRefuse any object larger than this many bytes. 0 or unset means no limit: an object above max_object_bytes is spooled to disk rather than held in memory, so file size is bounded by free disk, not by RAM.min 0
connection.max_keys_per_pageintegerNoMaximum objects requested per provider list API callmin 1, max 1000200
connection.max_object_bytesintegerNoMaximum bytes of one object held in memory. Larger objects are streamed to a temporary file (or read by byte range where the provider supports it), so this bounds memory rather than the size of file that can be scanned. See max_file_bytes to refuse large objects outright.min 10245242880
connection.region_namestringNoRegion (recommended for AWS; required by some S3-compatible providers)
connection.request_timeout_secondsnumberNoNetwork timeout in seconds for list/download operationsmin 1, max 30030
connection.verify_sslbooleanNoTLS certificate verification toggletrue
scopeobjectNoObject scope and filtering controls.no extra properties
scope.exclude_extensionsarrayNoOptional extension denylist
scope.exclude_extensions[]stringNo
scope.include_content_previewbooleanNoDownload object bytes to infer MIME and extract detector-ready text previewstrue
scope.include_empty_objectsbooleanNoInclude zero-byte objects in extraction resultsfalse
scope.include_extensionsarrayNoOptional extension allowlist (for example, .pdf, .csv, .parquet)
scope.include_extensions[]stringNo
scope.include_object_metadatabooleanNoAttach provider metadata (etag, size, content-type hints, timestamps) to asset checksumstrue
scope.prefixstringNoObject key prefix filter (for example, exports/2026/)
Last updated on