Hugging Face
Hugging Face
Stream files from a Hugging Face dataset, model, or space repository without cloning it.
- Category
- Warehouse & Lakehouse
- Source type
- HUGGING_FACE
- Produces
- fileimageaudiovideoarchive
Training data is data. A dataset repository your team published — or one a model was fine-tuned on — can carry customer records, internal documents, or credentials in a config file, and nothing about being on the Hub changes that.
What you need to connect
A repository id (namespace/name) and its kind: dataset, model, or
space. Public repositories need no credential; private and gated ones need a
read-scoped access token. Self-hosted Enterprise Hubs are supported by setting
the endpoint.
What Classifyre reads
Files inside the repository, at whichever revision you choose — a branch, a tag, or a commit SHA. The repository is never cloned: files are listed, then streamed one at a time, so scanning a repository with terabytes of weights in it costs nothing if you filter the weights out.
Narrow the scan with folder paths, glob allow/deny patterns, and extension
filters. Excluding .safetensors, .bin and .gguf is usually the first thing
you want.
Shared behaviour · File and object sources
Every object in the repository becomes one asset. What it is — a PDF, a spreadsheet, a screenshot, a video — is worked out from the bytes themselves, not from the file name, because storage systems routinely label everything as generic binary data. The full list of readable types is on File Formats.
Files hidden inside other files are pulled out and scanned in their own right: images embedded in a document, and every member of a ZIP, TAR, 7z or RAR archive. Images, audio and video need OCR and transcription switched on before their content can be read.
Large objects are streamed rather than loaded whole, so a multi-gigabyte file costs disk rather than memory, and only the part a sampling window asks for is read.
Files in this repository are read with the shared file pipeline: see Supported File Formats for everything it can open, and OCR & Transcription for reading text out of images, audio and video.
Metadata on every asset
Asset kind · file
| Field | Type | Always present | What it is |
|---|---|---|---|
| size_bytes | integer | Yes | Raw byte size of the content |
| mime_type | string | Yes | Resolved MIME type |
| parse_error | string | No | Set when content extraction failed |
| image_width | integer | No | Width in pixels |
| image_height | integer | No | Height in pixels |
| page_count | integer | No | Number of pages (pdf) |
| paragraph_count | integer | No | Number of paragraphs (docx) |
| table_count | integer | No | Number of tables (docx) |
| row_count | integer | No | Number of data rows |
| columns | object[] | No | Columns as {name, type} objects (type may be empty for csv/xlsx) |
| encoding | string | No | Detected character encoding |
| json_root_type | string | No | Root JSON type: object, array, or scalar |
| top_level_keys | integer | No | Number of top-level keys when the root is an object |
| array_length | integer | No | Length when the root is an array |
| provider | string | Yes | Storage provider label |
| object_key | string | Yes | File path relative to the repository root |
| repo_id | string | Yes | Repository the file belongs to (namespace/name) |
| repo_type | string | Yes | Repository kind the file was read from: model, dataset or space |
| revision | string | No | Commit SHA the file was read at (resolved from the configured revision) |
| etag | string | No | Content digest of the file: the Git LFS SHA-256 when present, otherwise the Git blob OID |
| blob_id | string | No | Git object ID (OID) of the file blob |
| lfs_sha256 | string | No | SHA-256 recorded in the Git LFS pointer, for files stored in LFS |
| web_url | string | No | Browser-accessible Hugging Face URL for the file |
| source_hash | string | No | Hash of the parent asset (embedded files and archive members only) |
| location | string | No | Location within the parent (embedded file location or archive member path) |
Asset kind · image
| Field | Type | Always present | What it is |
|---|---|---|---|
| size_bytes | integer | Yes | Raw byte size of the content |
| mime_type | string | Yes | Resolved MIME type |
| parse_error | string | No | Set when content extraction failed |
| image_width | integer | No | Width in pixels |
| image_height | integer | No | Height in pixels |
| provider | string | No | Storage provider label |
| object_key | string | No | File path relative to the repository root |
| repo_id | string | No | Repository the file belongs to (namespace/name) |
| repo_type | string | No | Repository kind the file was read from: model, dataset or space |
| revision | string | No | Commit SHA the file was read at (resolved from the configured revision) |
| etag | string | No | Content digest of the file: the Git LFS SHA-256 when present, otherwise the Git blob OID |
| blob_id | string | No | Git object ID (OID) of the file blob |
| lfs_sha256 | string | No | SHA-256 recorded in the Git LFS pointer, for files stored in LFS |
| web_url | string | No | Browser-accessible Hugging Face URL for the file |
| source_hash | string | No | Hash of the parent asset (embedded files and archive members only) |
| location | string | No | Location within the parent (embedded file location or archive member path) |
Asset kind · audio
| Field | Type | Always present | What it is |
|---|---|---|---|
| size_bytes | integer | Yes | Raw byte size of the content |
| mime_type | string | Yes | Resolved MIME type |
| parse_error | string | No | Set when content extraction failed |
| provider | string | Yes | Storage provider label |
| object_key | string | Yes | File path relative to the repository root |
| repo_id | string | Yes | Repository the file belongs to (namespace/name) |
| repo_type | string | Yes | Repository kind the file was read from: model, dataset or space |
| revision | string | No | Commit SHA the file was read at (resolved from the configured revision) |
| etag | string | No | Content digest of the file: the Git LFS SHA-256 when present, otherwise the Git blob OID |
| blob_id | string | No | Git object ID (OID) of the file blob |
| lfs_sha256 | string | No | SHA-256 recorded in the Git LFS pointer, for files stored in LFS |
| web_url | string | No | Browser-accessible Hugging Face URL for the file |
| source_hash | string | No | Hash of the parent asset (embedded files and archive members only) |
| location | string | No | Location within the parent (embedded file location or archive member path) |
Asset kind · video
| Field | Type | Always present | What it is |
|---|---|---|---|
| size_bytes | integer | Yes | Raw byte size of the content |
| mime_type | string | Yes | Resolved MIME type |
| parse_error | string | No | Set when content extraction failed |
| provider | string | Yes | Storage provider label |
| object_key | string | Yes | File path relative to the repository root |
| repo_id | string | Yes | Repository the file belongs to (namespace/name) |
| repo_type | string | Yes | Repository kind the file was read from: model, dataset or space |
| revision | string | No | Commit SHA the file was read at (resolved from the configured revision) |
| etag | string | No | Content digest of the file: the Git LFS SHA-256 when present, otherwise the Git blob OID |
| blob_id | string | No | Git object ID (OID) of the file blob |
| lfs_sha256 | string | No | SHA-256 recorded in the Git LFS pointer, for files stored in LFS |
| web_url | string | No | Browser-accessible Hugging Face URL for the file |
| source_hash | string | No | Hash of the parent asset (embedded files and archive members only) |
| location | string | No | Location within the parent (embedded file location or archive member path) |
Asset kind · archive
| Field | Type | Always present | What it is |
|---|---|---|---|
| size_bytes | integer | Yes | Raw byte size of the content |
| mime_type | string | Yes | Resolved MIME type |
| parse_error | string | No | Set when content extraction failed |
| provider | string | Yes | Storage provider label |
| object_key | string | Yes | File path relative to the repository root |
| repo_id | string | Yes | Repository the file belongs to (namespace/name) |
| repo_type | string | Yes | Repository kind the file was read from: model, dataset or space |
| revision | string | No | Commit SHA the file was read at (resolved from the configured revision) |
| etag | string | No | Content digest of the file: the Git LFS SHA-256 when present, otherwise the Git blob OID |
| blob_id | string | No | Git object ID (OID) of the file blob |
| lfs_sha256 | string | No | SHA-256 recorded in the Git LFS pointer, for files stored in LFS |
| web_url | string | No | Browser-accessible Hugging Face URL for the file |
| source_hash | string | No | Hash of the parent asset (embedded files and archive members only) |
| location | string | No | Location within the parent (embedded file location or archive member path) |
Lineage
Lineage
This source records no lineage. Nothing in the system it reads describes data moving from one place to another, so no FLOW edges are produced. Related items are still linked — see Lineage & Relationships for what those links mean and how they differ from lineage.
Worth knowing
- Parquet data files are read as tables, so a dataset’s rows are what detectors see — not an opaque binary blob.
- Last-commit dates cost extra requests. They’re what
LatestandAutomaticsampling order by; leave them off if ordering doesn’t matter to you and listing is slow. - Very large files are read by byte range, so a sampling window reads only the part it needs.
Configuration
Beyond the fields below, every source also has the settings shared by all of them: the sampling strategy, the detectors to run, the scan schedule, and the compute limits for its scan jobs.
Required
Without these, the source will not save.
| Field | Type | Required | What it does | Default |
|---|---|---|---|---|
| required | object | Yes | —no extra properties | — |
| repo_id | string | Yes | Repository to scan, as namespace/name (for example openai/gsm8k or my-org/internal-corpus) | — |
| repo_type | enum | Yes | Repository kind on the Hub. Datasets hold the data files (parquet, csv, images, audio); models hold weights and configuration; spaces hold app source code. Allowed: dataset, model, space | dataset |
Secrets
Stored encrypted and never shown again after you save them. See Configuration & Fields.
| Field | Type | Required | What it does | Default |
|---|---|---|---|---|
| masked | object | Yes | —no extra properties | — |
| token | string | Yes | Hugging Face user access token (starts with hf_). Create one under Settings → Access Tokens with at least read permission on the repository. The token is always passed explicitly — no environment variable or locally cached login is ever used. | — |
Optional
Everything you can tune. Sensible defaults apply when you leave them alone.
| Field | Type | Required | What it does | Default |
|---|---|---|---|---|
| optional | object | No | —no extra properties | — |
| connection | object | No | Network and resource controls for Hub requests. Files are streamed one at a time and capped, so a large repository never has to fit in memory or on disk.no extra properties | — |
| connection.endpoint | string | No | Hub endpoint to use. Defaults to https://huggingface.co; set this for a self-hosted Enterprise Hub. | — |
| connection.max_archive_member_bytes | integer | No | Maximum uncompressed bytes read from a single archive membermin 1024 | 10485760 |
| connection.max_archive_members | integer | No | Maximum member files expanded from one archive into child assets. Bounds fan-out, and is the guard against a zip bomb.min 1, max 10000 | 200 |
| connection.max_archive_total_bytes | integer | No | Maximum uncompressed bytes read across all members of one archive. The decompression-ratio ceiling: a zip bomb hits this before it hits memory.min 1024 | 104857600 |
| connection.max_embedded_files | integer | No | Maximum embedded files (parquet image/audio columns, office media) expanded from one container into child assets per runmin 1, max 10000 | 200 |
| connection.max_file_bytes | integer | No | Refuse any object larger than this many bytes. 0 or unset means no limit: an object above max_object_bytes is spooled to disk rather than held in memory, so file size is bounded by free disk, not by RAM.min 0 | — |
| connection.max_object_bytes | integer | No | Maximum bytes of one object held in memory. Larger objects are streamed to a temporary file (or read by byte range where the provider supports it), so this bounds memory rather than the size of file that can be scanned. See max_file_bytes to refuse large objects outright.min 1024 | 26214400 |
| connection.max_retries | integer | No | Maximum retries on transient Hub errors (5xx and rate limits)min 0, max 10 | 3 |
| connection.request_timeout_seconds | number | No | Network timeout in seconds for list and download requestsmin 1, max 300 | 60 |
| scope | object | No | Which files inside the repository are listed and read. The repository is never cloned or fully downloaded: the file tree is listed first, then each selected file is streamed one at a time.no extra properties | — |
| scope.allow_patterns | array | No | Glob allowlist applied to file paths, for example data/*.parquet or **/*.csv. A file is kept when it matches at least one pattern. Empty means every file is kept. | — |
| scope.allow_patterns[] | string | No | — | — |
| scope.exclude_extensions | array | No | Optional extension denylist (for example, .safetensors, .bin, .gguf) | — |
| scope.exclude_extensions[] | string | No | — | — |
| scope.ignore_patterns | array | No | Glob denylist applied to file paths, for example *.safetensors or **/checkpoints/*. Files matching any pattern are skipped. | — |
| scope.ignore_patterns[] | string | No | — | — |
| scope.include_content_preview | boolean | No | Stream file bytes to infer MIME and extract detector-ready text previews. Turn off for a metadata-only inventory of the repository. | true |
| scope.include_empty_objects | boolean | No | Include zero-byte files in extraction results | false |
| scope.include_extensions | array | No | Optional extension allowlist (for example, .parquet, .csv, .png) | — |
| scope.include_extensions[] | string | No | — | — |
| scope.include_last_commit | boolean | No | Fetch each file's last-commit date while listing. Needed for accurate LATEST and AUTOMATIC ordering, but makes listing noticeably slower on large repositories. When off, every file inherits the repository's last-modified date. | false |
| scope.include_object_metadata | boolean | No | Attach Hub metadata (content digest, size, timestamps) to asset checksums | true |
| scope.paths | array | No | Folders inside the repository to list, for example data/ or data/train. Each folder is walked recursively. Leave empty to list the whole repository. | — |
| scope.paths[] | string | No | — | — |
| scope.revision | string | No | Branch, tag or commit SHA to read (for example main, refs/convert/parquet, or v1.0). Defaults to the repository's default branch. | — |