Skip to content
SupportLogin
Create a dataset from file upload, HuggingFace, Kaggle, or Google Sheets

Create a dataset from file upload, HuggingFace, Kaggle, or Google Sheets

POST/api/v1/datasets

Unified ingest endpoint. Pass source.url to start an async import — the provider is inferred from the URL host (HuggingFace, Kaggle, or Google Sheets). Omit it to get upload instructions for a presigned S3 PUT. Adapt the result with POST /datasets/{dataset_id}/run, or send it straight to training with POST /autoscientist when it already holds prompt and completion columns.

Body ParametersJSONExpand Collapse
source: object { url, files, sheet_names, 6 more }

Dataset source. Pass url to import from HuggingFace, Kaggle, or Google Sheets, or name + file_format to upload a file. For Sheets, optional sheet_names selects tabs (omit or [] for all).

url: optional string

HuggingFace, Kaggle, or Google Sheets URL. Provide it to import from that provider; omit it to upload a local file instead.

files: optional array of string

File paths to download from the source. Required with HuggingFace/Kaggle url. Not used for Google Sheets — use sheet_names (or sheet_name) instead.

sheet_names: optional array of string

Google Sheets tab titles to import. Only for Google Sheets url. Omit or pass [] to import every tab; pass one or more titles to import only those tabs.

sheet_name: optional string

Legacy single Google Sheets tab title. Ignored when sheet_names is present (including []).

access: optional "oauth" or "public"

Google Sheets only. public imports a link-shared spreadsheet without a connected Google account. oauth (default) uses the connected account and may be unavailable on some accounts.

One of the following:
"oauth"
"public"
name: optional string

Name for the dataset. Required for file uploads, where it may double as a filename (“sales-data.csv”) to supply the format; derived from the URL for provider imports.

file_format: optional "csv" or "json" or "jsonl" or 8 more

Format of the file being uploaded. Optional when name ends in a supported extension — it is inferred from there, and an explicit value overrides it.

One of the following:
"csv"
"json"
"jsonl"
"parquet"
"pdf"
"docx"
"pptx"
"xlsx"
"html"
"zip"
"txt"
processing_mode: optional "adapt" or "raw"

How the data is ingested, for both uploads and provider imports. adapt (default) runs it through the adaptation pipeline. raw materializes the data straight to a trainable dataset with no augmentation, and needs a column_mapping plus tabular input — file_format for an upload, or files for an import. Imported files are read as one dataset, so they must share a single format and schema. Either can then be fine-tuned or run through AutoScientist.

One of the following:
"adapt"
"raw"
column_mapping: optional object { prompt, completion, chosen, 2 more }

Required when processing_mode is raw: which columns to canonicalize as the prompt and the answer — either a completion, or a chosen / rejected pair. Ignored when the dataset is adapted.

prompt: string

Name of the raw column holding the prompt / input text

completion: optional string

Name of the raw column holding the completion / target text. Required unless chosen and rejected are given.

chosen: optional string

Name of the raw column holding the preferred response. Give it together with rejected to ingest a preference dataset.

rejected: optional string

Name of the raw column holding the dispreferred response. Give it together with chosen.

context: optional array of string

Names of raw columns to fold in as context alongside the prompt. Each is validated against the uploaded file at ingestion; an unknown name fails the request listing the available columns. The columns themselves are kept in the dataset as well as folded.

defer_adaption: optional boolean

When true, the upload is stored but Adaptive Data does not start. The dataset stays in awaiting_preprocessing until POST /datasets/:id/start-adaption.

ReturnsExpand Collapse
dataset_id: string

ID of the newly created dataset

status: string

Current dataset status

upload_instructions: optional object { url, method, s3_key }

Upload instructions for file sources. PUT your file to the provided URL.

url: string

Pre-signed URL for uploading the file

method: string

HTTP method to use

s3_key: string

S3 object key — pass this back in the complete request if needed for verification

Create a dataset from file upload, HuggingFace, Kaggle, or Google Sheets

curl https://api.prod.adaptionlabs.ai/api/v1/datasets \
    -H 'Content-Type: application/json' \
    -H "Authorization: Bearer $ADAPTION_API_KEY" \
    -d '{
          "source": {}
        }'
{
  "dataset_id": "dataset_id",
  "status": "status",
  "upload_instructions": {
    "url": "https://s3.amazonaws.com/bucket/key?X-Amz-Signature=...",
    "method": "PUT",
    "s3_key": "s3_key"
  }
}
Returns Examples
{
  "dataset_id": "dataset_id",
  "status": "status",
  "upload_instructions": {
    "url": "https://s3.amazonaws.com/bucket/key?X-Amz-Signature=...",
    "method": "PUT",
    "s3_key": "s3_key"
  }
}