Skip to content
SupportLogin
Adaptive Data
Edit on GitHub

Create a dataset

Upload a local file or import a dataset from Hugging Face or Kaggle with the Python SDK.

Import data from a local file, Hugging Face, or Kaggle without writing a conversion script. Each method returns a dataset_id that you pass to datasets.adapt.

The examples assume you have installed the Python SDK and set ADAPTION_API_KEY:

from adaption import Adaption
client = Adaption()
1. Upload a local file

The SDK helper creates the dataset, uploads the file, and confirms the upload.

Supported extensions: .csv, .json, .jsonl, and .parquet.

dataset = client.datasets.upload_file("training_data.csv")
dataset_id = dataset.dataset_id
print(dataset_id)
2. Import from Hugging Face

Point at a Hugging Face dataset URL and the file(s) to import.

dataset = client.datasets.create(
source={
"url": "https://huggingface.co/datasets/tatsu-lab/alpaca",
"files": [
"data/train-00000-of-00001-a09b74b3ef9c3b56.parquet"
],
},
)
dataset_id = dataset.dataset_id
print(dataset_id)
3. Import from Kaggle

Use the Kaggle dataset page URL and the files to pull.

dataset = client.datasets.create(
source={
"url": "https://www.kaggle.com/datasets/uciml/sms-spam-collection-dataset",
"files": ["spam.csv"],
},
)
dataset_id = dataset.dataset_id
print(dataset_id)

Kaggle credentials must be registered in Adaption: open API keys settings (sign in if prompted) and add your Kaggle API credentials there before importing.

Imports run asynchronously. Poll the dataset status until ingestion has populated the dataset before starting an adaptation run:

import time
while True:
status = client.datasets.get_status(dataset_id)
if status.status == "failed":
error = status.error_data
message = (error and error.message) or "unknown error"
raise RuntimeError(f"Dataset import failed: {message}")
if status.row_count is not None:
print(f"Imported {status.row_count} rows")
break
time.sleep(5)

Next, map the imported columns, then configure the adaptation run. For endpoint details, see the create and get_status references. The datasets.upload_file helper handles the create, upload, and completion steps for local files.