Skip to content
SupportLogin

Add rows to a dataset from the curated pool

POST/api/v1/datasets/{dataset_id}/augment

Retrieves additional rows matching this dataset and returns them as a new dataset, leaving this one unchanged. The new dataset carries this dataset’s rows alongside the retrieved ones, so it can be trained on or downloaded directly. Poll GET /datasets/{dataset_id}/status on the returned id until it reaches succeeded. Set estimate=true to price the request without starting a run.

Path ParametersExpand Collapse
dataset_id: string
Body ParametersJSONExpand Collapse
domain_rows: optional number

How many additional rows to retrieve from topics matching this dataset. Defaults to 0.

minimum0
maximum100000
general_rows: optional number

How many additional rows to retrieve from other topics, to broaden the dataset. Defaults to 0.

minimum0
maximum100000
training_type: optional "instruction_dataset" or "preference_pairs"

Shape of the rows to retrieve. instruction_dataset (default) retrieves prompt/completion pairs for SFT; preference_pairs retrieves chosen/rejected pairs for DPO.

One of the following:
"instruction_dataset"
"preference_pairs"
estimate: optional boolean

When true, validates the request and returns the estimated credit cost without starting a run or creating a dataset.

idempotency_key: optional string

Opaque key that makes a retried request safe. A second request carrying the same key returns the dataset the first one created instead of starting and charging for a second run.

maxLength255
ReturnsExpand Collapse
dataset_id: string

The augmented dataset. It carries this dataset’s rows plus the retrieved ones, and is ready to download once its status reaches succeeded. Null for an estimate, which creates nothing.

status: "pending" or "running" or "awaiting_input" or 2 more

Status of the augmented dataset at the moment this response was sent, in the same vocabulary GET /datasets/{dataset_id}/status reports. Poll that endpoint until it reaches succeeded or failed. An idempotent replay of a finished run returns its terminal status here.

One of the following:
"pending"
"running"
"awaiting_input"
"succeeded"
"failed"
estimated_credits_consumed: number

Credits this run consumes, charged against the requested rows.

estimate: boolean

Whether this was an estimate-only request (no run started).

Add rows to a dataset from the curated pool

curl https://api.prod.adaptionlabs.ai/api/v1/datasets/$DATASET_ID/augment \
    -H 'Content-Type: application/json' \
    -H "Authorization: Bearer $ADAPTION_API_KEY" \
    -d '{
          "domain_rows": 3000,
          "general_rows": 1000,
          "training_type": "instruction_dataset"
        }'
{
  "dataset_id": "550e8400-e29b-41d4-a716-446655440000",
  "status": "running",
  "estimated_credits_consumed": 40,
  "estimate": true
}
Returns Examples
{
  "dataset_id": "550e8400-e29b-41d4-a716-446655440000",
  "status": "running",
  "estimated_credits_consumed": 40,
  "estimate": true
}