## Add rows to a dataset from the curated pool

**post** `/api/v1/datasets/{dataset_id}/augment`

Retrieves additional rows matching this dataset and returns them as a new dataset, leaving this one unchanged. The new dataset carries this dataset’s rows alongside the retrieved ones, so it can be trained on or downloaded directly. Poll GET /datasets/{dataset_id}/status on the returned id until it reaches `succeeded`. Set estimate=true to price the request without starting a run.

### Path Parameters

- `dataset_id: string`

### Body Parameters

- `domain_rows: optional number`

  How many additional rows to retrieve from topics matching this dataset. Defaults to 0.

- `general_rows: optional number`

  How many additional rows to retrieve from other topics, to broaden the dataset. Defaults to 0.

- `training_type: optional "instruction_dataset" or "preference_pairs"`

  Shape of the rows to retrieve. `instruction_dataset` (default) retrieves prompt/completion pairs for SFT; `preference_pairs` retrieves chosen/rejected pairs for DPO.

  - `"instruction_dataset"`

  - `"preference_pairs"`

- `estimate: optional boolean`

  When true, validates the request and returns the estimated credit cost without starting a run or creating a dataset.

- `idempotency_key: optional string`

  Opaque key that makes a retried request safe. A second request carrying the same key returns the dataset the first one created instead of starting and charging for a second run.

### Returns

- `dataset_id: string`

  The augmented dataset. It carries this dataset’s rows plus the retrieved ones, and is ready to download once its status reaches `succeeded`. Null for an estimate, which creates nothing.

- `status: "pending" or "running" or "awaiting_input" or 2 more`

  Status of the augmented dataset at the moment this response was sent, in the same vocabulary GET /datasets/{dataset_id}/status reports. Poll that endpoint until it reaches `succeeded` or `failed`. An idempotent replay of a finished run returns its terminal status here.

  - `"pending"`

  - `"running"`

  - `"awaiting_input"`

  - `"succeeded"`

  - `"failed"`

- `estimated_credits_consumed: number`

  Credits this run consumes, charged against the requested rows.

- `estimate: boolean`

  Whether this was an estimate-only request (no run started).

### Example

```http
curl https://api.prod.adaptionlabs.ai/api/v1/datasets/$DATASET_ID/augment \
    -H 'Content-Type: application/json' \
    -H "Authorization: Bearer $ADAPTION_API_KEY" \
    -d '{
          "domain_rows": 3000,
          "general_rows": 1000,
          "training_type": "instruction_dataset"
        }'
```

#### Response

```json
{
  "dataset_id": "550e8400-e29b-41d4-a716-446655440000",
  "status": "running",
  "estimated_credits_consumed": 40,
  "estimate": true
}
```
