## Adapt a dataset (or estimate cost)

**post** `/api/v1/datasets/{dataset_id}/adapt`

Validates column mapping and recipe configuration, reserves credits, and starts the adaptation pipeline, which improves the rows already in the dataset without changing how many there are. To add rows instead, use POST /datasets/{dataset_id}/augment. Set estimate=true to validate and get a cost quote without starting a run. When the mapped image column is also in context_columns, output rows are billed at 10 credits per 100 rows (1–100 rows cost 10 credits; see multimodalPricingApplied and creditMultiplier on the response). Once the run succeeds, train on the result with [POST /autoscientist](/api/python/resources/autoscientist/methods/create).

### Path Parameters

- `dataset_id: string`

### Body Parameters

- `column_mapping: optional object { prompt, completion, chat, 3 more }`

  Column role assignments for adaptation. Required for real runs, optional for estimate-only requests.

  - `prompt: optional string`

    Column to use as the prompt/instruction field. Trigger for prompt mode (with optional `completion` and `context`); cannot be combined with `chat` or `universal_prompt`.

  - `completion: optional string`

    Column to use as the completion/response field. Optional in prompt mode and universal-prompt mode; not allowed in chat mode.

  - `chat: optional string`

    Column containing chat/conversation data. Chat mode is exclusive: `prompt`, `completion`, `context` and `universal_prompt` must be omitted. An `image` column may accompany it — chat supplies the text framing an image column needs.

  - `context: optional array of string`

    Columns to include as context. Optional in prompt mode; required (at least one entry) in universal-prompt mode; not allowed in chat mode.

  - `universal_prompt: optional string`

    Dataset-wide instruction folded into every row alongside the context columns. Use when there is no per-row prompt column but rows share a common task framing. Requires a per-row differentiator, so that the folded prompt is not identical on every row: at least one entry in `context`, or an `image` column.

  - `image: optional string`

    Column containing per-row images (URLs, file paths, or encoded bytes) to fold into the adaptation as multimodal context. Allowed in any mode (prompt, universal-prompt, or chat) as long as text framing is also present — it cannot be the only mapping. The column is automatically added to `context` if omitted. Note: opting a dataset into multimodal context disqualifies it from finetuning.

- `recipe_specification: optional object { version, recipes }`

  Adaptation recipe configuration. Omitted recipes use backend defaults.

  - `version: optional string`

    Reserved. Accepted and stored for forward compatibility, but not yet read — recipe options are currently versioned by release rather than by this field. Omit it.

  - `recipes: optional object { prompt_rephrase, deduplication, reasoning_traces }`

    Adaptation recipe toggles. Omitted recipes use backend defaults.

    - `prompt_rephrase: optional boolean`

      Rephrase prompts for variety and clarity

    - `deduplication: optional boolean`

      Remove near-duplicate rows

    - `reasoning_traces: optional boolean`

      Add reasoning traces (chain-of-thought) to completions

- `job_specification: optional object { max_rows, idempotency_key }`

  Job execution parameters

  - `max_rows: optional number`

    Maximum number of rows to process in this run

  - `idempotency_key: optional string`

    Client-generated idempotency key for safe retries. If a launch with the same key already exists, the original response is returned.

- `brand_controls: optional object { length, safety_categories, hallucination_mitigation, blueprint }`

  Brand and quality controls for generated completions. Covers response length, content safety categories, web-search grounding, and a freeform blueprint system prompt.

  - `length: optional "minimal" or "concise" or "detailed" or "extensive"`

    Target response length. Controls verbosity of generated completions.

    - `"minimal"`

    - `"concise"`

    - `"detailed"`

    - `"extensive"`

  - `safety_categories: optional array of string`

    Turns on content-safety analysis for the generated completions. Supply one or more of `violence`, `self_harm`, `harassment`, `hate`, `sexual`. Any non-empty list enables the analysis, and every row is then scored against all five categories — the list does not narrow what is checked. Results arrive as the `prompt_safety_issues` and `response_safety_issues` columns on the adapted dataset. Rows are annotated, never removed.

  - `hallucination_mitigation: optional boolean`

    Enable web-search grounding to reduce hallucinations in generated completions.

  - `blueprint: optional string`

    Freeform brand/style instructions injected as a system prompt for every generated completion. Use this to enforce tone, language, persona, or any guideline that does not fit the structured length/safety/grounding controls.

- `training_type: optional "instruction_dataset" or "preference_pairs"`

  How to adapt the dataset. `instruction_dataset` (default) produces enhanced prompt/completion pairs for SFT; `preference_pairs` generates chosen/rejected pairs for DPO. Source columns are the same in both cases — `column_mapping.prompt` (and optionally `completion`) — the chosen/rejected fields are produced by the adaptation pipeline.

  - `"instruction_dataset"`

  - `"preference_pairs"`

- `estimate: optional boolean`

  When true, validates the request and returns the estimated credit cost without starting a run.

- `language_expansion: optional object { type, languages, pairs, sample_rate }`

  Translation/localization expansion. When set, the pipeline produces additional rows in the requested languages or country/language variants. Output row count ≈ input × (1 + sample_rate × target_count); credits are billed on output rows.

  Three-state field:

  - **omitted** — leaves any prior expansion config on the dataset unchanged (no-op). A wizard-saved expansion is preserved across SDK-driven /adapt calls that do not pass `language_expansion`.
  - **`language_expansion: null`** — explicitly clears any prior expansion config on this dataset (overrides wizard state).
  - **object** — overwrites the prior expansion config with the given spec.

  - `type: "translate" or "localize"`

    Expansion mode. `translate` produces one new row variant per target language; `localize` produces one new row variant per country/language pair.

    - `"translate"`

    - `"localize"`

  - `languages: optional array of string`

    Target language codes (required when type=translate, forbidden when type=localize). Mostly ISO 639-1 two-letter codes such as `es`, with three-letter codes where no two-letter code exists (`fil`, `yue`) and script-tagged codes for Chinese (`zh-Hans`, `zh-Hant`). Matching is case-insensitive, and bare `zh` resolves to `zh-Hans`. Validated against the supported language list at request time; unknown values return 400 with a sample of supported codes.

  - `pairs: optional array of object { country, language }`

    Country/language pairs (required when type=localize, forbidden when type=translate). Each pair is validated against the supported country/language pair list at request time.

    - `country: string`

      ISO 3166-1 alpha-2 country code.

    - `language: string`

      Language code, as ISO 639-1 — optionally with a script subtag. Supported pairs include codes that are not two letters (`zh-Hant` for Taiwan, `yue` for Hong Kong), so the length is not pinned at two here. Membership is decided against the supported-pair list, which names an unsupported entry in its own 400.

  - `sample_rate: number`

    Fraction (0.01–1) of input rows that are expanded per target. Output rows ≈ input × (1 + sample_rate × target_count). Required. Credits are billed on the expanded output row count.

### Returns

- `run_id: optional string`

  Unique identifier for this pipeline run. Null for estimate-only requests.

- `estimatedMinutes: number`

  Estimated processing time in minutes

- `estimatedCreditsConsumed: number`

  Estimated number of credits that will be consumed by this run

- `estimate: boolean`

  Whether this was an estimate-only request (no run started)

- `multimodalPricingApplied: boolean`

  True when an image column is mapped and also listed in context_columns; each output row is billed at a higher rate.

- `creditMultiplier: optional number`

  10 credits per 100 output rows when multimodalPricingApplied is true. Omitted for text-only pricing.

### Example

```http
curl https://api.prod.adaptionlabs.ai/api/v1/datasets/$DATASET_ID/adapt \
    -H 'Content-Type: application/json' \
    -H "Authorization: Bearer $ADAPTION_API_KEY" \
    -d '{
          "training_type": "instruction_dataset"
        }'
```

#### Response

```json
{
  "run_id": "dataset-550e8400-e29b-41d4-a716-446655440000-1712234567890",
  "estimatedMinutes": 0,
  "estimatedCreditsConsumed": 0,
  "estimate": true,
  "multimodalPricingApplied": true,
  "creditMultiplier": 0
}
```
