Skip to content
SupportLogin
Adapt a dataset (deprecated)

Adapt a dataset (deprecated)

Deprecated
POST/api/v1/datasets/{dataset_id}/run

Deprecated alias of POST /datasets/{dataset_id}/adapt, kept for existing clients. Use POST /datasets/{dataset_id}/adapt in new code.

Path ParametersExpand Collapse
dataset_id: string
Body ParametersJSONExpand Collapse
column_mapping: optional object { prompt, completion, chat, 3 more }

Column role assignments for adaptation. Required for real runs, optional for estimate-only requests.

prompt: optional string

Column to use as the prompt/instruction field. Trigger for prompt mode (with optional completion and context); cannot be combined with chat or universal_prompt.

completion: optional string

Column to use as the completion/response field. Optional in prompt mode and universal-prompt mode; not allowed in chat mode.

chat: optional string

Column containing chat/conversation data. Chat mode is exclusive: prompt, completion, context and universal_prompt must be omitted. An image column may accompany it — chat supplies the text framing an image column needs.

context: optional array of string

Columns to include as context. Optional in prompt mode; required (at least one entry) in universal-prompt mode; not allowed in chat mode.

universal_prompt: optional string

Dataset-wide instruction folded into every row alongside the context columns. Use when there is no per-row prompt column but rows share a common task framing. Requires a per-row differentiator, so that the folded prompt is not identical on every row: at least one entry in context, or an image column.

image: optional string

Column containing per-row images (URLs, file paths, or encoded bytes) to fold into the adaptation as multimodal context. Allowed in any mode (prompt, universal-prompt, or chat) as long as text framing is also present — it cannot be the only mapping. The column is automatically added to context if omitted. Note: opting a dataset into multimodal context disqualifies it from finetuning.

recipe_specification: optional object { version, recipes }

Adaptation recipe configuration. Omitted recipes use backend defaults.

version: optional string

Reserved. Accepted and stored for forward compatibility, but not yet read — recipe options are currently versioned by release rather than by this field. Omit it.

recipes: optional object { prompt_rephrase, deduplication, reasoning_traces }

Adaptation recipe toggles. Omitted recipes use backend defaults.

prompt_rephrase: optional boolean

Rephrase prompts for variety and clarity

deduplication: optional boolean

Remove near-duplicate rows

reasoning_traces: optional boolean

Add reasoning traces (chain-of-thought) to completions

job_specification: optional object { max_rows, idempotency_key }

Job execution parameters

max_rows: optional number

Maximum number of rows to process in this run

minimum1
idempotency_key: optional string

Client-generated idempotency key for safe retries. If a launch with the same key already exists, the original response is returned.

brand_controls: optional object { length, safety_categories, hallucination_mitigation, blueprint }

Brand and quality controls for generated completions. Covers response length, content safety categories, web-search grounding, and a freeform blueprint system prompt.

length: optional "minimal" or "concise" or "detailed" or "extensive"

Target response length. Controls verbosity of generated completions.

One of the following:
"minimal"
"concise"
"detailed"
"extensive"
safety_categories: optional array of string

Turns on content-safety analysis for the generated completions. Supply one or more of violence, self_harm, harassment, hate, sexual. Any non-empty list enables the analysis, and every row is then scored against all five categories — the list does not narrow what is checked. Results arrive as the prompt_safety_issues and response_safety_issues columns on the adapted dataset. Rows are annotated, never removed.

hallucination_mitigation: optional boolean

Enable web-search grounding to reduce hallucinations in generated completions.

blueprint: optional string

Freeform brand/style instructions injected as a system prompt for every generated completion. Use this to enforce tone, language, persona, or any guideline that does not fit the structured length/safety/grounding controls.

training_type: optional "instruction_dataset" or "preference_pairs"

How to adapt the dataset. instruction_dataset (default) produces enhanced prompt/completion pairs for SFT; preference_pairs generates chosen/rejected pairs for DPO. Source columns are the same in both cases — column_mapping.prompt (and optionally completion) — the chosen/rejected fields are produced by the adaptation pipeline.

One of the following:
"instruction_dataset"
"preference_pairs"
estimate: optional boolean

When true, validates the request and returns the estimated credit cost without starting a run.

language_expansion: optional object { type, languages, pairs, sample_rate }

Translation/localization expansion. When set, the pipeline produces additional rows in the requested languages or country/language variants. Output row count ≈ input × (1 + sample_rate × target_count); credits are billed on output rows.

Three-state field:

  • omitted — leaves any prior expansion config on the dataset unchanged (no-op). A wizard-saved expansion is preserved across SDK-driven /adapt calls that do not pass language_expansion.
  • language_expansion: null — explicitly clears any prior expansion config on this dataset (overrides wizard state).
  • object — overwrites the prior expansion config with the given spec.
type: "translate" or "localize"

Expansion mode. translate produces one new row variant per target language; localize produces one new row variant per country/language pair.

One of the following:
"translate"
"localize"
languages: optional array of string

Target language codes (required when type=translate, forbidden when type=localize). Mostly ISO 639-1 two-letter codes such as es, with three-letter codes where no two-letter code exists (fil, yue) and script-tagged codes for Chinese (zh-Hans, zh-Hant). Matching is case-insensitive, and bare zh resolves to zh-Hans. Validated against the supported language list at request time; unknown values return 400 with a sample of supported codes.

pairs: optional array of object { country, language }

Country/language pairs (required when type=localize, forbidden when type=translate). Each pair is validated against the supported country/language pair list at request time.

country: string

ISO 3166-1 alpha-2 country code.

language: string

Language code, as ISO 639-1 — optionally with a script subtag. Supported pairs include codes that are not two letters (zh-Hant for Taiwan, yue for Hong Kong), so the length is not pinned at two here. Membership is decided against the supported-pair list, which names an unsupported entry in its own 400.

sample_rate: number

Fraction (0.01–1) of input rows that are expanded per target. Output rows ≈ input × (1 + sample_rate × target_count). Required. Credits are billed on the expanded output row count.

minimum0.01
maximum1
ReturnsExpand Collapse
run_id: optional string

Unique identifier for this pipeline run. Null for estimate-only requests.

estimatedMinutes: number

Estimated processing time in minutes

estimatedCreditsConsumed: number

Estimated number of credits that will be consumed by this run

estimate: boolean

Whether this was an estimate-only request (no run started)

multimodalPricingApplied: boolean

True when an image column is mapped and also listed in context_columns; each output row is billed at a higher rate.

creditMultiplier: optional number

10 credits per 100 output rows when multimodalPricingApplied is true. Omitted for text-only pricing.

Adapt a dataset (deprecated)

curl https://api.prod.adaptionlabs.ai/api/v1/datasets/$DATASET_ID/run \
    -H 'Content-Type: application/json' \
    -H "Authorization: Bearer $ADAPTION_API_KEY" \
    -d '{
          "training_type": "instruction_dataset"
        }'
{
  "run_id": "dataset-550e8400-e29b-41d4-a716-446655440000-1712234567890",
  "estimatedMinutes": 0,
  "estimatedCreditsConsumed": 0,
  "estimate": true,
  "multimodalPricingApplied": true,
  "creditMultiplier": 0
}
Returns Examples
{
  "run_id": "dataset-550e8400-e29b-41d4-a716-446655440000-1712234567890",
  "estimatedMinutes": 0,
  "estimatedCreditsConsumed": 0,
  "estimate": true,
  "multimodalPricingApplied": true,
  "creditMultiplier": 0
}