Skip to content
SupportLogin
Edit on GitHub

Configure Adaptive Data

Generate instruction or preference data, apply recipes, limit rows, and control adapted outputs.

After you map your columns, configure how Adaptive Data transforms them. You can choose its training format, apply recipes, limit the run, and define output requirements.

The SDK examples below assume you have configured the Python SDK and set ADAPTION_DATASET_ID to an imported dataset:

import os
from adaption import Adaption
client = Adaption()
dataset_id = os.environ["ADAPTION_DATASET_ID"]
column_mapping = {
"prompt": "instruction",
"completion": "response",
}

A run always applies your settings in the same order. This means the order you pass your arguments to datasets.adapt has no effect.

  1. Map columns. column_mapping selects which source columns are read as the prompt, completion, chat, and context.
  2. Limit rows. If job_specification.max_rows is set, a sample of that many rows is selected: a random sample for text datasets, the first rows for image datasets. Later steps see only those rows.
  3. Deduplicate. Exact duplicate rows are removed. This happens after sampling, so a run with max_rows=500 can return fewer than 500 rows.
  4. Expand languages. If language_expansion is set, a share of the remaining rows is copied into each target language and added as new rows. Expansion comes before rephrasing and generation, so translated rows are rephrased and completed in their target language. The completions are written in that language, not translated afterwards.
  5. Rephrase prompts. If prompt_rephrase is on, each prompt is rewritten, including the prompts on rows added in step 4. Datasets mapped through a chat column are not rephrased. length is applied here as guidance attached to the rewritten prompt, so it takes effect only when prompt_rephrase is on.
  6. Generate completions. Each row receives its adapted completion. training_type decides what this step produces: one completion for instruction_dataset, or a chosen and rejected pair for preference_pairs. Two brand controls apply here:
    • blueprint is given to the model as a system instruction.
    • hallucination_mitigation runs a web search first, for prompts that need current information, and passes the results to the model. It applies to text datasets only.
  7. Capture reasoning traces. If reasoning_traces is on, every row is completed by reasoning models and the reasoning behind each completion is saved as it is generated. It is not a separate pass over the data.
  8. Review safety. If safety_categories is set, each finished completion is reviewed after generation.

Recipes are optional transformations that operate on the mapped data:

  • Prompt deduplication removes duplicate prompts to keep examples diverse.
  • Prompt rephrasing rewrites prompts for clarity and alignment with the task.
  • Reasoning traces provides associated reasoning for adapted completions.

In the SDK, recipes belong under recipe_specification.recipes. You can enable more than one recipe in the same run:

run = client.datasets.adapt(
dataset_id,
column_mapping=column_mapping,
recipe_specification={
"recipes": {
"deduplication": True,
"prompt_rephrase": True,
"reasoning_traces": True,
},
},
)
print(f"Run started: {run.run_id}")

Use reasoning traces if you want to fine-tune a model that supports chain-of-thought reasoning.

By default, Adaptive Data produces an instruction dataset for supervised fine-tuning. Set training_type to preference_pairs when you instead need chosen and rejected responses for preference training such as DPO.

The source mapping stays the same: provide a prompt column and, optionally, an existing completion column. Adaptive Data generates the preference-pair fields in the output.

preference_mapping = {
"prompt": "instruction",
"completion": "response",
}
estimate = client.datasets.adapt(
dataset_id,
column_mapping=preference_mapping,
training_type="preference_pairs",
estimate=True,
)
print(f"Estimated credits: {estimate.estimated_credits_consumed}")
run = client.datasets.adapt(
dataset_id,
column_mapping=preference_mapping,
training_type="preference_pairs",
)
print(f"Run started: {run.run_id}")

Use training_type="instruction_dataset"—or omit training_type—to produce enhanced prompt/completion pairs instead. Language expansion, recipes, job limits, and brand controls can be included in the same preference-pair run. See training_type on the datasets.adapt reference for the current schema.

Use job_specification.max_rows to test mappings and controls on a subset before processing a large dataset. Request an estimate with datasets.adapt first to validate the same configuration and receive a credit quote without starting adaptation:

job_specification = {"max_rows": 500}
estimate = client.datasets.adapt(
dataset_id,
column_mapping=column_mapping,
job_specification=job_specification,
estimate=True,
)
print(f"Estimated credits: {estimate.estimated_credits_consumed}")
run = client.datasets.adapt(
dataset_id,
column_mapping=column_mapping,
job_specification=job_specification,
)
print(f"Run started: {run.run_id}")

Omit max_rows when you want to process the full dataset. See JobSpecification on the datasets.adapt reference for all job-level options.

brand_controls makes output requirements part of the adaptation run:

  • length controls target verbosity: minimal, concise, detailed, or extensive.
  • safety_categories identifies content-safety categories to enforce.
  • hallucination_mitigation grounds generations with web search where applicable.
  • blueprint supplies freeform instructions for tone, persona, audience, language, or product-specific wording.

Hallucination mitigation steers generated completions toward verifiable information. Enable it for fact-sensitive data such as customer support, grounded question answering, or compliance workflows:

run = client.datasets.adapt(
dataset_id,
column_mapping=column_mapping,
brand_controls={"hallucination_mitigation": True},
)

Grounding can affect latency and credit use. Use estimate=True to compare the run before scaling it.

Length and safety constraints compose in the same brand_controls object. Use the datasets.adapt reference as the source of truth for supported safety category names:

run = client.datasets.adapt(
dataset_id,
column_mapping=column_mapping,
brand_controls={
"length": "concise",
"safety_categories": ["harassment", "hate"],
},
)

Use blueprint for qualitative requirements that do not fit a structured field. It acts as a system instruction for generated completions and applies consistently across the run:

run = client.datasets.adapt(
dataset_id,
column_mapping=column_mapping,
brand_controls={
"blueprint": (
"You are a customer-success assistant for Acme. "
"Answer in British English with a warm, direct tone. "
"Do not recommend third-party tools."
),
},
)

Structured controls and Blueprint instructions can be used together:

brand_controls = {
"length": "concise",
"safety_categories": ["harassment", "hate"],
"hallucination_mitigation": True,
"blueprint": "Answer in British English with a warm, direct tone.",
}
estimate = client.datasets.adapt(
dataset_id,
column_mapping=column_mapping,
brand_controls=brand_controls,
job_specification={"max_rows": 500},
estimate=True,
)
print(f"Estimated credits: {estimate.estimated_credits_consumed}")
run = client.datasets.adapt(
dataset_id,
column_mapping=column_mapping,
brand_controls=brand_controls,
job_specification={"max_rows": 500},
)
print(f"Run started: {run.run_id}")

After the run completes, evaluate dataset quality and download the adapted rows.