Configure Adaptive Data
Generate instruction or preference data, apply recipes, limit rows, and control adapted outputs.
After you map your columns, configure how Adaptive Data transforms them. You can choose its training format, apply recipes, limit the run, and define output requirements.
The SDK examples below assume you have configured the Python SDK and set ADAPTION_DATASET_ID to an imported dataset:
import os
from adaption import Adaption
client = Adaption()dataset_id = os.environ["ADAPTION_DATASET_ID"]column_mapping = { "prompt": "instruction", "completion": "response",}How a run is processed
Section titled “How a run is processed”A run always applies your settings in the same order. This means the order you pass your arguments to datasets.adapt has no effect.
- Map columns.
column_mappingselects which source columns are read as the prompt, completion, chat, and context. - Limit rows. If
job_specification.max_rowsis set, a sample of that many rows is selected: a random sample for text datasets, the first rows for image datasets. Later steps see only those rows. - Deduplicate. Exact duplicate rows are removed. This happens after sampling, so a run with
max_rows=500can return fewer than 500 rows. - Expand languages. If
language_expansionis set, a share of the remaining rows is copied into each target language and added as new rows. Expansion comes before rephrasing and generation, so translated rows are rephrased and completed in their target language. The completions are written in that language, not translated afterwards. - Rephrase prompts. If
prompt_rephraseis on, each prompt is rewritten, including the prompts on rows added in step 4. Datasets mapped through achatcolumn are not rephrased.lengthis applied here as guidance attached to the rewritten prompt, so it takes effect only whenprompt_rephraseis on. - Generate completions. Each row receives its adapted completion.
training_typedecides what this step produces: one completion forinstruction_dataset, or a chosen and rejected pair forpreference_pairs. Two brand controls apply here:blueprintis given to the model as a system instruction.hallucination_mitigationruns a web search first, for prompts that need current information, and passes the results to the model. It applies to text datasets only.
- Capture reasoning traces. If
reasoning_tracesis on, every row is completed by reasoning models and the reasoning behind each completion is saved as it is generated. It is not a separate pass over the data. - Review safety. If
safety_categoriesis set, each finished completion is reviewed after generation.
Apply recipes
Section titled “Apply recipes”Recipes are optional transformations that operate on the mapped data:
- Prompt deduplication removes duplicate prompts to keep examples diverse.
- Prompt rephrasing rewrites prompts for clarity and alignment with the task.
- Reasoning traces provides associated reasoning for adapted completions.
In the SDK, recipes belong under recipe_specification.recipes. You can enable more than one recipe in the same run:
run = client.datasets.adapt( dataset_id, column_mapping=column_mapping, recipe_specification={ "recipes": { "deduplication": True, "prompt_rephrase": True, "reasoning_traces": True, }, },)print(f"Run started: {run.run_id}")Use reasoning traces if you want to fine-tune a model that supports chain-of-thought reasoning.
Generate preference pairs
Section titled “Generate preference pairs”By default, Adaptive Data produces an instruction dataset for supervised fine-tuning. Set training_type to preference_pairs when you instead need chosen and rejected responses for preference training such as DPO.
The source mapping stays the same: provide a prompt column and, optionally, an existing completion column. Adaptive Data generates the preference-pair fields in the output.
preference_mapping = { "prompt": "instruction", "completion": "response",}
estimate = client.datasets.adapt( dataset_id, column_mapping=preference_mapping, training_type="preference_pairs", estimate=True,)print(f"Estimated credits: {estimate.estimated_credits_consumed}")
run = client.datasets.adapt( dataset_id, column_mapping=preference_mapping, training_type="preference_pairs",)print(f"Run started: {run.run_id}")Use training_type="instruction_dataset"—or omit training_type—to produce enhanced prompt/completion pairs instead. Language expansion, recipes, job limits, and brand controls can be included in the same preference-pair run. See training_type on the datasets.adapt reference for the current schema.
Limit rows and estimate a run
Section titled “Limit rows and estimate a run”Use job_specification.max_rows to test mappings and controls on a subset before processing a large dataset. Request an estimate with datasets.adapt first to validate the same configuration and receive a credit quote without starting adaptation:
job_specification = {"max_rows": 500}
estimate = client.datasets.adapt( dataset_id, column_mapping=column_mapping, job_specification=job_specification, estimate=True,)print(f"Estimated credits: {estimate.estimated_credits_consumed}")
run = client.datasets.adapt( dataset_id, column_mapping=column_mapping, job_specification=job_specification,)print(f"Run started: {run.run_id}")Omit max_rows when you want to process the full dataset. See JobSpecification on the datasets.adapt reference for all job-level options.
Configure brand controls
Section titled “Configure brand controls”brand_controls makes output requirements part of the adaptation run:
lengthcontrols target verbosity:minimal,concise,detailed, orextensive.safety_categoriesidentifies content-safety categories to enforce.hallucination_mitigationgrounds generations with web search where applicable.blueprintsupplies freeform instructions for tone, persona, audience, language, or product-specific wording.
Reduce hallucinations
Section titled “Reduce hallucinations”Hallucination mitigation steers generated completions toward verifiable information. Enable it for fact-sensitive data such as customer support, grounded question answering, or compliance workflows:
run = client.datasets.adapt( dataset_id, column_mapping=column_mapping, brand_controls={"hallucination_mitigation": True},)Grounding can affect latency and credit use. Use estimate=True to compare the run before scaling it.
Control length and safety
Section titled “Control length and safety”Length and safety constraints compose in the same brand_controls object. Use the datasets.adapt reference as the source of truth for supported safety category names:
run = client.datasets.adapt( dataset_id, column_mapping=column_mapping, brand_controls={ "length": "concise", "safety_categories": ["harassment", "hate"], },)Set a Blueprint
Section titled “Set a Blueprint”Use blueprint for qualitative requirements that do not fit a structured field. It acts as a system instruction for generated completions and applies consistently across the run:
run = client.datasets.adapt( dataset_id, column_mapping=column_mapping, brand_controls={ "blueprint": ( "You are a customer-success assistant for Acme. " "Answer in British English with a warm, direct tone. " "Do not recommend third-party tools." ), },)Combine brand controls
Section titled “Combine brand controls”Structured controls and Blueprint instructions can be used together:
brand_controls = { "length": "concise", "safety_categories": ["harassment", "hate"], "hallucination_mitigation": True, "blueprint": "Answer in British English with a warm, direct tone.",}
estimate = client.datasets.adapt( dataset_id, column_mapping=column_mapping, brand_controls=brand_controls, job_specification={"max_rows": 500}, estimate=True,)print(f"Estimated credits: {estimate.estimated_credits_consumed}")
run = client.datasets.adapt( dataset_id, column_mapping=column_mapping, brand_controls=brand_controls, job_specification={"max_rows": 500},)print(f"Run started: {run.run_id}")After the run completes, evaluate dataset quality and download the adapted rows.