Skip to content
SupportLogin
Edit on GitHub

Run on non-adapted data

Upload training-ready prompt/completion data for supervised fine-tuning.

Use raw dataset processing when a tabular file already contains training-ready prompt and completion columns and you want AutoScientist to train on those rows as-is with SFT. This skips Adaptive Data’s adaptation pipeline.

Raw processing supports CSV, JSON, JSONL, and Parquet files. Set processing_mode to raw and map both required source columns when creating the dataset.

This example uploads a CSV whose instruction and response columns contain the training pairs:

from adaption import Adaption
client = Adaption()
dataset = client.datasets.upload_file(
"training_data.csv",
processing_mode="raw",
column_mapping={
"prompt": "instruction",
"completion": "response",
},
)
client.datasets.wait_for_completion(dataset.dataset_id)
run = client.autoscientist.create(dataset_id=dataset.dataset_id)
run = client.autoscientist.wait_for_completion(run.id)
if run.status == "failed":
raise RuntimeError(f"failed: {run.error}")
if run.status != "succeeded":
raise RuntimeError(f"Run ended with status: {run.status}")
print(run.id, run.best_win_rate)

Do not call client.datasets.adapt for this workflow: raw processing makes the uploaded dataset ready for AutoScientist directly. The column_mapping above refers to the original file headers; an AutoScientist column_mapping, if supplied later, refers to the processed dataset schema.

See Running AutoScientist for polling, cancellation, and status details. For endpoint details, see the datasets.create and autoscientist.create references.