Datasets
Create a dataset from file upload, HuggingFace, Kaggle, or Google Sheets
Get a dataset by ID
List datasets
Get the processing status of a dataset
Download the processed dataset
Publish a dataset to an external platform
Adapt a dataset (or estimate cost)
Add rows to a dataset from the curated pool
Add translated rows to a dataset
Add localized rows to a dataset
Get evaluation results for a dataset
Generate a dataset from scratch
List the domains and subdomains available for generation
Delete a dataset
Poll saved URLs for successful dataset exports (Hugging Face, Kaggle, Google Sheets)
Best fine-tune launch config for re-launch parity
ModelsExpand Collapse
Dataset object { dataset_id, kind, source_dataset_id, 12 more }
kind: "uploaded" or "combined" or "invented" or 3 moreHow this dataset came about. uploaded was supplied by you, combined merges several datasets, invented was generated from a prompt, augmented adds rows retrieved from the curated pool to another dataset, and translated / localized add translated copies of another dataset’s rows.
How this dataset came about. uploaded was supplied by you, combined merges several datasets, invented was generated from a prompt, augmented adds rows retrieved from the curated pool to another dataset, and translated / localized add translated copies of another dataset’s rows.
The dataset this one was derived from. Set when kind is augmented, translated or localized. Null for uploaded and invented, and for combined, which has several sources rather than one.
status: "pending" or "running" or "succeeded" or "failed"Lifecycle status: pending, running, succeeded, or failed
Lifecycle status: pending, running, succeeded, or failed
configured_column_mapping: object { prompt, completion, chat, 2 more } User-configured column mapping. Null if not yet configured.
User-configured column mapping. Null if not yet configured.
evaluation_summary: object { grade_before, grade_after, score_before, 2 more } Compact evaluation summary. Null if evaluation has not completed.
Compact evaluation summary. Null if evaluation has not completed.
progress: object { percent, processed_rows, total_rows } Processing progress. Null when no run is active.
Processing progress. Null when no run is active.
image_column_formats: map["embedded_bytes" or "url" or "file_reference"]Per-column export encoding for detected image columns (column name → format). Use with GET /datasets/{dataset_id}/download: look up the active image column (mapped image column that is also in configured_column_mapping.context) to determine how each row’s original_image is encoded. Null or empty when no image columns were detected.
Per-column export encoding for detected image columns (column name → format). Use with GET /datasets/{dataset_id}/download: look up the active image column (mapped image column that is also in configured_column_mapping.context) to determine how each row’s original_image is encoded. Null or empty when no image columns were detected.
DatasetGetStatusResponse object { dataset_id, status, row_count, 2 more }
status: "pending" or "running" or "awaiting_input" or 2 moreCurrent processing status. awaiting_input means the dataset is uploaded but waiting for you to start adaptation via datasets.adapt(column_mapping=…) — it will not progress on its own.
Current processing status. awaiting_input means the dataset is uploaded but waiting for you to start adaptation via datasets.adapt(column_mapping=…) — it will not progress on its own.
DatasetAugmentResponse object { dataset_id, status, estimated_credits_consumed, estimate }
The augmented dataset. It carries this dataset’s rows plus the retrieved ones, and is ready to download once its status reaches succeeded. Null for an estimate, which creates nothing.
status: "pending" or "running" or "awaiting_input" or 2 moreStatus of the augmented dataset at the moment this response was sent, in the same vocabulary GET /datasets/{dataset_id}/status reports. Poll that endpoint until it reaches succeeded or failed. An idempotent replay of a finished run returns its terminal status here.
Status of the augmented dataset at the moment this response was sent, in the same vocabulary GET /datasets/{dataset_id}/status reports. Poll that endpoint until it reaches succeeded or failed. An idempotent replay of a finished run returns its terminal status here.
DatasetTranslateResponse object { dataset_id, status, estimated_credits_consumed, 2 more }
The expanded dataset. It carries this dataset’s rows plus the new ones, and is ready to download once its status reaches succeeded. Null for an estimate, which creates nothing.
status: "pending" or "running" or "awaiting_input" or 2 moreStatus of the expanded dataset at the moment this response was sent, in the same vocabulary GET /datasets/{dataset_id}/status reports. Poll that endpoint until it reaches succeeded or failed. An idempotent replay of a finished run returns its terminal status here.
Status of the expanded dataset at the moment this response was sent, in the same vocabulary GET /datasets/{dataset_id}/status reports. Poll that endpoint until it reaches succeeded or failed. An idempotent replay of a finished run returns its terminal status here.
DatasetLocalizeResponse object { dataset_id, status, estimated_credits_consumed, 2 more }
The expanded dataset. It carries this dataset’s rows plus the new ones, and is ready to download once its status reaches succeeded. Null for an estimate, which creates nothing.
status: "pending" or "running" or "awaiting_input" or 2 moreStatus of the expanded dataset at the moment this response was sent, in the same vocabulary GET /datasets/{dataset_id}/status reports. Poll that endpoint until it reaches succeeded or failed. An idempotent replay of a finished run returns its terminal status here.
Status of the expanded dataset at the moment this response was sent, in the same vocabulary GET /datasets/{dataset_id}/status reports. Poll that endpoint until it reaches succeeded or failed. An idempotent replay of a finished run returns its terminal status here.
DatasetInventResponse object { estimate, id, name, 8 more }
Dataset id, or null on an estimate. Pass it straight to finetune_jobs.create or autoscientist.create once the status is succeeded.
training_type: "instruction_dataset" or "preference_pairs"Shape of the generated data, echoed from the request.
Shape of the generated data, echoed from the request.
Domain codes the generation covers, echoed from the request — including any inferred from subdomains.
status: "pending" or "running" or "awaiting_input" or 2 moreLifecycle status, or null on an estimate. A fresh generation is always running; poll GET /datasets/{dataset_id} until it reports succeeded or failed.
Lifecycle status, or null on an estimate. A fresh generation is always running; poll GET /datasets/{dataset_id} until it reports succeeded or failed.
DatasetGetBestLaunchConfigResponse object { best_job_config }
best_job_config: object { finetune_job_id, training_experiment_id, original_model_name, 5 more } Launch-parity snapshot for the experiment best job (terminal experiment with best_finetune_job_id) or, when no experiment exists, the newest succeeded standalone job (training_experiment_id null). Null while an experiment is non-terminal, when no best job was chosen yet, or when no qualifying job exists.
Launch-parity snapshot for the experiment best job (terminal experiment with best_finetune_job_id) or, when no experiment exists, the newest succeeded standalone job (training_experiment_id null). Null while an experiment is non-terminal, when no best job was chosen yet, or when no qualifying job exists.
Fine-tune job whose config is shown (experiment best or standalone).
Training experiment when this snapshot is the AutoScientist best job; null for a standalone job.
Output label / suffix for the trained model. Taken from the value recorded at launch when present; otherwise derived from the current dataset name, in which case it can differ from the label the job was submitted with if the dataset was renamed. Null only when no label was recorded and the dataset is unavailable.
DatasetsUpload
Initiate a dataset upload
Complete a dataset upload and trigger processing
Complete a file upload and trigger processing
Initiate a batch upload
Complete a batch upload and trigger processing
ModelsExpand Collapse
DatasetsCombine
Combine multiple Datasets into a single new merged Dataset
Validate that a set of Datasets can be combined
ModelsExpand Collapse
CombineValidateResponse object { compatible, training_type, total_estimated_rows, issues }
training_type: optional "instruction_dataset" or "preference_pairs"Resolved training type for a compatible merge. Homogeneous source sets inherit their shared type; mixed instruction + preference sets resolve to instruction_dataset.
Resolved training type for a compatible merge. Homogeneous source sets inherit their shared type; mixed instruction + preference sets resolve to instruction_dataset.
Sum of processed_rows across all sources for a compatible merge.
issues: array of object { code, level, message, 9 more } Validation diagnostics. Errors block the merge; warnings describe reconciliation the merge performs automatically.
Validation diagnostics. Errors block the merge; warnings describe reconciliation the merge performs automatically.
code: "too_few_sources" or "too_many_sources" or "duplicate_source_ids" or 9 moreStable machine-readable validation issue code.
Stable machine-readable validation issue code.