Skip to content
SupportLogin
Datasets

Datasets

Get a dataset by ID
GET/api/v1/datasets/{dataset_id}
List datasets
GET/api/v1/datasets
Get the processing status of a dataset
GET/api/v1/datasets/{dataset_id}/status
Download the processed dataset
GET/api/v1/datasets/{dataset_id}/download
Publish a dataset to an external platform
POST/api/v1/datasets/{dataset_id}/publish
Adapt a dataset (deprecated)
Deprecated
POST/api/v1/datasets/{dataset_id}/run
Adapt a dataset (or estimate cost)
POST/api/v1/datasets/{dataset_id}/adapt
Add rows to a dataset from the curated pool
POST/api/v1/datasets/{dataset_id}/augment
Add translated rows to a dataset
POST/api/v1/datasets/{dataset_id}/translate
Add localized rows to a dataset
POST/api/v1/datasets/{dataset_id}/localize
Get evaluation results for a dataset
GET/api/v1/datasets/{dataset_id}/evaluation
Generate a dataset from scratch
POST/api/v1/datasets/invent
Delete a dataset
DELETE/api/v1/datasets/{dataset_id}
Best fine-tune launch config for re-launch parity
GET/api/v1/datasets/{dataset_id}/finetune/best-launch-config
ModelsExpand Collapse
Dataset object { dataset_id, kind, source_dataset_id, 12 more }
dataset_id: string

Unique dataset identifier

kind: "uploaded" or "combined" or "invented" or 3 more

How this dataset came about. uploaded was supplied by you, combined merges several datasets, invented was generated from a prompt, augmented adds rows retrieved from the curated pool to another dataset, and translated / localized add translated copies of another dataset’s rows.

One of the following:
"uploaded"
"combined"
"invented"
"augmented"
"translated"
"localized"
source_dataset_id: string

The dataset this one was derived from. Set when kind is augmented, translated or localized. Null for uploaded and invented, and for combined, which has several sources rather than one.

name: string

Human-readable name for the dataset

status: "pending" or "running" or "succeeded" or "failed"

Lifecycle status: pending, running, succeeded, or failed

One of the following:
"pending"
"running"
"succeeded"
"failed"
created_at: string

Timestamp when the dataset was created

formatdate-time
created_by_user_id: string

User who created the dataset

formatuuid
updated_at: string

Timestamp of the last update

formatdate-time
row_count: number

Total number of rows in the dataset

configured_column_mapping: object { prompt, completion, chat, 2 more }

User-configured column mapping. Null if not yet configured.

prompt: string
completion: string
chat: string
context: array of string
image: string
evaluation_summary: object { grade_before, grade_after, score_before, 2 more }

Compact evaluation summary. Null if evaluation has not completed.

grade_before: string

Letter grade (A-E) before adaptation

grade_after: string

Letter grade (A-E) after adaptation

score_before: number

Quality score before adaptation

score_after: number

Quality score after adaptation

improvement_percent: number

Relative improvement percentage

run_id: string

ID of the currently active run

progress: object { percent, processed_rows, total_rows }

Processing progress. Null when no run is active.

percent: number

Progress percentage (0-100)

processed_rows: number

Number of rows processed so far

total_rows: number

Total rows to process (samples_to_process or row_count)

error_data: object { message, code, level }

Error details if the dataset failed. Null otherwise.

message: string

Error message

code: string

Stable error code when the failure was structured (e.g. E0100)

level: "error" or "warning"

Severity when known

One of the following:
"error"
"warning"
image_column_formats: map["embedded_bytes" or "url" or "file_reference"]

Per-column export encoding for detected image columns (column name → format). Use with GET /datasets/{dataset_id}/download: look up the active image column (mapped image column that is also in configured_column_mapping.context) to determine how each row’s original_image is encoded. Null or empty when no image columns were detected.

One of the following:
"embedded_bytes"
"url"
"file_reference"
DatasetCreateResponse object { dataset_id, status, upload_instructions }
dataset_id: string

ID of the newly created dataset

status: string

Current dataset status

upload_instructions: optional object { url, method, s3_key }

Upload instructions for file sources. PUT your file to the provided URL.

url: string

Pre-signed URL for uploading the file

method: string

HTTP method to use

s3_key: string

S3 object key — pass this back in the complete request if needed for verification

DatasetListResponse object { dataset_id, name, status, 4 more }
dataset_id: string

Dataset ID

name: optional string

Dataset name

status: "pending" or "running" or "awaiting_input" or 2 more

Dataset status

One of the following:
"pending"
"running"
"awaiting_input"
"succeeded"
"failed"
updated_at: string

Last updated timestamp

formatdate-time
created_at: string

Timestamp when the dataset was created

formatdate-time
row_count: optional number

Total number of rows

description: optional string

Auto-generated description of the dataset contents

DatasetGetStatusResponse object { dataset_id, status, row_count, 2 more }
dataset_id: string

Dataset ID

status: "pending" or "running" or "awaiting_input" or 2 more

Current processing status. awaiting_input means the dataset is uploaded but waiting for you to start adaptation via datasets.adapt(column_mapping=…) — it will not progress on its own.

One of the following:
"pending"
"running"
"awaiting_input"
"succeeded"
"failed"
row_count: number

Number of rows in the dataset

progress: object { percent, processed_rows, total_rows }

Processing progress. Null when no run is active.

percent: number

Progress percentage (0-100)

processed_rows: number

Number of rows processed so far

total_rows: number

Total rows to process (samples_to_process or row_count)

error_data: object { message, code, level }

Error details if the dataset failed. Null otherwise.

message: string

Error message

code: string

Stable error code when the failure was structured (e.g. E0100)

level: "error" or "warning"

Severity when known

One of the following:
"error"
"warning"
DatasetPublishResponse object { publish_id, status, message }
publish_id: string

Unique identifier for the publish job

status: string

Status of the publish job

message: optional string

Additional information about the publish request

DatasetRunResponse object { run_id, estimatedMinutes, estimatedCreditsConsumed, 3 more }
run_id: optional string

Unique identifier for this pipeline run. Null for estimate-only requests.

estimatedMinutes: number

Estimated processing time in minutes

estimatedCreditsConsumed: number

Estimated number of credits that will be consumed by this run

estimate: boolean

Whether this was an estimate-only request (no run started)

multimodalPricingApplied: boolean

True when an image column is mapped and also listed in context_columns; each output row is billed at a higher rate.

creditMultiplier: optional number

10 credits per 100 output rows when multimodalPricingApplied is true. Omitted for text-only pricing.

DatasetAdaptResponse object { run_id, estimatedMinutes, estimatedCreditsConsumed, 3 more }
run_id: optional string

Unique identifier for this pipeline run. Null for estimate-only requests.

estimatedMinutes: number

Estimated processing time in minutes

estimatedCreditsConsumed: number

Estimated number of credits that will be consumed by this run

estimate: boolean

Whether this was an estimate-only request (no run started)

multimodalPricingApplied: boolean

True when an image column is mapped and also listed in context_columns; each output row is billed at a higher rate.

creditMultiplier: optional number

10 credits per 100 output rows when multimodalPricingApplied is true. Omitted for text-only pricing.

DatasetAugmentResponse object { dataset_id, status, estimated_credits_consumed, estimate }
dataset_id: string

The augmented dataset. It carries this dataset’s rows plus the retrieved ones, and is ready to download once its status reaches succeeded. Null for an estimate, which creates nothing.

status: "pending" or "running" or "awaiting_input" or 2 more

Status of the augmented dataset at the moment this response was sent, in the same vocabulary GET /datasets/{dataset_id}/status reports. Poll that endpoint until it reaches succeeded or failed. An idempotent replay of a finished run returns its terminal status here.

One of the following:
"pending"
"running"
"awaiting_input"
"succeeded"
"failed"
estimated_credits_consumed: number

Credits this run consumes, charged against the requested rows.

estimate: boolean

Whether this was an estimate-only request (no run started).

DatasetTranslateResponse object { dataset_id, status, estimated_credits_consumed, 2 more }
dataset_id: string

The expanded dataset. It carries this dataset’s rows plus the new ones, and is ready to download once its status reaches succeeded. Null for an estimate, which creates nothing.

status: "pending" or "running" or "awaiting_input" or 2 more

Status of the expanded dataset at the moment this response was sent, in the same vocabulary GET /datasets/{dataset_id}/status reports. Poll that endpoint until it reaches succeeded or failed. An idempotent replay of a finished run returns its terminal status here.

One of the following:
"pending"
"running"
"awaiting_input"
"succeeded"
"failed"
estimated_credits_consumed: number

Credits this run consumes, charged against the rows it adds. Zero on an idempotent replay: the earlier run reserved the credits, and this call charges nothing.

estimate: boolean

Whether this was an estimate-only request (no run started).

estimated_new_rows: number

Rows this run adds, which is what the credit cost is charged against. Zero on an idempotent replay, which adds none.

DatasetLocalizeResponse object { dataset_id, status, estimated_credits_consumed, 2 more }
dataset_id: string

The expanded dataset. It carries this dataset’s rows plus the new ones, and is ready to download once its status reaches succeeded. Null for an estimate, which creates nothing.

status: "pending" or "running" or "awaiting_input" or 2 more

Status of the expanded dataset at the moment this response was sent, in the same vocabulary GET /datasets/{dataset_id}/status reports. Poll that endpoint until it reaches succeeded or failed. An idempotent replay of a finished run returns its terminal status here.

One of the following:
"pending"
"running"
"awaiting_input"
"succeeded"
"failed"
estimated_credits_consumed: number

Credits this run consumes, charged against the rows it adds. Zero on an idempotent replay: the earlier run reserved the credits, and this call charges nothing.

estimate: boolean

Whether this was an estimate-only request (no run started).

estimated_new_rows: number

Rows this run adds, which is what the credit cost is charged against. Zero on an idempotent replay, which adds none.

DatasetGetEvaluationResponse object { dataset_id, status, quality, raw_results }
dataset_id: string

Dataset ID

status: string

Evaluation pipeline status: pending | running | succeeded | failed | skipped

quality: object { grade_before, grade_after, score_before, 3 more }

Structured quality metrics. Null until evaluation completes.

grade_before: string

Letter grade (A-E) before adaptation

grade_after: string

Letter grade (A-E) after adaptation

score_before: number

Quality score (0-10) before adaptation

score_after: number

Quality score (0-10) after adaptation

percentile_after: number

Percentile rank (0-100) after adaptation

improvement_percent: number

Relative quality improvement as a percentage

raw_results: map[unknown]

Raw evaluation results payload for advanced use. Null until evaluation completes.

DatasetInventResponse object { estimate, id, name, 8 more }
estimate: boolean

True when this response priced the request without running it.

id: string

Dataset id, or null on an estimate. Pass it straight to finetune_jobs.create or autoscientist.create once the status is succeeded.

name: string

Display name, echoed from the request or the default applied for you.

training_type: "instruction_dataset" or "preference_pairs"

Shape of the generated data, echoed from the request.

One of the following:
"instruction_dataset"
"preference_pairs"
domains: array of string

Domain codes the generation covers, echoed from the request — including any inferred from subdomains.

subdomains: array of string

Subdomain codes the generation is narrowed to, if any.

rows: number

Number of rows requested.

status: "pending" or "running" or "awaiting_input" or 2 more

Lifecycle status, or null on an estimate. A fresh generation is always running; poll GET /datasets/{dataset_id} until it reports succeeded or failed.

One of the following:
"pending"
"running"
"awaiting_input"
"succeeded"
"failed"
created_at: string

When the dataset was created, or null on an estimate.

formatdate-time
estimated_credits: number

Credits this generation consumes. Billed on the post-expansion row count, so a translated request quotes above its rows. Invent pricing is still provisional, so treat the figure as indicative.

available_credits: number

Credits currently available to your team.

DatasetInventDomainsResponse object { domains }
domains: array of object { code, title, subdomains }
code: string

Code to send in domains.

title: string

Human-readable name.

subdomains: array of object { code, title }

Subdomains available within this domain. Empty when the domain has no subdivision — send the domain code alone.

code: string

Qualified code to send in subdomains — copy it as-is rather than assembling it.

title: string

Human-readable name.

DatasetDeleteResponse object { message }
message: string
DatasetGetBestLaunchConfigResponse object { best_job_config }
best_job_config: object { finetune_job_id, training_experiment_id, original_model_name, 5 more }

Launch-parity snapshot for the experiment best job (terminal experiment with best_finetune_job_id) or, when no experiment exists, the newest succeeded standalone job (training_experiment_id null). Null while an experiment is non-terminal, when no best job was chosen yet, or when no qualifying job exists.

finetune_job_id: string

Fine-tune job whose config is shown (experiment best or standalone).

formatuuid
training_experiment_id: string

Training experiment when this snapshot is the AutoScientist best job; null for a standalone job.

formatuuid
original_model_name: string

Base model id the job was launched with.

trained_model_name: string

Output label / suffix for the trained model. Taken from the value recorded at launch when present; otherwise derived from the current dataset name, in which case it can differ from the label the job was submitted with if the dataset was renamed. Null only when no label was recorded and the dataset is unavailable.

training_method: "sft" or "dpo"
One of the following:
"sft"
"dpo"
training_type: "lora" or "full"
One of the following:
"lora"
"full"
data_format: "chat" or "instruction" or "preference"
One of the following:
"chat"
"instruction"
"preference"
hyperparams: map[unknown]

Hyperparameters the job was launched with.

DatasetsUpload

Initiate a dataset upload
POST/api/v1/datasets/upload/initiate
Complete a dataset upload and trigger processing
POST/api/v1/datasets/upload/complete
Complete a file upload and trigger processing
POST/api/v1/datasets/{dataset_id}/upload/complete
Initiate a batch upload
POST/api/v1/datasets/upload/initiate-batch
Complete a batch upload and trigger processing
POST/api/v1/datasets/upload/complete-batch
ModelsExpand Collapse
UploadInitiateResponse object { upload_url }
upload_url: string

Pre-signed S3 URL — upload the file directly to this URL via HTTP PUT

UploadCompleteResponse object { dataset_id }
dataset_id: string

ID of the newly created dataset

UploadCompleteByIDResponse object { dataset_id, status }
dataset_id: string

ID of the dataset

status: string

Current status of the dataset after completing upload

UploadInitiateBatchResponse object { dataset_id, uploads }
dataset_id: string

Dataset ID

uploads: array of object { file_name, upload_url, s3_key }
file_name: string

Original file name

upload_url: string

Presigned S3 upload URL

s3_key: string

S3 key for the uploaded file

UploadCompleteBatchResponse object { dataset_id, status }
dataset_id: string

Dataset ID

status: string

Current dataset status

DatasetsCombine

ModelsExpand Collapse
CombineCreateResponse object { dataset_id }
dataset_id: string

The newly created merged Dataset’s ID. Returns existing ID on idempotent replay.

CombineValidateResponse object { compatible, training_type, total_estimated_rows, issues }
compatible: boolean

Whether the proposed merge can proceed.

training_type: optional "instruction_dataset" or "preference_pairs"

Resolved training type for a compatible merge. Homogeneous source sets inherit their shared type; mixed instruction + preference sets resolve to instruction_dataset.

One of the following:
"instruction_dataset"
"preference_pairs"
total_estimated_rows: optional number

Sum of processed_rows across all sources for a compatible merge.

issues: array of object { code, level, message, 9 more }

Validation diagnostics. Errors block the merge; warnings describe reconciliation the merge performs automatically.

code: "too_few_sources" or "too_many_sources" or "duplicate_source_ids" or 9 more

Stable machine-readable validation issue code.

One of the following:
"too_few_sources"
"too_many_sources"
"duplicate_source_ids"
"dataset_not_found"
"wrong_organization"
"dataset_not_ready"
"missing_run_id"
"mixed_multimodal_sources"
"all_preference_sources_missing_dpo_columns"
"incompatible_column_types"
"preference_source_missing_ranked_generations"
"mixed_training_types_resolved_to_sft"
level: "error" or "warning"

error blocks the merge; warning describes automatic reconciliation.

One of the following:
"error"
"warning"
message: string

Preformed user-facing message.

count: optional number
limit: optional number
dataset_id: optional string
status: optional string
dataset_ids: optional array of string
multimodal_dataset_ids: optional array of string
text_dataset_ids: optional array of string
columns: optional array of object { column, types, cast_to }
column: string

Column whose source types do not agree.

types: array of object { dataset_id, type }
dataset_id: string

Source Dataset ID.

type: string

DuckDB type observed for this Dataset.

cast_to: string

Common DuckDB type used by the merge.

types: optional array of object { dataset_id, training_type }
dataset_id: string

Source Dataset ID.

training_type: "instruction_dataset" or "preference_pairs"

Training type resolved for this source Dataset.

One of the following:
"instruction_dataset"
"preference_pairs"