Skip to content
SupportLogin

Custom Evals

Start a custom-rubric evaluation for an adapted dataset
datasets.custom_evals.create(strdataset_id, CustomEvalCreateParams**kwargs) -> CustomEval
POST/api/v1/datasets/{dataset_id}/custom-evals
List custom-rubric evaluations for a dataset, newest first
datasets.custom_evals.list(strdataset_id) -> CustomEvalListResponse
GET/api/v1/datasets/{dataset_id}/custom-evals
Get a single custom-rubric evaluation
datasets.custom_evals.get(strcustom_eval_id, CustomEvalGetParams**kwargs) -> CustomEval
GET/api/v1/datasets/{dataset_id}/custom-evals/{custom_eval_id}
Draft a judge rubric tailored to an adapted dataset
datasets.custom_evals.prepare(strdataset_id) -> CustomEvalPrepareResponse
POST/api/v1/datasets/{dataset_id}/custom-evals/prepare
ModelsExpand Collapse
class CustomEval:
custom_eval_id: str
dataset_id: str
status: str

pending | running | succeeded | failed

judge_prompt: str
name: Optional[str]
score_before: Optional[float]

Mean rubric score on the original pairs. Null until the eval succeeds.

score_after: Optional[float]

Mean rubric score on the adapted pairs. Null until the eval succeeds.

improvement_percent: Optional[float]

Percentage change from score_before to score_after. Null until the eval succeeds.

scored_rows: Optional[int]

Rows the judge scored. A rubric score is an average over this sample, not the whole dataset.

error_message: Optional[str]
created_at: datetime
formatdate-time
completed_at: Optional[datetime]
formatdate-time
class CustomEvalListResponse:
items: List[CustomEval]
custom_eval_id: str
dataset_id: str
status: str

pending | running | succeeded | failed

judge_prompt: str
name: Optional[str]
score_before: Optional[float]

Mean rubric score on the original pairs. Null until the eval succeeds.

score_after: Optional[float]

Mean rubric score on the adapted pairs. Null until the eval succeeds.

improvement_percent: Optional[float]

Percentage change from score_before to score_after. Null until the eval succeeds.

scored_rows: Optional[int]

Rows the judge scored. A rubric score is an average over this sample, not the whole dataset.

error_message: Optional[str]
created_at: datetime
formatdate-time
completed_at: Optional[datetime]
formatdate-time
class CustomEvalPrepareResponse:
judge_prompt: str

A ready-to-edit rubric, accepted as-is by judge_prompt on POST /datasets/{dataset_id}/custom-evals.

adapted: bool

False when the stock rubric came back untouched — the sampled values did not fill both halves of the before/after contrast, or the adaptation model was unreachable. The prompt is still valid, just not tailored to this dataset.

sampled_rows: int

How many of the dataset’s own examples the rubric was written against — half original, half adapted. Lower than the usual target on a thin dataset; a dataset too thin to write from at all is rejected with a 400.