# Custom Evals

## Start a custom-rubric evaluation for an adapted dataset

**post** `/api/v1/datasets/{dataset_id}/custom-evals`

Start a custom-rubric evaluation for an adapted dataset

### Path Parameters

- `dataset_id: string`

### Body Parameters

- `judge_prompt: string`

  The judge rubric. Scored per sample against the original and adapted pairs.

- `name: optional string`

  Optional label so a history of rubrics stays legible.

### Returns

- `CustomEval object { custom_eval_id, dataset_id, status, 9 more }`

  - `custom_eval_id: string`

  - `dataset_id: string`

  - `status: string`

    pending | running | succeeded | failed

  - `judge_prompt: string`

  - `name: string`

  - `score_before: number`

    Mean rubric score on the original pairs. Null until the eval succeeds.

  - `score_after: number`

    Mean rubric score on the adapted pairs. Null until the eval succeeds.

  - `improvement_percent: number`

    Percentage change from score_before to score_after. Null until the eval succeeds.

  - `scored_rows: number`

    Rows the judge scored. A rubric score is an average over this sample, not the whole dataset.

  - `error_message: string`

  - `created_at: string`

  - `completed_at: string`

### Example

```http
curl https://api.prod.adaptionlabs.ai/api/v1/datasets/$DATASET_ID/custom-evals \
    -H 'Content-Type: application/json' \
    -H "Authorization: Bearer $ADAPTION_API_KEY" \
    -d '{
          "judge_prompt": "Rate how factually accurate the response is, from 0 (fabricated) to 10 (fully supported)."
        }'
```

#### Response

```json
{
  "custom_eval_id": "custom_eval_id",
  "dataset_id": "dataset_id",
  "status": "running",
  "judge_prompt": "judge_prompt",
  "name": "name",
  "score_before": 0,
  "score_after": 0,
  "improvement_percent": 0,
  "scored_rows": 0,
  "error_message": "error_message",
  "created_at": "2019-12-27T18:11:19.117Z",
  "completed_at": "2019-12-27T18:11:19.117Z"
}
```

## List custom-rubric evaluations for a dataset, newest first

**get** `/api/v1/datasets/{dataset_id}/custom-evals`

List custom-rubric evaluations for a dataset, newest first

### Path Parameters

- `dataset_id: string`

### Returns

- `items: array of CustomEval`

  - `custom_eval_id: string`

  - `dataset_id: string`

  - `status: string`

    pending | running | succeeded | failed

  - `judge_prompt: string`

  - `name: string`

  - `score_before: number`

    Mean rubric score on the original pairs. Null until the eval succeeds.

  - `score_after: number`

    Mean rubric score on the adapted pairs. Null until the eval succeeds.

  - `improvement_percent: number`

    Percentage change from score_before to score_after. Null until the eval succeeds.

  - `scored_rows: number`

    Rows the judge scored. A rubric score is an average over this sample, not the whole dataset.

  - `error_message: string`

  - `created_at: string`

  - `completed_at: string`

### Example

```http
curl https://api.prod.adaptionlabs.ai/api/v1/datasets/$DATASET_ID/custom-evals \
    -H "Authorization: Bearer $ADAPTION_API_KEY"
```

#### Response

```json
{
  "items": [
    {
      "custom_eval_id": "custom_eval_id",
      "dataset_id": "dataset_id",
      "status": "running",
      "judge_prompt": "judge_prompt",
      "name": "name",
      "score_before": 0,
      "score_after": 0,
      "improvement_percent": 0,
      "scored_rows": 0,
      "error_message": "error_message",
      "created_at": "2019-12-27T18:11:19.117Z",
      "completed_at": "2019-12-27T18:11:19.117Z"
    }
  ]
}
```

## Get a single custom-rubric evaluation

**get** `/api/v1/datasets/{dataset_id}/custom-evals/{custom_eval_id}`

Get a single custom-rubric evaluation

### Path Parameters

- `dataset_id: string`

- `custom_eval_id: string`

### Returns

- `CustomEval object { custom_eval_id, dataset_id, status, 9 more }`

  - `custom_eval_id: string`

  - `dataset_id: string`

  - `status: string`

    pending | running | succeeded | failed

  - `judge_prompt: string`

  - `name: string`

  - `score_before: number`

    Mean rubric score on the original pairs. Null until the eval succeeds.

  - `score_after: number`

    Mean rubric score on the adapted pairs. Null until the eval succeeds.

  - `improvement_percent: number`

    Percentage change from score_before to score_after. Null until the eval succeeds.

  - `scored_rows: number`

    Rows the judge scored. A rubric score is an average over this sample, not the whole dataset.

  - `error_message: string`

  - `created_at: string`

  - `completed_at: string`

### Example

```http
curl https://api.prod.adaptionlabs.ai/api/v1/datasets/$DATASET_ID/custom-evals/$CUSTOM_EVAL_ID \
    -H "Authorization: Bearer $ADAPTION_API_KEY"
```

#### Response

```json
{
  "custom_eval_id": "custom_eval_id",
  "dataset_id": "dataset_id",
  "status": "running",
  "judge_prompt": "judge_prompt",
  "name": "name",
  "score_before": 0,
  "score_after": 0,
  "improvement_percent": 0,
  "scored_rows": 0,
  "error_message": "error_message",
  "created_at": "2019-12-27T18:11:19.117Z",
  "completed_at": "2019-12-27T18:11:19.117Z"
}
```

## Draft a judge rubric tailored to an adapted dataset

**post** `/api/v1/datasets/{dataset_id}/custom-evals/prepare`

Takes no body: samples the dataset's own before/after pairs and rewrites the stock rubric around them, sizing the sample to what the dataset yields. Nothing is persisted and no judge tokens are spent — the returned `judge_prompt` is a suggestion to review and then POST to `/datasets/{dataset_id}/custom-evals`. Blocks for the full model round-trip; set a client timeout of several minutes.

### Path Parameters

- `dataset_id: string`

### Returns

- `judge_prompt: string`

  A ready-to-edit rubric, accepted as-is by `judge_prompt` on POST /datasets/{dataset_id}/custom-evals.

- `adapted: boolean`

  False when the stock rubric came back untouched — the sampled values did not fill both halves of the before/after contrast, or the adaptation model was unreachable. The prompt is still valid, just not tailored to this dataset.

- `sampled_rows: number`

  How many of the dataset's own examples the rubric was written against — half original, half adapted. Lower than the usual target on a thin dataset; a dataset too thin to write from at all is rejected with a 400.

### Example

```http
curl https://api.prod.adaptionlabs.ai/api/v1/datasets/$DATASET_ID/custom-evals/prepare \
    -X POST \
    -H "Authorization: Bearer $ADAPTION_API_KEY"
```

#### Response

```json
{
  "judge_prompt": "judge_prompt",
  "adapted": true,
  "sampled_rows": 0
}
```

## Domain Types

### Custom Eval

- `CustomEval object { custom_eval_id, dataset_id, status, 9 more }`

  - `custom_eval_id: string`

  - `dataset_id: string`

  - `status: string`

    pending | running | succeeded | failed

  - `judge_prompt: string`

  - `name: string`

  - `score_before: number`

    Mean rubric score on the original pairs. Null until the eval succeeds.

  - `score_after: number`

    Mean rubric score on the adapted pairs. Null until the eval succeeds.

  - `improvement_percent: number`

    Percentage change from score_before to score_after. Null until the eval succeeds.

  - `scored_rows: number`

    Rows the judge scored. A rubric score is an average over this sample, not the whole dataset.

  - `error_message: string`

  - `created_at: string`

  - `completed_at: string`

### Custom Eval List Response

- `CustomEvalListResponse object { items }`

  - `items: array of CustomEval`

    - `custom_eval_id: string`

    - `dataset_id: string`

    - `status: string`

      pending | running | succeeded | failed

    - `judge_prompt: string`

    - `name: string`

    - `score_before: number`

      Mean rubric score on the original pairs. Null until the eval succeeds.

    - `score_after: number`

      Mean rubric score on the adapted pairs. Null until the eval succeeds.

    - `improvement_percent: number`

      Percentage change from score_before to score_after. Null until the eval succeeds.

    - `scored_rows: number`

      Rows the judge scored. A rubric score is an average over this sample, not the whole dataset.

    - `error_message: string`

    - `created_at: string`

    - `completed_at: string`

### Custom Eval Prepare Response

- `CustomEvalPrepareResponse object { judge_prompt, adapted, sampled_rows }`

  - `judge_prompt: string`

    A ready-to-edit rubric, accepted as-is by `judge_prompt` on POST /datasets/{dataset_id}/custom-evals.

  - `adapted: boolean`

    False when the stock rubric came back untouched — the sampled values did not fill both halves of the before/after contrast, or the adaptation model was unreachable. The prompt is still valid, just not tailored to this dataset.

  - `sampled_rows: number`

    How many of the dataset's own examples the rubric was written against — half original, half adapted. Lower than the usual target on a thin dataset; a dataset too thin to write from at all is rejected with a 400.
