Skip to content
SupportLogin
AutoScientist
Edit on GitHub

Hyperparameters

Understand the training recipe AutoScientist selects and iterates on automatically.

AutoScientist proposes a complete training recipe before training, then evaluates and adjusts the recipe across iterations. For most runs, keep the recommended values unchanged. AutoScientist is optimized to select these values together and improve them based on measured results.

training_method is a top-level create argument (instruction or alignment). The remaining fields below are keys inside hyperparams.

These hyperparameters apply to both SFT and DPO runs.

HyperparameterDefinitionValid Values
learning_ratePeak learning rate for the optimizer. Controls the step size during weight updates.Float in [1e-6, 5e-4]
n_epochsNumber of complete passes training makes over the dataset.Integer, 1–5
batch_sizeNumber of samples per training step. In SFT, must be "max". In DPO, a power-of-2 integer or "max"."max" or power-of-2 integer ≥ 1
lr_scheduler_typeHow the learning rate changes during training. A linear schedule decays steadily; a cosine schedule slows its decay near the end."linear" | "cosine"
min_lr_ratioFloor learning rate as a fraction of the peak. The learning rate will not decay below learning_rate × min_lr_ratio.Float, 0.0–1.0
scheduler_num_cyclesNumber or fraction of cosine decay cycles across training. Only affects cosine schedules.Float > 0.0 (default 0.5)
warmup_ratioFraction of total training steps spent gradually increasing the learning rate from zero. Smooths the transition into the target learning rate.Float, 0.0–1.0 (typically 0.03–0.1)
max_grad_normMaximum gradient norm for gradient clipping. Caps gradient magnitude to reduce exploding gradients and unstable updates. Set to 0.0 to disable.Float ≥ 0.0
weight_decayL2 regularization penalty. Discourages excessively large weights, which helps prevent overfitting and improves generalization to unseen data.Float ≥ 0.0
training_typeTraining strategy. "lora" inserts low-rank adapters; "full" updates all model weights. Full fine-tuning is only available for supported models."lora" | "full"

These hyperparameters configure Low-Rank Adaptation and are only applicable when training_type is "lora". In alignment (SFT → DPO) pipelines, LoRA settings are frozen from the winning SFT iteration and cannot be changed during DPO.

HyperparameterDefinitionValid Values
lora_rRank of the LoRA adapter matrices. Controls adapter capacity — higher ranks can represent more complex adaptations but require more memory and compute.Power-of-2 integer, 1–64
lora_alphaScaling factor for the LoRA adapter’s contribution to the base model. Must be 1× or 2× the value of lora_r.Power-of-2 integer (1× or 2× lora_r)
lora_dropoutDropout probability applied to LoRA layers during training. Adds regularization to reduce overfitting.Float, 0.0–1.0 (typically 0.0–0.1)
lora_trainable_modulesWhich model layers receive LoRA adapters. "all-linear" targets every linear layer; alternatively, specify a comma-separated list of module names."all-linear" or comma-separated list (e.g. "q_proj,v_proj,o_proj")
HyperparameterDefinitionValid Values
train_on_inputsWhether training loss is computed over prompt tokens in addition to completion tokens. Setting to false masks the prompt so the model only learns to generate completions.true | false

These hyperparameters apply only to preference-based (DPO) training. They are inert for SFT runs.

HyperparameterDefinitionValid Values
dpo_betaControls how far the model is allowed to drift from the reference model. Lower values (around 0.1) update more aggressively toward the preferred output; higher values (around 0.7) stay closer to the reference.Float, 0.05–0.9
dpo_normalize_logratios_by_lengthNormalizes log ratios by sample length during loss calculation. Automatically set to true when simpo_gamma is above 0. Changing this flag alters the effective scale of dpo_beta; treat a change as a major intervention.true | false
rpo_alphaWeight of an auxiliary negative log-likelihood (NLL) loss on preferred responses, added on top of the standard DPO loss (Regularized Preference Optimization). Counteracts likelihood displacement — the failure mode where chosen-response probability drifts downward even as the reward margin grows. Set to 0.0 to disable. Cannot be combined with simpo_gamma.Float ≥ 0.0 (typical 0.1–1.0)
simpo_gammaTarget reward margin from Simple Preference Optimization. Above 0.0, the objective becomes reference-free and length-normalized: the chosen response must beat the rejected response by at least this margin. Raise dpo_beta alongside it since SimPO typically calls for a substantially larger beta. Set to 0.0 to disable. Cannot be combined with rpo_alpha.Float ≥ 0.0 (typical 0.3–1.6)

The proposed fields are editable, but the default workflow is to review and accept them:

  1. Review the model, algorithm, and generated recipe.
  2. Keep the recommended values unless a known constraint requires an override.
  3. Inspect the final AutoScientist Config JSON before confirming.
  4. Let AutoScientist evaluate and refine the recipe across iterations.
  5. After training, compare win rate and diagnostics as described in Interpreting results.

API overrides are an advanced escape hatch for controlled experiments or hard requirements. The SDK accepts the training objective through top-level training_method, the base model through top-level model, and recipe overrides through hyperparams. training_type is a field inside hyperparams, not a top-level argument.

If an override is necessary, avoid unsupported key names: copy a valid hyperparams object from the current app configuration or construct it from the current create API schema, save it as hyperparams.json, and pass it unchanged:

import json
from pathlib import Path
from adaption import Adaption
client = Adaption()
dataset_id = "dataset_abc123"
hyperparams = json.loads(Path("hyperparams.json").read_text())
run = client.autoscientist.create(
dataset_id=dataset_id,
training_method="alignment",
hyperparams={
**hyperparams,
"training_type": "lora",
"dpo_beta": 0.1,
},
)
print(run.id, run.status)

Unset hyperparameter keys remain under platform control. To enforce a specific model, review the supported models, list the currently available IDs, and pass one through model:

response = client.autoscientist.list_models()
for model in response.models:
print(model.id)

Prefer leaving model and hyperparams unset so AutoScientist can choose and iterate on them automatically. See Running AutoScientist for the complete request and status workflow.