Supported models
Base models available to AutoScientist and the per-model limits for context length, batch size, and LoRA rank.
Every model below can be passed as the model field when creating an AutoScientist run. The runtime model list is authoritative; use this page as a planning reference.
response = client.autoscientist.list_models()for model in response.models: print(model.id, model.api_model_size, model.methods, model.training_types)See the autoscientist.list_models API reference for the current response schema.
Available models
Section titled “Available models”| Model | model ID | Size | Methods | Supported training_type | Supports reasoning training | Min. rows |
|---|---|---|---|---|---|---|
| SFT / Alignment | ||||||
| Gemma 3 (4B Instruct) | google/gemma-3-4b-it | 4B | SFT, DPO | lora, full | false | 1,000 |
| Gemma 4 (31B Instruct) | google/gemma-4-31B-it | 31B | SFT, DPO | lora, full | true | 10,000 |
| gpt-oss (20B) | openai/gpt-oss-20b | 20B | SFT, DPO | lora | true | 1,000 |
| gpt-oss (120B) | openai/gpt-oss-120b | 120B | SFT, DPO | lora | true | 1,000 |
| Llama 3.3 (70B Instruct) | meta-llama/Llama-3.3-70B-Instruct-Reference | 70B | SFT, DPO | lora, full | false | 1,000 |
| Llama 4 Scout (17B-16E Instruct) | meta-llama/Llama-4-Scout-17B-16E-Instruct | 109B | SFT, DPO | lora | false | 1,000 |
| Meta Llama 3.2 (3B Instruct) | meta-llama/Llama-3.2-3B-Instruct | 3B | SFT, DPO | lora, full | false | 1,000 |
| Mistral (7B) Instruct v0.2 | mistralai/Mistral-7B-Instruct-v0.2 | 7B | SFT, DPO | lora | false | 1,000 |
| Mixtral 8x7B Instruct (v0.1) | mistralai/Mixtral-8x7B-Instruct-v0.1 | 46.7B | SFT, DPO | lora, full | false | 1,000 |
| NVIDIA Nemotron 3 Nano Omni 30B A3B Reasoning BF16 | nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 | 30B | SFT, DPO | lora | true | 10,000 |
| NVIDIA Nemotron 3 Super 120B A12B BF16 | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | 120B | SFT, DPO | lora | false | 10,000 |
| Qwen 3.5 (0.8B) | Qwen/Qwen3.5-0.8B | 0.8B | SFT, DPO | lora | false | 1,000 |
| Qwen 3.5 122B A10B | Qwen/Qwen3.5-122B-A10B | 122B | SFT, DPO | lora | true | 10,000 |
| Qwen 3.6 35B A3B | Qwen/Qwen3.6-35B-A3B | 35B | SFT, DPO | lora | true | 10,000 |
| SFT / Alignment (multimodal) | ||||||
| Gemma 3 (4B) VLM | google/gemma-3-4b-it-VLM | 4B | SFT, DPO | lora | false | 1,000 |
| Gemma 3 (27B) VLM | google/gemma-3-27b-it-VLM | 27B | SFT, DPO | lora | false | 1,000 |
| Gemma 4 (31B) VLM | google/gemma-4-31B-it-VLM | 31B | SFT, DPO | lora | false | 1,000 |
| Qwen 3.5 (4B) | Qwen/Qwen3.5-4B | 4B | SFT, DPO | lora | true | 1,000 |
| Qwen 3.5 (9B) | Qwen/Qwen3.5-9B | 9B | SFT, DPO | lora | false | 1,000 |
Min. rows is the minimum effective dataset size (base rows plus augmentation) required to select that model. DPO training requires at least 12,000 rows regardless of model; see Row count.
A full value for training_type means full fine-tuning is available in addition to lora.
Per-model training limits
Section titled “Per-model training limits”These limits are properties of the model and training hardware. AutoScientist accounts for them automatically; they matter primarily when you override its recommended hyperparameters.
model ID | Context (SFT) | Context (DPO) | Max batch (SFT) | Max batch (DPO) | Gradient accumulation | Maximum lora_r |
|---|---|---|---|---|---|---|
| SFT / Alignment | ||||||
google/gemma-3-4b-it | 131072 | 65536 | 8 | 8 | 1 | 64 |
google/gemma-4-31B-it | 49152 | 24576 | 4 | 4 | 2 | 64 |
openai/gpt-oss-20b | 131072 | 65536 | 1 | 1 | 8 | 64 |
openai/gpt-oss-120b | 65536 | 32768 | 2 | 2 | 8 | 64 |
meta-llama/Llama-3.3-70B-Instruct-Reference | 24576 | 12288 | 8 | 8 | 1 | 64 |
meta-llama/Llama-4-Scout-17B-16E-Instruct | 65536 | 12288 | 8 | 8 | 1 | 64 |
meta-llama/Llama-3.2-3B-Instruct | 131072 | 65536 | 8 | 8 | 1 | 64 |
mistralai/Mistral-7B-Instruct-v0.2 | 32768 | 32768 | 16 | 8 | 1 | 64 |
mistralai/Mixtral-8x7B-Instruct-v0.1 | 32768 | 16384 | 8 | 8 | 1 | 64 |
nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 | 65536 | 32768 | 8 | 8 | 1 | 64 |
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | 49152 | 24576 | 2 | 2 | 4 | 64 |
Qwen/Qwen3.5-0.8B | 131072 | 131072 | 8 | 8 | 1 | 64 |
Qwen/Qwen3.5-122B-A10B | 65536 | 32768 | 16 | 16 | 1 | 64 |
Qwen/Qwen3.6-35B-A3B | 65536 | 32768 | 8 | 8 | 1 | 64 |
| SFT / Alignment (multimodal) | ||||||
google/gemma-3-4b-it-VLM | 32768 | 32768 | 8 | 8 | 1 | 64 |
google/gemma-3-27b-it-VLM | 32768 | 24576 | 8 | 8 | 1 | 64 |
google/gemma-4-31B-it-VLM | 24576 | 12288 | 8 | 8 | 1 | 64 |
Qwen/Qwen3.5-4B | 131072 | 65536 | 8 | 8 | 1 | 64 |
Qwen/Qwen3.5-9B | 65536 | 49152 | 8 | 8 | 1 | 64 |
Batch size
Section titled “Batch size”Each model has one valid batch size. batch_size defaults to "max", which resolves to that model-specific value. Leaving it unchanged is the right choice for almost every run. Supplying another integer causes the launch to fail after the job is accepted.
Models use gradient accumulation when a larger effective batch is useful.
# Recommended: let AutoScientist resolve the model and batch size.run = client.autoscientist.create(dataset_id=dataset_id)
# Advanced: pin a model and its required batch size.run = client.autoscientist.create( dataset_id=dataset_id, model="openai/gpt-oss-20b", hyperparams={"batch_size": 1},)LoRA rank
Section titled “LoRA rank”lora_r accepts values from 1 through 64 for every model on this page. lora_alpha must be exactly one or two times lora_r; the API cross-checks the values when both are supplied in the same request.
Row count
Section titled “Row count”AutoScientist enforces the following minimum effective dataset sizes (base rows + augmentation):
| Training method | Minimum rows | Applies to |
|---|---|---|
| SFT (instruction) | 1,000 | All models |
| DPO (preference) | 12,000 | All models |
| Any method | 10,000 | Gemma (≥ 26B), Qwen (≥ 35B), NVIDIA |
Text datasets can use augmentation to reach a minimum. Multimodal datasets cannot currently be augmented, so they must already contain enough processed rows. AutoScientist treats 20,000 rows as a recommended quality target, not a hard launch requirement.
Context length and your data
Section titled “Context length and your data”Training context is the maximum sequence length used during fine-tuning. It is shorter than the serving context. Rows longer than the training limit are truncated rather than rejected.
DPO context can be shorter than SFT context. Before pinning a model for DPO, confirm that DPO appears in its runtime methods list and use the DPO context column above.