Skip to main content

Finetuning Configuration

Finetuning allows you to adapt pre-trained models to specific tasks or domains with minimal computational overhead. The finetuning process leverages existing model knowledge while updating parameters to optimize for your specific use case.

FinetuningConfig

The main configuration class for finetuning experiments, building on the ExperimentConfig structure.
str
required
Path to the pre-trained model checkpoint or Hugging Face model identifier to finetune from
DataConfig or List[DataConfig]
required
Data configuration(s) for finetuning tasks. Supports multi-task finetuning scenarios.
FinetuningOptimizationConfig
required
Finetuning-specific optimization parameters with typically lower learning rates
ModelConfig
required
Model architecture configuration. Must match the base model architecture.
MetaConfig
default:"Uses MetaConfig defaults"
Metadata and run-specific parameters for the finetuning experiment

MetaConfig

Configuration for experiment metadata and checkpointing behavior.
str
default:"trial-run"
Name identifier for this finetuning experimental run.
int
default:"42"
Random seed for reproducible finetuning.
str
default:"current working directory / run_name"
Directory path for saving finetuned model checkpoints.
int
default:"-1"
Frequency (in steps) for saving model checkpoints. Set to -1 to save only at the end of finetuning.
int
default:"-1"
Maximum number of model checkpoints to retain. Set to -1 for no limit.
WandbConfig or None
default:"None"
Weights & Biases logging configuration for finetuning experiment tracking and visualization

WandbConfig

Configuration for Weights & Biases experiment tracking and logging during finetuning.
str
required
Weights & Biases project name for organizing finetuning experiments
str or None
default:"None"
Weights & Biases team/organization name. If None, uses the default entity associated with your API key.
str or None
default:"None"
Custom run name for the finetuning experiment. If None, uses the MetaConfig name or auto-generates one.
List[str] or None
default:"None"
List of tags to associate with the finetuning run for easy filtering and organization
str or None
default:"None"
Optional notes or description for the finetuning experiment run
bool
default:"True"
Whether to log the finetuned model as a Weights & Biases artifact for version control. Defaults to True for finetuning.
int
default:"50"
Frequency (in steps) for logging metrics to Weights & Biases. Lower default for finetuning due to fewer total steps.
bool
default:"False"
Whether to log gradient histograms (can impact performance)
bool
default:"False"
Whether to log parameter histograms (can impact performance)
str or None
default:"None"
Model watching mode for logging gradients and parameters:
  • "gradients": Log gradient histograms
  • "parameters": Log parameter histograms
  • "all": Log both gradients and parameters
  • None: Disable model watching
dict or None
default:"None"
Additional configuration dictionary to log to Weights & Biases
bool
default:"True"
Whether to log LoRA adapter weights as artifacts when using LoRA finetuning

FinetuningOptimizationConfig

Specialized optimization configuration for finetuning with recommended parameter ranges.
int
required
Total number of finetuning steps. Typically much lower than full training (500-5000 steps).
float
required
Maximum learning rate for finetuning. Recommended range: 1e-5 to 5e-4 (lower than full training).
int
required
Global batch size for finetuning. Can be smaller than full training due to fewer steps.
str or callable
default:"linear"
Learning rate scheduling strategy for finetuning:
  • "linear": Linear decay (recommended for finetuning)
  • "cosine": Cosine annealing
  • "constant": Constant learning rate
  • Custom function with signature: (learning_rate, current_step, total_steps) → decayed_rate
int
default:"50"
Number of learning rate warmup steps. Typically 5-10% of total finetuning steps.
float
default:"0.01"
L2 regularization coefficient. Important for preventing overfitting in finetuning.
int
default:"1"
Number of steps to accumulate gradients before updating. Useful for effective larger batch sizes.
List[str] or None
default:"None"
List of layer patterns to freeze during finetuning. Example: ["embeddings", "layer.0", "layer.1"]
LoRAConfig or None
default:"None"
Low-Rank Adaptation configuration for parameter-efficient finetuning
str
default:"AdamW"
Optimizer algorithm. Options: "AdamW", "Adam", "SGD"
float
default:"1.0"
Gradient clipping threshold. Important for stability in finetuning.

LoRAConfig

Configuration for Low-Rank Adaptation (LoRA) parameter-efficient finetuning.
int
default:"16"
Rank of the adaptation matrices. Higher rank = more parameters but better expressiveness.
float
default:"32"
LoRA scaling parameter. Controls the magnitude of the adaptation.
float
default:"0.1"
Dropout probability for LoRA layers.
List[str] or None
default:"Auto-detected"
List of module names to apply LoRA to. If None, automatically targets attention and MLP layers.
str
default:"none"
Bias handling strategy:
  • "none": No bias adaptation
  • "all": Adapt all biases
  • "lora_only": Only adapt LoRA biases

FinetuningDataConfig

Extended data configuration with finetuning-specific options.
str or List[str]
required
Path(s) to finetuning data files. Should be formatted according to your task type.
str
default:"text_generation"
Type of finetuning task:
  • "text_generation": Generative language modeling
  • "classification": Text classification
  • "instruction_following": Instruction-tuning
  • "code_generation": Code completion/generation
  • "time_series_forecasting": Time series tasks
int
default:"512"
Maximum sequence length for finetuning examples. Shorter than training can speed up finetuning.
float
default:"0.1"
Portion of data reserved for validation during finetuning.
dict or None
default:"None"
Task-specific preprocessing options:
  • For instruction tuning: {"format": "alpaca", "prompt_template": "..."}
  • For classification: {"label_column": "label", "text_column": "text"}

Example Configurations

Advanced Features

Multi-Task Finetuning

Finetune on multiple related tasks simultaneously for better generalization:

Curriculum Learning

Gradually increase task complexity during finetuning:
Curriculum learning support is planned for future releases.

convert_finetuned_to_hf()

Convert finetuned models to Hugging Face format, preserving both base model and adaptations.
str
required
Path to the finetuned checkpoint directory
str
required
Path to the finetuning configuration YAML file
str
required
Destination directory for the converted Hugging Face model
bool
default:"False"
Whether to merge LoRA weights into the base model. If False, saves LoRA adapters separately.
bool
default:"False"
Whether to directly upload the converted model to Hugging Face Hub