EvaluationConfig
str
required
Path to the trained model checkpoint directory (e.g.,
/path/to/checkpoint/global_step_XXXXX)DataConfig
required
Configuration for evaluation data. Similar to training data config but typically with
validation_split=1.0str | List[str] | callable
default:"Auto-selected based on training objective"
Evaluation metrics to compute:
- For text/code models:
"perplexity","accuracy","bleu","rouge" - For time series:
"mse","mae","mape","smape","quantile_loss" - Custom callable functions with signature:
(predictions, targets) → metric_value
int
default:"32"
Batch size for evaluation.
bool
default:"False"
Whether to save predictions to file.
str | None
default:"model_path + '/evaluation'"
Directory to save evaluation results and predictions.
int | None
default:"None"
Maximum number of evaluation steps. Set to
None for full dataset evaluation.InferenceConfig
int
default:"1"
Batch size for inference.
int
default:"512"
Maximum number of new tokens to generate (for generative models).
float
default:"1.0"
Sampling temperature for text generation. Higher values increase randomness.
float
default:"1.0"
Nucleus sampling parameter. Only consider tokens with cumulative probability up to this value.
int | None
default:"None"
Only consider the k most likely tokens at each step.
bool
default:"True"
Whether to use sampling for generation. If False, uses greedy decoding.
float
default:"1.0"
Penalty for token repetition. Values > 1.0 discourage repetition.
float
default:"1.0"
Penalty for sequence length. Values > 1.0 encourage longer sequences.
str
default:"auto"
Device for inference (‘cuda’, ‘cpu’, ‘auto’).

