Skip to main content

Evaluation & Inference

The evaluation system in Nolano.AI provides comprehensive tools for assessing model performance across different modalities and tasks. The platform supports both built-in evaluation metrics and custom evaluation functions.

EvaluationConfig

Configuration class for model evaluation and inference settings.
str
required
Path to the trained model checkpoint directory (e.g., /path/to/checkpoint/global_step_XXXXX)
DataConfig
required
Configuration for evaluation data. Similar to training data config but typically with validation_split=1.0
str, List[str], or callable
default:"Auto-selected based on training objective"
Evaluation metrics to compute:
  • Text/code models: "perplexity", "accuracy", "bleu", "rouge"
  • Time series: "mse", "mae", "mape", "smape", "quantile_loss"
  • Custom callable functions with signature: (predictions, targets) → metric_value
int
default:"32"
Batch size for evaluation.
bool
default:"False"
Whether to save predictions to file.
str or None
default:"model_path + '/evaluation'"
Directory to save evaluation results and predictions.
int or None
default:"None"
Maximum number of evaluation steps. Set to None for full dataset evaluation.

Built-in Evaluation Metrics

  • Perplexity: Measures how well the model predicts the next token
  • Accuracy: Token-level or sequence-level accuracy
  • BLEU: Bilingual Evaluation Understudy score for text generation quality
  • ROUGE: Recall-Oriented Understudy for Gisting Evaluation
  • CodeBLEU: Specialized BLEU variant for code generation

Evaluation Examples

Running Evaluation

Inference

InferenceConfig

Configuration class for model inference settings.
int
default:"1"
Batch size for inference.
int
default:"512"
Maximum number of new tokens to generate (for generative models).
float
default:"1.0"
Sampling temperature for text generation. Higher values increase randomness.
float
default:"1.0"
Nucleus sampling parameter. Only consider tokens with cumulative probability up to this value.
int or None
default:"None"
Only consider the k most likely tokens at each step.
bool
default:"True"
Whether to use sampling for generation. If False, uses greedy decoding.
float
default:"1.0"
Penalty for token repetition. Values > 1.0 discourage repetition.
float
default:"1.0"
Penalty for sequence length. Values > 1.0 encourage longer sequences.
str
default:"auto"
Device for inference (‘cuda’, ‘cpu’, ‘auto’).

Inference Examples

Evaluation Output

Evaluation results are saved in JSON format with the following structure:
This comprehensive evaluation system enables thorough assessment of model performance and supports iterative improvement of your foundation models.