Plato's evaluator subsystem adds structured benchmark outputs on top of the trainer's normal scalar test metric.
A trainer's TestingStrategy still returns a single value such as accuracy or perplexity. When [evaluation] is configured, Plato then runs an evaluator and stores richer benchmark results in the trainer context for server-side logging.
Runtime flow
The evaluation path is:
TestingStrategy.test_model(...) computes the trainer's scalar test metric.
run_configured_evaluation(...) treats evaluator failures as non-fatal by default.
If evaluation.fail_on_error = false or omitted, Plato logs the exception and continues without structured evaluation metrics.
If evaluation.fail_on_error = true, the exception is raised and the run stops.
These settings govern exceptions raised while evaluating. Unknown or retired
evaluator selections fail during resolution so a misspelled or unsupported
backend cannot silently produce an unevaluated run.
Lighteval-specific behaviour
Plato's Lighteval adapter adds several integration details on top of upstream Lighteval:
task aliases are mapped to the concrete upstream ids:
arc_easy → arc:easy
arc_challenge → arc:challenge
piqa → custom piqa_hf
summary metrics are normalized into stable Plato names:
ifeval_avg
hellaswag
arc_easy
arc_challenge
arc_avg
piqa
detailed per-task metrics are also exported to the CSV, for example:
evaluation_ifeval_prompt_level_strict_acc
evaluation_arc_easy_acc
evaluation_arc_challenge_acc_stderr
safe runtime defaults are used for server-side evaluation:
batch_size = 1
model_parallel = false
device = Config.device()
dtype inferred from trainer.bf16 / trainer.fp16
when trainer.max_concurrency spawns a subprocess for testing, Plato persists evaluator state so the parent server process still logs the structured metrics.
The adapter also exposes the preset smollm_round_fast, used by the SmolLM2 server-side evaluation example.
Task scoring follows upstream Lighteval metric definitions. Plato maps those
outputs to stable summary names and detailed CSV columns; normalization does
not change the upstream scoring rules.
Current-model exports and configured local fallbacks receive fresh response
caches per evaluation. Local fallback directories are independent full copies,
with cleanup on success and failure; source artifacts remain untouched. See
model artifacts and response caches
for storage costs and concurrent-writer limits.
Optional runtime qualification
Provision the locked extra, test dependencies, and the exact required NLTK
resources before running the complete optional suite:
The strict profile requires its complete reviewed case inventory and successful
execution, with no skips, xfails, or xpasses. Missing dependencies or resources
fail during preflight; preflight does not download models or resources.
--collect-only checks prerequisites and inventory, not runtime behavior.
Filters such as -k, -m, and collection exclusions are incompatible with
complete qualification. Use an explicit filesystem path without the profile
for a focused run:
Focused runs retain prerequisite and strict outcome checks but do not qualify
the full inventory. Native MLX and Lighteval optional tests run separately;
optional --pyargs selection and cross-profile targets are unsupported.
For combined core and Lighteval qualification, the retained model-search tests
also need the test-model-search group, which includes test and ptflops:
Ordinary core, base, mandatory, and mlx-native requests exclude tests/llm_eval before importing its modules.
Existing unit tests in tests/evaluators remain core tests and do not establish
actual Lighteval runtime qualification.
The real runtime regression fixtures use a tiny local GPT-2 model, local parquet
tasks, and controlled one-token generation on CPU with Python 3.13. They exercise
the actual registry, runner, Pipeline, metrics, cache freshness, and cleanup.
Their scores demonstrate regression behavior, not public benchmark accuracy,
performance, downloaded Hub-model quality, GPU, or distributed qualification.
Server logging contract
Evaluator metrics appear in the runtime CSV under the evaluation_ prefix.
Every key in EvaluationResult.metrics becomes evaluation_<metric>.
Lighteval additionally exports detailed task-level metrics from the evaluator metadata.
The CSV schema expands automatically when new evaluation columns appear.
See Evaluation for configuration details and Results for logging behaviour.
Extending Plato with new evaluators
A good custom evaluator should:
Accept a lightweight config object.
Use request.context.state for temporary coordination instead of mutating the trainer directly.
Return a normalized EvaluationResult with small summary metrics in metrics.
Store heavier nested details inside metadata if they are still useful for downstream inspection.
Choose a stable primary_metric so CSV dashboards and automated comparisons have a clear headline number.
For larger integrations, pair the evaluator with a dedicated documentation example under docs/docs/examples/ and a smoke test under tests/evaluators/.