This installs the optional runtime and the exact NLTK tokenizer resources used
by its registry. Provision the resources before offline use. The ssl extra
requires a separate environment. Keep --extra llm_eval on every syncing
command so shared dependencies stay compatible; see
Installation.
The run trains SmolLM2 on HuggingFaceTB/smol-smoltalk and evaluates the
aggregated global model on the server after each round. The supplied config
selects evaluation.device = "cuda:0" and uses external model and dataset
artifacts. For CPU evaluation, copy the config, set that field to "cpu", and
add --cpu to the command for trainer placement. The local runtime regression
fixtures do not qualify this downloaded-model training workload or its GPU
performance. See runtime qualification
for the bounded checks and their limits.
Plato normalizes their outputs into the following summary metrics:
ifeval_avg
hellaswag
arc_easy
arc_challenge
arc_avg
piqa
Sampling behaviour
max_samples = 32 is a per-task cap, not a global cap.
That means the run evaluates up to 32 examples from each task, using a deterministic shuffled subset inside Lighteval. This is useful for fast round-by-round feedback, but it produces partial benchmark numbers rather than full leaderboard-style scores.
Result logging
The runtime CSV is the canonical result log. Summary columns can be declared up front:
Plato also appends detailed Lighteval metrics automatically when they appear, for example:
evaluation_ifeval_prompt_level_strict_acc
evaluation_ifeval_prompt_level_loose_acc
evaluation_ifeval_inst_level_loose_acc
evaluation_arc_easy_acc
evaluation_arc_challenge_acc_stderr
evaluation_hellaswag_em
evaluation_piqa_em
So you can start with the summary columns above and still get the detailed task metrics in the same CSV file.
Runtime notes
batch_size = 1 and model_parallel = false are conservative defaults intended for reliable server-side evaluation on multi-GPU hosts.
device = "cuda:0" keeps the evaluator on one explicit GPU instead of relying on broad multi-GPU auto-detection.
If you want evaluator failures to stop the run, add:
[evaluation]fail_on_error=true
Otherwise Plato logs the evaluator exception and continues the training run without structured benchmark outputs for that round.
The current server model and tokenizer are exported afresh for evaluation each
round. Configured local fallback directories instead require independent full
copies, which can use substantial disk space and time for large checkpoints.
Temporary artifacts and response caches are cleaned up on success and failure;
source directories remain unchanged. See
artifact and cache behavior
for details, including the requirement to keep source checkpoints stable while
copying.