Skip to content

Qwen3 Federated LoRA

This reference adapts Qwen/Qwen3-0.6B-Base through Plato’s existing Hugging Face trainer, LoRA datasource, and fedavg_lora algorithm. Clients exchange only PEFT adapter weights; the pretrained base weights remain frozen. The small local text fixture exercises training and held-out evaluation. Its metrics are an integration check, not a benchmark of model quality or training convergence.

Model and data provenance

The model and tokenizer are both pinned to da87bfb608c14b7cf20ba1ce41287e8de496c0cd. The official checkpoint is a 596,049,920-parameter, text-only base causal language model using native Qwen3ForCausalLM support in Transformers. It needs no remote custom code or chat template. This is a small 2025 Qwen3 reference, not the newest Qwen release. The model’s Apache-2.0 license is separate from the local data’s provenance.

tests/fixtures/qwen3/train.json contains five original prose records, and validation.json contains two disjoint records. Their origin and license are recorded in tests/fixtures/qwen3/PROVENANCE.txt. This original fixture text is distributed under Plato’s Apache-2.0 repository license. No upstream training corpus or benchmark dataset is bundled with this scenario.

Run the reference

Use Python 3.13, the default full qualification and CI target. From the repository root:

uv sync --python 3.13
uv run python plato.py --config configs/HuggingFace/fedavg_qwen3_06b_lora.toml --cpu

The first run downloads the pinned model and tokenizer from Hugging Face. Subsequent runs reuse the cache. A missing revision or unavailable artifact is an error; the loader does not replace the pinned checkpoint with random weights.

The configuration uses two clients, one round, one local epoch, and max_concurrency = 1 to serialize client training. --cpu selects CPU execution; model_dtype = "float32" selects FP32. Batch size is 1, maximum token length is 64, and gradient accumulation is 1. LoRA targets q_proj and v_proj with rank 4, alpha 8, and zero dropout; AdamW uses a learning rate of 0.0001. trainer.random_seed = 17 seeds adapter initialization; local training resets Torch’s global RNG to 17 + client_id. The IID sampler uses a separate generator seeded by data.random_seed = 17. Server selection, data partitioning, and data shuffling also use configured seeds of 17.

The stored BF16 weights occupy roughly 1.2 GB, and FP32 parameters alone require roughly 2.4 GB. Peak process memory is higher because model loading, server residency, activations, and adapter training need additional memory. These parameter-size estimates are not a measured peak-memory guarantee.

Server testing evaluates held-out causal-language-model loss and perplexity through the normal Hugging Face testing strategy. This reference omits the optional [evaluation] block. For structured benchmark tasks, see the separate Lighteval guide.

Validation scope

The offline test tests/integration/test_qwen3_lora.py uses a tiny locally initialized Qwen3 model and local tokenizer. It checks architecture/runtime behavior without downloading the official pretrained checkpoint. Passing that test does not qualify the 0.6B pretrained weights.

The explicit pretrained qualification command is:

uv run python examples/huggingface/qualify_qwen3.py --output /tmp/qwen3-proof

It is separate from offline CI and uses the official pinned artifacts. The qualification trains two logical clients serially with unequal sample counts (2 and 3), then uses the production server aggregation path. It shares one frozen base model in one process; it does not launch the socket-based client processes of plato.py. Its peak RSS therefore measures that qualification process, not a full deployment. Keep its receipt with the run: revision, dependency versions, device and dtype, seeds, trainable parameters, adapter changes, held-out metrics, elapsed time, and peak RSS define what was actually exercised. A short qualification run does not establish downstream quality, convergence, or GPU performance.

The October 2026 pretrained qualification passed on macOS ARM64 with Python 3.13.16, Torch 2.14.1, Transformers 5.18.0, PEFT 0.21.2, and Datasets 5.0.1, using four CPU threads. It performed five optimizer updates with 573,440 trainable adapter parameters, verified every tensor against weighted aggregation, preserved the frozen base hash, and restored matching evaluation logits. The final seeded run took 14.04 seconds and peaked at 4,259,299,328 bytes (about 3.97 GiB) RSS in the shared-base qualification process. Those figures exclude initial artifact acquisition and do not measure the full plato.py client-process deployment. The tiny held-out fixture produced finite perplexity; its value is not a model-quality result. See the corrected implementation’s qualification receipt for the recorded execution evidence.

The normal plato.py launcher was checked at implementation revision d3569ec with the reference configuration, --cpu, cached official weights, and isolated runtime and port overrides. This run exercised two configured clients with serial training, aggregation, held-out evaluation, and checkpoint saving, with socket coordination and actual adapter payload transfers (comm_simulation = false). It stopped naturally at the one-round budget with exit code 0 and no surviving processes after a 10-second descendant shutdown grace period; the watchdog did not intervene.

This separate launcher run took 48.71 seconds. Sampling every 0.5 seconds observed a maximum combined RSS of 11,497,668,608 bytes (about 10.71 GiB) across a process group containing at most four processes. This sampled aggregate measurement has a different scope from the shared-base qualifier and is not a hardware memory guarantee. At shutdown, Python’s resource_tracker emitted one warning about a semaphore left for cleanup. The run completed and left no surviving processes; the warning’s cause has not been established. The original launcher receipt records the command, measurements, and log for that revision.

Adapter checkpoints

For this reference, the Hugging Face trainer exports adapters through PEFT save_pretrained(save_embedding_layers=False) and stores the adapter tensors in Plato’s single-file .safetensors checkpoint. The accompanying <checkpoint>.hf JSON metadata records the exact base-model and tokenizer identities and native PEFT adapter configuration. A .pkl sidecar stores Plato run history. Keep the checkpoint and its sidecars together and retain access to the pinned base checkpoint.

Reload uses set_peft_model_state_dict with the existing PEFT wrapper. Adapter checkpoints contain no replacement copy of the base weights and do not imply a full optimizer-state resume. When transferring adapters to another application, use the recorded base revision and PEFT configuration; do not load them against an arbitrary later model revision.

Customization

Set trainer.model_revision and trainer.tokenizer_revision when changing the model. When both repositories are the same, an omitted tokenizer revision inherits the model revision. A different tokenizer repository needs its own explicit revision for reproducibility.

For local text data, keep data.dataset_name = "json" and change the train and validation paths under [data.data_files]. Run from the repository root for the reference paths. Keep the splits disjoint, and record the source and license of replacement data. For a Hub dataset, data.dataset_revision can pin its revision.

See Federated LoRA Fine-Tuning for the general adapter configuration.