Build a PEFT LoRA Adapter
Instead of merging, you can export the trained adapter on its own in PEFT format. This is the preferred path for vLLM and SGLang, where one base model stays loaded and lightweight adapters are attached to it, and it is the only practical option when you serve many fine-tunes of the same base.
| Merged model | PEFT adapter | |
|---|---|---|
| Size | Full model (GBs) | Just the LoRA matrices (MBs) |
| Deployment | Load like any HF model | Load base model + attach adapter |
| Multi-adapter | One model per adapter | One base + many adapters |
| Use with | Any framework | vLLM --lora-modules, SGLang --lora-paths, PEFT |
Convert
Download the checkpoint (see Deployment Basics), then call build_lora_adapter:
from tinker_cookbook import weights
adapter_dir = weights.download(
tinker_path="tinker://<run_id>/sampler_weights/final",
output_dir="./adapter",
)
weights.build_lora_adapter(
base_model="Qwen/Qwen3.5-4B",
adapter_path=adapter_dir,
output_path="./peft_adapter",
)
build_lora_adapter remaps Tinker's internal adapter keys to the HuggingFace parameter names that serving frameworks expect. It only reads the base model's config, not its weights, so it is fast and needs little memory. output_path must not already exist.
It raises WeightsAdapterError for model families that vLLM and SGLang cannot serve with LoRA (currently DeepSeek V3.1). For those, merge into a full model instead. See build_lora_adapter for parameters.
What you get
Two files:
adapter_config.json— PEFT metadata: base model, rankr,lora_alpha, andtarget_modulesadapter_model.safetensors— the LoRA weight matrices
Load or serve it
With PEFT and transformers:
from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B")
model = PeftModel.from_pretrained(base, "./peft_adapter")
With vLLM (one base, many adapters):
With SGLang:
Requests then select the adapter by name (my_adapter) in the model field.
Next steps
- Merge into a HuggingFace Model — produce a standalone model instead
- Publish to HuggingFace Hub — upload the adapter with a generated model card