Skip to content
View as Markdown

Build a PEFT LoRA Adapter

Instead of merging, you can export the trained adapter on its own in PEFT format. This is the preferred path for vLLM and SGLang, where one base model stays loaded and lightweight adapters are attached to it, and it is the only practical option when you serve many fine-tunes of the same base.

Merged model PEFT adapter
Size Full model (GBs) Just the LoRA matrices (MBs)
Deployment Load like any HF model Load base model + attach adapter
Multi-adapter One model per adapter One base + many adapters
Use with Any framework vLLM --lora-modules, SGLang --lora-paths, PEFT

Convert

Download the checkpoint (see Deployment Basics), then call build_lora_adapter:

from tinker_cookbook import weights

adapter_dir = weights.download(
    tinker_path="tinker://<run_id>/sampler_weights/final",
    output_dir="./adapter",
)

weights.build_lora_adapter(
    base_model="Qwen/Qwen3.5-4B",
    adapter_path=adapter_dir,
    output_path="./peft_adapter",
)

build_lora_adapter remaps Tinker's internal adapter keys to the HuggingFace parameter names that serving frameworks expect. It only reads the base model's config, not its weights, so it is fast and needs little memory. output_path must not already exist.

It raises WeightsAdapterError for model families that vLLM and SGLang cannot serve with LoRA (currently DeepSeek V3.1). For those, merge into a full model instead. See build_lora_adapter for parameters.

What you get

Two files:

  • adapter_config.json — PEFT metadata: base model, rank r, lora_alpha, and target_modules
  • adapter_model.safetensors — the LoRA weight matrices

Load or serve it

With PEFT and transformers:

from peft import PeftModel
from transformers import AutoModelForCausalLM

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B")
model = PeftModel.from_pretrained(base, "./peft_adapter")

With vLLM (one base, many adapters):

vllm serve Qwen/Qwen3.5-4B \
    --enable-lora \
    --lora-modules my_adapter=./peft_adapter

With SGLang:

python -m sglang.launch_server \
    --model Qwen/Qwen3.5-4B \
    --lora-paths my_adapter=./peft_adapter

Requests then select the adapter by name (my_adapter) in the model field.

Next steps