# Build a PEFT LoRA Adapter Instead of merging, you can export the trained adapter on its own in **PEFT format**. This is the preferred path for vLLM and SGLang, where one base model stays loaded and lightweight adapters are attached to it, and it is the only practical option when you serve many fine-tunes of the same base. | | Merged model | PEFT adapter | | ----------------- | ---------------------- | -------------------------------------------------- | | **Size** | Full model (GBs) | Just the LoRA matrices (MBs) | | **Deployment** | Load like any HF model | Load base model + attach adapter | | **Multi-adapter** | One model per adapter | One base + many adapters | | **Use with** | Any framework | vLLM `--lora-modules`, SGLang `--lora-paths`, PEFT | ## Convert Download the checkpoint (see [Deployment Basics](https://tinker-docs.thinkingmachines.ai/cookbook/deployment/basics/#download-the-checkpoint)), then call `build_lora_adapter`: ```python from tinker_cookbook import weights adapter_dir = weights.download( tinker_path="tinker:///sampler_weights/final", output_dir="./adapter", ) weights.build_lora_adapter( base_model="Qwen/Qwen3.5-4B", adapter_path=adapter_dir, output_path="./peft_adapter", ) ``` `build_lora_adapter` remaps Tinker's internal adapter keys to the HuggingFace parameter names that serving frameworks expect. It only reads the base model's config, not its weights, so it is fast and needs little memory. `output_path` must not already exist. It raises `WeightsAdapterError` for model families that vLLM and SGLang cannot serve with LoRA (currently DeepSeek V3.1). For those, [merge into a full model](https://tinker-docs.thinkingmachines.ai/cookbook/deployment/export-hf/index.md) instead. See [`build_lora_adapter`](https://tinker-docs.thinkingmachines.ai/cookbook/api-reference/weights/build_lora_adapter/index.md) for parameters. ## What you get Two files: - `adapter_config.json` — PEFT metadata: base model, rank `r`, `lora_alpha`, and `target_modules` - `adapter_model.safetensors` — the LoRA weight matrices ## Load or serve it With PEFT and `transformers`: ```python from peft import PeftModel from transformers import AutoModelForCausalLM base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B") model = PeftModel.from_pretrained(base, "./peft_adapter") ``` With vLLM (one base, many adapters): ```bash vllm serve Qwen/Qwen3.5-4B \ --enable-lora \ --lora-modules my_adapter=./peft_adapter ``` With SGLang: ```bash python -m sglang.launch_server \ --model Qwen/Qwen3.5-4B \ --lora-paths my_adapter=./peft_adapter ``` Requests then select the adapter by name (`my_adapter`) in the `model` field. ## Next steps - [Merge into a HuggingFace Model](https://tinker-docs.thinkingmachines.ai/cookbook/deployment/export-hf/index.md) — produce a standalone model instead - [Publish to HuggingFace Hub](https://tinker-docs.thinkingmachines.ai/cookbook/deployment/publish-hub/index.md) — upload the adapter with a generated model card