# Merge into a HuggingFace Model Merging folds a LoRA adapter back into its base model, producing a **standalone HuggingFace model directory** with no LoRA dependency. Use it when you want a single artifact that any HuggingFace-compatible framework can load. During LoRA training only the low-rank matrices change; the base weights stay frozen. Merging adds the adapter deltas back into the base weights, `W_merged = W_base + (B @ A) * (alpha / rank)`, so the result is an ordinary model: ```text Tinker checkpoint Merged HuggingFace model +-------------------+ +---------------------------+ | adapter weights | --> | model shards (.safetensors)| | adapter config | --> | config.json | +-------------------+ | tokenizer files ... | + base model +---------------------------+ (from HF Hub) ``` ## Merge Download the checkpoint (see [Deployment Basics](https://tinker-docs.thinkingmachines.ai/cookbook/deployment/basics/#download-the-checkpoint)), then call `build_hf_model` with the same base model the checkpoint was trained from: ```python from tinker_cookbook import weights adapter_dir = weights.download( tinker_path="tinker:///sampler_weights/final", output_dir="./adapter", ) weights.build_hf_model( base_model="Qwen/Qwen3.5-4B", adapter_path=adapter_dir, output_path="./merged_model", ) ``` `build_hf_model` downloads the base model from HuggingFace Hub, applies the deltas shard by shard, and writes the result. `output_path` must not already exist. Options worth knowing: - `merge_strategy` — `"auto"` (default) merges shard by shard, so peak memory is about one shard rather than the whole model. `"full"` loads the entire model; use it with `dtype` when you need a specific output dtype. - `trust_remote_code` — some architectures need it to load. Defaults to the `HF_TRUST_REMOTE_CODE` environment variable, then `False`. - `quantize="experts-fp8"` with `serving_format="vllm"` — quantize routed MoE expert weights to FP8 for vLLM. See [`build_hf_model`](https://tinker-docs.thinkingmachines.ai/cookbook/api-reference/weights/build_hf_model/index.md) for the full parameter list. ## What you get The output directory is a standard HuggingFace model: `config.json`, tokenizer files, and safetensors shards with an index. ```text chat_template.jinja config.json model-00001-of-00002.safetensors model-00002-of-00002.safetensors model.safetensors.index.json tokenizer.json tokenizer_config.json ... ``` ## Load or serve it With `transformers`: ```python from transformers import AutoModelForCausalLM, AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("./merged_model") model = AutoModelForCausalLM.from_pretrained("./merged_model", device_map="auto") inputs = tokenizer("The capital of France is", return_tensors="pt").to(model.device) output = model.generate(**inputs, max_new_tokens=20) print(tokenizer.decode(output[0], skip_special_tokens=True)) ``` With vLLM: ```bash vllm serve ./merged_model ``` ## Next steps - [Build a PEFT LoRA Adapter](https://tinker-docs.thinkingmachines.ai/cookbook/deployment/lora-adapter/index.md) — skip the merge and serve the adapter on top of a shared base model - [Publish to HuggingFace Hub](https://tinker-docs.thinkingmachines.ai/cookbook/deployment/publish-hub/index.md) — upload the merged model with a generated model card