Skip to content
View as Markdown

Merge into a HuggingFace Model

Merging folds a LoRA adapter back into its base model, producing a standalone HuggingFace model directory with no LoRA dependency. Use it when you want a single artifact that any HuggingFace-compatible framework can load.

During LoRA training only the low-rank matrices change; the base weights stay frozen. Merging adds the adapter deltas back into the base weights, W_merged = W_base + (B @ A) * (alpha / rank), so the result is an ordinary model:

Tinker checkpoint          Merged HuggingFace model
+-------------------+      +---------------------------+
| adapter weights   |  -->  | model shards (.safetensors)|
| adapter config    |  -->  | config.json               |
+-------------------+      | tokenizer files ...        |
      + base model         +---------------------------+
      (from HF Hub)

Merge

Download the checkpoint (see Deployment Basics), then call build_hf_model with the same base model the checkpoint was trained from:

from tinker_cookbook import weights

adapter_dir = weights.download(
    tinker_path="tinker://<run_id>/sampler_weights/final",
    output_dir="./adapter",
)

weights.build_hf_model(
    base_model="Qwen/Qwen3.5-4B",
    adapter_path=adapter_dir,
    output_path="./merged_model",
)

build_hf_model downloads the base model from HuggingFace Hub, applies the deltas shard by shard, and writes the result. output_path must not already exist.

Options worth knowing:

  • merge_strategy — "auto" (default) merges shard by shard, so peak memory is about one shard rather than the whole model. "full" loads the entire model; use it with dtype when you need a specific output dtype.
  • trust_remote_code — some architectures need it to load. Defaults to the HF_TRUST_REMOTE_CODE environment variable, then False.
  • quantize="experts-fp8" with serving_format="vllm" — quantize routed MoE expert weights to FP8 for vLLM.

See build_hf_model for the full parameter list.

What you get

The output directory is a standard HuggingFace model: config.json, tokenizer files, and safetensors shards with an index.

chat_template.jinja
config.json
model-00001-of-00002.safetensors
model-00002-of-00002.safetensors
model.safetensors.index.json
tokenizer.json
tokenizer_config.json
...

Load or serve it

With transformers:

from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("./merged_model")
model = AutoModelForCausalLM.from_pretrained("./merged_model", device_map="auto")

inputs = tokenizer("The capital of France is", return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=20)
print(tokenizer.decode(output[0], skip_special_tokens=True))

With vLLM:

vllm serve ./merged_model

Next steps