Manage Checkpoints
TrainingClients can save checkpoints for resuming training later, or serving the checkpoint for sampling.
| Method | Use it to | Path |
|---|---|---|
save_state |
Resume training | tinker://<run_id>/weights/<name> |
save_weights_for_sampler |
Sample, open in the Playground, or download | tinker://<run_id>/sampler_weights/<name> |
Training checkpoints
Save the full training state, including Adam's moment estimates, with save_state:
checkpoint = training_client.save_state("step-100", user_metadata={ "step": "100", "recipe": "example" }).result()
print(checkpoint.path) # tinker://<run_id>/weights/step-100
# load without optimizer state
training_client = service_client.create_training_client_from_state(checkpoint.path).result()
# load with optimizer state
training_client = service_client.create_training_client_from_state_with_optimizer(checkpoint.path).result()
# view checkpoint in Tinker console
print(checkpoint.get_console_url())
Sampler checkpoints
save_weights_for_sampler saves only the weights and not gradient or optimizer state for sampling in one of the following ways:
- Load using
SamplingClient - Use with
OpenAIorAnthropiccompatible APIs - Chat with in the Tinker Playground
sampler_checkpoint = training_client.save_weights_for_sampler("step-100", user_metadata={ "step": "100", "recipe": "example" }).result()
print(sampler_checkpoint.path) # tinker://<run_id>/sampler_weights/step-100
# load the checkpoint into a sampling client
sampling_client = service_client.create_sampling_client(model_path=sampler_checkpoint.path)
# chat with checkpoint in Tinker Playground
print(sampler_checkpoint.get_playground_url())
# manage checkpoint in Tinker console
print(sampler_checkpoint.get_console_url())
Expiration
Checkpoint storage is charged at $0.10 per GB per month. To ensure checkpoints are not kept forever, you can set a TTL (time to live) when saving a checkpoint.
training_client.save_state("step-100", ttl_seconds=24 * 3600)
training_client.save_weights_for_sampler("step-100", ttl_seconds=24 * 3600)
Without a TTL, a checkpoint is kept until you delete it. To change the TTL later, see Set a TTL below.
Manage with RestClient
Use the RestClient to manage checkpoints after they are saved.
List checkpoints
list_checkpoints lists the checkpoints of a single training run. list_user_checkpoints lists them across all of your runs, newest first:
response = rest_client.list_checkpoints("<run_id>").result()
for checkpoint in response.checkpoints:
print(checkpoint.checkpoint_type, checkpoint.tinker_path)
recent = rest_client.list_user_checkpoints(limit=10).result()
Set a TTL
Change or remove a checkpoint's TTL after it is saved:
# Expire in 7 days
rest_client.set_checkpoint_ttl_from_tinker_path(path, ttl_seconds=7 * 24 * 3600).result()
# Keep indefinitely
rest_client.set_checkpoint_ttl_from_tinker_path(path, ttl_seconds=None).result()
Publish a checkpoint
Publishing a checkpoint lets other Tinker users load it from its path. Only the owner of the training run can publish or unpublish:
rest_client.publish_checkpoint_from_tinker_path(path).result()
rest_client.unpublish_checkpoint_from_tinker_path(path).result()
Download sampler weights
Sampler checkpoints can be downloaded to local disk as a tar file:
import tarfile
import urllib.request
# create a signed url for the checkpoint
archive = rest_client.get_checkpoint_archive_url_from_tinker_path(sampler_checkpoint.path).result()
# download the checkpoint
urllib.request.urlretrieve(archive.url, "checkpoint.tar")
# extract the checkpoint
with tarfile.open("checkpoint.tar") as tar:
tar.extractall("./adapter", filter="data")
See tinker checkpoint download to download from the CLI.
Weights format
The archive holds a LoRA adapter, not a full model. It contains:
adapter_model.safetensors: the trained LoRAAandBmatrices for each adapted layer, as a safetensors file.adapter_config.json: the LoRA hyperparameters, including the rankrandlora_alpha. The update applied to each weight is scaled bylora_alpha / r.
To use the weights outside Tinker, you need either the adapter merged into the base model's weights or the keys remapped to the PEFT format.
The Tinker Cookbook's weights module provides utilities to download and deploy checkpoints from Tinker. See Tinker Cookbook Deployment:
- Merge into a HuggingFace Model: one standalone model you can load with
transformers, vLLM, or SGLang - Build a PEFT LoRA Adapter: a lightweight adapter to serve on top of the base model in vLLM or SGLang
- Publish to HuggingFace Hub: upload either one to the Hub