# Manage Checkpoints TrainingClients can save checkpoints for resuming training later, or serving the checkpoint for sampling. | Method | Use it to | Path | | ----------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------- | ------------------------------------------ | | [`save_state`](https://tinker-docs.thinkingmachines.ai/tinker/api-reference/trainingclient/#save_state) | Resume training | `tinker:///weights/` | | [`save_weights_for_sampler`](https://tinker-docs.thinkingmachines.ai/tinker/api-reference/trainingclient/#save_weights_for_sampler) | Sample, open in the Playground, or download | `tinker:///sampler_weights/` | ## Training checkpoints Save the full training state, including Adam's moment estimates, with `save_state`: ```python checkpoint = training_client.save_state("step-100", user_metadata={ "step": "100", "recipe": "example" }).result() print(checkpoint.path) # tinker:///weights/step-100 # load without optimizer state training_client = service_client.create_training_client_from_state(checkpoint.path).result() # load with optimizer state training_client = service_client.create_training_client_from_state_with_optimizer(checkpoint.path).result() # view checkpoint in Tinker console print(checkpoint.get_console_url()) ``` ## Sampler checkpoints `save_weights_for_sampler` saves only the weights and not gradient or optimizer state for sampling in one of the following ways: - Load using [`SamplingClient`](https://tinker-docs.thinkingmachines.ai/tinker/api-reference/samplingclient/index.md) - Use with [`OpenAI`](https://tinker-docs.thinkingmachines.ai/tinker/compatible-apis/openai/index.md) or [`Anthropic`](https://tinker-docs.thinkingmachines.ai/tinker/compatible-apis/anthropic/index.md) compatible APIs - Chat with in the [Tinker Playground](https://tinker.thinkingmachines.ai/playground) ```python sampler_checkpoint = training_client.save_weights_for_sampler("step-100", user_metadata={ "step": "100", "recipe": "example" }).result() print(sampler_checkpoint.path) # tinker:///sampler_weights/step-100 # load the checkpoint into a sampling client sampling_client = service_client.create_sampling_client(model_path=sampler_checkpoint.path) # chat with checkpoint in Tinker Playground print(sampler_checkpoint.get_playground_url()) # manage checkpoint in Tinker console print(sampler_checkpoint.get_console_url()) ``` ## Expiration Checkpoint storage is charged at $0.10 per GB per month. To ensure checkpoints are not kept forever, you can set a TTL (time to live) when saving a checkpoint. ```python training_client.save_state("step-100", ttl_seconds=24 * 3600) training_client.save_weights_for_sampler("step-100", ttl_seconds=24 * 3600) ``` Without a TTL, a checkpoint is kept until you delete it. To change the TTL later, see [Set a TTL](#set-a-ttl) below. ## Manage with RestClient Use the [`RestClient`](https://tinker-docs.thinkingmachines.ai/tinker/api-reference/restclient/index.md) to manage checkpoints after they are saved. ```python rest_client = service_client.create_rest_client() ``` ### List checkpoints `list_checkpoints` lists the checkpoints of a single training run. `list_user_checkpoints` lists them across all of your runs, newest first: ```python response = rest_client.list_checkpoints("").result() for checkpoint in response.checkpoints: print(checkpoint.checkpoint_type, checkpoint.tinker_path) recent = rest_client.list_user_checkpoints(limit=10).result() ``` ### Set a TTL Change or remove a checkpoint's TTL after it is saved: ```python # Expire in 7 days rest_client.set_checkpoint_ttl_from_tinker_path(path, ttl_seconds=7 * 24 * 3600).result() # Keep indefinitely rest_client.set_checkpoint_ttl_from_tinker_path(path, ttl_seconds=None).result() ``` ### Publish a checkpoint Publishing a checkpoint lets other Tinker users load it from its path. Only the owner of the training run can publish or unpublish: ```python rest_client.publish_checkpoint_from_tinker_path(path).result() rest_client.unpublish_checkpoint_from_tinker_path(path).result() ``` ### Download sampler weights Sampler checkpoints can be downloaded to local disk as a tar file: ```python import tarfile import urllib.request # create a signed url for the checkpoint archive = rest_client.get_checkpoint_archive_url_from_tinker_path(sampler_checkpoint.path).result() # download the checkpoint urllib.request.urlretrieve(archive.url, "checkpoint.tar") # extract the checkpoint with tarfile.open("checkpoint.tar") as tar: tar.extractall("./adapter", filter="data") ``` See [`tinker checkpoint download`](https://tinker-docs.thinkingmachines.ai/tinker/cli/checkpoint/#checkpoint-download) to download from the CLI. #### Weights format The archive holds a LoRA adapter, not a full model. It contains: - `adapter_model.safetensors`: the trained LoRA `A` and `B` matrices for each adapted layer, as a [safetensors](https://huggingface.co/docs/safetensors) file. - `adapter_config.json`: the LoRA hyperparameters, including the rank `r` and `lora_alpha`. The update applied to each weight is scaled by `lora_alpha / r`. To use the weights outside Tinker, you need either the adapter merged into the base model's weights or the keys remapped to the PEFT format. The Tinker Cookbook's [`weights`](https://tinker-docs.thinkingmachines.ai/cookbook/api-reference/weights/index.md) module provides utilities to download and deploy checkpoints from Tinker. See [Tinker Cookbook Deployment](https://tinker-docs.thinkingmachines.ai/cookbook/deployment/basics/index.md): - [Merge into a HuggingFace Model](https://tinker-docs.thinkingmachines.ai/cookbook/deployment/export-hf/index.md): one standalone model you can load with `transformers`, vLLM, or SGLang - [Build a PEFT LoRA Adapter](https://tinker-docs.thinkingmachines.ai/cookbook/deployment/lora-adapter/index.md): a lightweight adapter to serve on top of the base model in vLLM or SGLang - [Publish to HuggingFace Hub](https://tinker-docs.thinkingmachines.ai/cookbook/deployment/publish-hub/index.md): upload either one to the Hub