Safety & ops
What is Checkpoint?
A checkpoint is a saved state an agent run can resume from after failure or pause.
Checkpoints make long workflows recoverable. They should store artifacts, not hidden thoughts.
Why it matters
A checkpoint saves the state of an agent run so it can be resumed after failure, pause, or timeout. Without checkpoints, long-running workflows must restart from scratch when anything goes wrong.
Checkpoints store artifacts and state — not hidden model thoughts. They should be inspectable and reproducible.
Key takeaways
- 1Checkpoints enable resumable, recoverable agent workflows.
- 2They store artifacts and state, not model internals.
- 3Essential for long-running workflows where restart cost is high.
Common mistakes
- ✕Not implementing checkpoints for workflows that take more than a few minutes.
- ✕Storing model-internal state instead of observable artifacts.
Related terms
Concept neighborhood
Terms linked from Checkpoint in the glossary graph.
- Checkpoint
- Agentic Workflow
- Trace