What changed
Together AI introduced preemptible GPU nodes for its GPU Clusters in a September 10, 2026 public-preview announcement. The nodes use spare accelerator capacity and are billed at a fixed 50% of the equivalent on-demand rate rather than a fluctuating spot-market price. They can be added to an existing cluster alongside standard nodes, metered in one- to two-minute increments and replenished automatically toward a requested capacity target after reclamation. A reclaim notice starts a drain window of at most five minutes; workloads must checkpoint or finish before removal. The launch announcement focused on Kubernetes, while the current first-party documentation also describes public-preview support for Slurm with job requeue. The lower unit price trades away availability guarantees and predictable uninterrupted execution.
Why it matters
GPU spend can dominate the economics of training, evaluation, fine-tuning and batch inference, particularly for smaller AI teams. A fixed half-price tier makes the trade-off easier to model than a variable spot market, but the relevant metric is completed-job cost after checkpoints, retries and idle waiting, not GPU price alone. Builders can mix reliable coordinator and serving capacity with cheaper interruptible workers, using the same cluster control plane rather than managing a second pool manually. The preview's no-minimum-lifetime rule means jobs without restartable state can still lose money and time.
Half-price capacity joins ordinary clusters
Preemptible nodes are a compute class within Together GPU Clusters rather than a separate cluster product. Together says they cost 50% of the on-demand rate, with usage metered sub-hourly every one to two minutes. Customers set a desired preemptible GPU count, but actual allocated capacity can remain below that target when spare capacity is scarce. Each cluster requires at least one standard node; existing nodes cannot be converted between the two classes in place.
Five minutes is a maximum, not a promise to finish
When a node is reclaimed, Together starts a drain sequence that ends within five minutes. Kubernetes nodes are cordoned, an event is emitted and pods receive SIGTERM; pods can request up to 300 seconds of graceful termination. Slurm drains the node and can requeue interrupted jobs submitted with `--requeue`. Teams should checkpoint to shared storage before interruption and assume a node can disappear at any time during preview.
Keep the control plane on reliable capacity
Together labels Kubernetes preemptible nodes `together.ai/compute-class=preemptible` so operators can schedule suitable workers there while pinning coordinators, user-facing replicas and stateful anchors to standard nodes. The platform refills the requested preemptible target as capacity becomes available, but that does not guarantee an immediate replacement. This architecture favours batch jobs, evaluations and checkpointable experiments rather than strict-latency serving without fallback.
Price savings need workload-level measurement
The headline 50% reduction applies to GPU-node rates, not necessarily the total cost of a successful job. Checkpoint writes, failed partial work, retry loops, orchestration overhead and any persistent standard nodes still cost money. Compare cost per completed evaluation, training run or batch against on-demand rather than assuming the full workload bill is halved.
Documentation now covers Kubernetes and Slurm
Together's September announcement described Kubernetes availability and Slurm as planned. Its current documentation now says both Kubernetes and Slurm are in public preview and documents Slurm preemptible constraints and requeue behaviour. Operators should rely on the live documentation and confirm availability for their specific region and accelerator rather than the original launch wording alone.