# Together AI cuts interruptible GPU cluster compute to half price with a five-minute exit window

Together AI now offers preemptible GPU nodes at a fixed 50% discount to on-demand compute, with sub-hour billing, automatic refill and a maximum five-minute drain before reclaim. Current documentation covers Kubernetes and Slurm, but capacity and node lifetime are not guaranteed.

Training experiments and batch inference can use Together AI's discounted preemptible GPUs in existing clusters. Workloads must checkpoint or requeue on interruption, and at least one standard node is required.

- Status: Active
- Published: 2026-10-09T06:22:30+13:00
- Updated: 2026-10-09T06:22:30+13:00
- Categories: SaaS, Cloud & Infrastructure, AI SaaS, Compute & AI Infrastructure
- Tags: Cloud pricing, GPU inference, Together AI
- Canonical HTML: https://beyondthe.news/dossiers/together-ai-preemptible-gpu-clusters-half-price-drain-window

## What changed

Together AI introduced preemptible GPU nodes for its GPU Clusters in a September 10, 2026 public-preview announcement. The nodes use spare accelerator capacity and are billed at a fixed 50% of the equivalent on-demand rate rather than a fluctuating spot-market price. They can be added to an existing cluster alongside standard nodes, metered in one- to two-minute increments and replenished automatically toward a requested capacity target after reclamation. A reclaim notice starts a drain window of at most five minutes; workloads must checkpoint or finish before removal. The launch announcement focused on Kubernetes, while the current first-party documentation also describes public-preview support for Slurm with job requeue. The lower unit price trades away availability guarantees and predictable uninterrupted execution.

## Why it matters

GPU spend can dominate the economics of training, evaluation, fine-tuning and batch inference, particularly for smaller AI teams. A fixed half-price tier makes the trade-off easier to model than a variable spot market, but the relevant metric is completed-job cost after checkpoints, retries and idle waiting, not GPU price alone. Builders can mix reliable coordinator and serving capacity with cheaper interruptible workers, using the same cluster control plane rather than managing a second pool manually. The preview's no-minimum-lifetime rule means jobs without restartable state can still lose money and time.

## Half-price capacity joins ordinary clusters

Preemptible nodes are a compute class within Together GPU Clusters rather than a separate cluster product. Together says they cost 50% of the on-demand rate, with usage metered sub-hourly every one to two minutes. Customers set a desired preemptible GPU count, but actual allocated capacity can remain below that target when spare capacity is scarce. Each cluster requires at least one standard node; existing nodes cannot be converted between the two classes in place.

## Five minutes is a maximum, not a promise to finish

When a node is reclaimed, Together starts a drain sequence that ends within five minutes. Kubernetes nodes are cordoned, an event is emitted and pods receive SIGTERM; pods can request up to 300 seconds of graceful termination. Slurm drains the node and can requeue interrupted jobs submitted with `--requeue`. Teams should checkpoint to shared storage before interruption and assume a node can disappear at any time during preview.

## Keep the control plane on reliable capacity

Together labels Kubernetes preemptible nodes `together.ai/compute-class=preemptible` so operators can schedule suitable workers there while pinning coordinators, user-facing replicas and stateful anchors to standard nodes. The platform refills the requested preemptible target as capacity becomes available, but that does not guarantee an immediate replacement. This architecture favours batch jobs, evaluations and checkpointable experiments rather than strict-latency serving without fallback.

## Price savings need workload-level measurement

The headline 50% reduction applies to GPU-node rates, not necessarily the total cost of a successful job. Checkpoint writes, failed partial work, retry loops, orchestration overhead and any persistent standard nodes still cost money. Compare cost per completed evaluation, training run or batch against on-demand rather than assuming the full workload bill is halved.

## Documentation now covers Kubernetes and Slurm

Together's September announcement described Kubernetes availability and Slurm as planned. Its current documentation now says both Kubernetes and Slurm are in public preview and documents Slurm preemptible constraints and requeue behaviour. Operators should rely on the live documentation and confirm availability for their specific region and accelerator rather than the original launch wording alone.

## Key details

- September 10, 2026 announcement introduced preemptible compute in Together GPU Clusters.
- Fixed 50% discount relative to the equivalent on-demand GPU rate, not variable spot bidding.
- Billing is sub-hourly, metered approximately every one to two minutes.
- Reclamation drain window is at most five minutes, with SIGTERM on Kubernetes.
- Current documentation covers Kubernetes and Slurm public previews.
- Requested preemptible GPU capacity is a target, not a guarantee of allocation.
- At least one standard node is required; node classes cannot be converted in place.
- Together attempts automatic refill after reclamation as spare capacity returns.
- No minimum preemptible-node lifetime is guaranteed during preview.

## Builder takeaways

- Use preemptible capacity for retryable batch inference, evaluations and checkpointable fine-tuning, not critical coordinators or single-replica live serving.
- Set Kubernetes graceful termination to support checkpointing within 300 seconds; use Slurm requeue and shared checkpoints where appropriate.
- Track desired versus allocated preemptible GPUs and plan for periods with no spare capacity.
- Benchmark total cost per completed job after retries, checkpoints and required standard nodes before claiming 50% end-to-end savings.
- Verify live region, accelerator and scheduler support because the September launch and current documentation differ on Slurm availability.

## What to watch

- General availability and guaranteed limits for Kubernetes and Slurm preemptible compute.
- Actual preemption frequency and time-to-refill by GPU type and region.
- Changes to the flat 50% discount or minimum standard-node requirement.
- Workload-level independent measurements of completed-job cost and reliability.

## Uncertainties

- Together does not guarantee a minimum node lifetime or availability of spare GPU capacity.
- The 50% figure is a vendor-published rate discount, not independently measured end-to-end savings.
- Documentation now lists Slurm preview support, but the date that support began is not established by the September announcement.
- Actual checkpoint/retry overhead varies by workload and storage performance.

## Sources

- [Introducing preemptible compute: the same compute, half the price](https://www.together.ai/blog/introducing-preemptible-compute-the-same-compute-half-the-price) — Together AI · primary announcement · 2026-09-10T00:00:00+12:00. Fixed 50% rate, billing cadence, five-minute drain, auto-refill and initial Kubernetes availability.
- [Preemptible compute — GPU Clusters](https://docs.together.ai/docs/preemptible-compute) — Together AI Docs · primary documentation · 2026-10-09T00:00:00+13:00. Current Kubernetes and Slurm preview support, no minimum lifetime, labels, scheduling, reclamation and limitations.

