# Fireworks is raising on-demand GPU prices up to 30% on September 1

Fireworks' current pricing table shows September 1 increases across H100, H200, B200, B300 and GB300 on-demand deployments, with B200 rising from $10 to $13 per GPU-hour. Dedicated training jobs that use the same GPU-hour rates inherit the new economics.

The price changes are not uniform: H100/H200 rise about 14%, B200 30%, B300 25% and GB300 about 11%. Builders using dedicated inference or training should re-run workload economics before assuming newer accelerators remain the cheapest route per completed task.

- Status: Active
- Published: 2026-08-28T22:59:29+12:00
- Updated: 2026-08-28T22:59:29+12:00
- Categories: Artificial Intelligence, Cloud & Infrastructure, Compute & AI Infrastructure, Inference & APIs
- Tags: Cloud pricing, Fireworks AI, GPU inference, GPU training
- Canonical HTML: https://beyondthe.news/dossiers/fireworks-ai-on-demand-gpu-price-increase-september-2026

## What changed

Fireworks' current pricing page publishes a two-column transition for its on-demand GPU deployments: prices through August 31 and new rates from September 1, 2026. H100 80GB and H200 141GB increase from $7 to $8 per GPU-hour; B200 180GB rises from $10 to $13; B300 288GB from $12 to $15; and GB300 288GB from $18 to $20. Fireworks bills on-demand deployments per GPU second with no startup charge, and its Dedicated Training API uses the same GPU-hour pricing, so the change affects both dedicated serving and GPU-hour training jobs. Region-restricted deployments continue to carry a 1.5× premium.

## Why it matters

Dedicated GPU economics can dominate the unit cost of high-volume inference, fine-tuning and reinforcement-learning workloads. These increases range from roughly 11% to 30%, large enough to change break-even points between accelerator generations, serverless APIs and competing GPU clouds. The biggest increase is on B200, where a workload consuming 1,000 GPU-hours per month moves from $10,000 to $13,000 before regional premiums or other services. Teams should therefore benchmark completed-task throughput rather than choosing hardware on hourly price or peak specifications alone.

## The September increase varies materially by accelerator

H100 and H200 move from $7 to $8 per hour, about a 14.3% increase. B200 rises 30% from $10 to $13, B300 rises 25% from $12 to $15 and GB300 rises about 11.1% from $18 to $20. The uneven changes can alter the relative cost ranking of GPUs even when performance remains unchanged.

## Training inherits the same GPU-hour table

Fireworks states that Dedicated Training API jobs are priced per GPU hour using the On-Demand Pricing section. Teams running reinforcement fine-tuning or dedicated training should therefore update both serving and training forecasts, not only inference deployment budgets.

## Per-second billing limits idle waste but not rate exposure

On-demand deployments are billed by the second and do not charge extra startup time, which remains useful for bursty workloads. But sustained fleets still absorb the full hourly increase, and a region-restricted deployment costs 1.5 times the listed rate.

## Hourly price is not total inference cost

A more expensive accelerator can still win if it materially improves throughput, latency or batch efficiency. Builders should compare cost per successful request, training run or completed agent task under the exact model, quantization, context and concurrency they use.

## Key details

- The new Fireworks on-demand GPU rates begin September 1, 2026.
- H100 80GB: $7 → $8 per GPU-hour.
- H200 141GB: $7 → $8 per GPU-hour.
- B200 180GB: $10 → $13 per GPU-hour.
- B300 288GB: $12 → $15 per GPU-hour.
- GB300 288GB: $18 → $20 per GPU-hour.
- On-demand deployments are billed per GPU second with no extra startup charges.
- Region-restricted deployments are priced at a 1.5× premium.
- Dedicated Training API jobs use the same on-demand GPU-hour pricing.

## Builder takeaways

- Update September budgets now for any persistent Fireworks deployment; the B200 and B300 changes are large enough to invalidate old monthly forecasts.
- Benchmark cost per completed workload across H100/H200/B200/B300 rather than assuming a newer GPU is automatically cheaper after the rate changes.
- If workload residency requires region-restricted deployment, apply the 1.5× premium after the new base rate rather than modeling only the published headline number.
- For training, recalculate long-running jobs and experiment budgets because Dedicated Training uses the same GPU-hour table.
- Compare dedicated deployments with Fireworks serverless and competing providers after the increase; usage shape can change which model is economically preferable.

## What to watch

- Whether Fireworks changes serverless model token prices alongside the dedicated GPU increases.
- Customer-facing explanations or contract options for the September pricing change.
- Reserved or committed-capacity discounts that offset the higher on-demand rates.
- Independent throughput-per-dollar comparisons on B200/B300/GB300 after the new rates take effect.

## Uncertainties

- The pricing page publishes the dated rate transition but does not explain the commercial rationale for each GPU-specific increase.
- Effective cost varies significantly with model serving efficiency, batching, quantization, context length and utilization.
- Enterprise or negotiated contracts may differ from public on-demand pricing.

## Sources

- [Fireworks Pricing](https://fireworks.ai/pricing) — Fireworks AI · primary pricing documentation. Current dated pricing table for on-demand deployments, September 1 rates, region-restricted premium and Dedicated Training API pricing linkage.

