Artificial IntelligenceCloud & InfrastructureCompute & AI InfrastructureInference & APIsActive dossier
NVIDIA Groq 3 LPX enters full production with 3,431-token/s long-context inference
Groq 3 LPX is moving from architecture announcement to manufactured infrastructure. Artificial Analysis measured about 3,400 output tokens/s at both 10K and 100K context on an NVIDIA-hosted private endpoint, but the single-concurrency benchmark does not yet establish public-cloud price, multi-tenant throughput or end-to-end agent speed.