- vendor
- Groq
- whatItIs
- Custom LPU (Language Processing Unit) inference; delivers 500-1500 tok/sec on open-weight models with day-zero coverage of new releases.
- hosting
- managed-only
- pricingModel
- per-token (rates undisclosed publicly; enterprise custom)
- modelProviders
- open-weight (Llama, Mixtral, Qwen, Kimi, GPT-OSS)
- stateModel
- stateless
- toolModel
- OpenAI-compatible API
- includesBrowser
- false
- includesMemory
- false
- longRunning
- per-request
- hitlSupport
- false
- observability
- cost + latency metrics
- maturity
- ga
- launched
- 2016 (LPU); GroqCloud 2023
- notes
- Nvidia licensed Groq 3 LPU for DGX Cloud in 2026. Serves 3M+ developers; $750M Series C in 2025-Q3.