- vendor
- Fireworks Research
- whatItIs
- Ultra-low-latency LLM inference via FireAttention; processes 30T+ tokens/day at 180k req/sec.
- hosting
- managed-only
- pricingModel
- per-token ($0.20/1M for 8B, $0.90/1M for 70B as of 2026-04)
- modelProviders
- open-weight (DeepSeek, GLM, Qwen, MiniMax, Gemma, Kimi, Flux)
- stateModel
- stateless
- toolModel
- OpenAI-compatible API
- includesBrowser
- false
- includesMemory
- false
- longRunning
- per-request
- hitlSupport
- false
- observability
- latency + request logging
- maturity
- ga
- launched
- 2023
- notes
- $254M Series B at $4B valuation; Azure Foundry integration in 2026-Q1. Founded by ex-PyTorch core.