- vendor
- Together Compute
- whatItIs
- Full-stack AI cloud for inference, fine-tuning, and GPU compute on open-weight models with 2x faster throughput via proprietary kernels.
- hosting
- managed-only
- pricingModel
- per-token + dedicated GPU ($6.49/hr H100, $11.95/hr B200)
- modelProviders
- open-weight (Llama, Mistral, DeepSeek, Qwen, Gemma, MiniMax)
- stateModel
- stateless
- toolModel
- OpenAI-compatible API + Together API
- includesBrowser
- false
- includesMemory
- false
- longRunning
- batch API for high-volume jobs
- hitlSupport
- false
- observability
- cost tracking + request logging
- maturity
- ga
- launched
- 2022
- notes
- 30B tokens/day batch capacity. ISO 27001 certified. Llama 3.3 70B at $0.88/1M tokens.