Together AI
Together AI is an AI cloud platform headquartered in San Francisco, California, that specialises in inference, fine-tuning, and deployment of open-source and custom AI models at scale. Earlier backers include …
What Together AI Does
Together AI is an AI cloud platform headquartered in San Francisco, California, that specialises in inference, fine-tuning, and deployment of open-source and custom AI models at scale. Earlier backers include NVIDIA, Salesforce Ventures, General Catalyst, Kleiner Perkins, Coatue Management, and Lux Capital.
On 1 July 2026 the company closed $800 million led by Aramco Ventures at an $8.3 billion valuation, more than doubling the $3.3 billion mark set by its $305 million round roughly sixteen months earlier. Annualised revenue is estimated near $1 billion, tripling since mid-2025.
Together AI provides a serverless inference API supporting more than 200 open-source models including Llama, Mistral, DeepSeek, and Qwen, alongside dedicated fine-tuning infrastructure for teams training custom models on proprietary datasets. Its compute platform gives developers access to NVIDIA H100 and H200 GPU clusters without the overhead of managing infrastructure, with per-token pricing for inference and per-GPU-hour pricing for training.
Key differentiators include industry-leading inference speed (Together AI claims top throughput benchmarks on Llama models), model library breadth, and built-in support for RLHF and LoRA fine-tuning workflows. The platform is widely used by AI startups, enterprise engineering teams, and researchers who want the capability of frontier open-source models without the capital expense of building their own GPU infrastructure.
Together AI also publishes open research on efficient inference architectures and model quantisation techniques. Together AI is the odd one out in this category and should not be price-compared with the rest without adjustment: its serverless tier sells tokens, not GPU-hours, so what you are buying is a managed inference and fine-tuning layer with the capacity question handled for you, not a machine you operate.
That makes it complementary to CoreWeave or Lambda rather than an alternative — many teams prototype against its API and migrate to dedicated GPUs only when volume makes per-token economics worse than owning utilisation. Its strategic bet is that open-weight models close enough to frontier quality will win on cost and control, which is why breadth of the model library and speed of adding new releases matter more here than GPU inventory.
Revenue figures circulating for the company are third-party estimates rather than reported results, and should be treated as such. Buyer caveats: per-token pricing is excellent for spiky or early-stage workloads and can become the most expensive option at sustained high throughput, so model the crossover before committing; and because the draw is hosted open-weight models, the practical switching question is which models are available where, not which cloud holds the contract.
Best fit: teams shipping products on open-weight models who want inference and fine-tuning as a service rather than infrastructure to run.
Sign in with your company email to claim and enrich this profile.
How Together AI compares in its category
Read our independently researched buyer's guides to see where Together AI sits against the other leading vendors, how the category works, and what to check before shortlisting.