What it does
Together AI is a cloud platform for running, fine-tuning and serving open-source AI models. Developers call a large library of chat, code, image, video and voice models through an OpenAI-compatible API, with serverless, batch, reserved-throughput and dedicated deployment options. The same platform rents GPU clusters for training, offers fine-tuning on your own data, code sandboxes and managed storage, and Together Link lets you use its open models inside the coding harness you already work in.
Official source 1 ↗ · Official source 2 ↗Features
- Serverless inference on open models. Run open-source models on demand through an OpenAI-compatible API with no infrastructure to manage and no long-term commitment. The catalogue spans chat, vision, image, audio, video, transcription, embeddings, rerank and moderation models. Official source 1 ↗ · Official source 2 ↗ · Official source 3 ↗
- Batch inference. Process large workloads asynchronously at lower cost, up to 30 billion tokens per model, with any serverless model or a private deployment. Official source 1 ↗
- Provisioned Throughput. Reserve committed inference capacity in throughput units (PTUs) with token-based pricing and a 99% uptime SLA, keeping drop-in API compatibility. Official source 1 ↗ · Official source 2 ↗
- Dedicated model and container inference. Deploy models on single-tenant GPUs with guaranteed performance, custom-model support and autoscaling. Dedicated Container Inference targets video, audio and image models. Official source 1 ↗ · Official source 2 ↗
- GPU clusters. Rent NVIDIA GPU capacity from self-serve instant clusters up to thousands of GPUs, including H100 and B200 clusters with attached storage for training or large batch jobs, on preemptible, on-demand or reserved terms. Official source 1 ↗ · Official source 2 ↗
- Fine-tuning. Fine-tune open-source models on your own data with supervised fine-tuning or Direct Preference Optimization, using LoRA or full fine-tuning, then deploy the result for inference. Official source 1 ↗ · Official source 2 ↗ · Official source 3 ↗
- Together Link. Run Together's open models inside the coding harness you already use, with one install command, automatic routing and per-session cost receipts. It is marked beta in the docs. Official source 1 ↗ · Official source 2 ↗
- Code sandboxes and managed storage. Secure code sandboxes provide development environments for AI apps and agents, a Code Interpreter API runs LLM-generated code, and managed object storage and parallel filesystems have zero egress fees. Official source 1 ↗ · Official source 2 ↗
- Data-use controls. Together AI says it does not train its models on your data without explicit opt-in, and account privacy settings can stop it retaining training data, prompts or model responses. The docs also cover zero data retention and single sign-on. Official source 1 ↗ · Official source 2 ↗
Pricing & free limits
Each model has its own input and output token price, often with a cheaper cached-input rate and a separate Batch API price. Examples: gpt-oss-120B $0.15 / $0.60, MiniMax M3 $0.30 / $1.20 (cached input $0.06), Llama 3.3 70B $1.04 / $1.04. Official source 1 ↗
Media models are priced per output unit: GPT Image 2 at $0.053 per image, ByteDance Seedance 2.5 at $0.115 per video, Whisper Large v3 transcription at $0.0015 per audio minute. Text-to-speech is priced per million characters, for example Kokoro-82M at $4.00. Official source 1 ↗
You reserve a number of PTUs, each a fixed slice of capacity whose tokens per minute depend on the model. The pricing page includes a cost estimator; estimates assume 24/7 provisioning. Comes with a 99% uptime SLA. Official source 1 ↗ · Official source 2 ↗
Single-tenant GPU endpoints. H200, B300 and GB200/GB300 NVL72 hardware and all reserved capacity are quoted by sales. Official source 1 ↗
H100 runs $1.99 preemptible, $3.99 on demand, and $3.69 to $3.19 for 3- to 180-day reservations; longer reservations and NVL72 systems are by quote. Other on-demand rates: H200 $5.99, B200 $8.19, B300 $9.99. Official source 1 ↗
Supervised fine-tuning and DPO are priced per model, from $0.34 (SFT) and $0.84 (DPO) per million tokens on small Qwen models up to $40 and $100 on GLM-5.2. Each job has a minimum charge, from $4.00. Official source 1 ↗
Code Sandbox VMs cost $0.0446 per vCPU and $0.0149 per GiB of RAM per hour. Code Interpreter sessions of 60 minutes cost $0.03. Shared filesystem storage is $0.16 per GiB per month. Official source 1 ↗
Billing details
Fine-tuning cost is based on the tokens in the training dataset times the number of epochs, plus any evaluation-set tokens. Official source 1 ↗
Listed image and video prices are for the lowest resolution and duration settings; actual prices can be higher. Official source 1 ↗
Free access & limits
Free access and its limits were not established by the checked official sources.
Best for & limitations
Building an app or agent on open-source models through an OpenAI-compatible API without running inference infrastructure. Official source 1 ↗ · Official source 2 ↗
Customising an open model with your own data and serving the fine-tuned model in production. Official source 1 ↗
Renting H100 or B200 GPU clusters for model training or large batch jobs. Official source 1 ↗
Using cheaper open models inside an existing coding harness through Together Link. Official source 1 ↗
- Newest GPU systems and reservations by quote. H200, B300 and GB200/GB300 NVL72 dedicated endpoints, NVL72 clusters and long reservations have no public price and require contacting sales. Official source 1 ↗
- Together Link is in beta. Together Link is labelled beta in the documentation. Official source 1 ↗
- Media prices are minimums. Listed image and video model prices apply to the lowest resolution and duration settings; actual costs vary. Official source 1 ↗