inference cloud
Together AI
Models, pricing and feature compatibility available through Together AI.
Deployments
43
Models
42
Labs
8
With pricing
37
Models on Together AI
Follow any model to compare this deployment with the same model through other providers.
| Model | Provider model ID | Mode | Region | Input | Output | Context | Tools |
|---|---|---|---|---|---|---|---|
| Baai Bge Base En V1 5 | BAAI/bge-base-en-v1.5 | Embedding | global | $0.0080 / 1M | $0.0000 / 1M | 512 | not_applicable |
| Baai Bge Base En V1 5 | baai/bge-base-en-v1.5 | Embedding | global | $0.0080 / 1M | $0.0000 / 1M | 512 | not_applicable |
| DeepSeek-Ai Deepseek R1 | deepseek-ai/DeepSeek-R1 | Chat | global | $3.00 / 1M | $7.00 / 1M | 128K | supported |
| DeepSeek-Ai Deepseek R1 0528 Tput | deepseek-ai/DeepSeek-R1-0528-tput | Chat | global | $0.55 / 1M | $2.19 / 1M | 128K | supported |
| DeepSeek-Ai Deepseek | deepseek-ai/DeepSeek-V3 | Chat | global | $1.25 / 1M | $1.25 / 1M | 65.5K | supported |
| DeepSeek-Ai Deepseek V3 1 | deepseek-ai/DeepSeek-V3.1 | Chat | global | $0.60 / 1M | $1.70 / 1M | 128K | supported |
| Meta Llama Llama 3 2 3B Instruct Turbo | meta-llama/Llama-3.2-3B-Instruct-Turbo | Chat | global | — | — | — | supported |
| Meta Llama Llama 3 3 70B Instruct Turbo | meta-llama/Llama-3.3-70B-Instruct-Turbo | Chat | global | $0.88 / 1M | $0.88 / 1M | — | supported |
| Meta Llama Llama 3 3 70B Instruct Turbo Free | meta-llama/Llama-3.3-70B-Instruct-Turbo-Free | Chat | global | $0.0000 / 1M | $0.0000 / 1M | — | supported |
| Meta Llama Llama 4 Maverick 17B 128E Instruct FP8 | meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 | Chat | global | $0.27 / 1M | $0.85 / 1M | — | supported |
| Meta Llama Llama 4 Scout 17B 16E Instruct | meta-llama/Llama-4-Scout-17B-16E-Instruct | Chat | global | $0.18 / 1M | $0.59 / 1M | — | supported |
| Meta Llama Meta Llama 3 1 405B Instruct Turbo | meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo | Chat | global | $3.50 / 1M | $3.50 / 1M | — | supported |
| Meta Llama Meta Llama 3 1 70B Instruct Turbo | meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo | Chat | global | $0.88 / 1M | $0.88 / 1M | — | supported |
| Meta Llama Meta Llama 3 1 8B Instruct Turbo | meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo | Chat | global | $0.18 / 1M | $0.18 / 1M | — | supported |
| Mistralai Mistral 7B Instruct V0 1 | mistralai/Mistral-7B-Instruct-v0.1 | Chat | global | — | — | — | supported |
| Mistralai Mistral Small 24B Instruct 2501 | mistralai/Mistral-Small-24B-Instruct-2501 | Chat | global | — | — | — | supported |
| Mistralai Mixtral 8x7B Instruct V0 1 | mistralai/Mixtral-8x7B-Instruct-v0.1 | Chat | global | $0.60 / 1M | $0.60 / 1M | — | supported |
| Moonshotai Kimi K2 5 | moonshotai/Kimi-K2.5 | Chat | global | $0.50 / 1M | $2.80 / 1M | 256K | supported |
| Moonshotai Kimi K2 Instruct | moonshotai/Kimi-K2-Instruct | Chat | global | $1.00 / 1M | $3.00 / 1M | — | supported |
| Moonshotai Kimi K2 Instruct 0905 | moonshotai/Kimi-K2-Instruct-0905 | Chat | global | $1.00 / 1M | $3.00 / 1M | 262.1K | supported |
| gpt-oss-120b | openai/gpt-oss-120b | Chat | global | $0.15 / 1M | $0.60 / 1M | 131.1K | supported |
| gpt-oss-20b | openai/gpt-oss-20b | Chat | global | $0.05 / 1M | $0.20 / 1M | 128K | supported |
| Qwen Qwen2 5 72B Instruct Turbo | Qwen/Qwen2.5-72B-Instruct-Turbo | Chat | global | — | — | — | supported |
| Qwen Qwen2 5 7B Instruct Turbo | Qwen/Qwen2.5-7B-Instruct-Turbo | Chat | global | — | — | — | supported |
| Qwen Qwen3 235B A22B FP8 Tput | Qwen/Qwen3-235B-A22B-fp8-tput | Chat | global | $0.20 / 1M | $0.60 / 1M | 40K | unsupported |
| Qwen Qwen3 235B A22B Instruct 2507 Tput | Qwen/Qwen3-235B-A22B-Instruct-2507-tput | Chat | global | $0.20 / 1M | $6.00 / 1M | 262K | supported |
| Qwen Qwen3 235B A22B Thinking 2507 | Qwen/Qwen3-235B-A22B-Thinking-2507 | Chat | global | $0.65 / 1M | $3.00 / 1M | 256K | supported |
| Qwen Qwen3 5 397B A17B | Qwen/Qwen3.5-397B-A17B | Chat | global | $0.60 / 1M | $3.60 / 1M | 262.1K | supported |
| Qwen Qwen3 Coder 480B A35B Instruct FP8 | Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 | Chat | global | $2.00 / 1M | $2.00 / 1M | 256K | supported |
| Qwen Qwen3 Next 80B A3B Instruct | Qwen/Qwen3-Next-80B-A3B-Instruct | Chat | global | $0.15 / 1M | $1.50 / 1M | 262.1K | supported |
| Qwen Qwen3 Next 80B A3B Thinking | Qwen/Qwen3-Next-80B-A3B-Thinking | Chat | global | $0.15 / 1M | $1.50 / 1M | 262.1K | supported |
| Together Ai 21 1B 41B | together-ai-21.1b-41b | Chat | global | $0.80 / 1M | $0.80 / 1M | — | unknown |
| Together Ai 4 1B 8B | together-ai-4.1b-8b | Chat | global | $0.20 / 1M | $0.20 / 1M | — | unknown |
| Together Ai 41 1B 80B | together-ai-41.1b-80b | Chat | global | $0.90 / 1M | $0.90 / 1M | — | unknown |
| Together Ai 8 1B 21B | together-ai-8.1b-21b | Chat | global | $0.30 / 1M | $0.30 / 1M | 1K | unknown |
| Together Ai 81 1B 110B | together-ai-81.1b-110b | Chat | global | $1.80 / 1M | $1.80 / 1M | — | unknown |
| Together Ai Embedding 151m To 350m | together-ai-embedding-151m-to-350m | Embedding | global | $0.02 / 1M | $0.0000 / 1M | — | not_applicable |
| Together Ai Embedding Up To 150m | together-ai-embedding-up-to-150m | Embedding | global | $0.0080 / 1M | $0.0000 / 1M | — | not_applicable |
| Together Ai Up To 4B | together-ai-up-to-4b | Chat | global | $0.10 / 1M | $0.10 / 1M | — | unknown |
| Togethercomputer Codellama 34B Instruct | togethercomputer/CodeLlama-34b-Instruct | Chat | global | — | — | — | supported |
| Zai Org Glm 4 5 Air FP8 | zai-org/GLM-4.5-Air-FP8 | Chat | global | $0.20 / 1M | $1.10 / 1M | 128K | supported |
| Zai Org Glm 4 6 | zai-org/GLM-4.6 | Chat | global | $0.60 / 1M | $2.20 / 1M | 200K | supported |
| Zai Org Glm 4 7 | zai-org/GLM-4.7 | Chat | global | $0.45 / 1M | $2.00 / 1M | 200K | supported |