Skip to main content
In most cases, you can use Together AI serverless inference without encountering rate limits. Rate limiting may occasionally occur with high request volumes or large bursts of traffic.

Demand and performance

Response times can vary depending on the model and current demand. Together aims to maintain high performance for serverless requests, including:
  • Tokens per second (TPS): The speed at which a model generates output tokens.
  • Time to first token (TTFT): The time between sending a request and receiving the first output token.
When demand is high, Together may limit requests to maintain these performance goals. You may receive a 429 Too Many Requests or 503 Service Unavailable response. Serverless performance is best-effort. For committed throughput and reliability guarantees, see provisioned throughput.

Handle errors during high demand

The response code tells you how to adjust your requests:
  • 429 Too Many Requests: Reduce your request rate. Spread requests out over time and avoid sending large bursts. Use exponential backoff when retrying.
  • 503 Service Unavailable: Wait briefly, then retry. Use exponential backoff if the error continues.
Limit retries to fit your application’s latency needs. For workloads that need committed throughput and reliability, consider provisioned throughput.

Get throughput guarantees

If your workload needs committed throughput and reliability guarantees, provisioned throughput reserves capacity for a selected model or model family with a defined service level agreement (SLA). Contact sales to discuss your workload and capacity requirements.

Contract and model-specific limits

  • Enterprise and Scale contracts: If you have an active Enterprise or Scale contract, your purchased rate limits stay in place until the contract expires. Nothing changes during your current term.
  • Model-specific limits: When a model is in especially high demand, Together may apply custom rate limits or access restrictions to that model.