Demand and performance
Response times can vary depending on the model and current demand. Together aims to maintain high performance for serverless requests, including:- Tokens per second (TPS): The speed at which a model generates output tokens.
- Time to first token (TTFT): The time between sending a request and receiving the first output token.
429 Too Many Requests or 503 Service Unavailable response.
Serverless performance is best-effort. For committed throughput and reliability guarantees, see provisioned throughput.
Handle errors during high demand
The response code tells you how to adjust your requests:429 Too Many Requests: Reduce your request rate. Spread requests out over time and avoid sending large bursts. Use exponential backoff when retrying.503 Service Unavailable: Wait briefly, then retry. Use exponential backoff if the error continues.
Get throughput guarantees
If your workload needs committed throughput and reliability guarantees, provisioned throughput reserves capacity for a selected model or model family with a defined service level agreement (SLA). Contact sales to discuss your workload and capacity requirements.Contract and model-specific limits
- Enterprise and Scale contracts: If you have an active Enterprise or Scale contract, your purchased rate limits stay in place until the contract expires. Nothing changes during your current term.
- Model-specific limits: When a model is in especially high demand, Together may apply custom rate limits or access restrictions to that model.