Remove 240s request timeout on Inference API and make docs on parameters like reasoning effort clearer

Currently the 240s server-side timeout on all requests makes it difficult to seriously use the gateway especially for models with a low throughput, and I’m not certain if the gateway forwards parameters like effort for Anthropic & OpenAI models either?

Please authenticate to join the conversation.

Upvoters
Status

In Review

Board
πŸ’‘

Ideas & Suggestions

Date

11 days ago

Author

An Anonymous User

Subscribe to post

Get notified by email when there are changes.