Limits what a client may ask of the model: how large the input may be, how much the model may generate, and which models it may use. A request over the input limit or for a model that is not allowed is rejected with an error in the format of the configured provider; the output limit is written into the request instead of rejecting it. See tutorials/llm-gateway/openai/10-Basic-LLM-Gateway.yaml.
Rejects a request whose input exceeds this number of tokens. The gateway estimates the input from the size of the request before it is forwarded, so the number the provider counts may differ.
100
maxOutputTokens
false
0 (unlimited)
Caps how many tokens the model may generate. The gateway lowers the limit a client asked for and sets one where the client asked for none, but the provider decides what it honours.
200
models
false
(no restriction)
The models the gateway accepts. A request asking for any other model is rejected.