What knobs can I configure on a Dedicated Endpoint?

Last updated: August 6, 2026

When deploying a Dedicated Endpoint, you can configure a wide range of performance and scaling options to optimize its performance.

These include online quantization (8-bit to 4-bit), host KV cache offloading, request queuing, draft-model and n-gram speculative decoding, and autoscaling settings such as replica count and cooldown periods. You can also configure engine-level options, including maximum context length, special tokens, maximum batch size, and logging.

See our documentation for details:

Try deploying your first Dedicated Endpoint on Friendli Suite.