How much does it cost to deploy or run a model?

Last updated: July 31, 2026

It depends on the product. Current rates are at friendli.ai/pricing.

  • Model API: pay per token; the rate varies by model.

  • Dedicated Endpoints (on-demand): pay per GPU-hour, billed per second. The rate varies by GPU type.

Check more details on our docs.