
How are serving endpoints billed?
Databricks Model Serving uses serverless compute. The billing model depends on the endpoint and model type, including pay-per-token, CPU or GPU, and provisioned throughput.
Serving endpoint usage appears in the billing system tables under MODEL_SERVING. A DBU badge is a usage signal. The SKU and regional price determine the cost.
Published GPU Serving DBU rates
Public Model Serving rates in DBUs per hour. Multiply the DBU rate by the applicable regional DBU price to estimate a charge.
| GPU configuration | DBUs per hour | Hourly cost at $0.070 / DBU |
|---|---|---|
| T4 or equivalent | 10.48 | $0.73/hour |
| A10G x 1 GPU or equivalent | 20.00 | $1.40/hour |
| A10G x 4 GPU or equivalent | 112.00 | $7.84/hour |
| A10G x 8 GPU or equivalent | 290.80 | $20.36/hour |
| L40S x 1 GPU or equivalent | 44.86 | $3.14/hour |
| A100 40GB x 8 GPU or equivalent | 538.40 | $37.69/hour |
| A100 80GB x 8 GPU or equivalent | 628.00 | $43.96/hour |
| H100 x 1 GPU or equivalent | 100.00 | $7.00/hour |
| H100 x 8 GPU or equivalent | 800.00 | $56.00/hour |
Published rates checked October 6, 2026. Pay-per-token, CPU and provisioned throughput endpoints use different pricing.
Prices shown for US East (AWS us-east-1)View official Databricks pricing.
How to review endpoint usage
- Connect FinOpsWay and open Serving.
- Read the DBU badge beside an endpoint and check its reporting period.
- Compare endpoints with the same type and workload. Review scale-to-zero settings and traffic before acting on a change.
Common questions
Does scale to zero remove every cost?
It can reduce idle serving usage, but restarting an endpoint creates launch usage. Compare the startup pattern with latency and availability needs.
Can one rate price every endpoint?
No. Model type, serving mode, workload and region affect the billed SKU. Use your billing data for the applicable price.
Billing data updates every few hours. Check the reported period and freshness. How FinOpsWay calculates costs.