Home/Library/Bedrock Pricing
Explainer · AI & GPU · Updated June 2026

Amazon Bedrock Pricing Explained: On-Demand vs Provisioned Throughput

Amazon Bedrock gives you two ways to pay for the same models, and picking the wrong one can double your bill. Here is how on-demand per-token billing and provisioned throughput actually work, and the utilization break-even that tells you which to use.

TL;DR · Key takeaways

Amazon Bedrock charges two ways. On-demand bills per token with no commitment, which is cheapest for variable or low-volume traffic. Provisioned throughput reserves dedicated model units at an hourly rate, with one-month or six-month commitments that lower the rate, and becomes cheaper once sustained volume keeps those units busy. The decision is a utilization break-even: stay on-demand until your steady token volume is high enough that the reserved hourly cost per token falls below the on-demand per-token rate. Verify current rates on the AWS Bedrock pricing page, because they change as models are added.

Last updated: June 2026

Amazon Bedrock is AWS's managed service for calling foundation models through one API, and it offers two pricing modes for the same models. On-demand is pay-per-token with no commitment. Provisioned throughput reserves dedicated capacity, measured in model units, at an hourly price you can discount with a commitment. They are not tiers of the same plan; they are two cost structures, and the right one depends entirely on how steady your traffic is. Choosing well is one of the highest-leverage decisions in an AI bill, because inference runs continuously once a feature ships.

This article is part of our AI, GPU and ML cluster. For the full picture start with the complete guide to AI and GPU cost optimization, the pillar this piece links up to. If you are still deciding whether to use a managed API at all, read our companion guide on how to choose between self-hosted and API LLMs on cost.

How does Amazon Bedrock charge for on-demand usage?

On-demand Bedrock billing charges per token, with separate rates for input and output tokens that differ by model. You are billed only for what you process, with no reservation and no minimum, so cost scales directly with usage and drops to zero when no one calls the model. Output tokens usually cost more than input tokens, and image and embedding models bill on their own per-unit metrics rather than tokens. This mode suits prototypes, internal tools, and any workload whose volume is variable or hard to predict. Current per-model rates are on the AWS Bedrock pricing page.

What is provisioned throughput and how is it billed?

Provisioned throughput reserves dedicated model capacity, measured in model units, billed at an hourly rate whether or not you use it. Each model unit delivers a defined level of throughput, and you can buy units with no commitment, a one-month commitment, or a six-month commitment, with the longer commitments lowering the hourly rate. Provisioned throughput guarantees capacity, removes per-token variability, and is required to serve certain customized or fine-tuned models. Because you pay for the reserved hour regardless of traffic, it only makes financial sense when the units stay busy. The same commitment-versus-on-demand trade governs ordinary compute, which we cover across our commitment management work.

Decision criterionOn-demandProvisioned throughput
Billing unitPer input / output tokenPer model unit per hour
CommitmentNoneNone, 1 month, or 6 months
Pay when idle?No, scales to zeroYes, you pay the reserved hour
Best traffic shapeVariable, spiky, or lowSteady, high, predictable
Throughput guaranteeBest effort, sharedGuaranteed, dedicated
Custom modelsLimitedRequired for some

Verdict: start on-demand for anything variable or unproven, and switch to provisioned throughput only once sustained volume keeps the reserved units busy enough that cost per token drops below the on-demand rate.

When does provisioned throughput become cheaper than on-demand?

Provisioned throughput becomes cheaper at the point where sustained demand keeps the reserved units highly utilized. Because you pay the hourly model-unit rate around the clock, the effective cost per token is that hourly rate divided by the tokens you actually push through in the hour. At low utilization that number is high; as utilization rises it falls below the on-demand per-token rate, and beyond that break-even provisioned throughput wins. The trap is buying reserved capacity on optimistic forecasts: idle model units bill at full price and quietly inflate the AI line, exactly the kind of committed-but-unused waste described in what is GPU utilization and why idle accelerators cost so much.

Paying on-demand rates on a workload that should be reserved?

Our cost audit profiles your Bedrock token volume, finds the break-even between on-demand and provisioned throughput per model, and sizes commitments to your real demand curve so you never pay for idle model units. On the performance model, you pay only from realized savings. No savings, no fee.

Book a cloud cost audit →

How should I decide which Bedrock pricing mode to use?

Decide from your token volume curve, not a single average. Pull your input and output token usage per model over a representative period, look at how steady it is hour to hour, and model both options. If demand is flat and high, provisioned throughput with a commitment matched to your floor of guaranteed load is cheaper, with on-demand absorbing the spikes above it. If demand is spiky or still growing, stay fully on-demand until the pattern stabilizes. This hybrid, a committed base plus on-demand burst, is usually the lowest-cost shape, and sizing it correctly is the same forecasting discipline as forecasting AI and GPU infrastructure spend.

Go deeper · free guide

The AI and GPU Cost Control Guide includes our Bedrock break-even calculator for on-demand versus provisioned throughput, with the model-unit utilization formula. It is the downloadable companion to this article.

Frequently asked questions

How does Amazon Bedrock charge for on-demand usage?

On-demand Bedrock billing charges per token, with separate rates for input and output tokens that vary by model. You pay only for what you process and there is no commitment, which makes on-demand ideal for variable or low-volume workloads. Image and embedding models are billed on their own per-unit metrics.

What is provisioned throughput in Bedrock?

Provisioned throughput reserves dedicated model capacity in units called model units, billed at an hourly rate with optional one-month or six-month commitments that lower the rate. It guarantees throughput and is required for some customized models, and it becomes cheaper than on-demand once your sustained token volume is high enough to keep the reserved units busy.

When is provisioned throughput cheaper than on-demand?

Provisioned throughput is cheaper when sustained, predictable demand keeps the reserved model units highly utilized for most of every hour you pay for. If traffic is spiky or low, the reserved units sit idle and on-demand per-token billing costs less. The break-even is the volume at which the hourly reserved cost divided by your tokens drops below the on-demand per-token rate.

Where do I find current Bedrock prices?

Current Amazon Bedrock prices are on the official AWS Bedrock pricing page, which lists per-token rates per model for on-demand and the hourly model-unit rates and commitment discounts for provisioned throughput. Prices change as models are added, so verify against the live page before modeling cost.

The short version

Amazon Bedrock on-demand bills per token and is cheapest for variable or low traffic; provisioned throughput reserves model units by the hour and wins once steady volume keeps them busy. The right answer is usually a committed base sized to your guaranteed load plus on-demand for the spikes. When you want that break-even calculated and the commitment sized to your real demand, that is exactly what our FinOps implementation service delivers.

Written by Fredrik Filipsson and reviewed by Morten Andersen, applying the See, Cut, Lock, Run method. Independent and vendor neutral. Pricing structure verified against AWS Bedrock documentation as of June 2026; verify current rates before relying on them.

Primary sources & further reading

Cloud pricing and service behavior change frequently. Verify the specifics in this guide against the providers’ own current documentation and the FinOps Foundation: AWS pricing ↗, AWS documentation ↗ and FinOps Foundation Framework ↗. This article also reflects Cloud Cost Room’s hands-on, vendor-neutral engagement experience.

Co-founder of Cloud Cost Room and a FinOps Certified Practitioner, with 20 years in IT and cloud cost optimization across AWS, Azure, Google Cloud and OCI. More about Fredrik →

More from the AI, GPU & ML Cost cluster

See every guide in the AI, GPU & ML Cost cluster →

The Cloud Cost Brief

Cloud pricing moves. We tell you when it matters.

New commitment instruments, FOCUS changes, hyperscaler pricing shifts, and the plays that actually move a bill. No schedule, no filler.

Subscribe · Work email only