Operations8 min readOctober 3, 2026

Amazon Bedrock pricing for document Q&A: what a question about a PDF in S3 really costs

How Amazon Bedrock bills a question about a contract or invoice stored in S3: tokens per page, the on-demand, batch, Flex and Priority tiers, prompt caching, global versus regional inference profiles, and the extra lines that appear once you add guardrails or a knowledge base.

JeVaughn Ferguson
Founder, developer
The short version

The cost of a question about a PDF is mostly input tokens, roughly one to three thousand per page, more with citations or newer tokenizers. Keep interactive questions on Standard, move bulk jobs to batch or Flex for half price, cache documents that get follow-up questions, and use a global profile when residency allows. Add guardrails, knowledge bases and logging with their own meters in mind, and tag an application inference profile per app so you can see where the money goes.

"Amazon Bedrock pricing" is one of the most searched Bedrock topics, and most of the pricing page is a long table of per-model token rates. For a team that wants to ask questions about documents in S3, the rates matter less than the mechanics: how many tokens one page becomes, which tier the request runs on, and whether the same document is read twice.

This article explains those mechanics. It deliberately doesn't reproduce per-model rates, which change and differ by Region; take them from the Bedrock pricing page or the AWS Pricing Calculator.

The basic bill: tokens in, tokens out

On-demand Bedrock charges per input token and per output token, priced per model and per Region, with no commitment. For document Q&A the input is the document plus the question and any instructions, and the output is the answer. Input is usually the larger number by far, because a contract is long and an answer is short. Output tokens cost more each, so a request that asks for a long summary shifts the balance.

Reading the file from S3 adds a GET request at normal S3 rates, which is negligible next to the tokens.

How many tokens is a page?

Anthropic estimates 1,500 to 3,000 tokens per page of text for Claude, and there is no separate PDF fee. How the PDF is sent changes the count a lot. Through the Converse API without citations, Claude receives extracted text only, about 1,000 tokens for a three-page PDF in Anthropic's example. With citations turned on, each page is also processed as an image, about 7,000 tokens for the same file, and Claude can read charts and scanned pages.

Two caveats. These figures are documented for the Converse integration with earlier Claude models, and Anthropic notes its newest models use a tokenizer that produces roughly 30% more tokens for the same text. Measure with your own files: the Converse response returns inputTokens and outputTokens for every call.

Tiers: Standard, Flex, Priority, Reserved and batch

Bedrock now has service tiers chosen per request with the optional service_tier parameter (default, flex, priority or reserved), and the pricing page states the differences as percentages of Standard, the default. Not every model supports every tier.

A person waiting for an answer belongs on Standard. A nightly job that extracts the due date from every invoice in a prefix belongs on batch or Flex and costs half as much. The tier that actually served each request is shown in the response, in CloudTrail, and in CloudWatch as ResolvedServiceTier, so you can confirm the discount applied.

  • Standard: the normal on-demand price.
  • Flex: 50% cheaper, for work that can tolerate slower or deferred processing.
  • Priority: a 75% premium, for latency-sensitive traffic when demand is high.
  • Reserved: a fixed price for a 1 or 3 month reservation of tokens-per-minute capacity, arranged through your AWS account team.
  • Batch inference: a separate API at 50% below on-demand for select models, for jobs such as summarizing an entire folder overnight.

Prompt caching: don't pay full price to reread the same contract

If people ask several questions about the same document, prompt caching lets the repeated part of the prompt, the document, be read from cache instead of processed again. Cache reads are billed at a much lower rate than normal input, and cache writes can cost more than normal input. Anthropic lists a cache hit at a tenth of the base input price for most Claude models, and a write at 1.25 times for the 5-minute cache or 2 times for the 1-hour one; check Bedrock's own pricing for your model.

The default cache lifetime is 5 minutes, reset on every hit, and recent Claude models also accept a 1-hour TTL. Caching works with on-demand requests, not batch, and hits are not guaranteed. The usage block reports cacheReadInputTokens and cacheWriteInputTokens, so you can see whether it is working.

For a follow-up question asked a minute later, a cached document is the biggest single saving available. For a document opened once and asked one question, caching adds a write premium and saves nothing.

Global versus regional inference profiles

Newer Claude models on Bedrock are called through inference profiles. A geographic profile keeps processing within a geography such as the US or EU, at standard pricing. A global profile can route to any supported commercial Region and is about 10% cheaper, according to AWS. Pricing is based on the Region you call from, and CloudTrail records where each request was processed.

The choice is a data residency question first and a price question second. If client documents must stay in the EU, the 10% is not available to you.

The lines that appear later

Document Q&A rarely stays a bare model call. Each addition below brings its own meter.

  • Guardrails: billed per 1,000 text units (up to 1,000 characters each) for every policy you turn on, so cost grows with the size of the document you check against.
  • Knowledge bases: an index over the bucket brings vector store costs, plus embedding tokens for every file ingested and re-ingested. OpenSearch Serverless classic collections have a minimum capacity charge; S3 Vectors is aimed at infrequent queries.
  • Bedrock Data Automation: per-page charges for structured extraction from documents.
  • Logging: model invocation logs land in CloudWatch Logs or S3 at those services' rates, and they contain prompts, which means document text.

Seeing the cost per application

Bedrock charges show up as one service on the bill. To split them, create an application inference profile for each app or team on top of the model or system profile you use, tag it, activate the tags as cost allocation tags in Billing, and call Converse with the profile ARN. The price is the same as the underlying model; you just get a cost line you can attribute. Replace the source ARN below with the system inference profile you already use.

aws bedrock create-inference-profile \
  --inference-profile-name contracts-qa \
  --model-source copyFrom=arn:aws:bedrock:us-east-1::inference-profile/us.anthropic.example-model \
  --tags key=app,value=contracts-qa

How this compares with Amazon Q and BucketDesk

Amazon Q Business was priced per user plus index capacity, and it is closed to new customers; Amazon Quick, its replacement, is also per user with an account fee. With either, the token cost is hidden inside the subscription, and you pay for the index whether or not anyone asks a question.

BucketDesk Document AI keeps the Bedrock model calls in your AWS account, so you see them as ordinary Bedrock usage and pay AWS directly, with no index or capacity charge. Each conversation is scoped to one document, which keeps the input to the file someone actually opened, and answers come with a citation to the passage. Document AI is part of the Business plan.

Try it in BucketDesk

Starter is free. Deploy a scoped role with CloudFormation, sign in, and browse, without handing anyone an access key.

Connect a bucket

Primary sources

Discussion

0 comments · open to guests · moderated
Comments appear after a quick review.

Liked this? Get the next article by email. No schedule, no filler, one click to leave.

Keep reading

All writing →