What a document chat costs on Amazon Bedrock, by model, with prompt caching
We measured the same chat about one document on Amazon Nova Lite, Nova Pro and Claude Haiku 4.5. With Bedrock prompt caching, follow-up questions on Claude cost about a tenth of the first one, and Haiku ends up cheaper than Nova Pro from the second question.
With prompt caching, the price of a document chat depends on how many questions you ask, not just on the model's list price. On Claude Haiku 4.5, follow-ups cost about a tenth of the first question, which makes it cheaper than Nova Pro from the second question. On Nova, caching did not help for documents in our test, so Nova Lite and Nova Pro remain the cheaper choice for one-off questions.
When you ask a question about a document with the Amazon Bedrock Converse API, the whole document goes to the model with every question. A model has no memory between calls. Ask ten questions about a 30-page report and you pay for reading those 30 pages ten times.
Prompt caching changes that. Bedrock keeps the part of the request that does not change, such as the system prompt and the document, and the next question reads it from the cache at a fraction of the price. How much you save depends on the model you pick, and on one model we tested, caching saved nothing at all.
This article walks through what we measured on 2026-10-06 while adding prompt caching to BucketDesk's Document AI, and what it means for choosing a model.
The test
We sent the same two-question conversation about the same document to three models through their US inference profiles in us-east-1. The document was plain text, about 11,000 tokens long, which is roughly a 25 to 30 page text-heavy report. The system prompt and the questions were identical for every model.
The ten-question figures repeat the measured follow-up cost; in a longer chat each follow-up also carries the earlier turns, which adds a little. Costs below use the per-million-token prices in BucketDesk's model catalog: AWS's US prices for Nova, and Anthropic's list prices for Claude, so the Claude figures are close estimates rather than an invoice. Output was short in every case, so nearly all of the cost is the document.
Across a thousand ten-question chats, the costs below add up to about $7 on Nova Lite, $96 on Nova Pro, $114 on Haiku without caching, and $26 on Haiku with caching.
- Amazon Nova Lite: 0.07¢ for every question, about 0.7¢ for ten.
- Amazon Nova Pro: 0.96¢ for every question, about 9.6¢ for ten.
- Claude Haiku 4.5 without caching: 1.14¢ for every question, about 11.4¢ for ten.
- Claude Haiku 4.5 with caching: 1.42¢ for the first question, 0.13¢ for each follow-up, about 2.6¢ for ten.
Why the first Claude question costs more
Caching is not free. On Claude, writing the document into the cache is billed at 1.25 times the normal input rate, and reading it back is billed at 0.1 times. So the first question costs about a quarter more than it would without caching, and every follow-up within the cache window costs about a tenth.
Our Bedrock usage numbers show it directly. On the first question Haiku reported 11,238 tokens written to the cache and none read. On the second it reported the same 11,238 tokens read from the cache and none written, plus 47 new tokens for the question and the earlier turn.
The break-even point is the second question. A one-question chat on Claude costs slightly more with caching. A two-question chat already costs less, and the gap grows with every question after that.
Haiku overtakes Nova Pro from the second question
Claude Haiku 4.5 lists at $1.00 per million input tokens and Nova Pro at $0.80, so on paper Nova Pro is the cheaper of the two. With caching that only holds for the first question.
Nova Lite stays the cheapest by a wide margin at every point. It is a smaller model, so whether it is good enough is a question of answer quality rather than price.
- After one question: Haiku 1.42¢, Nova Pro 0.96¢.
- After two questions: Haiku 1.55¢, Nova Pro 1.92¢.
- After ten questions: Haiku about 2.6¢, Nova Pro about 9.6¢.
What happened on Nova
Nova models accept a cache point in the Converse API, and Bedrock documents implicit caching for Nova text prompts. In our test neither helped once a document was in the request. With an explicit cache point, Nova Lite and Nova Pro both reported about 11,900 tokens written to the cache on every question and none read back, so the follow-up cost the same as the first question. Without the cache point, they reported no cache activity at all.
Answers were correct either way. We left the cache point off for Nova in BucketDesk, since writing to a cache that is never read gains nothing. If AWS changes this, the table above will change with it.
Where the cache point goes
A cache point marks the end of the part of the request Bedrock should keep. Everything before it has to be identical from one request to the next, so the document has to come first in the conversation, not attached to the latest question.
Bedrock also refused a cache point placed directly after a document block, with a ValidationException saying a cache point cannot be inserted after the preceding content block. Putting a short fixed text block between the document and the cache point fixed it.
In the script below, the second question should print a non-zero "read from cache" number. If it prints zero, something before the cache point changed between the two calls, or the cache expired.
cat > cached_document_chat.py <<'EOF'
import boto3
MODEL_ID = "us.anthropic.claude-haiku-4-5-20251001-v1:0"
client = boto3.client("bedrock-runtime", region_name="us-east-1")
with open("report.pdf", "rb") as f:
document = {"format": "pdf", "name": "report", "source": {"bytes": f.read()}}
history = []
for question in ["What is the total budget?", "Which team owns the largest line item?"]:
history.append({"role": "user", "content": [{"text": question}]})
messages = [dict(m, content=list(m["content"])) for m in history]
# The document leads the first message, then a fixed line, then the cache point.
messages[0]["content"][:0] = [
{"document": document},
{"text": "The document is above."},
{"cachePoint": {"type": "default"}},
]
response = client.converse(
modelId=MODEL_ID,
system=[{"text": "Answer only from the document."}],
messages=messages,
)
answer = response["output"]["message"]["content"][0]["text"]
usage = response["usage"]
print(answer)
print(" written to cache:", usage.get("cacheWriteInputTokens", 0),
"read from cache:", usage.get("cacheReadInputTokens", 0))
history.append({"role": "assistant", "content": [{"text": answer}]})
EOF
python3 cached_document_chat.pyWhen caching does not help
- Short documents: Claude Haiku 4.5 only caches a prefix of at least 4,096 tokens, and Claude Sonnet 5 needs 1,024. Below that the request still works but nothing is cached. A one-page memo will cost the same either way.
- Long pauses: The default cache lives for 5 minutes, and each cache hit resets the timer. If someone reads for ten minutes before the next question, that question writes the cache again at 1.25 times the price.
- One question per document: The write premium makes a single question about a quarter more expensive on Claude.
- Changing the start of the request: Any change to the system prompt or the document, even a different file name, means a new prefix and a new cache write.
- Cross-region inference: AWS notes that inference profiles can route requests to other Regions at busy times, which can cause extra cache writes.
Which model to pick
- One quick question about a file: Nova Lite or Nova Pro. Caching never comes into play, and Nova's lower input price wins.
- A real back-and-forth on a long document: Claude Haiku 4.5. From the second question it is cheaper than Nova Pro and keeps getting cheaper.
- Short documents: pick on answer quality. Under Haiku's 4,096-token minimum there is no caching to factor in.
Doing this from the BucketDesk dashboard
BucketDesk's Document AI now puts the document first in every request and adds the cache point for Claude models, so follow-up questions in the same chat read the document from the cache automatically. There is nothing to set up. Bedrock usage is billed to your own AWS account, and the cache lives in Bedrock in that account for a few minutes. BucketDesk never stores the document.
Nova Pro remains the default model. Once Claude Haiku 4.5 shows as ready on your connection, you can switch to it in the chat's model picker when you expect more than a question or two.
Document AI on the features pagePlans and the Business trial
Starter is free. Deploy a scoped role with CloudFormation, sign in, and browse, without handing anyone an access key.
Primary sources
Discussion
0 comments · open to guests · moderatedLiked this? Get the next article by email. No schedule, no filler, one click to leave.