Which Amazon Bedrock model should answer questions about your documents?
Every model on Bedrock sees the same 4.5 MB document through the Converse API, so the choice comes down to citations, prompt caching minimums, data retention, Region routing, model lifecycle and price. Here is how Claude and Amazon Nova differ on each, as of October 2026.
For document questions on Bedrock the model choice is about citations, caching, retention, routing, lifecycle and price, not benchmarks. Claude gives you page citations, caches from 512 tokens on the newest models, offers a 1-hour cache and runs cheaper than Nova Pro over a conversation; Nova Lite is the cheapest single question; Nova's cache stops at 20,000 tokens; and every model's retention modes, Region routing and EOL date are on its card. Keep the model ID in configuration and re-check those six points when you swap.
Searches for "amazon bedrock models" and "amazon bedrock claude" have both climbed this year, and the question behind them is usually practical: which model should I put behind the thing that answers questions about our PDFs? Benchmarks do not settle it. The Bedrock Converse API gives every model the same document block with the same 4.5 MB limit, so the differences that matter for document work sit elsewhere.
This guide walks through those differences for the two families most teams compare, Anthropic Claude and Amazon Nova, with facts taken from the Bedrock model cards on 2026-10-06. Model lists and limits change, so treat it as a checklist to re-run rather than a ranking.
The models on offer today
Bedrock's catalogue groups models by provider. Anthropic currently lists three Claude generations: the 5.x line, including Claude Sonnet 5.5, Claude Opus 5.5 and Claude Fable 5.1, the 4.x line down to Claude Haiku 4.5, and two 3.x models still served. Amazon's own text models are Nova 2 Lite plus Nova Micro, Lite and Pro. OpenAI, Meta, Mistral, Google and others are on the same catalogue and work through the same Converse call.
For document questions the realistic shortlist is a small, cheap model for look-up questions and a larger one for judgement. On the Anthropic side that is Haiku 4.5 and a Sonnet; on the Amazon side, Nova Lite or Nova 2 Lite and Nova Pro.
Ask questions about S3 documents with Claude on Amazon Bedrock
- Anthropic: Claude 5.x (Sonnet 5.5, Opus 5.5, Fable 5.1 and others), 4.x (Opus 4.8 down to Haiku 4.5), 3.x (3.5 Haiku, 3 Haiku).
- Amazon: Nova 2 Lite, Nova Micro, Nova Lite, Nova Pro.
- All of them take the same Converse document block: pdf, csv, doc, docx, xls, xlsx, html, txt, md, up to 4.5 MB each and five per message.
Context window matters less than you think
Claude Haiku 4.5 has a 200K-token context window and Nova 2 Lite has 1M. Both numbers dwarf a single document: a 30-page text-heavy report is about 11,000 tokens, and the 4.5 MB document cap stops you long before either window fills.
Where the window does matter is a long conversation about several documents at once, or a chat that runs for fifty turns, because every earlier turn travels with each new question. If your product is one open file and a handful of questions, every model on the shortlist has room to spare. If it is a research assistant over a folder, the larger window is a real advantage.
- Single document under 4.5 MB: context window is not the constraint.
- Five documents plus a long chat: it starts to be.
Citations are a Claude feature
A document answer is only useful if someone can check it. Bedrock's document block has a citations setting, and when it is on the answer comes back as text plus a list of citations, each pointing at the page range the passage came from. AWS's documentation and Anthropic's describe this for Claude models.
If you are considering another provider, test for citationsContent in the response before you design a reader around it. An answer without a page reference is a claim; an answer with one is something a reviewer can verify in ten seconds.
Amazon Textract or Claude on Amazon Bedrock for documents in S3
Prompt caching: the minimums decide whether it works at all
Every question about a document re-sends the document. Prompt caching lets the model read it from a cache on follow-up questions at a fraction of the input price. On Bedrock, whether that helps depends on the model's minimum cacheable prefix.
The newest Claude models cache from 512 tokens, so almost any document qualifies. Claude Haiku 4.5, along with Opus 4.7 and 4.6, requires 4,096 tokens before a checkpoint takes effect, which means a short two-page letter is below the line and simply does not cache. Claude models support a 5-minute or 1-hour TTL and up to four checkpoints across system, messages and tools.
Amazon Nova works differently. Nova offers implicit caching for all text prompts without any configuration, and Nova 2 Lite also accepts explicit checkpoints from 1,000 tokens, but Nova caching is capped at 20,000 tokens and uses a 5-minute TTL only. A document longer than roughly 50 text pages will not fit in the Nova cache, while it caches fine on Claude. Writing to the cache can be billed above the normal input rate, so a single-question chat gains nothing; the saving starts on the second question.
What a document chat costs on Amazon Bedrock, by model, with prompt caching
- Claude Sonnet 5.5, Opus 5.5, Fable 5.1, Opus 5: 512-token minimum.
- Claude Sonnet 5, Opus 4.8, Sonnet 4.6: 1,024.
- Claude Haiku 4.5, Opus 4.7, Opus 4.6: 4,096.
- Nova 2 Lite: 1,000 minimum, 20,000 maximum, 5-minute TTL only.
- Claude: 5-minute or 1-hour TTL, four checkpoints.
Data retention and where the request runs
Bedrock lets you set a data retention mode per Region: none, default or aws_review. Some newer models are only offered under aws_review, and a request to them from an account set to none fails. Our data retention guide covers the modes; the point here is that the model and the mode have to be chosen together, and the model card tells you which modes a model allows.
Routing is the other quiet constraint. On the bedrock-runtime endpoint, Claude Haiku 4.5 and Nova 2 Lite are both served through geo or global inference profiles, which can route a request to another Region inside the geography or anywhere in the world. If a client contract requires single-Region processing, Haiku 4.5 is available in-Region through the bedrock-mantle endpoint in a handful of Regions, and you should check the card for any other model before promising residency.
Amazon Bedrock data retention modes for business documents
- Pick the model and the retention mode together; some models require aws_review.
- Geo and global inference profiles do not provide single-Region residency.
- Check the model card's in-Region column before committing to a Region.
Model lifecycle: plan the swap before you need it
Each model card carries an end-of-life floor. Claude Haiku 4.5, launched in October 2025, has an EOL no sooner than 16 October 2026 followed by a legacy period of at least six months. Nova 2 Lite, launched December 2025, has an EOL no sooner than 2 December 2026. Those are floors, not announcements, but they mean the model you pick this quarter may need replacing within a year or two.
The good news is that Converse makes the swap a configuration change: the document block, the question and the citations request are identical across models. What you have to re-test after a swap is the parts that differ above: whether citations still come back, whether the cache minimum still fits your typical document, and whether the retention mode still allows the new model.
- Store the model ID in configuration, not code.
- Re-run citation, caching and retention checks after every swap.
- Watch the "EOL no sooner than" date on the model card.
What it costs per question
We measured the same two-question chat about an 11,000-token document across three models while adding caching to BucketDesk's Document AI. Nova Lite cost about 0.07¢ a question. Nova Pro cost about 0.96¢. Claude Haiku 4.5 cost 1.14¢ without caching, and with caching 1.42¢ for the first question then 0.13¢ for each follow-up, which makes Haiku cheaper than Nova Pro from the second question onward.
One billing detail worth knowing: Claude is a third-party model sold through AWS Marketplace, so its charges appear in AWS Cost Explorer under the model provider rather than under Amazon Bedrock. If you allocate AI spend by application, tag the inference profile rather than searching the Bedrock line.
Amazon Bedrock pricing for document Q&AControl document AI spending in your AWS account
- Cheapest per question: Nova Lite.
- Cheapest for a conversation with follow-ups: Claude Haiku 4.5 with caching.
- Claude charges show under Anthropic in Cost Explorer, not under Bedrock.
A short decision list
Start from what the answer has to carry and how people will use it, then let the table above pick the model.
- Answers must show the page they came from: Claude, with citations enabled.
- Quick look-ups in a one-question session: Nova Lite, nothing to cache.
- Long conversations about one document: a Claude model with a 512 or 1,024-token minimum, or Haiku 4.5 if your documents are over about 4,000 tokens.
- Documents over 20,000 tokens with repeated questions: Claude, because Nova's cache will not hold them.
- Zero data retention required: a model whose card allows the none mode.
- Single-Region processing required: check the in-Region column; Haiku 4.5 via bedrock-mantle qualifies.
Choosing the model in BucketDesk
BucketDesk Document AI runs on Bedrock in your own AWS account, so all of the above applies to you rather than to us. It uses Amazon Nova Pro by default and lets a workspace admin choose a different Bedrock model in settings. Your account's retention mode applies, so a model that requires aws_review will not work while the account is set to none, and Bedrock usage is billed to you per request. Document AI is part of the Business plan, which has a 14-day trial.
Document AI on the features pagePlans and the Business trial
Starter is free. Deploy a scoped role with CloudFormation, sign in, and browse, without handing anyone an access key.
Primary sources
Discussion
0 comments · open to guests · moderatedLiked this? Get the next article by email. No schedule, no filler, one click to leave.