How to keep document AI spending under control in your own AWS account
BucketDesk document AI runs on Amazon Bedrock in your AWS account, so AWS bills the tokens to you. Here is what drives the bill, the limits BucketDesk already applies, and five controls you can set in AWS in under an hour.
Because document AI runs in your account, you hold every lever: an email budget for awareness, a budget action as a hard stop, anomaly alerts for spikes, CloudWatch for near-real-time numbers, and roles to decide who can ask. Most teams set up the first two in fifteen minutes and never think about it again.
When someone in your workspace asks BucketDesk a question about a document, the question and the document go to an AI model in Amazon Bedrock, running in your own AWS account through the separate document chat role you created. BucketDesk does not resell tokens or add a margin: AWS bills Bedrock usage to your account at its normal rates.
That keeps your documents in your account and the pricing transparent. It also means the spending controls live where the bill lives: in AWS. This guide shows what drives the cost and how to cap it.
What one question costs
Every question sends three things to the model: the document, your question, and the last few messages of the chat for context. The model then writes an answer. You pay per token for what goes in and what comes out.
The document is sent again with every question, because the model keeps no memory between calls. So the size of the document matters far more than the length of your question. A 20-page contract is roughly 15,000 input tokens; a 300-page manual can be well over 200,000.
A question about that 20-page contract with an 800-token answer costs about $0.0011 on Nova Lite, $0.015 on Nova Pro and $0.042 on GPT-6 Sol. The same question about the 300-page manual costs more than ten times as much. Check [Amazon Bedrock pricing](https://aws.amazon.com/bedrock/pricing/) for current rates; our [Bedrock pricing guide](/blog/amazon-bedrock-pricing-for-document-q-and-a) covers the details.
BucketDesk's models and their AWS list prices per million input and output tokens (US inference profiles, checked September 26, 2026):
- Amazon Nova Lite: $0.06 in, $0.24 out
- GPT-5.6 Luna: $0.22 in, $1.32 out
- Amazon Nova Pro: $0.80 in, $3.20 out
- Claude Haiku 4.5: $1.00 in, $5.00 out
- GPT-6 Sol: $2.20 in, $11.00 out
Limits BucketDesk already applies
BucketDesk applies these limits on its side. They stop runaway usage, but they are not a budget; the five steps after them are.
- Only some people can ask: Document AI is part of the Business plan, and only Owners, Admins and Operators can use it. Viewers cannot.
- Each person is rate limited: up to 12 questions a minute and 100 an hour per person in a workspace.
- Size caps: documents up to 10 MB, questions up to 4,000 characters, and answers up to 1,600 tokens.
- Chats end: a chat session expires after an hour, and nothing calls the model in the background.
- Only approved models: the document chat role our CloudFormation template creates can call only the models listed above, nothing else in Bedrock.
1. Set a Bedrock budget with email alerts
In the AWS Billing console, open Budgets, create a cost budget, and filter it to the Amazon Bedrock service. Set a monthly amount you are comfortable with and add email alerts at, say, 50%, 80% and 100% of actual spend, plus one on forecasted spend.
AWS updates billing data several times a day rather than instantly, so a budget warns you within hours, not seconds. That is fine for most teams; the next step adds a hard stop.
2. Add a budget action that pauses document AI
A budget can do more than send email. Add a budget action to the same budget that, at 100%, applies an IAM policy denying bedrock:InvokeModel and bedrock:InvokeModelWithResponseStream to the BucketDeskDocumentChat role (or whatever name you gave it in the template).
When the action runs, document AI in BucketDesk stops working and people see an error, while browsing, previews and sharing carry on. Next month, or after you raise the budget, remove the policy to switch it back on. You can choose whether the action runs automatically or waits for your approval. AWS Budgets needs permission to apply the policy; the console helps you create that role the first time.
3. Turn on Cost Anomaly Detection
Cost Anomaly Detection (in the Billing console) learns your normal spending and emails you when a service suddenly costs more than usual. Create a monitor for AWS services with an alert threshold that suits your size. It catches a spike, such as someone asking dozens of questions about a very large file, without you picking a number in advance.
4. Watch tokens in CloudWatch
Bedrock publishes Invocations, InputTokenCount and OutputTokenCount to CloudWatch under the AWS/Bedrock namespace, broken down by model. Put them on a dashboard, or create an alarm on InputTokenCount per hour. Unlike billing data, these metrics arrive within minutes.
5. Decide who and what can use AI
- Give the Operator role deliberately: anyone who only needs to read and share files can be a Viewer.
- Switch document chat off per connection: an admin can turn document chat off for a connection in its BucketDesk settings, for example on a bucket of large scanned archives.
- Prefer smaller models for big files: for long manuals and reports, Nova Lite or GPT-5.6 Luna answer most look-up questions at a fraction of the cost.
- Stop it entirely from AWS: removing the document chat role, or setting the template's document chat option to off, ends all AI calls from BucketDesk at once. Your files and the rest of BucketDesk are unaffected.
Starter is free. Deploy a scoped role with CloudFormation, sign in, and browse, without handing anyone an access key.
Primary sources
Discussion
0 comments · open to guests · moderatedLiked this? Get the next article by email. No schedule, no filler, one click to leave.