How to7 min readOctober 5, 2026

How to keep document AI spending under control in your own AWS account

BucketDesk document AI runs on Amazon Bedrock in your AWS account, so AWS bills the tokens to you. Here is what drives the bill, the limits BucketDesk already applies, and five controls you can set in AWS in under an hour.

JeVaughn Ferguson
Founder, developer
The short version

Because document AI runs in your account, you hold every lever: an email budget for awareness, a budget action as a hard stop, anomaly alerts for spikes, CloudWatch for near-real-time numbers, and roles to decide who can ask. Most teams set up the first two in fifteen minutes and never think about it again.

When someone in your workspace asks BucketDesk a question about a document, the question and the document go to an AI model in Amazon Bedrock, running in your own AWS account through the separate document chat role you created. BucketDesk does not resell tokens or add a margin: AWS bills Bedrock usage to your account at its normal rates.

That keeps your documents in your account and the pricing transparent. It also means the spending controls live where the bill lives: in AWS. This guide shows what drives the cost and how to cap it.

What one question costs

Every question sends three things to the model: the document, your question, and the last few messages of the chat for context. The model then writes an answer. You pay per token for what goes in and what comes out.

The document is sent again with every question, because the model keeps no memory between calls. So the size of the document matters far more than the length of your question. A 20-page contract is roughly 15,000 input tokens; a 300-page manual can be well over 200,000.

A question about that 20-page contract with an 800-token answer costs about $0.0011 on Nova Lite, $0.015 on Nova Pro and $0.042 on GPT-6 Sol. The same question about the 300-page manual costs more than ten times as much. Check [Amazon Bedrock pricing](https://aws.amazon.com/bedrock/pricing/) for current rates; our [Bedrock pricing guide](/blog/amazon-bedrock-pricing-for-document-q-and-a) covers the details.

BucketDesk's models and their AWS list prices per million input and output tokens (US inference profiles, checked September 26, 2026):

  • Amazon Nova Lite: $0.06 in, $0.24 out
  • GPT-5.6 Luna: $0.22 in, $1.32 out
  • Amazon Nova Pro: $0.80 in, $3.20 out
  • Claude Haiku 4.5: $1.00 in, $5.00 out
  • GPT-6 Sol: $2.20 in, $11.00 out

Limits BucketDesk already applies

BucketDesk applies these limits on its side. They stop runaway usage, but they are not a budget; the five steps after them are.

  • Only some people can ask: Document AI is part of the Business plan, and only Owners, Admins and Operators can use it. Viewers cannot.
  • Each person is rate limited: up to 12 questions a minute and 100 an hour per person in a workspace.
  • Size caps: documents up to 10 MB, questions up to 4,000 characters, and answers up to 1,600 tokens.
  • Chats end: a chat session expires after an hour, and nothing calls the model in the background.
  • Only approved models: the document chat role our CloudFormation template creates can call only the models listed above, nothing else in Bedrock.

1. Set a Bedrock budget with email alerts

In the AWS Billing console, open Budgets, create a cost budget, and filter it to the Amazon Bedrock service. Set a monthly amount you are comfortable with and add email alerts at, say, 50%, 80% and 100% of actual spend, plus one on forecasted spend.

AWS updates billing data several times a day rather than instantly, so a budget warns you within hours, not seconds. That is fine for most teams; the next step adds a hard stop.

2. Add a budget action that pauses document AI

A budget can do more than send email. Add a budget action to the same budget that, at 100%, applies an IAM policy denying bedrock:InvokeModel and bedrock:InvokeModelWithResponseStream to the BucketDeskDocumentChat role (or whatever name you gave it in the template).

When the action runs, document AI in BucketDesk stops working and people see an error, while browsing, previews and sharing carry on. Next month, or after you raise the budget, remove the policy to switch it back on. You can choose whether the action runs automatically or waits for your approval. AWS Budgets needs permission to apply the policy; the console helps you create that role the first time.

3. Turn on Cost Anomaly Detection

Cost Anomaly Detection (in the Billing console) learns your normal spending and emails you when a service suddenly costs more than usual. Create a monitor for AWS services with an alert threshold that suits your size. It catches a spike, such as someone asking dozens of questions about a very large file, without you picking a number in advance.

4. Watch tokens in CloudWatch

Bedrock publishes Invocations, InputTokenCount and OutputTokenCount to CloudWatch under the AWS/Bedrock namespace, broken down by model. Put them on a dashboard, or create an alarm on InputTokenCount per hour. Unlike billing data, these metrics arrive within minutes.

5. Decide who and what can use AI

  • Give the Operator role deliberately: anyone who only needs to read and share files can be a Viewer.
  • Switch document chat off per connection: an admin can turn document chat off for a connection in its BucketDesk settings, for example on a bucket of large scanned archives.
  • Prefer smaller models for big files: for long manuals and reports, Nova Lite or GPT-5.6 Luna answer most look-up questions at a fraction of the cost.
  • Stop it entirely from AWS: removing the document chat role, or setting the template's document chat option to off, ends all AI calls from BucketDesk at once. Your files and the rest of BucketDesk are unaffected.
Try it in BucketDesk

Starter is free. Deploy a scoped role with CloudFormation, sign in, and browse, without handing anyone an access key.

Connect a bucket

Primary sources

Discussion

0 comments · open to guests · moderated
Comments appear after a quick review.

Liked this? Get the next article by email. No schedule, no filler, one click to leave.

Keep reading

All writing →
Security8 min

Amazon Bedrock data retention modes: none, default or aws_review for business documents

When you ask AI on AWS about a company document, AWS now lets you choose whether that conversation may be kept. Here is what each choice means and who should make it.data_retention_mode none/default/aws_review/provider_data_share/inherit, project → account → model default resolution, allowed_modes and ValidationException, bedrock:PutAccountDataRetention with the DataRetentionMode condition key, and destination-Region storage under cross-Region inference.

How to8 min

S3 lifecycle rules explained: examples for versioned buckets, delete markers and storage class transitions

Lifecycle rules are how a bucket tidies itself: moving old files somewhere cheaper and clearing out copies nobody needs. Here is how to write them without deleting the wrong thing.Filter And blocks with prefix, tags and ObjectSizeGreaterThan, NoncurrentVersionExpiration with NewerNoncurrentVersions, ExpiredObjectDeleteMarker, AbortIncompleteMultipartUpload, conflict precedence, midnight-UTC rounding and transition request costs.

Product decisions8 min

Amazon Bedrock AgentCore for document agents on S3: what it runs, what it costs, and when you don't need it

AgentCore is AWS's toolkit for running your own AI agents. Here is what it would take to point one at your company's files, and when a simpler tool already answers the question.Runtime microVM sessions (8 hours, 100 MB payloads, 2 vCPU / 8 GB), execution roles that need your own s3:GetObject, Gateway turning Lambda and OpenAPI into MCP tools, vCPU-hour and GB-hour billing with free I/O wait, and Quick's MCP client.