Amazon Bedrock Data Automation for documents in S3: what it does, what it costs, and when to use something else
Bedrock Data Automation reads a file from S3 and writes structured JSON back to S3, either a standard layout of the document or the fields you define in a blueprint. Here are the limits, the API, the per-page pricing, and how it compares with Textract, Claude on Bedrock, Amazon Q and BucketDesk Document AI.
Bedrock Data Automation turns documents in S3 into JSON in S3: a standard layout at about $0.010 a page, or blueprint fields with confidence scores at $0.040 a page for up to 30 fields. It is asynchronous up to 500 MB and 3,000 pages, synchronous up to 50 MB and 10 pages, and always needs a cross-Region data automation profile. Reach for it when the same fields come out of many known document types; use Textract for raw OCR and a model, or BucketDesk Document AI, when a person has one question about one file.
Amazon Bedrock Data Automation, usually shortened to BDA, is AWS's answer to a common request: point a service at a folder of documents in S3 and get structured data back, without writing an OCR pipeline and a prompt for every document type. It reads documents, images, audio and video, and writes JSON to an S3 location you choose.
Search interest for it is still small but steady, mostly from people who have already used Textract or Claude on Bedrock and want to know whether BDA replaces either. It mostly sits between them. This guide covers what it returns, the limits and API, the pricing mechanics, and where it fits next to the tools a business team actually uses.
Standard output and custom output
BDA has two kinds of result. Standard output is what you get with no configuration: the document broken into pages, elements and words in reading order, with text as Markdown by default (or plain text, HTML, or CSV for tables) and a normalised bounding box for each element. You can switch on generative fields that add a 10-word and a 250-word summary of the document and captions for figures.
Custom output is driven by a blueprint. A blueprint is a list of fields you want, each with a name, a type, a plain-language description, and whether the value is read explicitly from the page or inferred. AWS ships catalog blueprints for common documents such as invoices and identity documents, and you can write your own. Custom output fields come back with a confidence score and the page they came from.
Blueprints live in a project. A project can hold up to 40 document blueprints and BDA picks the best match for each file, so one project can take a mixed folder of invoices, W-2s and bank statements and route each to the right field list.
- Standard output: layout, text, tables, figures, bounding boxes, optional summaries.
- Custom output: your fields, with confidence and page number.
- Blueprints: up to 100 fields for async calls, 15 for sync.
- Projects: up to 40 document blueprints, routed per file.
Limits that decide sync or async
The asynchronous API reads from S3 and writes to S3. It accepts PDF, TIFF, JPEG, PNG and DOCX up to 500 MB, and up to 3,000 pages with document splitting on. DOCX files are converted to PDF first, so page numbers in the output don't map back to the Word file. The console is stricter, at 200 MB and 20 pages.
The synchronous API, InvokeDataAutomation, returns results inline. It accepts PDF, TIFF, JPEG and PNG (not DOCX) up to 50 MB and 10 pages. It is meant for a single short document in a request path, not a backlog.
Document text is supported in English, German, Spanish, French, Italian and Portuguese. Vertical text is not.
Amazon Textract or Claude on Amazon Bedrock: which one should read the documents in your S3 bucket?
- Async: 500 MB, 3,000 pages, DOCX allowed, S3 in and out.
- Sync: 50 MB, 10 pages, no DOCX, inline result.
- Six languages, no vertical text.
Profiles and Regions
Every BDA call needs a data automation profile ARN, because BDA always uses cross-Region inference. There is no extra charge for it. In the US the profile is us.data-automation-v1, which spreads work across us-east-1, us-east-2 and us-west-2. There are equivalent eu, apac and na profiles, a us-gov profile, and an opt-in global profile available only from ap-southeast-1.
AWS says data stays stored in the source Region, while requests and results may be processed in the other Regions of the profile's geography. If your compliance review cares about where processing happens, not only where data rests, the profile is the setting to read.
Example: process a PDF from S3 and read the result
The call below sends one PDF to a project and polls until the job finishes. Results land under the output prefix. In production, set notificationConfiguration so BDA emits an EventBridge event when the job completes, instead of polling. The stage is LIVE or DEVELOPMENT, which lets you test a changed blueprint without touching the live one.
cat > bda_invoice.py <<'EOF'
import time
import boto3
REGION = "us-east-1"
ACCOUNT = boto3.client("sts").get_caller_identity()["Account"]
bda = boto3.client("bedrock-data-automation-runtime", region_name=REGION)
job = bda.invoke_data_automation_async(
inputConfiguration={"s3Uri": "s3://amzn-s3-demo-bucket/incoming/inv-10422.pdf"},
outputConfiguration={"s3Uri": "s3://amzn-s3-demo-bucket/bda-output/"},
dataAutomationConfiguration={
"dataAutomationProjectArn": f"arn:aws:bedrock:{REGION}:{ACCOUNT}:data-automation-project/invoices",
"stage": "LIVE",
},
dataAutomationProfileArn=f"arn:aws:bedrock:{REGION}:{ACCOUNT}:data-automation-profile/us.data-automation-v1",
)
while True:
status = bda.get_data_automation_status(invocationArn=job["invocationArn"])
if status["status"] not in ("Created", "InProgress"):
break
time.sleep(5)
print(status["status"], status.get("outputConfiguration", {}).get("s3Uri"))
EOF
python3 bda_invoice.pyS3 event notifications: run code when a file is uploaded to Amazon S3
What it costs
BDA bills per page for documents. The Bedrock pricing page gives its document prices through worked examples rather than a table: custom output is $0.040 a page for a blueprint with up to 30 fields, plus $0.0005 a page for each field beyond 30, so a 40-field blueprint is $0.045 a page. The Knowledge Bases example on the same page prices standard output at $0.010 a page. Images with custom output are $0.005 each, audio $0.006 a minute, and video standard output $0.050 a minute. Check the page before budgeting; these figures were read on October 8, 2026 and no Region is stated.
For comparison, Textract's Forms, Tables and Queries together run about $0.07 a page, and a single question to Claude Haiku about a 30-page report is around a cent. BDA is cheaper than stacking Textract features when you need many fields from every page, and more expensive than a model when a person only needs one answer from one file.
Amazon Bedrock pricing for document Q&A
- Custom output: $0.040 a page up to 30 fields, $0.0005 per extra field per page.
- Standard output: $0.010 a page in AWS's Knowledge Bases example.
- Paid per page processed, whether or not anyone reads the result.
BDA as a Knowledge Bases parser
Bedrock Knowledge Bases can use BDA as its parser instead of the default text parser or a foundation-model parser. That helps when the PDFs in your bucket carry their meaning in tables, charts and figures that a plain text extraction flattens. The catch is that choosing BDA applies it to every PDF in the data source, including the text-only ones, and each page is billed.
Amazon Bedrock Knowledge Bases on S3, or just ask the document?
- Good for chart- and table-heavy PDFs.
- Applies to every PDF in the data source.
- Adds a per-page cost to every sync.
When to use BDA, Textract or a model
Use BDA when you have many documents of a few known types and want the same fields out of each, with confidence scores, without writing a Textract-plus-prompt pipeline yourself. Mixed folders are where the project routing pays off.
Use Textract when you need raw OCR geometry, the specialised expense, ID or lending APIs, or the cheapest possible text detection at $0.0015 a page. Use Claude or another model on Bedrock directly when the question isn't known in advance: a person asking whether a contract allows early termination doesn't need a blueprint, they need an answer with a page citation.
Which Amazon Bedrock model should answer questions about your documents?
- BDA: repeatable fields from known document types, at volume.
- Textract: OCR geometry, specialised forms, lowest per-page text cost.
- A model: one person, one document, an open question.
Amazon Q, Amazon Quick and BucketDesk Document AI
None of the three services above is something a finance or legal team opens in a browser. At that level the comparison is between assistants. Amazon Q Business, which indexed S3 buckets behind a chat app, is closed to new customers, and AWS points new work to Amazon Quick. Both crawl and index the bucket before answering, which suits questions across a whole library and adds an index to pay for and keep in sync. Neither returns per-field confidence the way BDA does.
BucketDesk Document AI takes the per-document route. Someone opens a file in a connected bucket, asks in plain language, and gets an answer with a citation that highlights the supporting passage. The model runs on Bedrock in your own AWS account, nothing is indexed in advance, and BucketDesk keeps no file content after the session. It doesn't extract fields from ten thousand invoices; that is BDA's job. It replaces the script or the download someone would otherwise need to ask one contract one question.
Document AI on the features page
- Amazon Q Business: closed to new customers, indexes the bucket.
- Amazon Quick: the AWS successor, also index-first.
- BDA: batch extraction into JSON, per page.
- BucketDesk Document AI: one open document, cited answer, Bedrock in your account.
Starter is free. Deploy a scoped role with CloudFormation, sign in, and browse, without handing anyone an access key.
Primary sources
- What is Amazon Bedrock Data Automation? ↗
- How Bedrock Data Automation works ↗
- Bedrock Data Automation limits ↗
- Blueprints ↗
- Standard output for documents ↗
- Cross-Region support in Bedrock Data Automation ↗
- InvokeDataAutomationAsync API reference ↗
- GetDataAutomationStatus API reference ↗
- Advanced parsing options for Knowledge Bases ↗
- Amazon Bedrock pricing ↗
- Amazon Q Business availability change ↗
Discussion
0 comments · open to guests · moderatedLiked this? Get the next article by email. No schedule, no filler, one click to leave.