Cloud OCR Compared: Amazon Textract vs Google Cloud Vision vs Azure AI Document Intelligence
Invoices, receipts, scanned contracts and forms still arrive as images or PDFs. Before any AI model can summarize, classify or extract fields from them, you need Optical Character Recognition (OCR) to turn pixels into text. All three major clouds offer a managed OCR service, and the differences matter when you pick one for production.
This post compares Amazon Textract, Google Cloud Vision / Document AI and Azure AI Document Intelligence (formerly Form Recognizer) on the points that usually decide the choice.
The three services at a glance
| Amazon Textract | Google Cloud Vision / Document AI | Azure AI Document Intelligence | |
|---|---|---|---|
| Plain text OCR | DetectDocumentText | Vision TEXT_DETECTION / DOCUMENT_TEXT_DETECTION | prebuilt-read |
| Layout (tables, key-value) | AnalyzeDocument with TABLES, FORMS | Document AI Form Parser | prebuilt-layout |
| Prebuilt domain models | Invoices, receipts, IDs, lending | Invoice, receipt, ID, procurement processors | Invoice, receipt, ID, W-2, health insurance card, contracts |
| Custom models | Custom Queries / Adapters | Custom Document Extractor | Custom extraction and classification models |
| Handwriting | Yes (English focus) | Yes | Yes |
| Input | JPEG, PNG, PDF, TIFF (async for multi-page) | JPEG, PNG, PDF, TIFF, GIF | JPEG, PNG, PDF, TIFF, BMP, HEIF, Office documents |
Accuracy
On clean printed text all three are close to each other and mistakes are rare. The gaps show up on hard inputs:
- Low resolution scans and skewed photos: Azure
prebuilt-readand GoogleDOCUMENT_TEXT_DETECTIONtend to keep reading order and paragraph structure better than a raw text-detection call. - Handwriting: all three support it, but quality depends heavily on language. If you have handwritten non-English forms, run your own sample set through each service before committing.
- Tables: Textract
TABLESand Azureprebuilt-layoutreturn cell-level structures with row/column indexes. Google returns tables through Document AI, not through the plain Vision API.
The practical advice: build a small benchmark from your documents (20 to 50 real pages) and measure character error rate and field accuracy. Vendor benchmarks are not your documents.
Supported languages
- Textract supports a shorter list (English, Spanish, Italian, Portuguese, French, German) for text extraction.
- Google Cloud Vision covers the broadest set of printed languages and detects the language automatically.
- Azure Document Intelligence supports a wide printed-text list and a growing handwriting list; check the language support page for the exact model version you use.
If you process documents in many scripts (for example Devanagari or CJK), Google and Azure are usually the shortlist.
Pricing model
All three bill per page (or per image), with cheaper tiers for plain OCR and pricier ones for layout and prebuilt models.
- Plain text OCR is typically around $1.50 per 1,000 pages on each provider at the entry tier, dropping with volume.
- Layout, tables and forms cost several times more per page than plain OCR.
- Prebuilt domain models (invoice, receipt, ID) are the most expensive per page.
Two things people forget: multi-page PDFs are billed per page, not per file, and asynchronous jobs on AWS have separate request limits per region. Always price your page volume, not your document count.
Security and compliance
- All three offer encryption at rest and in transit, private endpoints and regional data residency.
- Azure lets you run Document Intelligence in a container on your own infrastructure for disconnected or regulated scenarios.
- AWS and Google both allow you to opt out of using your content to improve their services; make that setting explicit in your account before going to production.
- Check for HIPAA, SOC 2 and ISO 27001 coverage per service and per region; not every prebuilt model is available in every sovereign region.
Related reading on the Azure side: Things to consider before using Azure OpenAI in your organization.
Developer experience
- Textract has SDKs for every major language and integrates naturally with S3 and Step Functions for batch pipelines. Multi-page documents require the asynchronous
Start*/Get*calls. - Google gives you a single Vision call for quick OCR and a separate Document AI product for structured extraction; two products means two sets of quotas and pricing pages.
- Azure exposes everything through one REST API and SDK (
azure-ai-formrecognizer/azure-ai-documentintelligence) and has Document Intelligence Studio, a browser tool for testing and labeling that saves a lot of time when building custom models.
Which one should you choose?
- Already on AWS with S3-based ingestion and mostly English documents: Textract.
- Many languages, image-heavy inputs, or you need the broadest printed-language coverage: Google Cloud Vision / Document AI.
- Mixed Office and PDF inputs, need for on-premises containers, or an existing Azure estate: Azure AI Document Intelligence.
Whatever you pick, wrap it behind your own small interface. OCR is a commodity that improves every quarter, and switching providers should be a configuration change, not a rewrite.
