Post

Cloud OCR Compared: Amazon Textract vs Google Cloud Vision vs Azure AI Document Intelligence

Cloud OCR Compared: Amazon Textract vs Google Cloud Vision vs Azure AI Document Intelligence

Invoices, receipts, scanned contracts and forms still arrive as images or PDFs. Before any AI model can summarize, classify or extract fields from them, you need Optical Character Recognition (OCR) to turn pixels into text. All three major clouds offer a managed OCR service, and the differences matter when you pick one for production.

This post compares Amazon Textract, Google Cloud Vision / Document AI and Azure AI Document Intelligence (formerly Form Recognizer) on the points that usually decide the choice.

The three services at a glance

 Amazon TextractGoogle Cloud Vision / Document AIAzure AI Document Intelligence
Plain text OCRDetectDocumentTextVision TEXT_DETECTION / DOCUMENT_TEXT_DETECTIONprebuilt-read
Layout (tables, key-value)AnalyzeDocument with TABLES, FORMSDocument AI Form Parserprebuilt-layout
Prebuilt domain modelsInvoices, receipts, IDs, lendingInvoice, receipt, ID, procurement processorsInvoice, receipt, ID, W-2, health insurance card, contracts
Custom modelsCustom Queries / AdaptersCustom Document ExtractorCustom extraction and classification models
HandwritingYes (English focus)YesYes
InputJPEG, PNG, PDF, TIFF (async for multi-page)JPEG, PNG, PDF, TIFF, GIFJPEG, PNG, PDF, TIFF, BMP, HEIF, Office documents

Accuracy

On clean printed text all three are close to each other and mistakes are rare. The gaps show up on hard inputs:

  • Low resolution scans and skewed photos: Azure prebuilt-read and Google DOCUMENT_TEXT_DETECTION tend to keep reading order and paragraph structure better than a raw text-detection call.
  • Handwriting: all three support it, but quality depends heavily on language. If you have handwritten non-English forms, run your own sample set through each service before committing.
  • Tables: Textract TABLES and Azure prebuilt-layout return cell-level structures with row/column indexes. Google returns tables through Document AI, not through the plain Vision API.

The practical advice: build a small benchmark from your documents (20 to 50 real pages) and measure character error rate and field accuracy. Vendor benchmarks are not your documents.

Supported languages

  • Textract supports a shorter list (English, Spanish, Italian, Portuguese, French, German) for text extraction.
  • Google Cloud Vision covers the broadest set of printed languages and detects the language automatically.
  • Azure Document Intelligence supports a wide printed-text list and a growing handwriting list; check the language support page for the exact model version you use.

If you process documents in many scripts (for example Devanagari or CJK), Google and Azure are usually the shortlist.

Pricing model

All three bill per page (or per image), with cheaper tiers for plain OCR and pricier ones for layout and prebuilt models.

  • Plain text OCR is typically around $1.50 per 1,000 pages on each provider at the entry tier, dropping with volume.
  • Layout, tables and forms cost several times more per page than plain OCR.
  • Prebuilt domain models (invoice, receipt, ID) are the most expensive per page.

Two things people forget: multi-page PDFs are billed per page, not per file, and asynchronous jobs on AWS have separate request limits per region. Always price your page volume, not your document count.

Security and compliance

  • All three offer encryption at rest and in transit, private endpoints and regional data residency.
  • Azure lets you run Document Intelligence in a container on your own infrastructure for disconnected or regulated scenarios.
  • AWS and Google both allow you to opt out of using your content to improve their services; make that setting explicit in your account before going to production.
  • Check for HIPAA, SOC 2 and ISO 27001 coverage per service and per region; not every prebuilt model is available in every sovereign region.

Related reading on the Azure side: Things to consider before using Azure OpenAI in your organization.

Developer experience

  • Textract has SDKs for every major language and integrates naturally with S3 and Step Functions for batch pipelines. Multi-page documents require the asynchronous Start* / Get* calls.
  • Google gives you a single Vision call for quick OCR and a separate Document AI product for structured extraction; two products means two sets of quotas and pricing pages.
  • Azure exposes everything through one REST API and SDK (azure-ai-formrecognizer / azure-ai-documentintelligence) and has Document Intelligence Studio, a browser tool for testing and labeling that saves a lot of time when building custom models.

Which one should you choose?

  • Already on AWS with S3-based ingestion and mostly English documents: Textract.
  • Many languages, image-heavy inputs, or you need the broadest printed-language coverage: Google Cloud Vision / Document AI.
  • Mixed Office and PDF inputs, need for on-premises containers, or an existing Azure estate: Azure AI Document Intelligence.

Whatever you pick, wrap it behind your own small interface. OCR is a commodity that improves every quarter, and switching providers should be a configuration change, not a rewrite.

Further reading

This post is licensed under CC BY 4.0 by the author.