Google Cloud Vision
Google Cloud Vision provides visual model and visual recognition capabilities via API, targeting developers and enterprise workflows that can be integrated into applications.
Who it is for: Developers needing to integrate visual recognition capabilities into applications, and product and engineering teams needing to process images, documents, or videos
Core capabilities
- Pretrained Visual Recognition API: Based on pretrained computer vision models, it provides image tagging, face and landmark detection, OCR, and explicit content labeling via REST and RPC APIs.
- Image Text Detection: Identifies and extracts UTF-8 text from images.
- Object Localization and Safe Search: Provides general labels and bounding box annotations for multiple objects in an image, and returns likelihood scores for explicit content categories.
- Video Analysis: Processes and analyzes stored and streaming video, identifying objects, locations, and actions within them.
- Visual Captioning: Generates image descriptions using Imagen Visual Captioning, and supports metadata, automatic captions, and product or visual asset descriptions.
Pricing
New customers can receive up to $300 in free credit; Cloud Vision API has 1,000 feature units free per month; Video Intelligence API has the first 1,000 minutes free per month; the official website also lists pricing signals for Imagen multimodal embeddings at $0.0001 per input image and Imagen visual captioning at $0.0015 per image.
Updated 2026-08-18