AGENTHUBINTEL TERMINAL

Google Cloud Vision

Google Cloud Vision provides visual model and visual recognition capabilities via API, targeting developers and enterprise workflows that can be integrated into applications.

Who it is for: Developers needing to integrate visual recognition capabilities into applications, and product and engineering teams needing to process images, documents, or videos

Core capabilities

  • Pretrained Visual Recognition API: Based on pretrained computer vision models, it provides image tagging, face and landmark detection, OCR, and explicit content labeling via REST and RPC APIs.
  • Image Text Detection: Identifies and extracts UTF-8 text from images.
  • Object Localization and Safe Search: Provides general labels and bounding box annotations for multiple objects in an image, and returns likelihood scores for explicit content categories.
  • Video Analysis: Processes and analyzes stored and streaming video, identifying objects, locations, and actions within them.
  • Visual Captioning: Generates image descriptions using Imagen Visual Captioning, and supports metadata, automatic captions, and product or visual asset descriptions.

Pricing

New customers can receive up to $300 in free credit; Cloud Vision API has 1,000 feature units free per month; Video Intelligence API has the first 1,000 minutes free per month; the official website also lists pricing signals for Imagen multimodal embeddings at $0.0001 per input image and Imagen visual captioning at $0.0015 per image.

Updated 2026-08-18

Visit Google Cloud Vision