Computer Vision AI tools

Computer Vision platforms enable machines to understand and process visual information using artificial intelligence. These tools provide capabilities for image and video analysis, object detection, and visual data understanding across various applications and industries.

38 verified AI-first sites in Computer Vision.

Viso

Enterprise computer vision platform. Features no-code vision AI development, automated model training, and deployment tools. Includes real-time processing and edge deployment. Supports industrial applications.

Pricing: Custom enterprise pricing. Contact for solutions.

Voxel51

Open-source computer vision platform. Features dataset curation, model evaluation, and visual debugging tools. Includes data visualization and quality metrics. Supports ML model development and deployment workflows.

Pricing: Open-source core. Enterprise: Contact for pricing.

Plainsight

Vision AI that turns existing cameras into operational intelligence. Watches physical processes across stores, plants, and sites to measure execution, flag exceptions, and trigger follow-up without new hardware.

Pricing: Free trial available. Enterprise pricing based on usage

NetLux

IQGeo visual AI for telecom and utility field ops (formerly Deepomatic Lens). Real-time photo analysis for network asset capture, construction QA, and closeout. Deployed by major fiber and utility operators for first-time-right fieldwork.

Pricing: Contact IQGeo for pricing.

Remark Vision

Video analytics platform specializing in computer vision and facial recognition. Features include real-time video content analysis, crowd density monitoring, behavioral analytics, and anomaly detection. Provides enterprise-grade solutions for security, access control, and public safety applications with advanced identity recognition capabilities.

Pricing: Enterprise pricing with custom solutions.

Groundlight

Computer vision platform combining ML models with human-in-the-loop validation for real-time visual inspections. Features natural language queries, immediate deployment without pre-existing datasets, and automatic model tuning with 24/7 human annotation support.

Pricing: Contact for pricing. Free developer documentation available.

Picsellia

Enterprise MLOps platform for computer vision featuring comprehensive data management, model training, and deployment capabilities. Includes integrated labeling tools, experiment tracking, and model monitoring with ISO 27001:2022 certification.

Pricing: Community Edition available free. Contact for enterprise pricing.

Twelve Labs

Video understanding models and API. Marengo embeds visual, audio, speech, and on-screen text into one searchable index; Pegasus generates grounded summaries, answers, and descriptions from video. Used by media companies and enterprises to search and analyze large video archives. Raised $100M in 2026 from NEA, NAVER, and Amazon.

Pricing: Free tier (under 10 hours indexing). Developer: Under 10,000 hours indexing. Enterprise: Unlimited indexing with dedicated infrastructure.

Matrice

Vision AI Factory for orchestrating computer vision across any camera and compute environment. Features Streaming Gateway for universal camera ingestion, Inference Engine for low-latency AI processing, drift monitoring with automatic retraining, and priority-based orchestration. Privacy-first with on-premise deployment, GDPR and AI Act compliance. Serves energy, retail, manufacturing, and smart city applications.

Pricing: Enterprise pricing. Contact for solutions.

alwaysAI

Edge computer vision platform for enterprise deployment. Features data management, model training (BYOM or custom), edgeIQ Python library for development, and real-time inferencing with analytics console. Serves manufacturing, mining, energy, logistics, warehousing, and retail industries. Trusted by NVIDIA, Qualcomm, and Cognizant.

Pricing: Project-based pricing with unlimited inferencing. Contact for demo.

Accurately

End-to-end computer vision platform. Annotate, train, deploy in one place. AI-powered auto-labeling 10x faster; collaborative labeling, one-click model training. Deploy to web browser or cloud. Built by Datawow. Free tier available.

Pricing: Free tier. Contact for paid plans.

Ultralytics

YOLO ecosystem and tooling for training, validation, and deploying computer vision models. Provides open-source libraries, pretrained models, and enterprise offerings for detection, segmentation, classification, and pose across cloud and edge.

Pricing: Open-source core; Ultralytics HUB and enterprise plans; see site for current pricing.

Supervisely

End-to-end computer vision platform for curating data, AI-assisted labeling, and building production models. Supports images, video, 3D point clouds, and medical imagery with collaboration, custom models via SDK, and deployment workflows.

Pricing: Community and team tiers; enterprise options; see site for current plans.

Matroid

Enterprise no-code computer vision platform for automated visual inspection, defect detection, SOP verification, and live video analytics. Designed to work across existing camera infrastructure and industrial environments without requiring custom model coding for every use case.

Pricing: Enterprise pricing; request demo for deployment and licensing options.

Elorian

Frontier AI lab building native multimodal models for visual reasoning. Instead of translating images into text before reasoning, Elorian's models reason directly over visual, spatial, and physical information for tasks in engineering, robotics, design, and science. Founded in 2025 by former Google DeepMind and Apple researchers (Andrew Dai, Yinfei Yang); ~$55M seed from Menlo Ventures, Altimeter, and Striker.

Pricing: Early-stage research lab; no public product or API pricing yet.

Lightly

Computer vision data curation and labeling suite. Embedding-based selection, LightlyStudio for curation and QA, and LightlyTrain for self-supervised pretraining to cut annotation cost and improve dataset quality.

Pricing: Open-source core; cloud and enterprise plans; see site for pricing.

Amazon Rekognition

AWS managed vision API for image and video analysis. Object and scene detection, face analysis, text in image, content moderation, and custom labels without building a full CV stack from scratch.

Pricing: Pay-as-you-go API pricing; free tier available; see AWS pricing page.

Google Cloud Vision

Google Cloud Vision AI for image understanding. Label detection, OCR, landmark and logo recognition, SafeSearch, and AutoML custom models integrated with Vertex AI and Google Cloud workflows.

Pricing: Pay-as-you-go API pricing; free monthly units; see Google Cloud pricing.

Azure Vision

Microsoft Azure Vision (Foundry Tools) for image analysis, OCR, object detection, and spatial understanding via prebuilt APIs. Cloud vision service for apps and agents without training custom models from scratch.

Pricing: Pay-as-you-go Azure pricing; free tier available; see Microsoft pricing.

Ximilar

Visual AI platform for image recognition, visual search, and custom computer vision models. No-code training for classification and detection with REST APIs for retail, collecting, and industrial visual search.

Pricing: Free tier and paid API plans; see ximilar.com for current pricing.

Meta Segment Anything 3

Meta's Segment Anything Model 3 for promptable segmentation in images and video. Open research models and demos for object segmentation that underpin many modern computer vision labeling and editing workflows.

Pricing: Open research models; see Meta AI for licenses and access.

LandingLens

LandingAI deep-learning computer vision platform for industrial visual inspection. Train and deploy custom defect-detection models for manufacturing quality without traditional CV expertise.

Pricing: Cloud and enterprise plans; see LandingAI for current pricing.

Roboflow

End-to-end computer vision platform used by 1M+ developers: dataset versioning, labeling, training, and deployment from edge to cloud. Workflows, Inference server, and Universe models for custom detection and segmentation pipelines.

Pricing: Free tier; Core and Enterprise plans on site.

NVIDIA Metropolis

NVIDIA vision AI platform and developer stack for video analytics, perception, and smart spaces. SDKs, pretrained models, and reference workflows to build and deploy production computer vision on NVIDIA GPUs at edge and data center.

Pricing: Developer tools free; enterprise/GPU programs via NVIDIA.

Flock Safety

AI license-plate cameras and a searchable vehicle network for police and private sites. Machine-vision hits on plate, make, model, and color; OS Investigate adds English prompts over those logs plus records and commercial identity data.

Pricing: Agency and private-network contracts; contact Flock Safety.

Reka Vision

Visual intelligence platform for understanding and searching images and long video. Provides temporal recognition, multimodal reasoning, highlight extraction, and real-time monitoring through enterprise APIs.

Pricing: Enterprise and developer pricing; contact Reka.

Coactive

Visual data intelligence platform that turns unstructured image and video into searchable, labeled datasets. Helps media and AI teams discover content, curate training data, and build computer vision workflows.

Pricing: Usage-based and enterprise pricing; contact Coactive.

Spot

AI video platform that turns existing cameras into agents for security, safety, and operations. On-edge recording plus cloud search and real-time actions across 1,000+ customer sites without rip-and-replace hardware.

Pricing: Business and enterprise plans; see Spot.

Superb

Vision MLOps platform from labeling and curation through training, deployment, and video analytics. Used by Fortune 500 manufacturers and mobility teams; NVIDIA Physical AI partner in Korea with 100+ enterprise deployments.

Pricing: Enterprise platform; contact Superb.

Camio

AI video monitoring on standard cameras. Searches and alerts on people, vehicles, and events across multi-site camera networks without proprietary hardware lock-in.

Pricing: Subscription plans; see Camio.

Moondream

Open vision-language model for on-device and edge visual Q&A. Lightweight multimodal weights for describing images, answering visual questions, and embedding vision into apps without heavyweight cloud VLMs.

Pricing: Open weights and hosted API options; see Moondream.

CVEDIA

AI video analytics platform for detection, tracking, and situational awareness on camera streams. Privacy-forward vision models for smart cities, security, and industrial monitoring without heavy custom training.

Pricing: Commercial platform plans; contact CVEDIA.

Irisity

AI video analytics platform (including AgentVi heritage) for real-time surveillance intelligence. Detects people, vehicles, and behaviors across camera networks for security and operations teams.

Pricing: Enterprise video analytics; contact Irisity.

Vaidio

AI vision platform that turns camera video into searchable events and alerts. Object detection, recognition, and operational insights for enterprise security and facilities monitoring.

Pricing: Commercial plans; see Vaidio.

viisights

Behavioral AI video understanding for security and smart spaces. Recognizes complex activities and situations in live video beyond simple object detection.

Pricing: Enterprise licensing; contact viisights.

NVIDIA TAO

NVIDIA Train-Adapt-Optimize toolkit for building and fine-tuning vision AI models. Transfer learning, pretrained models, and deployment paths into NVIDIA Metropolis and edge inference stacks.

Pricing: Free toolkit; NVIDIA hardware and cloud usage may apply.

Valossa

Video analysis AI for video-to-text, search, captions, and clip discovery. Multimodal understanding that indexes spoken and visual content for media and archives.

Pricing: API and commercial plans; see Valossa.

Oosto

Real-time facial recognition and vision AI for security and access use cases. Camera-based identity and watchlist matching designed for live video environments.

Pricing: Enterprise licensing; contact Oosto.