How Vision System AIs Are Shifting from Classification to Understanding

Analytics from security cameras deliver value but they also generate complaints, with false alarms topping the list.
Published: July 7, 2026

Most commercial security cameras sold in the past three years include some form of on-board artificial intelligence analytics, including person/facial detection or vehicle classification.

Indeed, those on the market will often go beyond these to also include high-speed license plate recognition, crowd forming detection and behavior alerts among the range of heavy-duty AI feature sets that are integrated directly onto these edge devices.

The analytics deliver value but they also generate complaints, with false alarms topping the list. These aren’t necessarily false positives; they just lack context or situational awareness: a camera that can detect a person near a loading dock sends the same alert for a delivery driver on schedule as it would for an unauthorized individual in the middle of the night.

And a people-detection model at a retail entrance will produce an alert stream that, while technically accurate, is operationally useless without additional filtering logic.

SSI Newsletter

In short, object-level classification systems therefore generate high alert volumes, firing on every match – including the routine and irrelevant.

Adding Context to Security Cameras

Previously integrators and end users have been forced to close that gap themselves: building zone rules, tuning sensitivity, scheduling analytics by time of day, and re-tuning to minimize both false positives and false negatives.

This year’s ISC West, however, showed this limitation is being addressed and vision camera AIs are moving past this limitation with multiple vendors demonstrating natural-language video search and custom alert creation in working products.

What could be seen was a new category of on-device agentic AI that can interpret camera feeds from user descriptions written in plain English, and delivering results across hundreds of cameras in seconds: “Alert me if anyone lingers near the jewelry case for more than 90 seconds”; “Alert me if anyone goes around the back after 6pm”; “Alert me if a person falls”.

For these, the integrator did not draw detection zones, select an analytics module, or build time-based scheduling rules. It’s therefore no surprise that 87 percent of security professionals now consider agentic AI integration a priority.

Notably, some manufacturers went further by running generative AI models entirely on the camera’s edge processor, without any cloud dependency. This included fisheye cameras processing free-text detection queries on-device, accepting natural-language descriptions of what to monitor and generating real-time alerts when those conditions appeared in the scene.

This has been enabled through two key advances in processors and language models, with a number of models being produced (including by major LLM companies such as Meta) that are now small enough to run on both near- and far-edge systems. This means potentially sensitive data doesn’t leave the site, AI analysis can run without an internet connection, and latency is reduced to milliseconds.

Scaling

Context-aware systems that understand operational intent produce far fewer nuisance alerts than object-level classifiers. Fewer false alarms also mean fewer callbacks and higher client satisfaction.

Configuration through natural language also reduces commissioning time and allows for remote updates and service agreements built around ongoing tuning, capability updates, and performance reporting. Vertical diversification and scaling also gets easier, with a retail client, warehouse operation, and school district all able to use the same hardware, and differentiating only in the instructions applied at each site

Evaluating the Security Cameras

As we noted above, the evolutions in processor capabilities have been core to this shift from categorization to understanding. As such, it’s important to know that not every camera will support this and the resulting capabilities.

The processor needs enough headroom to run the vision language model (VLM) – typically up to 4 billion parameters for local on-camera processing – alongside standard video encoding and streaming within the camera’s thermal and power envelope.

This is no mean feat, especially given security cameras will often be using power over Ethernet (PoE) and integrators need to evaluate the specific camera’s on-device processing capability and what workloads the edge processor can sustain simultaneously.

Also, as noted above, on-device processing offers several benefits. However, cloud-based processing layered onto existing cameras is another viable option, giving access to more powerful servers and larger language models.

This approach also introduces trade-offs, including latency, bandwidth costs, data residency requirements and subscription-based pricing that can increase the total cost of ownership. So, while there isn’t a universal right approach, the decision needs to take into account the specific client’s operational and privacy requirements.

Realistically, most mid-to-large installations will land on a tiered architecture. What this means in practice will be on-device perception at each camera, with local server-based coordination across camera groups.

This will allow for not only correlated alerting and multi-camera tracking but also enable optional cloud connectivity for fleet-level reporting and model updates. And this ability to design and support multi-tier deployments will open doors to undertake the highest-value projects. 

Conclusion

Object-level classification systems have been effective at identifying when a specific event happens but have been terrible at understanding the context – driving users mad with too many alerts and requiring significant programming and sector-specific expertise that is hard to scale effectively (certainly cost effectively).

The shift to agentic AI will change this and is very much in demand, but not all cameras will be created equally. If the benefits it offers are to be accessed, system integrators need to not only evaluate the sensor, lens and operational constraints, but also pay particular attention to the processor.

Specifically, they must evaluate its ability to run these advanced VLMs right at the edge while still working within the limited power budgets afforded by PoE.

Jerome Gigot is vice president of product marketing at Ambarella.

Strategy & Planning Series
Strategy & Planning Series
Strategy & Planning Series
Strategy & Planning Series