https://www.asmag.com/showpost/36098.aspx
INSIGHTS
Security camera AI moves from seeing objects to understanding behavior
Security camera AI moves from seeing objects to understanding behavior
Comments from Eluviant and Genetec suggest the physical security industry is unlikely to converge on one type of AI model soon.

Security camera AI moves from seeing objects to understanding behavior

Date: 2026/08/24
Source: Prasanth Aby Thomas, Consultant Editor
Artificial intelligence in video surveillance is entering a phase in which some of the biggest gains may come from understanding how events develop over time rather than simply identifying what appears in individual images.
 
That change is beginning to alter how vendors think about video analytics and the infrastructure needed to run them. It also creates a practical issue for systems integrators. More capable models can extract more context from surveillance footage, but those gains have to be delivered across large camera estates without pushing hardware requirements beyond what customers can justify.
 
Comments from Eluviant and Genetec suggest the industry is unlikely to converge on one type of AI model over the next three to five years. Broader models with stronger reasoning capabilities could take on more complex interpretation, while specialized computer vision models remain useful when an application has tight performance or processing constraints.

From detecting objects to understanding events

Callum Wilson, co-founder and CEO of Eluviant, sees one of the most important advances already taking place in the move beyond analysis of individual images.

“The first big shift is already underway: the move from still-image analysis to genuine moving-image understanding,” Wilson said.
 
Eluviant launched Aurora Flow in July, describing it as a large-scale visual language model that understands video sequences and can place time-based events and behaviors in context.
 
The distinction matters because many security incidents are difficult to understand from a single frame.
 
Someone standing near an ATM may be engaged in normal activity. The sequence of actions around the machine may provide the context needed to recognize tampering. A still image in a retail store can show a person holding merchandise, but behavior associated with shoplifting is more apparent when movement is considered across a longer sequence.
 
Fights present a similar problem. The meaning lies in what people are doing over time, not simply in how they appear at one instant.
 
Wilson said analyzing the whole sequence allows Aurora Flow to identify what is happening with a higher degree of confidence.
 
For video surveillance, this changes the problem AI is being asked to solve. Conventional analytics have often focused on detecting an object or classifying something visible in a scene. A sequence-aware model can consider how activity develops and use that temporal context to distinguish situations that may look similar in an isolated image.
 
The shift could make analytics more useful in situations where intent or behavior cannot be inferred reliably from a snapshot. But greater analytical depth also raises the amount of processing needed, making architecture an increasingly important part of the AI discussion.

Foundation models will not replace specialized analytics

The move toward more capable models does not mean narrower computer vision models are likely to disappear.
 
Laurent Villeneuve, Senior Manager, Product and Industry Marketing at Genetec, expects foundation models and task-specific computer vision models to remain in use together.
 
“The two approaches are likely to coexist,” Villeneuve said.
Foundation models, he said, can expand the general reasoning and multimodal capabilities available to physical security systems. Their value becomes particularly relevant when a system needs to interpret information from video, audio, text and other sources together.
 
That creates room for AI systems that can reason across more context than a conventional analytic built around one predefined visual condition.
In a physical security environment, however, broader capability is only useful when it can be applied within the performance limits of the deployment.
 
Villeneuve said task-specific models will remain important where organizations require high accuracy or low latency. They will also continue to matter when computing resources are limited or when processing must take place at the edge.
 
“The most practical direction is a hybrid architecture that applies each type of model according to the requirements of the use case,” Villeneuve said.
 
For integrators, that makes the question of whether a platform “has AI” increasingly inadequate.
 
A more useful assessment is likely to focus on which model performs a required function, where that model runs and what infrastructure is needed to support it at the customer’s intended scale.
 
A general-purpose model may be appropriate when a security system has to interpret complex activity or combine multiple kinds of information. A smaller model may make more sense when a narrowly defined detection has to run quickly and consistently on constrained hardware.
Model size and sophistication therefore do not necessarily translate directly into operational value.

Compute becomes the practical bottleneck

This distinction becomes particularly important because Wilson argues that computing capacity, rather than the capabilities of current AI models, is one of the biggest constraints facing advanced video intelligence.
 
“The biggest limitation right now is the compute, not the models,” he said.
The problem is amplified in enterprise surveillance.
 
An AI application that performs well on a small set of video streams may be much harder to deploy economically when the same customer operates hundreds or thousands of cameras.
 
Wilson said a large part of Eluviant’s research and development work is focused on scalability, with the goal of making advanced video intelligence reliable across large deployments without creating what he described as “astronomical hardware costs.”
 
Sequence understanding can increase the processing burden because the model is considering movement over time instead of examining isolated frames. More demanding models can also require substantially more computational resources than simpler analytics.
 
Applying that level of processing continuously to every stream can make the infrastructure requirement a central part of the business case.
 
For systems integrators, this means an AI demonstration on a handful of feeds may say relatively little about the economics of a full deployment.
 
A proof of concept can show whether a system recognizes a particular behavior. It may not show what happens to processing requirements when the capability is extended across a large surveillance network.
 
That makes scalability testing increasingly important.
 
An analytics system can appear responsive when only a few streams are under analysis, yet an enterprise customer may need the same performance across hundreds of cameras. Integrators therefore have to consider not only whether an AI feature works, but how its infrastructure requirements change as more feeds are added.
 
The architecture around the model effectively becomes part of the capability.

Running deeper AI only when it is needed

Eluviant is addressing the compute problem by trying to avoid applying the most resource-intensive processing continuously.
 
Wilson said the company’s self-learning unusual behavior layer first surfaces moments that may matter. Deeper, more compute-hungry models are then run on relevant material instead of analyzing every frame from every stream at the same level.
 
“A large part of our answer is architectural,” Wilson said.
The approach treats compute as a resource that should be directed toward footage where additional interpretation is justified.
 
A lighter analytical layer performs the initial filtering, reducing the amount of video that reaches more demanding models.
 
The concept is particularly relevant to large camera deployments. If advanced reasoning is needed for only a fraction of recorded activity, applying the heaviest processing indiscriminately could waste computing resources without adding useful security information.
 
Unusual activity is not necessarily malicious activity. In this architecture, its role is to identify moments that warrant further analysis.
 
That also creates a balancing act. If an initial analytical layer passes too much footage onward, much of the computing advantage is lost. If it filters too aggressively, potentially relevant activity may never receive deeper interpretation.

The interview responses do not suggest that a single model will solve every part of this process. Instead, they point toward security systems in which different analytical layers perform different jobs depending on the requirements of the application and the resources available.

That principle aligns closely with Genetec’s view of hybrid AI. A foundation model could be called upon where more sophisticated interpretation is required, while task-specific analytics continue handling functions where efficiency, speed or predictable performance matters more. 

Edge AI will remain a question of trade-offs

The edge is one area where those decisions are likely to remain particularly visible.
Processing close to the camera or video source can reduce dependence on centralized infrastructure, but edge devices operate within tighter computing limits.
 
That means model selection remains an engineering decision rather than a contest to deploy the largest or most sophisticated AI system available.
 
A task-specific model may remain preferable for an edge application that has to carry out a narrow function with low latency. More computationally demanding interpretation can be reserved for infrastructure capable of supporting it when the additional context is justified.
 
For integrators, this could influence how AI is specified in future surveillance projects.
 
Customers may want to add more intelligence to existing camera estates without replacing large portions of the infrastructure. Whether that is viable will depend partly on how efficiently the analytics platform distributes workloads between cameras, edge devices and other computing resources.
 
It could also make comparisons between products more complicated.
Two systems may both claim advanced video understanding while making very different assumptions about when high-compute models operate.
 
A feature that looks similar at the application level could therefore have a significantly different infrastructure footprint when deployed across hundreds of cameras.
Integrators may need to look beneath the feature description and examine how often advanced models are invoked, where processing takes place and what happens to hardware requirements as camera counts rise.

Cheaper compute could accelerate adoption

The compute challenge also explains why progress in camera AI over the next few years may depend partly on improvements in hardware economics.
 
Wilson expects step changes in the availability of powerful yet affordable compute to unlock new possibilities for companies working on computer vision and accelerate mass adoption.
 
Cheaper processing could make demanding models practical across more video streams and increase the range of analytics that can be supported within a given infrastructure budget.
 
His comments nevertheless suggest vendors cannot simply wait for processor costs to decline.
 
Surveillance is a continuous workload, and enterprise environments can involve very large camera estates. Architectural efficiency will remain important even as hardware becomes more capable.
 
That is likely to keep hybrid approaches relevant. Some models can handle routine or narrowly defined tasks continuously. More computationally intensive systems can be reserved for events where their additional reasoning capability has value. The result may be less like a single AI engine replacing existing analytics and more like a hierarchy of models working together. 

Integrators will need to look beyond AI claims

The growth of foundation models creates another challenge for the security industry: separating useful capabilities from broad promises about what AI may eventually achieve.
 
Villeneuve said the industry needs to set realistic expectations about the technology and focus on practical operational benefits.
 
That caution becomes increasingly relevant as more general models make it possible to demonstrate forms of reasoning that were difficult for earlier generations of video analytics.
 
An impressive demonstration does not automatically establish an operational need.
Security teams still have to decide whether a system improves how an event is detected, understood or handled, and whether the infrastructure required to provide that capability is proportionate to its value.
 
Over the next several years, advances in moving-image understanding could make surveillance AI substantially better at interpreting behavior that cannot be captured in a single frame.
 
At the same time, task-specific models are unlikely to disappear. Their efficiency and predictable performance remain important in applications where latency, edge processing or computing limitations are decisive.
 
For security systems integrators, the more consequential change may therefore be how AI systems are evaluated.
 
The headline capability of a model is only one part of the equation. Integrators will increasingly have to understand how a platform allocates processing, when deeper reasoning is activated and whether the available compute can sustain the capability across the intended camera estate.
 
If affordable processing improves as Wilson expects, some of those limitations will ease.
 
Until then, the ability to use computing resources selectively could determine whether sophisticated video AI remains an impressive demonstration or becomes practical across real-world security networks.
 


https://www.asmag.com/project/the_manpower_survey/
Related Articles
Midwest Technology CEO on solving remote video security challenges with Luminys
Midwest Technology CEO on solving remote video security challenges with Luminys
AI video analytics raises new questions over evidence integrity and customer data
AI video analytics raises new questions over evidence integrity and customer data
Exclusive Interview: ONVIF Ambassador Roberto Licari on the new Profile V; open cloud ecosystems vs. vendor lock-in
Exclusive Interview: ONVIF Ambassador Roberto Licari on the new Profile V; open cloud ecosystems vs. vendor lock-in