Machine Learning

Why the Best AI Agents Are the Ones with Eyes

The idea that AI agents need ‘eyes’ – real visual information from the real world, and not just paperwork and dashboards, have become a marker of advancement in construction technology. Yet mostly what has been referred to as an “agent” today is not actually acting based on what it sees. It narrates and alerts while leaving the decisions to a human. Giving a camera does not allow the software to become an agent; it merely gives it more intelligent monitoring capabilities.

That difference matters more on a construction site than almost anywhere else AI is being deployed. A retail agent that detects a stockout and notifies it a few minutes later is a minor inefficiency. On a construction site, however, the conditions are different: there are safety risks that don’t wait for a review cycle, and the cost of a delayed response is much more than a missed metric – it is a person getting injured or even losing his life. So, the question worth asking isn’t whether an agent can see. It’s what that agent’s vision was actually built to do once it sees something.

Two Kinds of Vision

There are two different types of ‘vision’ that has been built into construction AI now, and the industry is overwhelmingly focused on one of them. The first one is ‘documentation vision’, that is, cameras and models trained to record a site, verify progress, and create a visual history. Most current investment and coverage sit here. The second one is ‘protective vision’, that is, systems built not to document a hazard after-the-fact, but to recognize it in the moment and trigger a response before it becomes an incident, like a worker without PPE near moving machinery, an unauthorized entry into a restricted zone, or a structural risk building in real-time.

Most of the industry has gotten very good at proving what already happened on a site. That’s valuable. But it’s a different job from stopping something before it happens, and we’ve been slow to admit those are two different problems.” – Gary Ng, CEO & Co-Founder, viAct

Why Protective Vision Has No “After”

The distinction is not about which agent has superior cameras; rather, it is about what the “eyes” are for. Documentation vision is built to clear up issues post-event: was the incident resolved, is the process closed out, or what has changed in the past week. Protective vision, on the other hand, has no after. It has to recognize risk and respond on the spot. A safety agent that detects the danger five minutes after the event has already failed to fulfil its task. This single difference in purpose alters everything about the way the agent has to be built. 

From Detection to Response

Spotting danger in a video is not the toughest part of the problem. Most of the computer vision systems nowadays can identify a case of missing hard hat or anyone unauthorized in a restricted area easily. What’s harder is what happens in the next second: does the system simply log the event for a report, or does it trigger an actual response to the right person, an automatic escalation, a barrier lockout while the risk is still live. The latter is the real design problem: connecting detections across PPE non-compliance, danger zone breach or human-machinery proximity to real-time response, instead of waiting for report at the end of the day.

A hazard that gets logged correctly but acted on late is still a failure. We don’t measure our systems by detection accuracy alone. We measure them by how fast that detection turns into a person doing something differently.” – Gary Ng

From APAC to the GCC

This distinction plays out concretely across live deployments. At a large construction site in Singapore, an integrated safety deployment, including PPE detection, danger-zone alerts, and machinery tracking, drove a 10x improvement in the site’s overall safety score, moving safety oversight from periodic audits to continuous, real-time monitoring. In Saudi Arabia, a combined AI video analytics and wearable monitoring deployment reduced on-site medical emergencies by 63% and prevented roughly 4,800 lost work hours tied to heat stress, a hazard that, by its nature, has no meaningful ‘after’.

These outcomes share a pattern worth noting: the value came from how quickly recognition turned into a live response, not from the visual record produced afterward. That’s the mechanism protective vision depends on, whether or not the system built on top of it is agentic, semi-automated, or still human-triggered at the response stage.

The Real Measure of Intelligence

Drawing on deployments like these across construction and industrial sites in APAC and the GCC, the real test of an AI agent isn’t camera count or model size, it’s whether the system was built to answer “what happened” or to answer “what do we do right now.” Both are useful. Only one of them can prevent an injury.

The industry keeps asking how smart these systems are. The better question is how fast they turn seeing into doing because on a live site, that gap is where people actually get hurt.” – Gary Ng

Gary Ng

Gary Ng is the CEO and Co-Founder of viAct, an AI-driven company he founded in 2016. With over 10 years of experience in bringing technological innovation to the construction industry, Gary transitioned from a building engineering background to becoming an AI entrepreneur. Before founding viAct, he served as Managing Director of EFI Optitex and was recognized as a top regional executive at Stratasys. He is also a visiting faculty member at The Hong Kong Polytechnic University and an active speaker promoting AI-driven sustainability in workplaces.

Recent Posts

Stellar Converter for EDB Review: The Best EDB to PST Conversion Tool

If you've ever faced the challenge of extracting Exchange mailboxes data from an offline EDB file, you know how painful…

2 months ago

How to Automate Data Validation and Find the Best Tools for Monitoring Research Integrity

As artificial intelligence (AI) and Internet of Things (IoT) accelerate the pace of discovery, research teams are grappling with an…

6 months ago

AI That Can Think Like a Human—How Close Are We?

Artificial Intelligence has come a long way in recent years. From chatbots that answer questions to AI systems that compose…

7 months ago

Protecting What Powers Progress – A Modern Look at Data Security

Modern progress runs on information. Every business, no matter the size or industry, depends on the constant movement of data…

10 months ago

AI Startups and the Legal Risk of Getting It Wrong

Artificial intelligence (AI) is exploding. There were 5,509 AI startups in the US between 2013 and 2023. And according to…

10 months ago

Top-Rated SaaS Financial Management Tools for K-12 Schools

Efficient and accountable financial management is nonnegotiable in today’s K-12 landscape. Outdated, traditional software packages can’t keep pace with the…

11 months ago