Edge AI + Multimodal Retrieval + Agentic Systems
Real-Time Video Intelligence on Edge Devices
Built an on-device system that watches live video in real time and also lets people search and ask questions across everything it has already seen, with both halves running on embedded hardware instead of in the cloud.
45%
reduction in manual monitoring effort
30 FPS
sustained real-time analytics on embedded hardware
Sub-second search across weeks of recorded video
Answers grounded in retrieved footage rather than model assertion
The problem
Continuous video monitoring does not scale with people. Someone has to watch the feed, attention degrades long before the shift ends, and the events that matter usually get found afterwards rather than as they happen. Pushing everything to the cloud does not solve it either: the bandwidth is not always there, the latency is too high for anything that has to act on what it sees, and in plenty of deployments the footage is not permitted to leave the site at all. So whatever runs has to run locally, on hardware with a fixed power and thermal budget, unattended, for a long time. And once a system has been watching for a few weeks, the recording becomes its own problem, because nobody is going to scrub through days of footage to find the one moment they need.
What I built
Built one system that covers both halves. A real-time path runs perception on the live stream on the device, decides what is worth acting on, and acts, all inside the frame budget. A retrieval path turns what the system has already seen into something searchable, so a person can ask an ordinary question and get back the moments that answer it with the supporting footage attached rather than a bare claim. Both paths work off the same on-device representation of the video, which is what makes it affordable to run them together on embedded hardware instead of standing up a separate indexing tier.
Technical approach
- Perception, decision, and action run as a single loop on the device, sized so the whole cycle fits inside the real-time frame budget instead of quietly degrading to best-effort under load
- Models are compiled and quantized for the target accelerator before deployment, and the video path uses hardware decode so GPU time goes to inference rather than to unpacking frames
- What the system has seen is kept as a compact representation rather than as raw footage, which is what keeps weeks of history searchable on a device with a fixed storage and memory budget
- Search runs in two stages, a fast pass to narrow the field and a slower one to rank what survives, so query latency stays roughly flat as the history grows
- Language model reasoning sits on top of retrieval rather than replacing it, so every answer points back at the footage it came from and a person can check it
- Built to run unattended. It recovers from stream drops, restarts, and power loss on its own, because there is nobody on site to restart it