Skip to content
All case studies

Edge AI + Multimodal Retrieval + Agentic Systems

Real-Time Video Intelligence on Edge Devices

Built an on-device system that watches live video in real time and also lets people search and ask questions across everything it has already seen, with both halves running on embedded hardware instead of in the cloud.

45%

reduction in manual monitoring effort

30 FPS

sustained real-time analytics on embedded hardware

Sub-second search across weeks of recorded video

Answers grounded in retrieved footage rather than model assertion

The problem

Continuous video monitoring does not scale with people. Someone has to watch the feed, attention degrades long before the shift ends, and the events that matter usually get found afterwards rather than as they happen. Pushing everything to the cloud does not solve it either: the bandwidth is not always there, the latency is too high for anything that has to act on what it sees, and in plenty of deployments the footage is not permitted to leave the site at all. So whatever runs has to run locally, on hardware with a fixed power and thermal budget, unattended, for a long time. And once a system has been watching for a few weeks, the recording becomes its own problem, because nobody is going to scrub through days of footage to find the one moment they need.

What I built

Built one system that covers both halves. A real-time path runs perception on the live stream on the device, decides what is worth acting on, and acts, all inside the frame budget. A retrieval path turns what the system has already seen into something searchable, so a person can ask an ordinary question and get back the moments that answer it with the supporting footage attached rather than a bare claim. Both paths work off the same on-device representation of the video, which is what makes it affordable to run them together on embedded hardware instead of standing up a separate indexing tier.

Technical approach

  • Perception, decision, and action run as a single loop on the device, sized so the whole cycle fits inside the real-time frame budget instead of quietly degrading to best-effort under load
  • Models are compiled and quantized for the target accelerator before deployment, and the video path uses hardware decode so GPU time goes to inference rather than to unpacking frames
  • What the system has seen is kept as a compact representation rather than as raw footage, which is what keeps weeks of history searchable on a device with a fixed storage and memory budget
  • Search runs in two stages, a fast pass to narrow the field and a slower one to rank what survives, so query latency stays roughly flat as the history grows
  • Language model reasoning sits on top of retrieval rather than replacing it, so every answer points back at the footage it came from and a person can check it
  • Built to run unattended. It recovers from stream drops, restarts, and power loss on its own, because there is nobody on site to restart it

Visuals

The two paths side by side: live perception loop and retrieval over history
Where the frame budget goes across decode, inference, and decision
Query to answer: narrow, rank, then reason over what came back
What stays on the device and what never leaves it