Skip to content
All case studies

ML Systems + Inference Optimization

High-Performance Inference and Deployment for Real-Time AI

Rebuilt the inference and deployment path for real-time AI workloads so latency became predictable, deployments became repeatable, and regressions got caught before they reached a device.

2x

inference throughput improvement

60%

latency reduction

50%

reduction in inter-process latency

30%

faster deployment turnaround

The problem

Real-time AI workloads fail differently from batch ones. Average latency can look perfectly healthy while the tail quietly breaks the product, memory creeps until something falls over days into a run, and a model that tested well behaves differently once it is running against live input on the actual hardware. Deployment made it worse. Builds were hard to reproduce, so it was never entirely clear which version of what was running where, which meant a performance regression could not reliably be traced back to the change that caused it.

What I built

Reworked the pipeline end to end. Models are compiled and quantized for the target hardware ahead of time rather than served through a general-purpose runtime. The video path moved onto hardware-accelerated decode and encode so the GPU spends its time on inference instead of moving frames around. Deployments were made reproducible and versioned, and validation runs against real hardware rather than a stand-in, so thermal and power behaviour shows up before a device in the field does something surprising.

Technical approach

  • Models compiled and quantized for the target accelerator instead of served through a general-purpose runtime, which is where most of the throughput came from
  • Hardware-accelerated decode and encode on the video path, so GPU time goes to inference rather than frame handling
  • Batching and scheduling tuned against tail latency rather than average, because a real-time system is judged on its worst frames and nobody notices a good average
  • Inter-process communication reworked to stop copying large frames between stages, removing a cost that scaled with resolution and got worse exactly as cameras got better
  • Deployments versioned and reproducible, so whatever is running on a device traces back to a known build
  • Validation runs against real hardware in the loop, since behaviour under thermal and power limits does not reproduce on a workstation

Visuals

Latency distribution before and after, with the tail called out
Where GPU time goes across decode, inference, and encode
Build to device, with the validation gate in the path
Throughput against resolution, before and after the copy was removed