ML Systems + Inference Optimization
High-Performance Inference and Deployment for Real-Time AI
Rebuilt the inference and deployment path for real-time AI workloads so latency became predictable, deployments became repeatable, and regressions got caught before they reached a device.
2x
inference throughput improvement
60%
latency reduction
50%
reduction in inter-process latency
30%
faster deployment turnaround
The problem
Real-time AI workloads fail differently from batch ones. Average latency can look perfectly healthy while the tail quietly breaks the product, memory creeps until something falls over days into a run, and a model that tested well behaves differently once it is running against live input on the actual hardware. Deployment made it worse. Builds were hard to reproduce, so it was never entirely clear which version of what was running where, which meant a performance regression could not reliably be traced back to the change that caused it.
What I built
Reworked the pipeline end to end. Models are compiled and quantized for the target hardware ahead of time rather than served through a general-purpose runtime. The video path moved onto hardware-accelerated decode and encode so the GPU spends its time on inference instead of moving frames around. Deployments were made reproducible and versioned, and validation runs against real hardware rather than a stand-in, so thermal and power behaviour shows up before a device in the field does something surprising.
Technical approach
- Models compiled and quantized for the target accelerator instead of served through a general-purpose runtime, which is where most of the throughput came from
- Hardware-accelerated decode and encode on the video path, so GPU time goes to inference rather than frame handling
- Batching and scheduling tuned against tail latency rather than average, because a real-time system is judged on its worst frames and nobody notices a good average
- Inter-process communication reworked to stop copying large frames between stages, removing a cost that scaled with resolution and got worse exactly as cameras got better
- Deployments versioned and reproducible, so whatever is running on a device traces back to a known build
- Validation runs against real hardware in the loop, since behaviour under thermal and power limits does not reproduce on a workstation