Edge Systems + Distributed Reliability + Python
Edge to Cloud Sync for Unattended Field Devices
Designed and built the service that moves data off unattended field devices and into the cloud reliably, over networks that drop constantly, without ever losing a file or uploading one twice.
Zero data loss across crashes, power loss, and network drops
No duplicate uploads under concurrent processing
No long-lived cloud credentials stored on any device
Months of unattended operation without intervention
The problem
Devices deployed in the field generate a large volume of data every day and have to get it to the cloud continuously, for months, with nobody on site. The network is the first problem, because links drop often enough that any design assuming a stable connection fails within a day. The second is that the device itself will crash, lose power, and restart, and every one of those events happens partway through something. The third is credentials. A device sitting in a shed cannot hold long-lived cloud keys, because physical access to that device would then mean access to the cloud account.
What I built
Built the service as a set of isolated worker processes that share no memory, only a durable local record of what has happened to each file. Because every state change is written down before it is acted on, a crash at any point leaves the system in a state it can reason about when it comes back, rather than one it has to guess at. The device holds no permanent cloud credentials. It proves its identity with a hardware-held certificate and receives short-lived credentials that expire on their own. A file is deleted locally only after the cloud side has independently confirmed it arrived and a retention window has passed.
Technical approach
- Work is split across isolated processes so a slow or stuck stage cannot block the others, which is what keeps throughput steady when the network degrades instead of stalling everything behind one retry
- Every state change is written durably before the action it describes, so recovering from a crash is a matter of reading the record rather than inspecting the filesystem and guessing
- Claiming a file for upload is atomic, so two processes racing on the same file resolve without coordinating and without uploading it twice
- A stalled transfer is reclaimed after a timeout rather than waited on forever, because the common failure here is a link that goes quiet, not one that returns an error
- Retries distinguish between failures worth retrying and failures that will never succeed, so one permanently bad file does not consume the retry budget indefinitely
- Local deletion is gated on independent confirmation from the cloud plus a retention window, so the device is never the only remaining copy at the moment it deletes
- Credentials are short-lived, issued against a hardware-held certificate, and never written to disk in a form anyone could reuse