Conversations about AI in medicine mostly concern data rather than the model. You can take an architecture off the shelf, fine-tune the weights, and squeeze the metrics. But what you can't get anywhere is 500 hours of labeled video from a real patient room, where a real nurse changes a dressing. You have to shoot it yourself.
This is the story of a system built to do exactly that. The goal was to create a clinical Edge AI model that could automatically detect and classify what was happening in a patient room in real time, and distinguish a routine check from an active procedure without requiring staff to do anything differently. Before that model could exist, the data to train it had to be collected. And collecting that data turned out to be its own engineering problem.
Before you can train a model, you have to collect the data to train it on. And that's where something curious surfaces: collecting the data is itself the problem.
Recording should start automatically when a nurse enters a patient room, begins a procedure, and stop when she leaves. No manual control should be needed since the staff's hands and attention are already occupied. That requirement sounds straightforward until you ask the obvious follow-up: how does the system know a procedure has started?
The first engineering answer looked like this. Bluetooth Low Energy (BLE) beacons, placed in the corners of every room, broadcasted a room identifier every 40 milliseconds. The nurse wore a receiver tag that listened, averaged the signal strength, and worked out which room it was in. A workstation took the tag's report, loaded the room configuration, drove the camera to a saved preset, and started recording. Once the tag dropped below the threshold, it stopped recording.
An honest, cheap solution that works. But it answers the wrong question.
What the project needed to know was whether a procedure had started. What the system actually determined was whether a signal level crossed a threshold. Between those two questions sits a chain of assumptions: strong signal means the tag is close, a close tag means the nurse is in the room, and a nurse in the room means she's performing a procedure.
Every link in that chain breaks in its own way. Radio passes through drywall, so a tag in the corridor or the room next door produces a confident, false trigger. A nurse steps in to restock supplies, adjust an IV, or answer a question, and the system records 40 minutes of footage about nothing. The camera looks at a preset saved on installation day, and the bed has since moved.
This is called a proxy label, or a measurement that correlates with your target event without actually being it. A proxy is cheap at collection time, but expensive at training time. You end up with a large dataset full of footage the model can't learn from.
The answer was to replace the inference chain with direct observation. Instead of asking where the tag was and guessing what that implied, the system could simply look at the room and describe what it saw.
A camera connects to the workstation, and a computer vision model on that workstation finds the person in frame and determines their posture: standing, sitting, or lying. The camera preset no longer has to be fixed. If the model found a person, it found the region of interest, which means the bed can be moved and the system still tracks it. The metadata becomes substantive too, evolving from "someone was in the room" to "2 people in frame, one of them horizontal."
The beacon was silent most of the day: no tag in the room, no data, no events, no recording, nothing beyond its own triggers. A camera produces a continuous stream, and after a few weeks that stream yields something the project never had before: a real dataset from a real environment.
That dataset opened a door the BLE approach never could.
The first version of the system could do exactly what was built into it by hand: BLE, thresholds, and person detection. It had no notion of what was actually happening in the room, and no way to acquire one. But once the first few thousand clips were labeled, there was enough material to train a classifier that separates:
The model went back onto the device as a new selection filter. It didn't need to be accurate to be useful, because at that stage, its job was to discard, not decide. The cost of an error is an extra clip in the labeling queue rather than a missed event.
Over the next month, the system collected data an order of magnitude cleaner than before. The next version trained on it. That version produced a better filter. And the better filter produced a cleaner dataset. The loop closed:

That loop worked. It also had a pathology built into it.
The filter passes what it already recognizes. The next version trains on precisely that. Over time, the system's picture of the room narrows to whatever the model handled in the previous iteration. However, this means rare situations fail to enter the sample because there was nothing to flag them. Metrics climb while real coverage stands still.
The fix is 2 channels that bypass the filter entirely: keep what the model is uncertain about, and hold back a small percentage of footage selected at random and labeled by hand. That random sample is the only way to measure what the system isn't seeing.
With that safeguard in place, the picture holds together. Meaning Edge AI turns out to be not only part of the product, but part of the process that builds it, and each new model improves the mechanism that collects the data for the next one.
The lesson here generalizes past this one clinic. Any medtech team building an AI-enabled connected medical device eventually hits the same wall: the model is rarely the hard part, the data that would make it good enough doesn't exist until the system that's supposed to use it starts collecting under real conditions.
That's an architecture decision as much as a machine learning one. It belongs at the same table as the hardware and workflow design, not handed off after the fact. Teams that treat data collection as a logistics problem and model development as the real work tend to discover, late, that they're inseparable.
If your team is working through a similar Edge AI challenge in healthcare, let’s connect.