Project
Perception
Turning cameras into a source of useful events rather than simply stored video — and being precise about which parts of that work are done.
- Frames uploaded anywhere
- 0
- Frames uploaded anywhere
- Physical camera on the vehicle
- 1
- Physical camera on the vehicle
Status
The pipeline runs. The model is not installed.
Everything in the pipeline is built: frames are read from the camera, motion is detected, objects are tracked across frames, and the result becomes a stream of structured events that the camera wall displays and a local API serves. It runs on the vehicle, and it works today.
What is missing is the AI model. A model is a downloaded artifact rather than code, and none is installed on this machine yet. Until it is, the system detects motion and tracks shapes — it does not identify people — and it reports itself as degraded when asked what it is doing.
What it will not do is pretend. A stand-in detector reports itself as a stand-in. That behaviour is the design goal rather than an accident, and it is the reason this page leads with the limitation instead of burying it.
Not implemented
-
Person detection
Requires a model that is not installed. Motion detection is what currently works.
-
Zones, line crossing, occupancy counting
You get “a shape moved through the frame”, not “a person entered the galley”.
-
Recording or retention
Events live in a bounded in-memory buffer and are gone when the service restarts.
-
Remote or phone notifications
Nothing is pushed anywhere. You have to go and look at the camera wall.
-
Multiple cameras
The code is written for many. One physical camera exists on this machine.
-
Face recognition or identity
The tracker knows “the same shape across frames”, not who it is. Tracking IDs are per-session.
Pipeline
Camera to event.
The path a frame takes between pointing a camera at a room and getting a structured record out the other side.
-
Stage 1 of 8: Camera
A live feed from the vehicle, currently one physical camera serving four view roles.
-
Stage 2 of 8: Perception
The pipeline runs locally on the vehicle and is built and running.
-
Stage 3 of 8: Motion detection
Current. Frame-difference detection finds movement without a model installed.
-
Stage 4 of 8: Tracking and event generation
Current. Objects are tracked and emitted as structured events.
-
Stage 5 of 8: Visualisation and API
Current. Events are visible in the camera wall and readable over the local API.
-
Stage 6 of 8: Model-based detection
In development. Requires an AI model that is not installed yet, so the system does not currently identify people.
-
Stage 7 of 8: Zones, line crossing, and occupancy
Planned. Depends on the model stage above.
-
Stage 8 of 8: Recording and retention
Planned. Events currently live in a bounded in-memory buffer and are lost on restart.
A measured limitation
A person can cross the frame and be missed.
Detection runs only when the motion gate fires, and the gate has a sensitivity threshold. There is a measured case where someone walking across the frame over three seconds changes roughly 0.83% of sampled pixels per frame — which never clears a 2% threshold.
That is a real person, present, and not detected. It is documented rather than tuned away, because lowering the threshold has a cost and the cost is not obvious. A written record is more useful than a quietly changed constant.
Why the motion gate exists
Running a detector on every frame of a quiet camera spends compute for no result. The gate is the cheap pre-check that decides whether the expensive stage is worth running at all, which is why it is tuned against false positives rather than tuned to catch everything.
Once a model is installed the trade-off is revisited, because a model-based stage can filter what the frame-difference stage flagged. That lets the gate get more sensitive without drowning the pipeline in noise.
This is published here rather than kept in an internal document because it is the kind of detail that tells you whether the rest of this page can be trusted.
Why local
No frames leave the vehicle.
Perception runs entirely on the head unit. There is no upload step, no cloud inference, and no third party in the path — which means the privacy properties follow from the architecture rather than from a policy that could be changed later.
It is also why this works where cell service does not, which is most of the places this platform goes.
Honest caveat
None of this has been reviewed by anyone outside this project, and no privacy claim here has been audited. It is an accurate description of how the software is built, not a compliance position.
Camera systems
Working on something that needs to see?
Perception is early and honest about it. If you are evaluating camera-based detection and want to know what this pipeline does and does not do, that conversation is worth having — the limits are documented well enough to be useful.