Vision & perception
AETHER sees through four tiers. They are not fallbacks for one another — they answer different questions, at different costs, on different hardware. A question goes to the tier that can actually answer it, and when no tier can, AETHER says so instead of guessing.
| Tier | Answers | Runs | Install | Measured |
|---|---|---|---|---|
| Camera + VLM | Open-ended description — “what do you see?” | Your vision provider | [vision] | One provider round trip |
| YOLO | Fixed 80-class boxes | Locally, on the robot | [yolo] | ~300 ms on a Raspberry Pi class host |
| MediaPipe hands | 21 landmarks, finger counts, gestures | Locally, CPU only | [hands] | 8.5–10 Hz continuous |
| SAM 3 | Open-vocabulary concepts, pixel masks, tracking | Hosted GPU endpoint | Endpoint URL + key | 2330 ms warm on an NVIDIA T4 |
The camera
One capture path, three backends: picamera2 on a Raspberry Pi, OpenCV on a laptop or desktop, and rpicam where neither fits. The device is probed, not assumed — the vision tools only appear in the tool registry when a camera is genuinely accessible, so a model is never offered a capability the hardware cannot deliver.
Frames are held in memory. Nothing on the detection path writes an image to disk; the single exception is capture_image, where a photograph is the thing the user asked for. aether doctor reports the camera by taking one real picture and the tiers by running real inference — never by checking that a package imports.
Local geometry — MediaPipe hands
A deterministic tier: 21 hand landmarks, from which finger extension and a small gesture vocabulary are derived by geometry. The same hand in the same pose returns the same integer every time, which is the difference between this and asking a language model to look at a photograph.
$ pip install 'aether-robotics[hands]'
The extra pins mediapipe<1.0 deliberately. MediaPipe 1.x aborts the process on macOS the first time it opens a Metal graph, so the tier declines to load it and names the version and the remedy rather than crashing. Both the older solutions API and the newer Tasks API are supported, because 0.10.22 and later ship only the latter.
Watching — continuous local sensing
When a goal is a condition a local sense can decide — “wait until I hold up four fingers” — AETHER does not iterate through a language model. It blocks on the sense at roughly 10 Hz and returns within one poll interval of the condition becoming true. A loop that asked an LLM to re-judge the same boolean would spend seconds per round on work the geometry already answered in milliseconds.
Two properties matter more than the speed. A watcher refuses to start when the detector cannot run, rather than examining a hundred frames with a broken one. And “I looked and saw nothing” is reported differently from “I could not look” — a detector failure is never rendered as a negative detection.
Local detection — YOLO
The fast, free path. A local detector over a fixed 80-class vocabulary, returning boxes, with no network call and no per-request cost. It stays the default for the classes it knows; SAM 3 is for the ones it does not.
Open-vocabulary detection — SAM 3
A hosted GPU endpoint running Meta's SAM 3. It takes a noun phrase rather than a class id — “the red mug”, “a chicken” — and returns a pixel mask, a bounding box, a centroid and a bearing per instance, plus identity across consecutive frames. Navigation needs a bearing to steer by and grasping needs a mask, neither of which a box provides.
It is bring-your-own-endpoint: set the URL and key and AETHER routes concept queries to it. deploy/sam3/ in the repository deploys one to Azure Container Apps on a serverless T4 that scales to zero, with the weights baked into the image so a cold start is an image pull and not a call to a model hub.
What it actually costs
Measured, on the deployment described above:
| Measurement | Value |
|---|---|
| Warm inference, NVIDIA T4, fp32 | 2330 ms |
| Model load, once per replica | 9.7 s |
| Image pulled per cold node | 6.42 GB |
| Idle cost at min-replicas 0 | Nothing |
Meta publishes ~30 ms per image for SAM 3, on an H200. That is a card roughly forty times the price of a T4, and the difference between fp32 on a 2018 inference card and half precision on a 2023 flagship accounts for the gap; there is no missing optimisation behind it. An fp16 measurement on the T4 is in progress and this figure will be updated with it.
So SAM 3 is not the fast tier and is not sold as one. It is the tier that answers questions the others cannot: an open vocabulary, a real mask, and identity that persists. Latency-sensitive work belongs on the local tiers.
Licence
SAM 3 is used under Meta's SAM License, which permits commercial hosted inference and prohibits military and weapons use. Because the deployment bakes the weights into the image, that image is a copy of SAM Materials and must stay in a private registry — never a public one, and never handed to a third party.
How a tier is chosen, and what is never claimed
Routing is by capability, not by preference. A tier is offered only when a real probe says it can run, so a question that no available tier can answer is refused with the reason — a missing camera, an unreachable endpoint, a detector that will not load — rather than answered approximately by something else.
Nothing about what the camera saw reaches analytics. A concept travels as a SHA-256 hash and nothing else, so the counter can answer “was the same thing looked for twice?” without recording what it was; labels, boxes, masks and centroids are never written to an event.
