MOTHERSHIP

Computer vision to help catch snakes.

Snakes kept getting onto a large site, right where people walk at night, and one missed snake can mean a bite. The site already had 70 cameras, but no one could watch every screen all night.

We built a computer vision system that watches every feed, confirms each sighting and alerts the site team as soon as it's spotted.

70
cameras on one server
12M
checks a day
3
layers before an alert
6–8 ms
per check

How we achieved this.

A snake is 0.1% of the picture.

On a 1920 by 1080 camera, a snake is about 70 by 40 pixels, often at night and half in shadow. A standard object detector downscales the frame to 640 pixels before inference, and the snake shrinks to a 23 by 13 smear. False positives are the opposite problem. Leaves, shadows and cables move all night, and at two frames a second across 70 cameras, the system screens about 12 million frames a day. If every leaf sends an alert, people stop trusting the alerts.

    A detection pipeline built for a 70-pixel target.

    Each feed is sampled twice a second. Motion detection runs first: background subtraction compares every frame with the camera's learned background and passes on only the regions that changed. Object detection comes next. The detector runs on each changed region cropped from the full-resolution frame, so the snake keeps every pixel, and a detection only counts once it holds for five consecutive frames. Validation is last. A vision-language model reviews the saved frame with the whole scene around the box and returns snake, not a snake or not sure. Only a snake verdict triggers an alert, sent to Telegram with the frame, camera, time, confidence and the model's reasoning.

    • Every confirmed event stored with its frame, camera and timestamp
    • Uncertain verdicts logged for review, never sent as alerts

      Seventy cameras on one server.

      Running all 70 cameras on a single GPU server was the last hard part. Seventy feeds at two frames a second is 140 frames a second to screen. Motion gating keeps the detector on changed regions only, at 6 to 8 milliseconds per inference on the GPU, against 13 to 15 on a CPU. Cameras are assigned round-robin across inference workers and motion analysis goes to whichever thread is free, so no worker is overloaded and no checks are dropped. Exclusion masks cover the timestamp burned into each feed, which background subtraction otherwise reads as constant motion. We tested it against the site's full camera setup.

      • No new cameras or cabling
      • Each camera runs on its own, so one bad feed can't stall the rest

        Beyond snakes.

        The same technology can be adapted to other use cases, like catching intruders in restricted areas at night and tracking objects across a factory floor. It runs on the cameras that factories, warehouses and construction sites already have.

          Parallel proof of work

          Explore how our core modules are deployed across different operational architectures.