|
Download README.md from sheneman/NanoDetector: direct link, hf CLI and curl.
- Browser
- Download file 11 kB
-
https://huggingface.co/sheneman/NanoDetector/resolve/main/README.md
- Command line
-
hf download hf://sheneman/NanoDetector/README.md
-
curl -L -o README.md https://huggingface.co/sheneman/NanoDetector/resolve/main/README.md
11 kB
| license: mit | |
| pipeline_tag: object-detection | |
| tags: | |
| - object-detection | |
| - wildlife | |
| - camera-trap | |
| - ecology | |
| - conservation | |
| - yolov5 | |
| - edge-ai | |
| - biodiversity | |
| - pytorch | |
| model-index: | |
| - name: NanoDetector_YOLOv5m_v1 | |
| results: | |
| - task: | |
| type: object-detection | |
| name: Object Detection | |
| dataset: | |
| type: camera-trap | |
| name: NanoDetector merged camera-trap corpus (IDFG, NorthAmerica, WCS, Caltech) | |
| metrics: | |
| - type: mAP@0.5 | |
| value: 0.968 | |
| name: mAP@0.5 (validation) | |
| - type: mAP@0.5:0.95 | |
| value: 0.902 | |
| name: mAP@0.5:0.95 (validation) | |
| - type: precision | |
| value: 0.947 | |
| name: Precision (validation) | |
| - type: recall | |
| value: 0.930 | |
| name: Recall (validation) | |
| # NanoDetector | |
| **A YOLOv5m animal detector and 25-class species classifier for North American camera-trap imagery, | |
| built to run on a Jetson Nano in a battery-powered field enclosure.** | |
| NanoDetector is one half of a system for monitoring wildlife where there is no power and no network. | |
| The other half is a device: an Arduino-scheduled, satellite-connected trail-camera unit that wakes on | |
| a timer, classifies imagery **on-device**, and transmits the conclusions over Iridium instead of the | |
| photographs. A season of imagery collapses from tens of gigabytes to a stream of small structured | |
| messages. | |
| | | | | |
| |---|---| | |
| | 🔧 **Code, device, docs** | [github.com/sheneman/NanoDetector](https://github.com/sheneman/NanoDetector) | | |
| | 🌐 **Project site** | [nanodetector.ai](https://nanodetector.ai) | | |
| | 🦌 **Parent project** | [WHISL — Wildlife & Humans in Shared Landscapes](https://whisl.org) | | |
| | 💵 **Funding** | NSF Award [#2317537](https://www.nsf.gov/awardsearch/showAward?AWD_ID=2317537) | | |
| --- | |
| ## Quick start | |
| NanoDetector is a standard YOLOv5 checkpoint trained on **stock, unmodified** | |
| [Ultralytics YOLOv5](https://github.com/ultralytics/yolov5) — no fork or patch is needed. | |
| ```python | |
| import torch | |
| from huggingface_hub import hf_hub_download | |
| weights = hf_hub_download("sheneman/NanoDetector", "NanoDetector_YOLOv5m_v1.pt") | |
| model = torch.hub.load("ultralytics/yolov5", "custom", path=weights) | |
| model.conf = 0.75 # threshold used in field deployment | |
| results = model("camera_trap_image.jpg") | |
| results.print() | |
| print(results.pandas().xyxy[0]) # xmin ymin xmax ymax confidence class name | |
| ``` | |
| Or with the YOLOv5 CLI: | |
| ```bash | |
| git clone https://github.com/ultralytics/yolov5.git && cd yolov5 | |
| pip install -r requirements.txt | |
| python detect.py --weights NanoDetector_YOLOv5m_v1.pt --source ./images --conf-thres 0.75 | |
| ``` | |
| Field deployments use **`--conf-thres 0.75`**, far above YOLOv5's 0.25 default: species labels feed a | |
| satellite summary, so precision matters more than recall. MegaDetector, not this model, is responsible | |
| for not missing animals. | |
| --- | |
| ## Model details | |
| | | | | |
| |---|---| | |
| | **File** | `NanoDetector_YOLOv5m_v1.pt` (169 MB) | | |
| | **sha256** | `397cd7d62fb38116cea7575bed3a7270b6c8301e434834f792e0bb7eb4081d4c` | | |
| | **Architecture** | YOLOv5m — 290 layers, 20,949,918 parameters | | |
| | **Input** | RGB, 640 × 640 | | |
| | **Classes** | 25 | | |
| | **Framework** | PyTorch / Ultralytics YOLOv5 @ [`c23a441`](https://github.com/ultralytics/yolov5/commit/c23a441c9df7ca9b1f275e8c8719c949269160d1) | | |
| | **Version** | v1 — `best.pt` snapshot at epoch 438 of a 500-epoch run | | |
| | **Developed by** | Luke Sheneman, Research Computing and Data Services, University of Idaho | | |
| | **License** | MIT | | |
| ## Training data | |
| **1,391,742** labelled images merged from five public camera-trap corpora distributed by | |
| [LILA BC](https://lila.science/) — IDFG, IDFG2, North American Camera Trap Images, WCS, and Caltech | |
| Camera Traps — harmonised into a single 25-class taxonomy | |
| ([mapping table](https://github.com/sheneman/NanoDetector/blob/main/model/data/class_map.tsv)). | |
| | Split | Images | | |
| |---|---:| | |
| | train | 1,177,647 | | |
| | val | 67,449 | | |
| | test | 146,646 | | |
| The corpus is heavily imbalanced — `cattle`, `elk` and `deer` dominate, while `badger` has 29 | |
| instances in the entire test split. | |
| **Training:** 500 epochs at 640×640, batch 200, SGD, `hyp.scratch-low`, 2 GPUs via DDP on the | |
| University of Idaho RCDS cluster. Full recipe, exact command, and archived `opt.yaml`/`hyp.yaml`: | |
| [docs/training.md](https://github.com/sheneman/NanoDetector/blob/main/docs/training.md). | |
| ## Classes | |
| | idx | class | idx | class | idx | class | | |
| |---:|---|---:|---|---:|---| | |
| | 0 | human | 9 | lagomorph | 18 | badger | | |
| | 1 | coyote | 10 | bobcat | 19 | unknown cervid ⚠️ | | |
| | 2 | fox | 11 | pronghorn ⚠️ | 20 | rabbit | | |
| | 3 | mountain lion | 12 | bighorn sheep ⚠️ | 21 | horse | | |
| | 4 | elk | 13 | unknown ungulate ⚠️ | 22 | unknown cat | | |
| | 5 | deer | 14 | vehicle | 23 | turkey | | |
| | 6 | unknown canid ⚠️ | 15 | wolf | 24 | domestic sheep | | |
| | 7 | domestic dog | 16 | moose | | | | |
| | 8 | bear | 17 | cattle | | | | |
| ⚠️ **unvalidated** — no instances (or a single instance) in the held-out splits. Predictions for these | |
| classes are unsupported by evidence. | |
| ## Evaluation | |
| Stock YOLOv5 `val.py` at 640×640 against this checkpoint. | |
| | Split | Images | Labels | P | R | mAP@0.5 | mAP@0.5:0.95 | | |
| |---|---:|---:|---:|---:|---:|---:| | |
| | val | 67,449 | 89,912 | 0.947 | 0.930 | 0.968 | 0.902 | | |
| | test | 146,646 | 193,326 | 0.849 | 0.835 | 0.872 | 0.809 | | |
| > **Read the test number carefully.** YOLOv5's `all` row is an *unweighted mean over classes*. The test | |
| > split contains exactly **one** instance each of `unknown canid` and `unknown cervid`; both score 0, | |
| > and two zeros in a 22-class mean cost ~9 points of mAP by themselves. Excluding those two degenerate | |
| > classes, test mAP@0.5 is **0.959** — consistent with validation. The model does not degrade on test. | |
| Selected per-class test results (full tables: | |
| [docs/evaluation.md](https://github.com/sheneman/NanoDetector/blob/main/docs/evaluation.md)): | |
| | Class | Labels | P | R | mAP@0.5 | | |
| |---|---:|---:|---:|---:| | |
| | deer | 29,495 | 0.968 | 0.963 | 0.989 | | |
| | elk | 45,625 | 0.941 | 0.963 | 0.987 | | |
| | cattle | 85,800 | 0.916 | 0.949 | 0.976 | | |
| | mountain lion | 2,465 | 0.950 | 0.965 | 0.989 | | |
| | bear | 3,024 | 0.943 | 0.935 | 0.982 | | |
| | wolf | 631 | 0.941 | 0.946 | 0.979 | | |
| | coyote | 4,175 | 0.936 | 0.950 | 0.962 | | |
| | moose | 3,617 | 0.936 | **0.770** | 0.870 | | |
| | turkey | 9,025 | **0.807** | 0.925 | 0.948 | | |
| | human | 199 | 0.908 | **0.693** | 0.869 | | |
| Speed: 4.0 ms/image inference, batch 32, datacenter GPU. Jetson Nano is substantially slower. | |
| ## Intended use | |
| **In scope** | |
| - Species-level classification of North American camera-trap imagery | |
| - Edge/offline inference where uploading imagery is impractical | |
| - Second-stage classifier behind a class-agnostic detector such as | |
| [MegaDetector](https://github.com/agentmorris/MegaDetector) | |
| - Triage and retrieval across large existing camera-trap archives | |
| **Out of scope** | |
| - **Person detection** — see limitations below | |
| - Regions outside western North America | |
| - Species not among the 25 classes (they will be forced into the nearest class, not rejected) | |
| - Individual identification, age, sex, health, or behaviour | |
| - Legal, enforcement, or regulatory determinations without human review | |
| ## Limitations | |
| 1. **Human recall is poor (0.69 test / 0.72 val).** The model misses roughly three people in ten. Only | |
| 199 human instances exist across the 193,326-label test split — humans are ~0.1% of a corpus drawn | |
| from wildlife surveys. **Do not use this model to detect people**, and do not use it as a privacy | |
| filter. The deployed device relies on MegaDetector for the `person` class precisely for this reason. | |
| 2. **Moose are undercounted** — recall 0.77 at precision 0.94, consistently across both splits. | |
| 3. **Four to five classes are unvalidated** (⚠️ above): no held-out evidence exists for them. | |
| 4. **Turkey over-predicts on test data** (precision 0.807 vs 0.954 on validation). | |
| 5. **Sparse-class metrics are fragile** — `badger` rests on 29 test instances. | |
| 6. **Geographic scope** — trained on western North American imagery; assume poor performance elsewhere. | |
| 7. **Split independence is unverified.** Camera-trap imagery is highly autocorrelated (bursts, fixed | |
| backgrounds). If splits were assigned per-image rather than per-camera-station, these metrics | |
| overstate accuracy at a genuinely new site. The split procedure is not recorded in surviving | |
| artifacts. | |
| 8. **No open-set rejection** — every detection is assigned one of 25 classes; novel species are | |
| misclassified rather than flagged. | |
| 9. **Night/IR imagery** was not separately evaluated, though it is common in camera-trap data. | |
| ## Ethical considerations | |
| **Human imagery.** Camera traps photograph people. This model has a `human` class, but its human recall | |
| is low, so it must **not** be used to find and redact people — roughly three in ten will get through. | |
| Use a purpose-built person detector, and handle imagery under an appropriate human-subjects and | |
| retention policy. | |
| **Sensitive species.** Near-real-time, georeferenced locations for wolves, mountain lions and bear can | |
| enable poaching or inflame conflict. Treat detection streams with the same care as any sensitive | |
| wildlife location data; aggregate or delay before public release. | |
| **Management decisions.** Errors concentrate in rare classes, which are often the classes of greatest | |
| conservation concern. Model counts should inform human judgement, not replace it. Decisions with legal | |
| or livelihood consequences need human verification of the underlying imagery. | |
| ## How it's used in the field | |
| The device runs **two** models in sequence. MegaDetector v5 finds *that* something is there | |
| (animal/person/vehicle) with high recall; NanoDetector says *what* it is. Outputs are fused by IoU so | |
| MegaDetector stays canonical for localisation and for person/vehicle calls while NanoDetector supplies | |
| species identity. Detections within a 30-second burst are then normalised by majority vote — never | |
| across the `human` or `vehicle` classes. | |
| Architecture, firmware, and the fusion pipeline: | |
| [github.com/sheneman/NanoDetector](https://github.com/sheneman/NanoDetector). | |
| ## Citation | |
| ```bibtex | |
| @software{sheneman_nanodetector, | |
| author = {Sheneman, Luke}, | |
| title = {NanoDetector: An edge AI device and species classifier | |
| for satellite-connected wildlife monitoring}, | |
| url = {https://github.com/sheneman/NanoDetector}, | |
| note = {Model weights: https://huggingface.co/sheneman/NanoDetector}, | |
| institution = {University of Idaho} | |
| } | |
| ``` | |
| ## Acknowledgments | |
| Supported by the National Science Foundation under Award | |
| [#2317537](https://www.nsf.gov/awardsearch/showAward?AWD_ID=2317537) and by | |
| [Research Computing and Data Services](https://hpc.uidaho.edu/) at the University of Idaho. Built on | |
| [Ultralytics YOLOv5](https://github.com/ultralytics/yolov5) and | |
| [MegaDetector](https://github.com/agentmorris/MegaDetector). Training imagery from | |
| [LILA BC](https://lila.science/). | |