Current work · UAV detection

Counter-UAV thermal imaging

Detecting small drones in thermal video when the camera itself is moving.

The abstract is quoted verbatim from the paper. Every number below was checked against the table it comes from in the paper itself. Figures are the paper’s own, with the paper’s own captions, except the one diagram labelled as drawn for this page.

Optical Flow-Enhanced Thermal UAV Detection Under Camera Ego-Motion for Real-Time Tactical C-UAV

J. Computational Vision and Imaging Systems 11(1) · pp. 122-127 · Proc. CVIS 2025 · journal record published 2026-03-01 · Sole first author · VIP Lab, University of Waterloo

Thermal frames I(t-1) AND I(t) Dense optical flow Farnebäck THREE-CHANNEL INPUT thermal appearance horizontal flow u vertical flow v OpticalFlowConv STEM two paths, kept apart: appearance · motion then learnable fusion YOLO11 detector mAP50 0.358 BoT-SORT + coherence filter OPTIONAL · NOT IN THE SCORES When the camera moves, everything in the frame moves. That destroys the motion cue a small drone would otherwise give away. The flow channels put that cue back, and the split stem stops the detector mixing motion and appearance in its very first layer.

Workflow drawn from the method described in the paper. Not a figure from the paper.

Abstract, verbatim

Thermal UAV detection from mobile platforms is difficult because camera ego-motion corrupts the motion cues needed to detect small airborne targets. We present an optical flow-enhanced YOLO detector that fuses thermal appearance with dense horizontal and vertical flow channels through a custom OpticalFlowConv stem. On the Anti-UAV thermal benchmark (233,667 frames across 205 sequences), using a sequential per-sequence split that preserves temporal order, the proposed detector reaches 35.8% mAP50 and 21.9% mAP50-95, outperforming a single-frame YOLO11 baseline by 11.3 mAP50 points and a frame-differencing baseline by 9.1 points. These results support motion-enhanced thermal detection as a practical sensing component for mobile tactical C-UAV systems under significant camera motion.

35.8%mAP50, against 24.5% for YOLO11 and 26.7% for frame differencing
+11.3mAP50 points over the single-frame baseline
233,667frames across 205 sequences in the Anti-UAV benchmark

The problem. A drone in a thermal frame is a handful of pixels. It has almost no appearance to recognise, so a detector leans on motion instead. But if the camera is handheld or vehicle-mounted, the whole scene moves (buildings, trees, horizon) and the drone’s own motion disappears into that. Frame differencing, the usual cheap fix, breaks for exactly this reason.

The method. Compute dense optical flow between consecutive thermal frames with Farnebäck’s method. Keep the horizontal and vertical components as two extra channels alongside the thermal frame. Feed all three to YOLO11, but replace its first convolution, which would blend the channels immediately, with a stem that processes appearance and motion on separate paths and only then fuses them, with the fusion weights learned.

The three input channels

Three-channel detector input: current thermal frame I_t, horizontal flow \tilde{u}_t, and vertical flow \tilde{v}_t.
Three-channel detector input: current thermal frame It, horizontal flow ŭt, and vertical flow ẽt.
Thermal frame and optical-flow channels with ground-truth annotation. The UAV target remains small in appearance space but exhibits a distinctive localized motion signature in the horizontal and vertical flow components.
Thermal frame and optical-flow channels with ground-truth annotation. The UAV target remains small in appearance space but exhibits a distinctive localized motion signature in the horizontal and vertical flow components.

Results on the Anti-UAV validation partition

Sequential per-sequence protocol, which keeps frames in temporal order rather than shuffling them.

MethodmAP50mAP50-95PrecisionRecall
YOLO110.2450.1560.3120.234
Frame differencing0.2670.1710.3280.251
Ours0.3580.2190.4010.324

What the split stem is worth

ConfigurationmAP50mAP50-95
Standard convolution0.3120.189
Separate branches0.3410.207
+ learnable fusion0.3580.219

The three-channel input alone gets 0.312. Keeping appearance and motion apart adds 0.029. Learning how to recombine them adds another 0.017.

The harder the camera shakes, the bigger the gap

Camera motionFrame differencingOurs
Low0.2890.387
Medium0.2510.342
High0.2040.315

Frame differencing loses 29% of its score from low to high motion. This method loses 19%, and its lead widens from 9.8 points to 11.1.

Keeping a track alive

Temporal-coherence filtering in an operational thermal scene. The true UAV track remains consistent with the predicted path, while thermally salient but motion-incoherent false alarms are rejected.
Temporal-coherence filtering in an operational thermal scene. The true UAV track remains consistent with the predicted path, while thermally salient but motion-incoherent false alarms are rejected.
What the numbers do and do not include

The paper states this itself: every benchmark figure above is measured at the detector stage, before the BoT-SORT tracker and the temporal-coherence filter are applied. Those two are part of the deployment stack, not part of the scores. The intent is to isolate what the motion-enhanced detector contributes from what downstream heuristics clean up afterwards.

BibTeX
@article{maser2025cvis,
  author  = {Maser, Bob and Srebrnjak Yang, Adam and Ramlal, Adrian and Zelek, John},
  title   = {Optical Flow-Enhanced Thermal UAV Detection Under Camera Ego-Motion
             for Real-Time Tactical C-UAV},
  journal = {Journal of Computational Vision and Imaging Systems},
  volume  = {11},
  number  = {1},
  pages   = {122--127},
  year    = {2025},
  url     = {https://openjournals.uwaterloo.ca/index.php/vsl/article/view/7216}
}