Active work at the University of Waterloo VIP Lab. No manuscript submitted, and no benchmark numbers are quoted on this page. What follows is the problem and the approach, not a results claim.
The problem
A camera on a moving platform sees a scene in which everything moves at once: the camera, the target, and the background. Recovering what the target is doing means separating three motions that arrive superimposed in a single 2-D image stream, with no stereo baseline and no active depth sensor to fall back on.
Rigid targets and non-rigid ones fail differently. A rigid body’s motion can be explained by one transform; a deforming body cannot, and a pipeline that assumes rigidity will quietly mis-attribute deformation to translation.
The approach
A multi-cue pipeline, each cue covering a different failure mode of the others:
| Cue | What it contributes |
|---|---|
| Dense optical flow (RAFT) | 2-D correspondence between consecutive frames |
| Optical expansion | Change in apparent size, the depth-rate signal flow alone cannot give |
| Monocular metric depth (Metric3D) | Absolute scale, so velocities come out in metres per second |
Fusing these yields, per target, an estimate of depth, velocity, time-to-collision, and a 3-D intercept solution, from one camera.
Why it matters
The same constraint drives the counter-UAV work: when the platform moves, the motion cue that makes a small target visible is exactly the cue that camera ego-motion destroys. That paper solved a narrow version of the problem: detection, in thermal, with flow channels. This is the general version, in 3-D.