CResMOT: Continuous RGB–Event Stateful Memory for Multi-Object Tracking
| dc.contributor.advisor | Habibi, Golnaz | |
| dc.contributor.author | McCulley, Evan | |
| dc.contributor.committeeMember | Xu, Shuozhi | |
| dc.contributor.committeeMember | Ghamarian, Iman | |
| dc.date.accessioned | 2026-09-08T19:15:25Z | |
| dc.date.embargoExpiration | ||
| dc.date.issued | 2026 | |
| dc.date.proquestAvailable | 01/01/2026 | |
| dc.date.updated | 2026-09-08T19:15:25Z | |
| dc.description.abstract | Recent advances in machine learning have made RGB-camera-only approaches increasingly viable for multi-object tracking. In the frame-based RGB setting considered here, however, detections arrive per frame at 20 Hz to 30 Hz and tracks are carried between frames by a motion model, so a track receives no new evidence between observations. Additional temporally resolved sensing is therefore needed if the tracker is to incorporate new measurements, rather than prediction alone, during those intervals. The usual remedy is a second sensor, but the established alternatives are themselves sampled. LiDAR, long-wave infrared, and radar all report at discrete instants tens of milliseconds apart. Continuous-time formulations model the motion between those instants, but they interpolate rather than observe. An event camera differs in kind rather than in rate. Each pixel reports independently when its log-luminance crosses a threshold, giving microsecond-timestamped output four to five orders of magnitude finer than the frame interval. It therefore supplies measurement during the interval, not only at its endpoints. This thesis investigates CResMOT (Continuous RGB–Event Stateful Memory for Multi-Object Tracking), whose defining requirement is that local event-derived state persist across RGB intervals and condition persistent per-object memory rather than being rebuilt as a frame-rate event image. A frozen YOLOX-m detector supplies RGB detections, a recurrent EVA-derived encoder supplies event features, and ROI-local fusion forms the observation. Each track carries a matrix-valued latent state propagated between observations by a Neural Ordinary Differential Equation and corrected at matched observations by an RWKV-7-style structured jump. The study establishes RGB and Speed-Invariant Frame ByteTrack baselines on DSEC-MOT, then evaluates detector-matched implementations of the event pathway. On a held-out sequence, an adapter acting on motion and covariance before association improves HOTA by 1.757, AssA by 4.421, and IDF1 by 2.574 over matched RGB ByteTrack. Memory added after candidate formation does not improve on it, and ground-truth oracles at that point are similarly limited. This locates the depth at which event state must enter. The complete architecture remains to be trained and evaluated as one system. | |
| dc.identifier.uri | https://shareok.org/handle/11244/342915 | |
| dc.language.iso | en | |
| dc.publisher | University of Oklahoma – Graduate College | |
| dc.subject | Artificial intelligence | |
| dc.subject | Computer science | |
| dc.subject | Mechanical engineering | |
| dc.subject | autonomous driving | |
| dc.subject | computer vision | |
| dc.subject | event camera | |
| dc.subject | multi-object tracking | |
| dc.subject | recurrent neural network | |
| dc.subject | RGB-event fusion | |
| dc.thesis.degree | M.S. | |
| dc.title | CResMOT: Continuous RGB–Event Stateful Memory for Multi-Object Tracking | |
| ou.group | Aerospace and Mechanical Engr: Engineering |