Vision Transformer model for classification of aerial targets using non-coherent X-band radar
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Item Statistics
- Total Views: 2
- Total Downloads: 0
- Views in the Last Month: 2
Abstract
Small unmanned aerial vehicles can exhibit radar cross sections and bulk radial velocities comparable to those of birds, limiting discrimination based on range, received amplitude, and translational Doppler alone. Micro-Doppler signatures provide complementary information because wing flapping and rotor motion produce target-dependent modulation in the time–frequency domain. This thesis presents three related contributions. First, a leakage-referenced coherent-on-receive (COR) processing framework is developed for intermediate-frequency data collected with a Furuno FAR-15x8 magnetron X-band marine radar. The framework combines adaptive intermediate-frequency estimation, conditional residual frequency-offset correction, leakage-window selection, rank-1 singular-value-decomposition reference estimation, pulse-wise phase compensation, matched filtering, and range-tracked time–frequency analysis. Phase-only-correlation timing estimates are retained as diagnostics; the reported products use pass-through timing. Representative Galveston bird and drone collections show range-localized range–Doppler and short- time Fourier-transform products after COR processing, but the scene-dependent workflow is not presented as an aggregate performance benchmark. Second, a kinematic and scattering simulation framework extends a two-segment bird-wing representation to a three-segment model with asymmetric stroke timing, lead–lag motion, body oscillation, optional intermittent-flight gating, aspect-dependent ellipsoidal scattering, and Reynolds-rule flock trajectories. The Herring Gull parameterization is a simulation design choice rather than a species-validation study. Third, a FAN-integrated Lightweight Hybrid Vision Transformer (FAN-LH-ViT) is evaluated on the separate public DIAT-µSAT benchmark. The model combines multiscale convolutional feature extraction, self-attention, and sine–cosine token projections motivated by periodic micro-motion. The retained checkpoint achieves 98.63% accuracy on a 1,455-sample stratified hold-out subset. The classifier is not trained or evaluated on the Furuno/Galveston collections; consequently, the COR, simulation, and benchmark-classification results are reported as separate contributions rather than as an end-to-end measured-data classification validation.