Accepted · PRICAI 2026 · Regular paper, oral presentation
Enroll Once, Identify Everywhere: Few-Shot Cross-Scenario Visual Speaker Recognition via Event-driven SlowFast Networks
Overview
This paper studies visual speaker recognition from lip motion captured by event cameras. Its few-shot enrollment protocol separates training and test identities and recognizes new speakers from three enrollment samples, including changes in illumination and viewpoint.
Method
A SlowFast network processes two-channel polarity-aware event count volumes. Complementary temporal pathways and fast-to-slow lateral fusion capture lip articulation. After training on known identities, a frozen feature extractor supports nearest-neighbor identification without retraining for new users.
Evaluation
On DVSpeaker, Rank-1 accuracy is 99.18% in matched frontal-bright conditions, 87.25% under low-light shift, 74.95% at 45°, and 50.20% at 90°. Ablations examine the contribution of temporal pathways and lateral fusion.
| Enrollment ↓ / Probe → | Frontal, bright | 45° | 90° | Frontal, low light |
|---|---|---|---|---|
| Frontal, bright | 99.18% | 74.95% | 50.20% | 87.25% |
| 45° | 77.45% | 97.32% | 58.75% | 70.05% |
| 90° | 50.80% | 53.10% | 98.04% | 35.75% |
| Frontal, low light | 90.90% | 65.90% | 39.00% | 96.34% |
- event cameras
- visual speaker recognition
- few-shot enrollment
- SlowFast networks
- biometrics
Citation
@inproceedings{yao2026enroll,
title={Enroll Once, Identify Everywhere: Few-Shot Cross-Scenario Visual Speaker Recognition via Event-driven SlowFast Networks},
author={Junguang Yao and Peichun Hua and Yue Zheng},
booktitle={The 23rd Pacific Rim International Conference on Artificial Intelligence (PRICAI)},
year={2026},
organization={Springer}
}