Eyes for the Blind: A Multi-Agent Assistant for the Visually Impaired
Principal Investigator: Shriyadita De, APL ASPIRE Intern in AMDS/A5B
Dates of Project Work: June - August 2026
References
Visual Impairment and Motivation
World Health Organization. (2019). World report on vision. https://www.who.int/publications/i/item/9789241516570
World Health Organization. (2026, February 10). Blindness and vision impairment. https://www.who.int/news-room/fact-sheets/detail/blindness-and-visual-impairment
Assistive Navigation Systems for Visually Impaired Users
Abidi, M. H., Siddiquee, A. N., Alkhalefah, H., & Srivastava, V. (2024). A comprehensive review of navigation systems for visually impaired individuals. Heliyon, 10(11), e31825. https://doi.org/10.1016/j.heliyon.2024.e31825
Ahmetovic, D., Gleason, C., Ruan, C., Kitani, K. M., Takagi, H., & Asakawa, C. (2016). NavCog: A navigational cognitive assistant for the blind. In Proceedings of the 18th International Conference on Human-Computer Interaction with Mobile Devices and Services (pp. 90–99). Association for Computing Machinery.
Messaoudi, M. D., Menelas, B.-A. J., & Mcheick, H. (2022). Review of navigation assistive tools and technologies for the visually impaired. Sensors, 22(20), 7888. https://doi.org/10.3390/s22207888
Sato, D., Oh, U., Naito, K., Takagi, H., Kitani, K. M., & Asakawa, C. (2017). NavCog3: An evaluation of a smartphone-based blind indoor navigation assistant with semantic features in a large-scale environment. In Proceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility (pp. 270–279). Association for Computing Machinery.
Computer Vision and Scene Understanding for Visually Impaired Users
Caraiman, S., Morar, A., Owczarek, M., Burlacu, A., Rzeszotarski, D., Botezatu, N., Herghelegiu, P., Moldoveanu, F., Strumillo, P., & Moldoveanu, A. (2017). Computer vision for the visually impaired: The Sound of Vision system. In Proceedings of the IEEE International Conference on Computer Vision Workshops (pp. 1480–1489).
Valipoor, M. M., & de Antonio, A. (2023). Recent trends in computer vision-driven scene understanding for VI/blind users: A systematic mapping. Universal Access in the Information Society, 22, 983–1005. https://doi.org/10.1007/s10209-022-00868-w
Object Detection and Deep Learning
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to-end object detection with transformers. In Computer Vision – ECCV 2020 (pp. 213–229). Springer. https://doi.org/10.1007/978-3-030-58452-8_13
Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 779–788).
Ultralytics. (2023). YOLOv8 documentation. https://docs.ultralytics.com/models/yolov8/
Large Language Models and Multi-Agent Systems
Guo, T., Chen, X., Wang, Y., Chang, R., Pei, S., Chawla, N. V., Wiest, O., & Zhang, X. (2024). Large language model based multi-agents: A survey of progress and challenges. arXiv. https://arxiv.org/abs/2402.01680
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., & Wang, C. (2024). AutoGen: Enabling next-gen LLM applications via multi-agent conversation. In Conference on Language Modeling.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations.
Edge AI and NVIDIA Jetson
NVIDIA. (2026). Jetson Orin Nano Developer Kit user guide. https://docs.nvidia.com/jetson/orin-nano-devkit/user-guide/index.html
NVIDIA. (2026). JetPack documentation. https://docs.nvidia.com/jetson/