Eyes for the Blind: A Multi-Agent Assistant for the Visually Impaired

Principal Investigator: Shriyadita De, APL ASPIRE Intern in AMDS/A5B

Dates of Project Work: June - August 2026

References

Visual Impairment and Motivation

World Health Organization. (2019). World report on vision. https://www.who.int/publications/i/item/9789241516570

World Health Organization. (2026, February 10). Blindness and vision impairment. https://www.who.int/news-room/fact-sheets/detail/blindness-and-visual-impairment

Assistive Navigation Systems for Visually Impaired Users

Abidi, M. H., Siddiquee, A. N., Alkhalefah, H., & Srivastava, V. (2024). A comprehensive review of navigation systems for visually impaired individuals. Heliyon, 10(11), e31825. https://doi.org/10.1016/j.heliyon.2024.e31825

Ahmetovic, D., Gleason, C., Ruan, C., Kitani, K. M., Takagi, H., & Asakawa, C. (2016). NavCog: A navigational cognitive assistant for the blind. In Proceedings of the 18th International Conference on Human-Computer Interaction with Mobile Devices and Services (pp. 90–99). Association for Computing Machinery.

Messaoudi, M. D., Menelas, B.-A. J., & Mcheick, H. (2022). Review of navigation assistive tools and technologies for the visually impaired. Sensors, 22(20), 7888. https://doi.org/10.3390/s22207888

Sato, D., Oh, U., Naito, K., Takagi, H., Kitani, K. M., & Asakawa, C. (2017). NavCog3: An evaluation of a smartphone-based blind indoor navigation assistant with semantic features in a large-scale environment. In Proceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility (pp. 270–279). Association for Computing Machinery.

Computer Vision and Scene Understanding for Visually Impaired Users

Caraiman, S., Morar, A., Owczarek, M., Burlacu, A., Rzeszotarski, D., Botezatu, N., Herghelegiu, P., Moldoveanu, F., Strumillo, P., & Moldoveanu, A. (2017). Computer vision for the visually impaired: The Sound of Vision system. In Proceedings of the IEEE International Conference on Computer Vision Workshops (pp. 1480–1489).

Valipoor, M. M., & de Antonio, A. (2023). Recent trends in computer vision-driven scene understanding for VI/blind users: A systematic mapping. Universal Access in the Information Society, 22, 983–1005. https://doi.org/10.1007/s10209-022-00868-w

Object Detection and Deep Learning

Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to-end object detection with transformers. In Computer Vision – ECCV 2020 (pp. 213–229). Springer. https://doi.org/10.1007/978-3-030-58452-8_13

Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 779–788).

Ultralytics. (2023). YOLOv8 documentation. https://docs.ultralytics.com/models/yolov8/

Large Language Models and Multi-Agent Systems

Guo, T., Chen, X., Wang, Y., Chang, R., Pei, S., Chawla, N. V., Wiest, O., & Zhang, X. (2024). Large language model based multi-agents: A survey of progress and challenges. arXiv. https://arxiv.org/abs/2402.01680

Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., & Wang, C. (2024). AutoGen: Enabling next-gen LLM applications via multi-agent conversation. In Conference on Language Modeling.

Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations.

Edge AI and NVIDIA Jetson

NVIDIA. (2026). Jetson Orin Nano Developer Kit user guide. https://docs.nvidia.com/jetson/orin-nano-devkit/user-guide/index.html

NVIDIA. (2026). JetPack documentation. https://docs.nvidia.com/jetson/