My research sits at the intersection of Health AI and multimodal deep learning: particularly building models that can reason jointly over video, sensor data, and text, and solving real clinical problems, especially in Neonatology. I did my PhD at the University of Wisconsin-Madison, where I was fortunate to be advised by Prof. Yin Li.
I have also gained valuable research experience through internships at Amazon Science, Microsoft Research and EMC.
My research is driven by problems in Health AI and Biomedicine, and by the multimodal deep learning needed to solve them. I build models that learn from sequential data such as video, sensor data (time series), and text, and I am especially interested in how Vision Language Models and Large Language Models can reason over these signals.
Mar 2026: Our paper "DUALVISION: RGB–Infrared Multimodal Large Language Models for Robust Visual Reasoning" was accepted at CVPR Findings 2026.
Mar 2026: Our paper "SimpleCall: A Lightweight Image Restoration Agent in Label-Free Environments with MLLM Perceptual Feedback" was accepted at CVPR Findings 2026.
Feb 2026: Our paper "Analgesia and sedation trends in a level IV NICU, 2014–2024: Opioid and dexmedetomidine use" was accepted at Journal of Perinatology.
Jan 2026: Our paper "Deep learning to assess laryngoscope insertion depth during neonatal intubation with video laryngoscopy" was accepted at Journal of Perinatology.
July 2025: Our paper "Agentic Prompt Optimization for Evidence-Grounded Clinical Question Answering" was accepted as an Oral presentation at BioNLP @ ACL 2025.
June 2025: Our team was awarded 2nd position in the Shared Task on grounded question answering (QA) from electronic health records (ArchEHR-QA 2025) at BioNLP@ACL 2025.
June 2025: Our paper "LETS Forecast: Learning Embedology for Time Series Forecasting" has been accepted at the International Conference on Machine Learning (ICML) 2025.
April 2025: Our team was awarded 1st place in the Machine Learning Challenge at the Pediatric Academic Societies (PAS) 2025 Conference, Honolulu, Hawaii.
Jan 2025: Our Poster "Deep learning to quantify care manipulation activities in neonatal intensive care units" won an Award for Best Innovation in Neonatology at the Cleveland Clinic Children's SHINE (Syposium on Health Innovation and Neonatal Excellence) Conference, Orlando, FL.
Nov 2024: Our paper "RICA2: Rubric-Informed, Calibrated Assessment of Actions" won the Best Poster Award at NSF Poster Competition at Purdue University, West Lafayette, IN.
A lightweight RGB-IR fusion module for multimodal large language models that enables robust visual reasoning under visual degradations like blur, low-light, and fog.
A novel time series forecasting method that combines principles from nonlinear dynamical systems with deep learning to model latent temporal structure for accurate forecasts.
Automatically quantify care manipulation activities in neonatal intensive care units (NICUs), while integrating physiological
signal data to monitor neonatal stress in NICUs.