Amol Harsh
M.Sc. Computer Vision @ MBZUAI

I am a Computer Vision M.Sc. student at MBZUAI, advised by Prof. Fahad Shahbaz Khan. My research is on 3D scene understanding and multimodal models, with applied work in healthcare imaging and robotics. I am currently a visiting researcher at Microsoft Research India, and previously worked at UC San Diego (MOSAIC Lab) and SUTD (MARVL Lab).
News
- Sep 2026Presenting Ground3D-LMM at ECCV 2026 in Malmö, with a main-conference poster on 12 September.
- Aug 2026Our MAP-CKD study was accepted to Kidney International Reports. The nailfold microvascular measurements reported in the paper were produced by the computer-vision pipeline I developed with the MOSAIC Lab at UC San Diego, under the guidance of Prof. Tauhidur Rahman and Dr. Rakesh Malhotra. The clinical team now uses it to assess microvascular change in chronic kidney disease.
- Jul 2026Ground3D-LMM accepted to ECCV 2026. The paper, code, and Ground3D dataset are now publicly available; the dataset has since passed 13,000 downloads on Hugging Face.
- Jun 2026Joined Microsoft Research India as a visiting researcher, working with Prof. Vineeth N. Balasubramanian on video world models.
Experience
Visiting Researcher · Microsoft Research India
Jun 2026 – PresentProf. Vineeth N. Balasubramanian · Bangalore, India
- Open-world evaluation and repair of video world models, characterizing object collapse for rare and unseen concepts under image-conditioned generation.
- Built a detect–repair–re-evaluate pipeline that reconditions the model without retraining, plus an object-centric robustness benchmark across common, rare, and novel concepts.
Visiting Scholar · MOSAIC Lab, UC San Diego
Aug 2024 – Dec 2024Prof. Tauhidur Rahman · San Diego, USA
- Led the AI nailfold-capillaroscopy pipeline covering detection, tracking, and microvascular pattern classification.
- Delivered a clinical-validation Streamlit app used by collaborators at UCSD Health and Maastricht University.
Visiting Researcher · MARVL Lab, Singapore University of Technology and Design
Jul 2023 – Aug 2023Prof. Malika Meghjani · Singapore
- Generalized SimMobility traffic simulation to any OSM-supported city, with a PostgreSQL-backed XML pipeline and C++ source modifications.
Publications
- [1]
Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM
A. Harsh, Z. Han, J. Lahoud, Y. Liu, R. M. Anwer, H. Cholakkal, S. Khan, F. S. Khan
European Conference on Computer Vision (ECCV)·2026·Accepted
- [2]
Associations of blood and urine mitochondrial DNA with kidney function and nailfold microvascular measures in chronic kidney disease: the MAP-CKD study
A. Ahmadi, M. Rahaman, A. Harsh, X. Li, J. Yang, B. Ghanim, S. Dasgupta, T. Rahman, A. J. H. M. Houben, M. Hepokoski, J. H. Ix, R. Malhotra
Kidney International Reports·2026·Accepted
- [3]
Time-Resolved Finger Nailfold Capillaroscopy for Dynamic Capillary Density Estimation
M. Rahaman*, A. Harsh*, A. Ahmadi, J. Yang, B. Ghanim, S. Dasgupta, P. Kotanko, R. N. Weinreb, A. J. H. M. Houben, J. H. Ix, T. Rahman, R. Malhotra
Scientific Reports (Nature Portfolio)·2026·Under Review·* equal contribution
- [4]
Enhancing Public Speaking Skills in Engineering Students Through AI
A. Harsh, B. Prince, S. Siddharth, D. R. P. Muthirayan, K. S. Bhalla, E. S. Gupta, S. Sahu
IEEE Frontiers in Education Conference (FIE), Full Paper Track·2025·Published
- [5]
‘The World of AI’: A Novel Approach to AI Literacy for First-Year Engineering Students
S. Siddharth, B. Prince, A. Harsh, S. Ramachandran
Artificial Intelligence in Education (CORE-A) 2025 — Springer CCIS, vol. 2591·2025·Published
- [6]
SANGO: Socially Aware Navigation through Grouped Obstacles
R. Malladi, A. Harsh, A. Sangwan, S. Chauhan, S. Manjanna
Indian Control Conference (ICC-10)·2024·Published
Talks & Presentations
Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM
12 September 2026, 15:00 CESTMain conference poster·ECCV 2026 — Poster Session 6, Multimodal, Video & Document Understanding
ExHall, poster #162 · Malmö, Sweden
Projects
Ground3D-LMM: Fine-Grained 3D Point Grounding & Spatial Reasoning
Sep 2025 – Mar 2026First author·Prof. Fahad Shahbaz Khan, MBZUAI·Accepted, ECCV 2026- Unified 3D LMM that, for the first time, jointly produces point-level 3D segmentation masks and metric-consistent numerical responses (size, distance, clearance) at object and part granularity.
- Proposed the 3D Grounded Measurement task and constructed Ground3D, a corpus of roughly 3M QA pairs across 2.5K ScanNet and ScanNet++ scenes with dense object and part annotations and multi-turn grounded dialogue.
- Achieves SOTA across object- and part-level grounding; +5.15% mIoU over Reason3D and ~20% improvement on ScanRefer instance grounding, with markedly lower metric error than image-only baselines.
- Released the Ground3D dataset publicly on Hugging Face, with over 13,000 downloads to date.
3D VisionLMMsGroundingSpatial ReasoningOpen-World Evaluation and Repair of Video World Models
Jun 2026 – PresentVisiting Researcher·Prof. Vineeth N. Balasubramanian, Microsoft Research India·Ongoing- Studying open-world failure modes in video world models, in particular object collapse for rare or unseen concepts under image-conditioned generation.
- Developing a detect–repair–re-evaluate pipeline that identifies collapsed generations, augments object references, and reconditions the model without retraining.
- Curating an object-centric benchmark spanning common, rare, and novel concepts to measure robustness across video world models.
Video GenerationWorld ModelsRobustnessBenchmarkingMicrocirculation Analysis with AI for Nailfold Capillaries
Aug 2024 – Dec 2024Lead developer·Prof. Tauhidur Rahman, MOSAIC Lab, UC San Diego·2 papers (KI Reports, Scientific Reports)- Built a deep-learning pipeline for nailfold capillary video analysis covering automated detection, tracking, and pattern classification linked to chronic kidney disease.
- 92% F1 using YOLOv11 + optical flow + graph-based methods, robust to artifacts and noisy clinical inputs.
- Shipped a Streamlit interface for real-time analysis, validated with collaborators at UCSD Health and Maastricht University.
Medical ImagingDetectionTracking
Education

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
2025 – 2027M.Sc. in Computer Vision·Advised by Prof. Fahad Shahbaz Khan
Abu Dhabi, UAE
- Coursework: Human and Computer Vision, Probabilistic & Statistical Inference, Visual Object Recognition and Detection.

Plaksha University
2021 – 2025B.Tech. in Computer Science and Artificial Intelligence
GPA 9.57 / 10 · India
- Dean's List 2021–2024 (top 5 students across all majors, three consecutive years).
- 2021–22 batch topper, with the highest CGPA across all majors.
- Coursework: Machine Learning & Pattern Recognition, Deep Learning, Reinforcement Learning, Data Science & AI.

Ashoka University
2020 – 2021Freshman year (liberal arts & sciences)
GPA 3.93 / 4.0 · India
- Dean's List, Monsoon 2020 and Spring 2021.
Skills
- Programming & Systems
- Python · C++ · Dart · Java · PostgreSQL · Ubuntu · Flutter · Firebase · Git
- ML & Deep Learning
- PyTorch · Reinforcement Learning (PPO) · Stable Baselines · Model Development
- Vision & Data
- OpenCV · NumPy · Pandas · Matplotlib · Statistical Modeling · Data Wrangling
- Other
- Scientific Writing · Simulation Design · Mobile App Development · Shell Scripting
Awards
- Fully funded merit scholarship for the M.Sc. in Computer Vision at MBZUAI (≈US$407K), 2025.
- Full merit scholarship for the B.Tech. in Computer Science and AI at Plaksha University (≈US$280K), 2021.
- 1st place in the SP Dutt Award for Innovation and Impact 2024, Plaksha University ($2,500).
- $1,000 annual funding from the US National Academy of Engineering for Yotta Expo.
- Member, New York Academy of Sciences (top 9% of 17,000 global applicants, 2018).