Ground3D-LMM accepted at ECCV 2026

Amol Harsh

M.Sc. Computer Vision @ MBZUAI

Amol Harsh

I am a Computer Vision M.Sc. student at MBZUAI, advised by Prof. Fahad Shahbaz Khan. My research is on 3D scene understanding and multimodal models, with applied work in healthcare imaging and robotics. I am currently a visiting researcher at Microsoft Research India, and previously worked at UC San Diego (MOSAIC Lab) and SUTD (MARVL Lab).

3D Vision & Scene UnderstandingMultimodal & Vision-Language ModelsEmbodied AI / Robot PerceptionMedical Image Analysis

News

  1. Sep 2026Presenting Ground3D-LMM at ECCV 2026 in Malmö, with a main-conference poster on 12 September.
  2. Aug 2026Our MAP-CKD study was accepted to Kidney International Reports. The nailfold microvascular measurements reported in the paper were produced by the computer-vision pipeline I developed with the MOSAIC Lab at UC San Diego, under the guidance of Prof. Tauhidur Rahman and Dr. Rakesh Malhotra. The clinical team now uses it to assess microvascular change in chronic kidney disease.
  3. Jul 2026Ground3D-LMM accepted to ECCV 2026. The paper, code, and Ground3D dataset are now publicly available; the dataset has since passed 13,000 downloads on Hugging Face.
  4. Jun 2026Joined Microsoft Research India as a visiting researcher, working with Prof. Vineeth N. Balasubramanian on video world models.
Research positions

Experience

  1. Microsoft Research India

    Visiting Researcher · Microsoft Research India

    Jun 2026 – Present

    Prof. Vineeth N. Balasubramanian · Bangalore, India

    • Open-world evaluation and repair of video world models, characterizing object collapse for rare and unseen concepts under image-conditioned generation.
    • Built a detect–repair–re-evaluate pipeline that reconditions the model without retraining, plus an object-centric robustness benchmark across common, rare, and novel concepts.
  2. MOSAIC Lab, UC San Diego

    Visiting Scholar · MOSAIC Lab, UC San Diego

    Aug 2024 – Dec 2024

    Prof. Tauhidur Rahman · San Diego, USA

    • Led the AI nailfold-capillaroscopy pipeline covering detection, tracking, and microvascular pattern classification.
    • Delivered a clinical-validation Streamlit app used by collaborators at UCSD Health and Maastricht University.
  3. MARVL Lab, Singapore University of Technology and Design

    Visiting Researcher · MARVL Lab, Singapore University of Technology and Design

    Jul 2023 – Aug 2023

    Prof. Malika Meghjani · Singapore

    • Generalized SimMobility traffic simulation to any OSM-supported city, with a PostgreSQL-backed XML pipeline and C++ source modifications.
Peer-reviewed & under review

Publications

  1. [1]

    Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM

    A. Harsh, Z. Han, J. Lahoud, Y. Liu, R. M. Anwer, H. Cholakkal, S. Khan, F. S. Khan

    European Conference on Computer Vision (ECCV)·2026·Accepted

  2. [2]

    Associations of blood and urine mitochondrial DNA with kidney function and nailfold microvascular measures in chronic kidney disease: the MAP-CKD study

    A. Ahmadi, M. Rahaman, A. Harsh, X. Li, J. Yang, B. Ghanim, S. Dasgupta, T. Rahman, A. J. H. M. Houben, M. Hepokoski, J. H. Ix, R. Malhotra

    Kidney International Reports·2026·Accepted

  3. [3]

    Time-Resolved Finger Nailfold Capillaroscopy for Dynamic Capillary Density Estimation

    M. Rahaman*, A. Harsh*, A. Ahmadi, J. Yang, B. Ghanim, S. Dasgupta, P. Kotanko, R. N. Weinreb, A. J. H. M. Houben, J. H. Ix, T. Rahman, R. Malhotra

    Scientific Reports (Nature Portfolio)·2026·Under Review·* equal contribution

  4. [4]

    Enhancing Public Speaking Skills in Engineering Students Through AI

    A. Harsh, B. Prince, S. Siddharth, D. R. P. Muthirayan, K. S. Bhalla, E. S. Gupta, S. Sahu

    IEEE Frontiers in Education Conference (FIE), Full Paper Track·2025·Published

  5. [5]

    ‘The World of AI’: A Novel Approach to AI Literacy for First-Year Engineering Students

    S. Siddharth, B. Prince, A. Harsh, S. Ramachandran

    Artificial Intelligence in Education (CORE-A) 2025 — Springer CCIS, vol. 2591·2025·Published

  6. [6]

    SANGO: Socially Aware Navigation through Grouped Obstacles

    R. Malladi, A. Harsh, A. Sangwan, S. Chauhan, S. Manjanna

    Indian Control Conference (ICC-10)·2024·Published

Conference presentations

Talks & Presentations

  1. Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM

    12 September 2026, 15:00 CEST

    Main conference poster·ECCV 2026 — Poster Session 6, Multimodal, Video & Document Understanding

    ExHall, poster #162 · Malmö, Sweden

Selected research

Projects

  1. Ground3D-LMM: Fine-Grained 3D Point Grounding & Spatial Reasoning

    Sep 2025 – Mar 2026
    First author·Prof. Fahad Shahbaz Khan, MBZUAI·Accepted, ECCV 2026
    • Unified 3D LMM that, for the first time, jointly produces point-level 3D segmentation masks and metric-consistent numerical responses (size, distance, clearance) at object and part granularity.
    • Proposed the 3D Grounded Measurement task and constructed Ground3D, a corpus of roughly 3M QA pairs across 2.5K ScanNet and ScanNet++ scenes with dense object and part annotations and multi-turn grounded dialogue.
    • Achieves SOTA across object- and part-level grounding; +5.15% mIoU over Reason3D and ~20% improvement on ScanRefer instance grounding, with markedly lower metric error than image-only baselines.
    • Released the Ground3D dataset publicly on Hugging Face, with over 13,000 downloads to date.
    3D VisionLMMsGroundingSpatial Reasoning
  2. Open-World Evaluation and Repair of Video World Models

    Jun 2026 – Present
    Visiting Researcher·Prof. Vineeth N. Balasubramanian, Microsoft Research India·Ongoing
    • Studying open-world failure modes in video world models, in particular object collapse for rare or unseen concepts under image-conditioned generation.
    • Developing a detect–repair–re-evaluate pipeline that identifies collapsed generations, augments object references, and reconditions the model without retraining.
    • Curating an object-centric benchmark spanning common, rare, and novel concepts to measure robustness across video world models.
    Video GenerationWorld ModelsRobustnessBenchmarking
  3. Microcirculation Analysis with AI for Nailfold Capillaries

    Aug 2024 – Dec 2024
    Lead developer·Prof. Tauhidur Rahman, MOSAIC Lab, UC San Diego·2 papers (KI Reports, Scientific Reports)
    • Built a deep-learning pipeline for nailfold capillary video analysis covering automated detection, tracking, and pattern classification linked to chronic kidney disease.
    • 92% F1 using YOLOv11 + optical flow + graph-based methods, robust to artifacts and noisy clinical inputs.
    • Shipped a Streamlit interface for real-time analysis, validated with collaborators at UCSD Health and Maastricht University.
    Medical ImagingDetectionTracking
Academic background

Education

  1. Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)

    Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)

    2025 – 2027

    M.Sc. in Computer Vision·Advised by Prof. Fahad Shahbaz Khan

    Abu Dhabi, UAE

    • Coursework: Human and Computer Vision, Probabilistic & Statistical Inference, Visual Object Recognition and Detection.
  2. Plaksha University

    Plaksha University

    2021 – 2025

    B.Tech. in Computer Science and Artificial Intelligence

    GPA 9.57 / 10 · India

    • Dean's List 2021–2024 (top 5 students across all majors, three consecutive years).
    • 2021–22 batch topper, with the highest CGPA across all majors.
    • Coursework: Machine Learning & Pattern Recognition, Deep Learning, Reinforcement Learning, Data Science & AI.
  3. Ashoka University

    Ashoka University

    2020 – 2021

    Freshman year (liberal arts & sciences)

    GPA 3.93 / 4.0 · India

    • Dean's List, Monsoon 2020 and Spring 2021.

Skills

Programming & Systems
Python · C++ · Dart · Java · PostgreSQL · Ubuntu · Flutter · Firebase · Git
ML & Deep Learning
PyTorch · Reinforcement Learning (PPO) · Stable Baselines · Model Development
Vision & Data
OpenCV · NumPy · Pandas · Matplotlib · Statistical Modeling · Data Wrangling
Other
Scientific Writing · Simulation Design · Mobile App Development · Shell Scripting

Awards

  • Fully funded merit scholarship for the M.Sc. in Computer Vision at MBZUAI (≈US$407K), 2025.
  • Full merit scholarship for the B.Tech. in Computer Science and AI at Plaksha University (≈US$280K), 2021.
  • 1st place in the SP Dutt Award for Innovation and Impact 2024, Plaksha University ($2,500).
  • $1,000 annual funding from the US National Academy of Engineering for Yotta Expo.
  • Member, New York Academy of Sciences (top 9% of 17,000 global applicants, 2018).