Data Scientist · Microsoft Ads

Elahe Vahdani

Multimodal retrieval and ranking for ads · Ph.D. in Computer Science, The City University of New York

I build vision-language systems that decide which images belong with which ads across the Microsoft Audience Network — retrieval, ranking and evaluation at the scale of billions of ads, shipped and measured in production.

Before Microsoft, I did a Ph.D. in computer vision at The City University of New York, working on action understanding in video, cross-modal retrieval and sign language recognition, with papers at CVPR and IEEE TPAMI.

Computer vision Vision-language models Retrieval & ranking Multimodal evaluation Large-scale ML pipelines Online experimentation
Portrait of Elahe Vahdani

Work

Microsoft Ads · Microsoft Audience Network · 2024–present
Travel ads

Matching travel ads to the right property images

Launched a multimodal pipeline that replaces generic stock images with the advertiser's own photos of each property, across billion-scale travel text ads. Ads are matched to properties through destination-URL entity resolution, and a fine-tuned scene classifier filters out unsuitable photos so that only representative hotel scenes are shown. Validated in an online A/B test.

Image ClassificationEntity resolutionA/B testing
Text ads

Image retrieval for text ads from landing-page content

Redesigned how images are chosen for text ads: landing-page content is embedded with a vision-language model and matched with DiskANN vector search, with a tuned retrieval-distance gate that suppresses poor matches. Built with LLM-as-judge evaluation in the loop; shipped through a backfill plus a daily pipeline and evaluated online.

Vision-language modelDiskANNLLM-as-judgeA/B testing
Cashback ads

Category-aware images for ads without a logo

Replaced random fallback images with imagery of what each merchant sells, for logo-less cashback ads at hundred-million scale. An offline pipeline classifies each ad into a fine-grained taxonomy with thousands of categories, retrieves candidate images per category via embedding retrieval, and shortlists them with a fine-tuned re-ranker. Validated in an online A/B test.

Multimodal retrievalRe-ranking
Ads quality

Incident investigation with multimodal retrieval

Built an embedding-based text-to-image and image-to-image retrieval system over the full ad corpus for incident triage, with automated reporting and mitigation workflows that cut investigation time from hours to minutes.

Multimodal retrievalVector searchAutomation

Earlier: Research Science Intern at Dataminr (fall 2021) and Data Science Intern at Expedia Group (summer 2021).

Publications

Hover a thumbnail for the paper figure · Google Scholar

* Equal contribution.

Background

Education

Ph.D. in Computer Science, The City University of New York — defended December 2023.

Dissertation: Deep Learning-Based Human Action Understanding in Videos. Advised by Professor Yingli Tian at the Media Lab of The City University of New York; research spanning action detection in untrimmed videos, cross-modal retrieval, sign language understanding, time-series analysis and vehicle re-identification.

Service

Reviewer for CVPR, AAAI, TMM, CVIU, MVAP, and JVCI.

Outside work

Hiking, yoga, strength training, and reading.

Timeline

  • 2024 – Present Data Scientist, Microsoft Ads — multimodal retrieval and ranking for the Microsoft Audience Network
  • Dec 2023Defended Ph.D. thesis, The City University of New York
  • Fall 2021Research Science Intern, Dataminr
  • Summer 2021Data Science Intern, Expedia Group
  • 2018Joined the Media Lab at The City University of New York as a Ph.D. student under the supervision of Professor Yingli Tian.

Contact

The fastest way to reach me is email: ellie.vahdani@gmail.com. I'm also on LinkedIn, GitHub and Google Scholar.