University of Maryland, College Park

Anubhav Gupta

PhD Candidate, Computer Science  ·  advised by Abhinav Shrivastava

I work on video understanding — what a representation of a long video should carry, and how it ought to change with the question being asked of it. Lately that has run from semantics inside learned representations to temporal structure in egocentric video, and now toward vision–language models.

My work sits in video understanding, and the question I keep coming back to is what a representation of a long video should actually carry. A representation is only ever good for something: the cut that answers “what step is happening now?” is not the one that answers “why did this go wrong?” That tension, between a fixed representation and a shifting question, is what most of my recent work circles.

One line of it asks what learned representations already encode. Latent-INR showed that implicit neural representations of video can carry discriminative semantics rather than only reconstructing pixels; LEIA found latents for 3D articulation that hold up across viewpoints; CSD pulled style out as a measurable, separable property of generated images. I have also worked on learning from noisy human and VLM preferences with TREND.

The other line is temporal. VidParse (ECCV 2026) treats online action understanding in egocentric video as graph-constrained inference: it finds semantic boundaries from manipulation-anchored features off frozen foundation models, then decodes with a beam search restricted to valid step transitions. It takes no gradient steps and is up to 10× more accurate at multi-step parsing than strong trained online baselines.

Where this is going. A parse fixed in advance is a strong assumption, and procedural task graphs are a rigid way to hold knowledge — they have to be induced or authored, and they do not transfer. I am starting to look at representations whose structure is not settled ahead of the question being asked. Vision–language models are the obvious way in, since they already carry a great deal of this knowledge implicitly. This is early; no results yet.

video understanding video representation learning egocentric & procedural video video–language models
Under review
In submission

Two Party Evaluation of the Open Visual World

Defining a paradigm for unbiased evaluation in open worlds.

I came to the PhD after eight years in industry: analyst roles at several banks first, then computer vision for automotive safety at Netradyne and geospatial machine learning at Swiggy. My four stints at Amazon are internships taken during the PhD — edge-based detection, visual geolocalization, and diffusion-based image editing. I did my B.Tech. in Electrical Engineering at IIT Delhi.

Jan 2021 — present
University of Maryland, College Park
Graduate Research Assistant · PhD, Computer Science
Video understanding — representation, temporal structure, and video–language models. Advised by Abhinav Shrivastava.
May — Aug 2024
Amazon Fashion · Sunnyvale, CA
Applied Scientist Intern
Diffusion-based multi-view image editing for indoor scenes; built a large-scale dataset starting from 3D-FRONT.
May — Aug 2023
Amazon Fashion · Sunnyvale, CA
Applied Scientist Intern
Keystroke-assisted human body editing (pose and shape) with Stable Diffusion 1.5 and ControlNets, aimed at 2D virtual try-on.
Oct — Dec 2022
Amazon
Applied Scientist Intern (Co-op)
Visual geolocalization.
May — Aug 2022
Amazon · Seattle, WA
Applied Scientist Intern
Edge-based object detection; took a prototype model to production for field trials.
Sep 2020 — Jan 2021
Swiggy · Bangalore, India
Data Scientist
Road network extraction from satellite imagery; building footprints and point-of-interest polygons at scale from OpenStreetMap.
Jun 2017 — Aug 2020
Netradyne Technologies · Bangalore, India
Senior Research Engineer
Object detection across geographies, model acceleration and compression for on-device analytics, and the in-house video data lake and model evaluation tooling.
  • VidParse is accepted to ECCV 2026. Online, training-free parsing of egocentric procedures.

  • Named an Outstanding Reviewer at CVPR 2025.

  • Three papers accepted to ECCV 2024 — two of them as shared first author.

  • Applied Scientist Intern at Amazon Fashion, on diffusion-based controllable indoor scene creation.

  • Applied Scientist Intern at Amazon Fashion.

Earlier news
  • Organized the workshop Dealing with Novelty in Open Worlds (DNOW) at WACV 2023.

  • Applied Scientist Intern (Co-op) at Amazon, working on visual geolocalization.

  • Organized ObjClsDisc: In-the-Wild Object Discovery Challenge, part of the Visual Perception and Learning in an Open World workshop at CVPR 2022.

  • Applied Scientist Intern at Amazon — edge-based object detection models, deployed for field trials.

  • Organized the workshop Dealing with Novelty in Open Worlds (DNOW) at WACV 2022.

  • PatchGame is accepted to NeurIPS 2021.

  • Began my MS in Computer Science at UMD.

  • Joined Swiggy as a Research Intern, then Data Scientist.

  • Built computer vision models and other cool things at Netradyne.

Reviewing

  • Outstanding Reviewer, CVPR 2025
  • Reviewer, ICRA 2025
  • Reviewer, CVPR 2024

Workshops organized

  • Open World Vision, CVPR 2023 · Vancouver
  • Dealing With Novelty in Open Worlds, WACV 2023 · Hawaii
  • Open World Vision, CVPR 2022 · New Orleans
  • Dealing With Novelty in Open Worlds, WACV 2022 · Hawaii

Education

  • PhD, Computer Science · University of Maryland, College Park
  • MS, Computer Science · University of Maryland, College Park
  • B.Tech., Electrical Engineering · IIT Delhi

Patent

Other

  • Advisor, Swasti — an NGO working in India (Nov 2024 – present)