Anubhav Gupta — Univ. of Maryland

§ 01

Research

I work on video understanding — mostly on what a model should hold on to from a long video, and how much of it can be worked out while the video is still playing rather than after the fact. Most of my thesis has gone into the streaming version of that, in egocentric procedural video: cooking, assembly, lab protocols, where you get one frame at a time, no future context and no second pass. That rules out most of what the field currently does well.

VidParse (ECCV 2026) does it without training anything. It segments the video using features anchored in hand–object interaction, then decodes the sequence of steps with a beam search that can only follow transitions the task graph permits. Everything runs off frozen foundation models, and on multi-step parsing it beats trained online baselines by up to 10×.

The task graph is the weak part: VidParse has to commit to one up front. What I’m working on now is using a VLM to build a hierarchical memory in a single streaming pass, structured so that different questions can be answered from it later.

Earlier work was about what learned representations already encode. Latent-INR found discriminative semantics inside implicit neural representations of video, which were supposed to be doing reconstruction. LEIA learned articulation latents that survive a change of viewpoint. CSD made style a measurable, separable property of diffusion output. TREND, from a robotics collaboration, is preference-based RL that holds up when the preferences — human or VLM — are noisy.

I spent eight years in industry before the PhD: analyst work at a few banks, then computer vision for automotive safety at Netradyne and geospatial ML at Swiggy, plus four summers at Amazon during it.

The direction I am trying to pull these threads toward is robotics and continual learning: systems that keep learning from what they watch, instead of being trained once and then frozen.

Video understanding Egocentric & procedural video Representation learning Vision–language models

§ 02

Publications

In submission

Two Party Evaluation of the Open Visual World

Defining a paradigm for unbiased evaluation in open worlds.

Under review

§ 03

Patent

US 10,782,654

Detection of driving actions that mitigate risk

Granted September 2020 · Assigned to Netradyne, Inc.

Detecting the driving manoeuvres that actively reduce risk, rather than only the ones that cause it, from vehicle-mounted video and telemetry analysed on-device. From my time building automotive-safety computer vision at Netradyne.

§ 04

Experience

Jan 2021 — present
University of Maryland, College Park
Graduate Research Assistant · PhD, Computer Science
Video understanding — representation, temporal structure, and vision–language models. Advised by Abhinav Shrivastava.
May — Aug 2024
Amazon Fashion · Sunnyvale, CA
Applied Scientist Intern
Diffusion-based multi-view image editing for indoor scenes; built a large-scale dataset starting from 3D-FRONT.
May — Aug 2023
Amazon Fashion · Sunnyvale, CA
Applied Scientist Intern
Keystroke-assisted human body editing (pose and shape) with Stable Diffusion 1.5 and ControlNets, aimed at 2D virtual try-on.
Oct — Dec 2022
Amazon
Applied Scientist Intern (Co-op)
Visual geolocalization.
May — Aug 2022
Amazon · Seattle, WA
Applied Scientist Intern
Edge-based object detection; took a prototype model to production for field trials.
Sep 2020 — Jan 2021
Swiggy · Bangalore, India
Data Scientist
Road network extraction from satellite imagery; building footprints and point-of-interest polygons at scale from OpenStreetMap.
Jun 2017 — Aug 2020
Netradyne Technologies · Bangalore, India
Senior Research Engineer
Object detection across geographies, model acceleration and compression for on-device analytics, and the in-house video data lake and model evaluation tooling.

§ 05

News

  • VidParse is accepted to ECCV 2026. Online, training-free parsing of egocentric procedures.

  • Named an Outstanding Reviewer at CVPR 2025.

  • Three papers accepted to ECCV 2024 — two of them as shared first author.

  • Applied Scientist Intern at Amazon Fashion, on diffusion-based controllable indoor scene creation.

  • Applied Scientist Intern at Amazon Fashion.

Earlier news
  • Organized the workshop Dealing with Novelty in Open Worlds (DNOW) at WACV 2023.

  • Applied Scientist Intern (Co-op) at Amazon, working on visual geolocalization.

  • Organized ObjClsDisc: In-the-Wild Object Discovery Challenge, part of the Visual Perception and Learning in an Open World workshop at CVPR 2022.

  • Applied Scientist Intern at Amazon — edge-based object detection models, deployed for field trials.

  • Organized the workshop Dealing with Novelty in Open Worlds (DNOW) at WACV 2022.

  • PatchGame is accepted to NeurIPS 2021.

  • Began my MS in Computer Science at UMD.

  • Joined Swiggy as a Research Intern, then Data Scientist.

  • Built computer vision models and other cool things at Netradyne.

§ 06

Service & more

Reviewing

  • Outstanding Reviewer, CVPR 2025
  • Reviewer, ICRA 2025
  • Reviewer, CVPR 2024

Workshops organized

  • Open World VisionCVPR 2023 · Vancouver
  • Dealing With Novelty in Open WorldsWACV 2023 · Hawaii
  • Open World VisionCVPR 2022 · New Orleans
  • Dealing With Novelty in Open WorldsWACV 2022 · Hawaii

Education

  • PhD, Computer ScienceUniversity of Maryland, College Park
  • MS, Computer ScienceUniversity of Maryland, College Park
  • B.Tech., Electrical EngineeringIIT Delhi

Other

  • Advisor, SwastiAn NGO working in India · Nov 2024 – present