This week, the international community of industry and academic leaders shaping the future of computer vision and machine learning will gather for the 19th European Conference on Computer Vision (ECCV) in Malmö, Sweden. Adobe Researchers and their collaborators will present groundbreaking papers during the event, sharing experimental work in generative video and image technology that could give users more intuitive and precise control over the things they create.
“ECCV remains the top computer vision venue in Europe,” says Niloy Mitra, Adobe Research Principal Scientist and co-author of one of the papers selected for this year’s conference. “We expect to see breakthrough research combining spatial intelligence, multimodal LLMs, and consistent video generation.”
The Adobe Research ECCV papers: Camera-angles, 3D object orchestration, and more precise image generation
Adobe researchers will present GimbalDiffusion: Gravity-Aware Camera Control for Video Generation. GimbalDiffusion provides a missing element in text-to-video generation: fine-grained control over camera motion, especially with extreme trajectories, such as looking up or down. Like a physical gimbal that supports an object as it rotates, this framework allows camera control grounded in real-world coordinates for accurate, interpretable control over camera parameters.
The technology can tackle a full sphere of viewpoints, including extreme pitch and roll. It also adds null-pitch conditioning, which prevents the model from overriding camera specifications in the case of conflicting prompt content, such as generating grass when the camera points to the sky. The research includes new benchmarks for evaluating gravity-aware, camera-controlled video generation. See the project page and full paper here.

Adobe researchers will also share LooseControlVideo: Directorial Video Control Using Spatial Blocking, a new framework that addresses the challenge of 3D spatial orchestration in text-to-video generation. Current methods, which use depth-conditioned models for directing the movement of objects in a video, are complex and labor intensive. But with LooseControlVideo (LCV), users can edit using simple 3D boxes, offering an intuitive, expressive interface for directing object trajectories, rotations, occlusions, and camera motion.
The method also allows for localized refinement with minimal disruption to the global scene context. Evaluations show that LCV significantly outperforms existing methods for controlling video objects in 3D space. See the project page and full paper here.

Another paper by Adobe researchers tackles the challenge of using complex, real-world scenes as reference images for image generation. Existing models tend to look at an entire image, so they cannot reliably generate a new image using attributes from a specific element of an image. To solve this problem, Adobe researchers developed RefDiT, a novel new framework for reference-guided image generation.
RefDiT uses a reference image, a text prompt, and optional user-provided guidance as its input so that users can direct image generation based on a specific local element in a reference image. RefDiT achieves a local attribute matching score of 0.88, significantly outperforming other state-of-the-art methods. See more details about RefDiT here.

Putting users in control of generative AI
Adobe Research’s papers at ECCV are part of a bigger mission: to develop tools that give visual artists and filmmakers more control over the things they create with generative AI. To learn more about the work Adobe Research is doing in this area, check out the latest AI assistant in Photoshop, the experimental technology behind MotionStream, and a new feature that lets users rotate 2D objects using 3D technology.
Wondering what else is happening in Adobe Research? Check out our latest news here.