Customer-obsessed science
Research areas
-
July 30, 20268 min readInstead of compromising among parameter updates dictated by different training objectives, ControlG allocates computational capacity to objectives sequentially and dynamically.
-
-
July 9, 202610 min read
-
Featured news
-
2026Multimodal Large Language Models (MLLMs) have recently demonstrated promising capabilities in multimodal coding tasks such as chart-to-code generation. However, existing methods primarily rely on supervised fine-tuning (SFT), which requires the model to learn code patterns through chart-code pairs but does not expose the model to a code execution environment. Moreover, while self-correction through execution
-
VLDB 20262026Accurate optimizer statistics are fundamental to query and ML-prediction performance in modern database systems, yet maintaining them poses a significant challenge for large-scale data warehouses. Traditional statistics collection relies on full table scans, which become prohibitively expensive as tables grow to billions of rows and beyond. This creates a critical tension: statistics must be kept current
-
CVPR 2026 Workshop on Personalization in Generative AI2026Makeup transfer models enable fun augmented reality (AR) experiences as well as virtual try-on (VTO) for online makeup shopping. While recent state-of-the-art diffusion-based solutions such as Stable-Makeup [45] dramatically improve the accuracy and realism of makeup transfer, they still face limitations in identity and skin color preservation, making production-level VTO for makeup shopping unrealistic
-
2026Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physicsbased perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses, which are unable to generalize to complex real-world settings. We introduce ∆YNAMICS, a visionlanguage framework that uses language as a unified representation of rigid-body
-
CVPR 2026 Workshop on Fine-Grained Visual Categorization2026Fine-grained visual recognition demands attention to subtle, localized differences that current multimodal large language models (MLLMs) often overlook when guided by generic prompts. We propose APO-Pair, a prompt-optimization framework that learns classification rules by contrasting image pairs. A multimodal agent views these pairs, judges whether they depict the same fine-grained class, and iteratively
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all