My name is Jiuxiang Gu (顾久祥). I am a Senior Research Scientist at Adobe Research in Seattle. I received my Ph.D. from Nanyang Technological University, Singapore (2016.1–2019.5), under the supervision of Prof. Jianfei Cai, Dr. Gang Wang, and Prof. Tsuhan Chen. My research journey began in hardware design. From 2010 to 2015, I worked as a hardware engineer (ASIC, FPGA, and PCB design). In 2015, I made the transition to Artificial Intelligence.
My current research focuses on multimodal foundation models, efficient AI, reasoning, and document intelligence.
Outside of research, I enjoy hiking and exploring the outdoors, as well as 3D printing, painting, and designing and building robots.
Open to collaborations and internships in the above areas.
Latest Updates
All updates-
LaViDa-R1, our reasoning model for unified multimodal diffusion, appears at ICML 2026.
-
Sparse-LaViDa, our efficient sparse multimodal diffusion model, appears at CVPR 2026.
-
FLARE, our framework for converting hybrid autoregressive models into diffusion language models, is now available.
-
LaViDa-O, our unified diffusion model for multimodal understanding and generation, appears at ICLR 2026.
-
FastCar, our work on efficient autoregressive video generation for edge devices, appears at ICLR 2026.
-
OIDA-QA, our multimodal benchmark for opioid-industry document analysis, appears at AAAI 2026.
Selected Publications
2026
-
CVPR 2026Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models
-
ICLR 2026LaViDa-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
2025
-
AAAI 2025Numerical pruning for efficient autoregressive models
2024
-
ICLR 2024Lrm: Large reconstruction model for single image to 3d
-
ICLR 2024ADoPD: A large-scale document page decomposition dataset
2021
-
NeurIPS 2021Unidoc: Unified pretraining framework for document understanding
2018
-
AAAI 2018Stack-Captioning: Coarse-to-Fine Learning for Image Captioning
-
CVPR 2018Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models
-
Pattern Recognition, 2018Recent advances in convolutional neural networks