ADOPD (2024, 2026)
Ongoing work on understanding and generating multimodal documents
This project explores treating documents as a first-class multimodal modality, bridging vision and language to both interpret and produce richly structured content. Rather than separating comprehension from creation, it investigates whether a single framework can reason over document layouts and semantics while also generating document-style visual content, extending the reach of general-purpose multimodal systems.
ADOPD: A Large-Scale Document Page Decomposition Dataset (ICLR 2024 poster).
Focus
- Unifying document understanding and generation within one modeling framework, rather than relying on task-specific components.
- Grounding the structural and layout properties that distinguish documents from ordinary images.
- Connecting the visual and textual dimensions of document content through shared reasoning pathways.
- Learning general-purpose document representations that transfer across a broad range of document types and tasks.
Demos
Parallel, streaming generation of document structure — bounding boxes and polygon masks over a full page.
Agentic detection with self-refinement: over-recalled regions are grouped and merged into grounded evidence.
This work is currently under review; more details will be shared once it is published.