Multimodal Garment Designer
This is the official repository for the paper "Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing". ICCV 2023
Professional software vendor delivering innovative solutions on the Softono platform. Specialized in both open-source and proprietary software development.
This is the official repository for the paper "Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing". ICCV 2023
ReflectiVA is a framework that extends multimodal large language models with self‑reflective tokens, enabling them to dynamically query external knowledge bases for visual‑question‑answering tasks that require factual information beyond their training data. The system supports retrieval‑augmented inference, integrates with Hugging Face models and datasets, and includes scripts for installing dependencies, training, and evaluating on datasets such as Infoseek and Encyclopedic‑VQA. Use cases include answering knowledge‑intensive visual questions, combining image understanding with external fact retrieval while preserving natural‑language fluency.