Home
Softono
w

wisconsinaivision

Professional software vendor delivering innovative solutions on the Softono platform. Specialized in both open-source and proprietary software development.

Total Products
1

Software by wisconsinaivision

ViP LLaVA
Open Source

ViP LLaVA

ViP-LLaVA is a CVPR 2024 accepted research project that enables Large Multimodal Models to understand arbitrary visual prompts. Developed by researchers from Cruise LLC and the University of Wisconsin-Madison, it introduces a technique where visual prompts such as bounding boxes or scribbles are directly overlaid onto the original image during visual instruction tuning. This approach allows the model to interpret user-defined regions and answer questions about specific image areas in a natural, zero-shot manner without requiring fine-tuning for every new prompt type. The framework supports modern LLM backbones including Llama-3-8B and Phi-3-mini. The project includes ViP-Bench, the first zero-shot region-level benchmark designed to evaluate large multimodal models on visual prompt understanding, along with an official evaluation server and leaderboard. ViP-LLaVA is integrated into Hugging Face Transformers and provides pre-trained weights for various model sizes. It is distributed under an Apache 2.0 code lic

ML Frameworks
339 Github Stars