OpenAgents
OpenAgents - AI Agent Networks for Open Collaboration
Professional software vendor delivering innovative solutions on the Softono platform. Specialized in both open-source and proprietary software development.
OpenAgents - AI Agent Networks for Open Collaboration
OSWorld is a comprehensive benchmark and environment framework designed to evaluate multimodal AI agents performing open-ended tasks in real computer operating systems. Presented at NeurIPS 2024, the platform empowers researchers to test agents on complex, multi-step workflows requiring visual perception and interaction with standard desktop interfaces. The system supports a diverse range of tasks across web browsing, software usage, and file management within realistic virtualized environments. Key features include support for multiple virtualization backends such as VMware, VirtualBox, and cloud platforms like AWS, Azure, and Azure. Recent updates introduce OSWorld-Verified, which offers enhanced benchmark signals, community bug fixes, and significantly reduced evaluation times through AWS parallelization. The suite is compatible with Docker for server deployments and includes a dedicated desktop environment package for easy installation. Researchers can access detailed documentation, evaluation examples, a
OSWorld-G is a research project presented as a NeurIPS 2025 Spotlight focusing on scaling computer-use grounding through UI decomposition and synthesis. The project provides the official repository for the OSWorld-G benchmark and the Jedi data collection pipeline. It includes access to pre-trained models, specifically Jedi-3B and Jedi-7B, which are designed to predict precise click coordinates on user interfaces based on natural language instructions. The system leverages the Qwen-2.5-VL computer use agent template and is optimized for high-resolution inputs. Key components of the release include the OSWorld-G benchmark containing original and refined instruction sets for pure grounding tasks, as well as the comprehensive Jedi dataset comprising over 4 million data points. The dataset pipeline covers icon recognition, component rendering, real-world augmentation for documents and spreadsheets, layout analysis, and refusal data collection. The software is built on Python 3.9 or higher with dependencies managed
CUA-Gym-Hub: mock web apps as reproducible RL training environments for computer-use agents