PointLLM
PointLLM is a multi-modal large language model designed to enable advanced understanding of colored 3D point clouds. By processing geometric structures, object types, and visual appearance directly, it overcomes common challenges in 3D perception such as ambiguous depth, occlusion, and viewpoint dependency. The system employs a two-stage training strategy supported by a proprietary dataset containing over 730,000 instruction-text pairs covering both simple and complex point cloud scenarios. PointLLM is capable of performing tasks like generative 3D object classification and automated 3D object captioning with high precision. To ensure rigorous assessment of its perceptual and generalization abilities, the developers established specific benchmarks evaluated through multiple methods. This research has achieved notable recognition, including selection as a Best Paper candidate at ECCV 2024 and acceptance in TPAMI 2025 with an improved version known as PointLLM-V2. The project provides open-source code for train