Home
Softono

ECCV2024 Papers With Code

Open source
2.3K
Stars
271
Forks
9
Issues
35
Watchers
2 years
Last Commit

 About ECCV2024 Papers With Code

ECCV 2024 论文和开源项目合集,同时欢迎各位大佬提交issue,分享ECCV 2024论文和开源项目

Platforms

Web Self-hosted

Need Help Installing ECCV2024 Papers With Code?

We provide expert installation service for this software. Our team will install, configure, and secure ECCV2024 Papers With Code on your server. plans start at just $30.

ECCV2024 Papers With Code

View on GitHub

ECCV 2024 论文和开源项目合集(Papers with Code)

ECCV 2024 decisions are now available!

注1:欢迎各位大佬提交issue,分享ECCV 2024论文和开源项目!

注2:关于往年CV顶会论文以及其他优质CV论文和大盘点,详见: https://github.com/amusi/daily-paper-computer-vision

想看ECCV 2024和最新最全的顶会工作,欢迎扫码加入【CVer学术交流群】,这是最大的计算机视觉AI知识星球!每日更新,第一时间分享最新最前沿的计算机视觉、深度学习、自动驾驶、医疗影像和AIGC等方向的学习资料,学起来!

【ECCV 2024 论文开源目录】

3DGS(Gaussian Splatting)

MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images

CityGaussian: Real-time High-quality Large-Scale Scene Rendering with Gaussians

FSGS: Real-Time Few-shot View Synthesis using Gaussian Splatting

Mamba / SSM

VideoMamba: State Space Model for Efficient Video Understanding

ZIGMA: A DiT-style Zigzag Mamba Diffusion Model

Avatars

Backbone

CLIP

MAE

Embodied AI

GAN

OCR

Bridging Synthetic and Real Worlds for Pre-training Scene Text Detectors

PosFormer: Recognizing Complex Handwritten Mathematical Expression with Position Forest Transformer

Occupancy

Fully Sparse 3D Occupancy Prediction

NeRF

NeRF-MAE: Masked AutoEncoders for Self-Supervised 3D Representation Learning for Neural Radiance Fields

DETR

Prompt

多模态大语言模型(MLLM)

SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant

ControlCap: Controllable Region-level Captioning

大语言模型(LLM)

NAS

ReID(重识别)

扩散模型(Diffusion Models)

ZIGMA: A DiT-style Zigzag Mamba Diffusion Model

Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation

The Lottery Ticket Hypothesis in Denoising: Towards Semantic-Driven Initialization

Vision Transformer

GiT: Towards Generalist Vision Transformer through Universal Language Interface

视觉和语言(Vision-Language)

GalLoP: Learning Global and Local Prompts for Vision-Language Models

  • Paper:https://arxiv.org/abs/2407.01400

目标检测(Object Detection)

Relation DETR: Exploring Explicit Position Relation Prior for Object Detection

Cross-Domain Few-Shot Object Detection via Enhanced Open-Set Object Detector

异常检测(Anomaly Detection)

目标跟踪(Object Tracking)

语义分割(Semantic Segmentation)

Context-Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation

医学图像(Medical Image)

Brain-ID: Learning Contrast-agnostic Anatomical Representations for Brain Imaging

FairDomain: Achieving Fairness in Cross-Domain Medical Image Segmentation and Classification

医学图像分割(Medical Image Segmentation)

ScribblePrompt: Fast and Flexible Interactive Segmentation for Any Biomedical Image

AnatoMask: Enhancing Medical Image Segmentation with Reconstruction-guided Self-masking

Representing Topological Self-Similarity Using Fractal Feature Maps for Accurate Segmentation of Tubular Structures

视频目标分割(Video Object Segmentation)

DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries

自动驾驶(Autonomous Driving)

Fully Sparse 3D Occupancy Prediction

milliFlow: Scene Flow Estimation on mmWave Radar Point Cloud for Human Motion Sensing

4D Contrastive Superflows are Dense 3D Representation Learners

3D点云(3D-Point-Cloud)

3D目标检测(3D Object Detection)

3D Small Object Detection with Dynamic Spatial Pruning

Ray Denoising: Depth-aware Hard Negative Sampling for Multi-view 3D Object Detection

3D语义分割(3D Semantic Segmentation)

图像编辑(Image Editing)

图像补全/图像修复(Image Inpainting)

BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion

视频编辑(Video Editing)

Low-level Vision

Restoring Images in Adverse Weather Conditions via Histogram Transformer

OneRestore: A Universal Restoration Framework for Composite Degradation

超分辨率(Super-Resolution)

去噪(Denoising)

图像去噪(Image Denoising)

3D人体姿态估计(3D Human Pose Estimation)

图像生成(Image Generation)

Object-Conditioned Energy-Based Attention Map Alignment in Text-to-Image Diffusion Models

Every Pixel Has its Moments: Ultra-High-Resolution Unpaired Image-to-Image Translation via Dense Normalization

ZIGMA: A DiT-style Zigzag Mamba Diffusion Model

Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation

视频生成(Video Generation)

VideoStudio: Generating Consistent-Content and Multi-Scene Videos

3D生成

视频理解(Video Understanding)

VideoMamba: State Space Model for Efficient Video Understanding

C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition

行为识别(Action Recognition)

SA-DVAE: Improving Zero-Shot Skeleton-Based Action Recognition by Disentangled Variational Autoencoders

知识蒸馏(Knowledge Distillation)

图像压缩(Image Compression)

Image Compression for Machine and Human Vision With Spatial-Frequency Adaptation

立体匹配(Stereo Matching)

场景图生成(Scene Graph Generation)

计数(Counting)

Zero-shot Object Counting with Good Exemplars

视频质量评价(Video Quality Assessment)

数据集(Datasets)

其他(Others)

Multi-branch Collaborative Learning Network for 3D Visual Grounding

PDiscoFormer: Relaxing Part Discovery Constraints with Vision Transformers

SPVLoc: Semantic Panoramic Viewport Matching for 6D Camera Localization in Unseen Environments

REFRAME: Reflective Surface Real-Time Rendering for Mobile Devices