ForCenNet: Foreground-Centric Network for Document Image Rectification Paper • 2507.19804 • Published Jul 26 • 11
Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval Paper • 2509.09118 • Published Sep 11 • 8
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning Paper • 2510.13515 • Published 13 days ago • 11
ORID: Organ-Regional Information Driven Framework for Radiology Report Generation Paper • 2411.13025 • Published Nov 20, 2024 • 2
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs Paper • 2411.15296 • Published Nov 22, 2024 • 21
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Paper • 2501.13826 • Published Jan 23 • 25
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Paper • 2509.23661 • Published 30 days ago • 44
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Paper • 2509.23661 • Published 30 days ago • 44
Seedream 4.0: Toward Next-generation Multimodal Image Generation Paper • 2509.20427 • Published Sep 24 • 75
Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation Paper • 2509.18824 • Published Sep 23 • 22
Region-based Cluster Discrimination for Visual Representation Learning Paper • 2507.20025 • Published Jul 26 • 19
MobileVOS: Real-Time Video Object Segmentation Contrastive Learning meets Knowledge Distillation Paper • 2303.07815 • Published Mar 14, 2023 • 1
$V_kD:$ Improving Knowledge Distillation using Orthogonal Projections Paper • 2403.06213 • Published Mar 10, 2024 • 2
Region-based Cluster Discrimination for Visual Representation Learning Paper • 2507.20025 • Published Jul 26 • 19
Long-tailed Instance Segmentation using Gumbel Optimized Loss Paper • 2207.10936 • Published Jul 22, 2022