OpenGVLab

community

https://github.com/opengvlab

opengvlab

OpenGVLab

Activity Feed Request to join this org

AI & ML interests

Computer Vision

Recent Activity

gulixin0922 updated a Space about 22 hours ago

OpenGVLab/InternVL

wzk1015 updated a model 1 day ago

OpenGVLab/PIIP

luotto authored a paper 2 days ago

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

View all activity

Organization Card

Community About org cards

OpenGVLab

Welcome to OpenGVLab! We are a research group from Shanghai AI Lab focused on Vision-Centric AI research. The GV in our name, OpenGVLab, means general vision, a general understanding of vision, so little effort is needed to adapt to new vision-based tasks.

Models

InternVL: a pioneering open-source alternative to GPT-4V.
InternImage: a large-scale vision foundation models with deformable convolutions.
InternVideo: large-scale video foundation models for multimodal understanding.
VideoChat: an end-to-end chat assistant for video comprehension.
All-Seeing-Project: towards panoptic visual recognition and understanding of the open world.

Datasets

ShareGPT4o: a groundbreaking large-scale resource that we plan to open-source with 200K meticulously annotated images, 10K videos with highly descriptive captions, and 10K audio files with detailed descriptions.
InternVid: a large-scale video-text dataset for multimodal understanding and generation.
MMPR: a high-quality, large-scale multimodal preference dataset.

Benchmarks

MVBench: a comprehensive benchmark for multimodal video understanding.
CRPE: a benchmark covering all elements of the relation triplets (subject, predicate, object), providing a systematic platform for the evaluation of relation comprehension ability.
MM-NIAH: a comprehensive benchmark for long multimodal documents comprehension.
GMAI-MMBench: a comprehensive multimodal evaluation benchmark towards general medical AI.

Collections 19

spaces 11

InternVL

VideoChat Flash

Hierarchical Compression for Long-Context Video Modeling

MVBench Leaderboard

Running on Zero

InternVideo2 Chat 8B HD

ControlLLM

Running on Zero

VideoMamba

models 140

OpenGVLab/PIIP

Object Detection • Updated 1 day ago • 4

OpenGVLab/VideoChat-Flash-Qwen2-7B_res224

Video-Text-to-Text • Updated 3 days ago • 124

OpenGVLab/VideoChat-Flash-Qwen2-7B_res448

Video-Text-to-Text • Updated 3 days ago • 462 • 4

OpenGVLab/VideoChat-Flash-Qwen2_5-2B_res448

Video-Text-to-Text • Updated 3 days ago • 301 • 5

OpenGVLab/VideoMAEv2-giant

Video Classification • Updated 4 days ago • 46 • 1

OpenGVLab/VideoMAEv2-Huge

Video Classification • Updated 4 days ago • 19

OpenGVLab/VideoMAEv2-Large

Video Classification • Updated 4 days ago • 19

OpenGVLab/VideoMAEv2-Base

Video Classification • Updated 4 days ago • 178 • 1

OpenGVLab/InternViT-300M-448px

Image Feature Extraction • Updated 10 days ago • 19.5k • 54

OpenGVLab/InternVL2_5-78B-MPO-AWQ

Image-Text-to-Text • Updated 12 days ago • 653 • 6

datasets 30

OpenGVLab/MMPR-v1.1

Preview • Updated 28 days ago • 1.05k • 35

OpenGVLab/MMPR

Preview • Updated 28 days ago • 457 • 44

OpenGVLab/GMAI-MMBench

Preview • Updated Dec 17, 2024 • 193 • 14

OpenGVLab/V2PE-Data

Preview • Updated Dec 14, 2024 • 728 • 5

OpenGVLab/InternVL-Domain-Adaptation-Data

Preview • Updated Dec 9, 2024 • 117 • 7

OpenGVLab/GUI-Odyssey

Viewer • Updated Nov 20, 2024 • 7.74k • 11.5k • 10

OpenGVLab/OmniCorpus-YT

Updated Nov 17, 2024 • 561 • 9

OpenGVLab/OmniCorpus-CC-210M

Viewer • Updated Nov 17, 2024 • 208M • 235 • 19

OpenGVLab/OmniCorpus-CC

Viewer • Updated Nov 17, 2024 • 986M • 24k • 12

OpenGVLab/MVBench

Viewer • Updated Oct 18, 2024 • 4k • 8.42k • 28