Expanding Performance Boundaries of Open-Source MLLM
OpenGVLab
community
AI & ML interests
Computer Vision
Organization Card
OpenGVLab
Welcome to OpenGVLab! We are a research group from Shanghai AI Lab focused on Vision-Centric AI research. The GV in our name, OpenGVLab, means general vision, a general understanding of vision, so little effort is needed to adapt to new vision-based tasks.
Models
- InternVL: a pioneering open-source alternative to GPT-4V.
- InternImage: a large-scale vision foundation models with deformable convolutions.
- InternVideo: large-scale video foundation models for multimodal understanding.
- VideoChat: an end-to-end chat assistant for video comprehension.
- All-Seeing-Project: towards panoptic visual recognition and understanding of the open world.
Datasets
- ShareGPT4o: a groundbreaking large-scale resource that we plan to open-source with 200K meticulously annotated images, 10K videos with highly descriptive captions, and 10K audio files with detailed descriptions.
- InternVid: a large-scale video-text dataset for multimodal understanding and generation.
Benchmarks
- MVBench: a comprehensive benchmark for multimodal video understanding.
- CRPE: a benchmark covering all elements of the relation triplets (subject, predicate, object), providing a systematic platform for the evaluation of relation comprehension ability.
- MM-NIAH: a comprehensive benchmark for long multimodal documents comprehension.
- GMAI-MMBench: a comprehensive multimodal evaluation benchmark towards general medical AI.
models
89
OpenGVLab/Mono-InternVL-2B
Image-Text-to-Text
•
Updated
•
17.1k
•
21
OpenGVLab/InternVL
Updated
•
20
OpenGVLab/InternVideo2-Chat-8B
Video-Text-to-Text
•
Updated
•
1.86k
•
17
OpenGVLab/VideoChat2_HD_stage4_Mistral_7B_hf
Updated
•
250
OpenGVLab/InternVL2-8B
Image-Text-to-Text
•
Updated
•
53.8k
•
143
OpenGVLab/InternVL2-4B
Image-Text-to-Text
•
Updated
•
92.3k
•
42
OpenGVLab/InternVL2-Llama3-76B-AWQ
Image-Text-to-Text
•
Updated
•
1.18k
•
24
OpenGVLab/InternVL2-40B-AWQ
Image-Text-to-Text
•
Updated
•
775
•
17
OpenGVLab/InternVL2-26B-AWQ
Image-Text-to-Text
•
Updated
•
468
•
18
OpenGVLab/InternVL2-8B-AWQ
Image-Text-to-Text
•
Updated
•
2.16k
•
12
datasets
29
OpenGVLab/OmniCorpus-CC
Viewer
•
Updated
•
986M
•
10.8k
•
7
OpenGVLab/InternVL-Domain-Adaptation-Data
Preview
•
Updated
•
171
•
1
OpenGVLab/MVBench
Viewer
•
Updated
•
4k
•
6.17k
•
27
OpenGVLab/GMAI-MMBench
Preview
•
Updated
•
500
•
13
OpenGVLab/OmniCorpus-YT
Updated
•
507
•
7
OpenGVLab/VideoMAEv2-TAL-Features
Preview
•
Updated
•
44
OpenGVLab/InternVL-SA-1B-Caption
Viewer
•
Updated
•
8.63M
•
99
•
7
OpenGVLab/InternVL-Chat-V1-2-SFT-Data
Viewer
•
Updated
•
573k
•
950
•
11
OpenGVLab/InternVL-LaionCOCO-OCR
Updated
•
46
•
2
OpenGVLab/InternVL-WuKong-OCR
Updated
•
45
•
2