xiangan's picture

xiangan

xiangan

·

https://anxiangsir.github.io/

anxiangsir

AI & ML interests

None yet

Recent Activity

liked a model 5 days ago

FreedomIntelligence/openPangu-VL-7B

liked a dataset 6 days ago

ByteDance/AncientDoc

upvoted a paper 8 days ago

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

View all activity

Organizations

authored 4 papers 3 months ago

ForCenNet: Foreground-Centric Network for Document Image Rectification

Paper • 2507.19804 • Published Jul 26, 2025 • 12

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

Paper • 2509.09118 • Published Sep 11, 2025 • 8

UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning

Paper • 2510.13515 • Published Oct 15, 2025 • 12

ORID: Organ-Regional Information Driven Framework for Radiology Report Generation

Paper • 2411.13025 • Published Nov 20, 2024 • 2

authored a paper 4 months ago

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Paper • 2509.23661 • Published Sep 28, 2025 • 48

authored a paper 6 months ago

Region-based Cluster Discrimination for Visual Representation Learning

Paper • 2507.20025 • Published Jul 26, 2025 • 19

authored a paper 11 months ago

RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm

Paper • 2502.12513 • Published Feb 18, 2025 • 16

authored a paper about 1 year ago

Croc: Pretraining Large Multimodal Models with Cross-Modal Comprehension

Paper • 2410.14332 • Published Oct 18, 2024 • 1