zhiqiulin
/

clip-flant5-xl

Text2Text Generation

text-generation-inference

Inference Endpoints

Model card Files Files and versions Community

Edit model card

CLIP-FlanT5-XL (VQAScore)

This model is a fine-tuned version of google/flan-t5-xl designed for image-text retrieval tasks, as presented in the VQAScore paper.

Model Description

Developed by: Zhiqiu Lin and collaborators
Model type: Vision-Language Generative Model
License: Apache-2.0
Finetuned from model: google/flan-t5-xxl

Model Sources [optional]

Repository: https://github.com/linzhiqiu/CLIP-FlanT5
Paper: https://arxiv.org/pdf/2404.01291
Demo: https://huggingface.co/spaces/zhiqiulin/VQAScore

Downloads last month: 4,836

Inference Examples

Text2Text Generation

This model does not have enough activity to be deployed to Inference API (serverless) yet. Increase its social visibility and check back later, or deploy to Inference Endpoints (dedicated) instead.

Model tree for zhiqiulin/clip-flant5-xl

Base model

google/flan-t5-xl

Finetuned

(19)

this model