Hi,
I want to train a Multi-modal using Image and Text for Multi-label classification.
Can you please help me to understand what latest multi-modal are available that takes image and text as an input and fine-tune on my classification task.
Looking forward to your reply.
thanks
Hi,
I want to train a Multi-modal using Image and Text for Multi-label classification.
Can you please help me to understand what latest multi-modal are available that takes image and text as an input and fine-tune on my classification task.
Looking forward to your reply.
thanks