Skip to content
#

image-text-alignment

Here are 8 public repositories matching this topic...

Language: All
Filter by language

Vision Transformer (ViT)-based pipeline for multimodal image knowledge extraction: fine-grained botanical taxonomy, cultural landmark recognition, and semantic object analysis. Combines pretrained ViTs, domain adapters, and generative language models to produce structured annotations, contextual metadata, and adaptive study resources. + RAG support

  • Updated Aug 8, 2026
  • Python

Add this topic to your repo

To associate your repository with the image-text-alignment topic, visit your repo's landing page and select "manage topics."

Learn more