Table of Contents
This is a repository made to demonstrate the cability and usefulness of Llama 3.2 vision, it is meant to show how the model can do in a professional use case, an example of how
it should do is on Blind-Spot which uses Google Gemini. The whole point is to test whether or not this model can act as a good replacement and if it exceeds the performance and usefulness of Gemini, the next phase of testing will occur after work in this repository is done which will be the use of an agent through one of the models or both, to test which is better and should be used in the for seeable future.
The tools used are:
The way the image description will be demonstrated is via a video on a basic Streamlit Application that will use the model to take in various images and show the description out to the user, the whole point is to mimic Blind-Spots and show if any changes can be made to the model to make better descriptions and help blind people better through the use of a new model.
Video Link TBD
As stated before the point is to create an agent using either Gemini or Ollama to provide better descriptions to the user while ensuring the best possible resources and model. Any future repository made based on this data will be listed below.
