llava
π LLaVA is a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding. Updated to version 1.6.
id model/llavavendor LLaVAfamily LLaVA
- Sizes
7b13b34b- Capabilities
- vision
- Model card
- https://ollama.com/library/llava β
Providers
| Provider | Model id | Get it | Tags | Pulls | Updated |
|---|---|---|---|---|---|
| Ollama β | llava | ollama pull llava | 98 | 14.8M | 2024-02-01 |
Same family
ollamalocalvision
Author LLaVA. Source ollama.com/library (original β). License unknown β the source states none; review before redistributing. Imported 2026-09-03. Description, capabilities, sizes, pulls, tag count and updated date as shown on the Ollama library page. Vendor, family and task are inferred from the model name; weights license is not published on the listing.