SkillsHooksPromptsAgentsPersonasModelsPoliciesToolsTemplatesBundlesCategoriesStart here
← Models
Mo modelmultimodalaistable

llava

πŸŒ‹ LLaVA is a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding. Updated to version 1.6.

id model/llavavendor LLaVAfamily LLaVA
Sizes
7b13b34b
Capabilities
vision
Model card
https://ollama.com/library/llava β†—

Providers

ProviderModel idGet itTagsPullsUpdated
Ollama β†—llavaollama pull llava9814.8M2024-02-01

Same family

ollamalocalvision
Author LLaVA. Source ollama.com/library (original β†—). License unknown β€” the source states none; review before redistributing. Imported 2026-09-03. Description, capabilities, sizes, pulls, tag count and updated date as shown on the Ollama library page. Vendor, family and task are inferred from the model name; weights license is not published on the listing.