Open-Source Community / LMSYS

LLaVA (Large Language and Vision Assistant)

Open-source instruction-following multimodal model.

Combines LLM with visual encoderInstruction-following capabilitiesOpen-source and adaptableGood for visual chat applications

Where it ranks today

Best for / Not great for

Best for
  • Visual question answering (VQA)
  • Image-based chatbots
  • Generating descriptions for images
  • Educational VQA tools
Not great for
  • Audio or video processing
  • Highly complex reasoning across many modalities
  • State-of-the-art performance in generation

Why it ranks here

LLaVA represents the power of open-source collaboration in creating capable multimodal models. It bridges the gap between large language models and visual understanding, making it a popular choice for research and specific applications.

30-day trend

Score breakdown

Search trends82
Benchmarks81
Developer buzz83
News mentions81

Pricing

API: $0.00 in · $0.00 out per 1M tokens · Consumer: $0.00/mo

Pricing plans

Popular
Open Source
Freely available for research and non-commercial use.
Free
  • Model weights and code
  • Customizable architecture
  • Active GitHub community
  • Instruction-tuned
Get LLaVA
Hosted Inference API
Easy-to-use API for LLaVA.
$0.05 /usage
  • Managed infrastructure
  • Pay-per-request
  • Quick integration
  • Supports various LLaVA versions
Try API
Compare with another modelHow is this score calculated? →Snapshot 2026-08-17