Microsoft Research

KOSMOS-2

Multimodal model for grounding language models to vision.

Vision-language groundingImage captioningVisual Question AnsweringOpen-source
Today's score
85.0
Try KOSMOS-2

Where it ranks today

Best for / Not great for

Best for
  • Describing image regions
  • Answering questions about images
  • Multimodal dialogue systems
  • Research in grounded AI
Not great for
  • Audio processing
  • Generating photorealistic images
  • Complex audio-visual reasoning
  • Large-scale real-time applications

Why it ranks here

KOSMOS-2 remains a valuable model for tasks requiring strong grounding between visual elements and language. Its open-source availability and focus on multimodal understanding contribute to its continued presence in the field.

30-day trend

Score breakdown

Search trends84
Benchmarks86
Developer buzz87
News mentions84

Pricing

API: $0.00 in · $0.00 out per 1M tokens · Consumer: $0.00/mo

Pricing plans

Popular
Open Source
Access model weights for research.
Free
  • Grounding capabilities
  • Image captioning
  • VQA support
  • Requires integration
Explore model
Compare with another modelHow is this score calculated? →Snapshot 2026-08-03