Salesforce Research

BLIP-2

Efficient vision-language pre-training.

State-of-the-art performanceEfficient pre-trainingImage captioningVisual Question Answering (VQA)
Today's score
82.0
Try BLIP-2

Where it ranks today

Best for / Not great for

Best for
  • Image captioning
  • Visual QA
  • Image-text retrieval
  • Foundation for multimodal research
Not great for
  • Audio or video processing
  • Generating complex, long-form text
  • Real-time conversational AI
  • Directly controlling robots

Why it ranks here

BLIP-2 offers a highly effective and efficient approach to vision-language pre-training, achieving strong results on various benchmarks. Its focus on efficiency makes it a popular choice for researchers and developers looking for solid multimodal capabilities without extreme computational cost.

30-day trend

Score breakdown

Search trends83
Benchmarks82
Developer buzz83
News mentions81

Pricing

API: $0.00 in · $0.00 out per 1M tokens · Consumer: $0.00/mo

Pricing plans

Popular
Open Source
Full access to code and weights.
Free
  • Pre-trained models
  • Codebase
  • Research focused
  • Permissive license
Get the code
Cloud Deployment (via platforms)
Managed service for BLIP-2.
$0.20 /usage
  • Easy deployment
  • Scalable
  • API access
  • Pay for usage
Use via API
Compare with another modelHow is this score calculated? →Snapshot 2026-08-15