Skillzwave

blip-2-vision-language

422 stars 30 forks Updated Dec 17, 2025
48.0

Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.

Marketplace
#errors Error#solutions#image#image captioning#Model
Also in: machine learning docker ci cd

Third-Party Skill: Review the code before installing. Skills execute in your AI assistant's environment and can access your files. Learn more about security

skilz install zechenzhangAGI_AI-research-SKILLs/blip-2-vision-language
skilz install zechenzhangAGI_AI-research-SKILLs/blip-2-vision-language --agent opencode
skilz install zechenzhangAGI_AI-research-SKILLs/blip-2-vision-language --agent codex
skilz install zechenzhangAGI_AI-research-SKILLs/blip-2-vision-language --agent gemini

First time? Install Skilz: pip install skilz

Works with 14 AI coding assistants

Cursor, Aider, Copilot, Windsurf, Qwen, Kimi, and more...

View All Agents
Download Skill ZIP

Extract and copy to ~/.claude/skills/ then restart Claude Desktop

1. Clone the repository:
git clone https://github.com/zechenzhangAGI/AI-research-SKILLs
2. Copy the skill directory:
cp -r AI-research-SKILLs/18-multimodal/blip-2 ~/.claude/skills/

Need detailed installation help? Check our platform-specific guides:

Related Skills

Details

Stars
422
Forks
30
Type
Technical
Meta-Domain
media
Primary Domain
image
Sub-Domain
images text
Skill Size
48.3 KB
Files
3
Quality Score
48.0

AI-Detected Topics

Extracted using NLP analysis

errors Error solutions image image captioning Model

Browse Category

More media skills

Report Security Issue

Found a security vulnerability in this skill?