Skillzwave

ai-multimodal

2 stars 1 forks Updated Nov 21, 2025
72.0

Process and generate multimedia content using Google Gemini API. Capabilities include analyze audio files (transcription with timestamps, summarization, speech understanding, music/sound analysis up to 9.5 hours), understand images (captioning, object detection, OCR, visual Q&A, segmentation), process videos (scene detection, Q&A, temporal analysis, YouTube URLs, up to 6 hours), extract from documents (PDF tables, forms, charts, diagrams, multi-page), generate images (text-to-image with Imagen 4

Commands Hooks
#ai multimodal description#multimedia ai#process audio images#analysis youtube processing#reference images ai
Also in: video github api

Third-Party Skill: Review the code before installing. Skills execute in your AI assistant's environment and can access your files. Learn more about security

skilz install The1Studio_theone-training-skills/ai-multimodal
skilz install The1Studio_theone-training-skills/ai-multimodal --agent opencode
skilz install The1Studio_theone-training-skills/ai-multimodal --agent codex
skilz install The1Studio_theone-training-skills/ai-multimodal --agent gemini

First time? Install Skilz: pip install skilz

Works with 14 AI coding assistants

Cursor, Aider, Copilot, Windsurf, Qwen, Kimi, and more...

View All Agents
Download Skill ZIP

Extract and copy to ~/.claude/skills/ then restart Claude Desktop

1. Clone the repository:
git clone https://github.com/The1Studio/theone-training-skills
2. Copy the skill directory:
cp -r theone-training-skills/.claude/skills/ai-multimodal ~/.claude/skills/

Need detailed installation help? Check our platform-specific guides:

Related Skills

Details

Stars
2
Forks
1
Type
Technical
Meta-Domain
media
Primary Domain
image
Sub-Domain
audio hours
Skill Size
344.8 KB
Files
26
Quality Score
72.0

AI-Detected Topics

Extracted using NLP analysis

ai multimodal description multimedia ai process audio images analysis youtube processing reference images ai

Browse Category

More media skills

Report Security Issue

Found a security vulnerability in this skill?