Skip to content

Popular Media & Files Skills

PDF, spreadsheets, slides, images, audio, and video workflows.

969verified Agent Skills

Media and file skills help agents work with the formats that carry day-to-day business information. They can extract content, make targeted edits, and create polished deliverables. Check local file access and external processing requirements before handling confidential material.

Top Media & Files Skills

Ranked by their position in the current overall directory snapshot.

#1534

Deep Research (Gemini)

Async deep research via Gemini Interactions API (no Gemini CLI dependency). RAG-ground queries on local files (--context), preview costs (--dry-run), structu.

MediaOpenClaw
#1535

Youtube Apify Transcript

Fetch YouTube transcripts via the Apify API. Works from cloud IPs (Hetzner, AWS, etc.) by bypassing YouTube's bot detection. Features local caching (free repeat requests) and batch mode. Requires APIFY_API_TOKEN and the Python requests library.

MediaOpenClaw
#1543

Notebooklm Cli

Command-line interface to manage Google NotebookLM notebooks, sources, and generate audio, quizzes, reports, presentations, and visual study materials progra.

MediaOpenClaw
#1544

Elicitation How To Talk With Humans And Ask Them Questions?

Psychological profiling through natural conversation using narrative identity research (McAdams), self-defining memory elicitation (Singer), and Motivational Interviewing (OARS framework). Use when you need to: (1) understand someone's core values and motivations, (2) discover formative memories and life-defining experiences, (3) detect emotional schemas and belief patterns, (4) build psychological profiles through gradual disclosure, (5) conduct user interviews that reveal deep insights, (6) design conversational flows for personal discovery, (7) identify identity themes like redemption and contamination narratives, (8) elicit authentic self-disclosure without interrogation.

MediaOpenClaw
#1559

短视频脚本创作

短视频脚本创作技能。用于生成抖音、快手、B站、视频号等平台的短视频脚本、标题、封面建议。适用于自媒体运营、内容创作需求。

MediaOpenClaw
#1571

Paper Recommendation

Automates discovery, parallel review, scoring, and briefing generation of AI research papers from arXiv, supporting daily updates and PDF analysis.

MediaOpenClaw
#1572

Docker Ctl

Inspect containers, logs, and images via podman

MediaOpenClaw
#1573

Image Handler

Read, analyze metadata, convert formats, resize, rotate, crop, compress, and batch process PNG, JPG, GIF, WebP, TIFF, BMP, HEIC, SVG, and ICO images.

MediaOpenClaw
#1574

Domain Authority Auditor

Use when auditing domain authority, trust, or citation credibility; runs a peer-relative 40-item CITE profile with evidence coverage and verified manipulatio.

MediaOpenClaw
#1576

Cinematic Video

AI cinematic video production powered by CellCog. Short films, music videos, brand films, widescreen cinematics. Consistent characters, cinematic lighting, visual storytelling — grand compositions from a single prompt.

MediaOpenClaw
#1581

Clawhub Skill

Full-stack AI marketing toolkit — scout X/Twitter and Reddit for trending topics, discover and deep-analyze competitors, find content gaps, publish SEO- and.

MediaOpenClaw
#1582

Content Generation

Generate high-quality content across multiple formats. Create articles, reports, social media posts, marketing copy, and other content types with professiona.

MediaOpenClaw
#1590

GitLab Agent Profile

Maintain the GitLab agent profile page and static contribution performance chart.

MediaOpenClaw
#1594

Music Generation

Generate AI music with optimized prompts, style control, and production-ready audio output.

MediaOpenClaw
#1604

Xpoz Setup

Set up and authenticate the Xpoz MCP server for social media intelligence. Required by all Xpoz skills. Handles server configuration, OAuth login, and connection verification with minimal user interaction.

MediaOpenClaw
#1607

Content Claim Navigator

A calm way for artists, creators, managers/teams and labels to get guidance or better understand/organize automated content claims across YouTube, META & TikTok, so nothing important gets missed. Verified official sources, estimated deadline timelines, plain words. *Not expert legal advice - a guide

MediaOpenClaw
#1609

Moltspaces

Join audio room spaces to talk and hang out with other agents and users on Moltspaces.

MediaOpenClaw
#1613

Youtube Data

Use when structured YouTube data is needed: pasted video/channel/playlist links, transcripts for analysis, video metadata, channel upload history, search res.

MediaOpenClaw
#1624

Flyworks Avatar Video

Generate videos using Flyworks (a.k.a HiFly) Digital Humans. Create talking photo videos from images, use public avatars with TTS, or clone voices for custom audio.

MediaOpenClaw
#1631

Context Anchor

Recover from context compaction by scanning memory files and surfacing where you left off. Use when waking up fresh, after compaction, or when you feel lost about what you were doing.

MediaOpenClaw
#1633

Soul Guardian

Drift detection + baseline integrity guard for agent workspace files with automatic alerting support

MediaOpenClaw
#1642

Baoyu Post To X

Posts content and articles to X (Twitter). Supports regular posts with images/videos and X Articles (long-form Markdown). In Codex, honor explicit requests f.

MediaOpenClaw
#1645

Agent Bom Scan

Open security scanner for agentic infrastructure — agents, MCP, packages, blast radius, runtime, and trust for package CVEs (OSV, NVD, EPSS, KEV), container images, provenance, filesystems, and SBOMs. Use when: "check package", "scan image", "verify", "is this safe", "scan dependencies", "CVE lookup", "blast radius".

MediaOpenClaw
#1647

XAI Grok Search

Search the web and X (Twitter) using xAI's Grok API with real-time access, citations, and image understanding

MediaOpenClaw
#1649

Moltchan

Image board for AI agents (4chan-style). Same auth as Moltbook; boards, threads, image posts, replies, upvotes.

MediaOpenClaw
#1666

PaddleOCR Text Recognition

Use this skill whenever the user wants text extracted from images, photos, scans, screenshots, or scanned PDFs. Returns exact machine-readable strings with l.

MediaOpenClaw
#1667

Transcribe Audio Files Via OpenRouter Using Audio Capable Models

Transcribe audio files via OpenRouter using audio-capable models (Gemini, GPT-4o-audio, etc).

MediaOpenClaw
#1668

Pitch Deck Visuals

Investor pitch deck structure with slide-by-slide framework, visual design rules, and data presentation. Covers the 12-slide framework, chart types, team sli.

MediaOpenClaw
#1679

Twitter Operations

Automate and manage Twitter/X accounts by posting, scheduling, replying, analyzing, tracking trends, managing followers, and handling media and analytics.

MediaOpenClaw
#1680

Documents

Build a personal document system for instant access to IDs, contracts, certificates, and important files.

MediaOpenClaw
#1681

Travel

Runs a traveler's standing system: dream list, passports and visas, Schengen day counts, bookings, points, budgets, and what broke last time. Use when someone names a destination they want to visit someday, when a passport, visa, ETA, or entry rule has to be checked before dates are fixed, when counting days already spent in a visa-limited region, when a reservation or cancellation deadline needs recording, when a flight is cancelled or delayed, a bag goes missing, a passport is stolen or a claim has to be filed, when deciding which destination to take next against a season and a budget, when travelling with children, a group, elderly parents or a pet, when a stay runs past a month or needs per-diem handling, or when points or elite status are about to expire. Not for building one trip's day-by-day itinerary (`travel-planning`), fare search (`flight`), accommodation search (`booking`), rental cars (`car-rental`), or moving abroad for good (`expat`).

MediaOpenClaw
#1684

Create Content

Thinking partner that transforms ideas into platform-optimized content

MediaOpenClaw
#1688

YouTube Video

AI YouTube content creation powered by CellCog. YouTube videos, Shorts, thumbnails, video scripts, tutorials, vlogs, educational videos, product reviews, video essays. From script to finished video with voiceover and music.

MediaOpenClaw
#1695

Baoyu Xhs Images

Generates infographic image card series with 12 visual styles, 8 layouts, and 3 color palettes. Breaks content into 1-10 cartoon-style image cards optimized.

MediaOpenClaw
#1697

文档整理技能 (document Organizer)

支持批量将旧版 Office 文档(.doc/.xls)高质量转换为 Markdown 并保持目录结构与格式完整。

MediaOpenClaw
#1700

Gno

Search local documents, files, notes, and knowledge bases. Index directories, search with BM25/vector/hybrid, get AI answers with citations. Use when user wa.

MediaOpenClaw
#1705

Video Transcript

Use when video content needs to be extracted as text: pasted YouTube links or IDs, requests to transcribe, summarize, quote, translate, convert video to text.

MediaOpenClaw
#1707

Code Docs Search By Exa

Search real code snippets and documentation from GitHub, docs sites, and Stack Overflow for accurate syntax and usage examples.

MediaOpenClaw
#1712

Baoyu Slide Deck

Generates professional slide deck images from content. Creates outlines with style instructions, then generates individual slide images. Use when user asks t.

MediaOpenClaw
#1714

VAP Media API Skill For Realistic Image Video Music

VAP Media API skill for image, video, music, and media editing through VAP. Uses VAP product keys with current Media API endpoints.

MediaOpenClaw
#1717

OCR With Python

Extract Chinese and English text from images and scanned PDFs, including documents like invoices and contracts, using PaddleOCR in Python.

MediaOpenClaw
#1719

Speech To Text

Transcribe audio to text with Whisper models via inference.sh CLI. Models: Fast Whisper Large V3, Whisper V3 Large. Capabilities: transcription, translation,.

MediaOpenClaw
#1727

Apple Docs

Query Apple Developer Documentation, APIs, and WWDC videos (2014-2025). Search SwiftUI, UIKit, Objective-C, Swift frameworks and watch sessions.

MediaOpenClaw
#1729

AudioPod

Use AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction, speech-to-text transcription, speaker separation, and media extraction. Use when the user needs to generate music/songs/rap from text, split a song into stems/vocals/instruments, generate speech from text, clean up noisy audio, transcribe audio/video, or extract audio from YouTube/URLs. Requires AUDIOPOD_API_KEY env var or pass api_key directly.

MediaOpenClaw
#1731

Gemini Yt Video Transcript

Create a verbatim transcript for a YouTube URL using Google Gemini (speaker labels, paragraph breaks; no time codes). Use when the user asks to transcribe a YouTube video or wants a clean transcript (no timestamps).

MediaOpenClaw
#1732

Follow Builders

AI builders digest — monitors top AI builders on X and YouTube podcasts, remixes their content into digestible summaries. Use when the user wants AI industry.

MediaOpenClaw
#1733

Theme Factory

Curated collection of professional color and typography themes for styling artifacts — slides, docs, reports, landing pages. Use when applying visual themes to presentations, generating themed content, or creating custom brand palettes. Triggers on theme, color palette, font pairing, slide styling, presentation theme, brand colors.

MediaOpenClaw
#1745

AWS | Amazon Web Services

Architects, debugs, secures, and cost-optimizes AWS infrastructure — EC2, Lambda, RDS, VPC, IAM, ECS, CloudFront. Use when deploying or reviewing anything on AWS, when a bill jumps or spend has to come down, when an AccessDenied, throttle, timeout, 502/503/504, or unreachable-database error has no obvious cause, when choosing between Lambda, Fargate, EC2, RDS, DynamoDB, SQS, or EventBridge, when hardening IAM policies, S3 exposure, security groups, or secrets, when writing Terraform/CloudFormation/CDK against AWS, when auditing an account you inherited, or when a service quota, cold start, connection limit, or failover is the thing that broke. Covers VPC and subnet design, NAT versus VPC endpoints, Organizations and cross-account roles, backups and disaster recovery, and CLI/SSO profiles. Not for object-storage patterns in depth (`s3`), DynamoDB key modeling (`dynamodb`), Kubernetes manifest authoring (`k8s`), or Terraform language mechanics (`terraform`).

MediaOpenClaw

Media Skills FAQ

What is a media agent skill?

It is a reusable instruction package that teaches an AI agent a focused media & files workflow, often including commands, checks, and supporting resources.

Which media skill should I try first?

Start with a narrow task you already understand. The current category leader is Ontology, but requirements and access scope matter more than rank alone.

Does a popular skill mean it is safe?

No. Popularity reflects adoption and interest, not a security guarantee. Read SKILL.md, review commands and dependencies, and test with minimal permissions.