Podcastfy Openclaw Skill
by @mr-11even
Convert text, images, PDFs, websites, or YouTube videos into multilingual AI-generated podcast audio using Podcastfy's open-source Python toolkit.
clawhub install podcastfy-openclaw-skillπ About This Skill
Podcastfy Skill for OpenClaw
> Transform content into engaging AI-generated podcasts using Podcastfy.
Overview
Podcastfy is an open-source Python package that transforms multi-modal content (text, images, websites, PDFs, YouTube videos) into engaging, multi-lingual audio conversations using GenAI.
It's like NotebookLM's podcast feature but open-source and programmatic.
Prerequisites
1. Install Podcastfy
pip install podcastfy
2. Install FFmpeg (required for audio processing)
# macOS
brew install ffmpegUbuntu/Debian
sudo apt install ffmpeg
3. Configure API Keys
Create a .env file with your API keys:
# Required: At least one LLM + one TTS
GEMINI_API_KEY=your_gemini_api_key # Default: transcript generation
OPENAI_API_KEY=your_openai_api_key # Default: TTSOptional (for other TTS models)
ELEVENLABS_API_KEY=your_elevenlabs_api_key
Supported TTS Models: | Model | Quality | API Key Required | |:---|:---|:---| | OpenAI (default) | Good | Yes | | ElevenLabs | Great (customizable) | Yes | | Google (gemini/geminimulti) | Best (English only) | Yes | | Edge | Basic | No |
Recommended Setup:
Usage
Basic Examples
#### From URLs (websites, articles)
python -m podcastfy.client --url https://example.com/article1 --url https://example.com/article2
#### From YouTube Video
python -m podcastfy.client --url https://youtube.com/watch?v=xxx
#### From PDF
python -m podcastfy.client --file path/to/document.pdf
#### From Images
python -m podcastfy.client --image path/to/image1.jpg --image path/to/image2.png
#### From Text
python -m podcastfy.client --text "Your raw text content here"
#### Longform (30+ minutes)
python -m podcastfy.client --url https://example.com/article --longform
#### Only Generate Transcript (no audio)
python -m podcastfy.client --url https://example.com/article --transcript-only
Python API Usage
from podcastfy.client import generate_podcastFrom URLs
audio_file = generate_podcast(
urls=["https://example.com/article1", "https://example.com/article2"]
)From local files
audio_file = generate_podcast(
file="path/to/urls.txt"
)From images
audio_file = generate_podcast(
image=["path/to/image1.jpg", "path/to/image2.png"]
)Longform
audio_file = generate_podcast(
urls=["https://example.com/article"],
longform=True
)
Advanced Options
# Custom conversation config
python -m podcastfy.client --url https://example.com/article --conversation-config config.yamlUse specific TTS model
python -m podcastfy.client --url https://example.com/article --tts-model elevenlabsUse local LLM for privacy
python -m podcastfy.client --url https://example.com/article --transcript-only --localOutput to specific directory
python -m podcastfy.client --url https://example.com/article --output-dir ./my-podcasts
Integration with OpenClaw
Method 1: Direct CLI Execution
Use OpenClaw's exec tool:
python -m podcastfy.client --url https://example.com/article
Method 2: Python Script Generation
Generate a Python script and run it:
# Script will be created by the skill
from podcastfy.client import generate_podcastaudio = generate_podcast(
urls=[""],
output_dir="./podcasts"
)
Use Cases for Dental/Medical Content
Transform Articles into Podcasts
# Convert dental research articles to audio
python -m podcastfy.client --url https://www.dental.com/research/article
Convert Course Materials
# Turn course PDFs into listenable content
python -m podcastfy.client --file ./dental-course/lecture1.pdf
Create Patient Education
# Transform patient info sheets to audio
python -m podcastfy.client --text "Dental implant procedure explained..."
Summarize YouTube Dental Videos
# Turn dental education videos into podcasts
python -m podcastfy.client --url https://youtube.com/watch?v=dental-education-video
Output
--transcript-only or alongside audio)./podcastfy_output/Troubleshooting
"ffmpeg not found"
# Install ffmpeg
brew install ffmpeg # macOS
sudo apt install ffmpeg # Linux
"API key not found"
Make sure your .env file is in the correct directory and API keys are set.
"Multi-speaker voices not available"
Google's multi-speaker TTS requires allowlisting. Use OpenAI or ElevenLabs instead.
Resources
Note: Requires Python 3.11+ and FFmpeg.
π‘ Examples
Basic Examples
#### From URLs (websites, articles)
python -m podcastfy.client --url https://example.com/article1 --url https://example.com/article2
#### From YouTube Video
python -m podcastfy.client --url https://youtube.com/watch?v=xxx
#### From PDF
python -m podcastfy.client --file path/to/document.pdf
#### From Images
python -m podcastfy.client --image path/to/image1.jpg --image path/to/image2.png
#### From Text
python -m podcastfy.client --text "Your raw text content here"
#### Longform (30+ minutes)
python -m podcastfy.client --url https://example.com/article --longform
#### Only Generate Transcript (no audio)
python -m podcastfy.client --url https://example.com/article --transcript-only
Python API Usage
from podcastfy.client import generate_podcastFrom URLs
audio_file = generate_podcast(
urls=["https://example.com/article1", "https://example.com/article2"]
)From local files
audio_file = generate_podcast(
file="path/to/urls.txt"
)From images
audio_file = generate_podcast(
image=["path/to/image1.jpg", "path/to/image2.png"]
)Longform
audio_file = generate_podcast(
urls=["https://example.com/article"],
longform=True
)
Advanced Options
# Custom conversation config
python -m podcastfy.client --url https://example.com/article --conversation-config config.yamlUse specific TTS model
python -m podcastfy.client --url https://example.com/article --tts-model elevenlabsUse local LLM for privacy
python -m podcastfy.client --url https://example.com/article --transcript-only --localOutput to specific directory
python -m podcastfy.client --url https://example.com/article --output-dir ./my-podcasts
βοΈ Configuration
1. Install Podcastfy
pip install podcastfy
2. Install FFmpeg (required for audio processing)
# macOS
brew install ffmpegUbuntu/Debian
sudo apt install ffmpeg
3. Configure API Keys
Create a .env file with your API keys:
# Required: At least one LLM + one TTS
GEMINI_API_KEY=your_gemini_api_key # Default: transcript generation
OPENAI_API_KEY=your_openai_api_key # Default: TTSOptional (for other TTS models)
ELEVENLABS_API_KEY=your_elevenlabs_api_key
Supported TTS Models: | Model | Quality | API Key Required | |:---|:---|:---| | OpenAI (default) | Good | Yes | | ElevenLabs | Great (customizable) | Yes | | Google (gemini/geminimulti) | Best (English only) | Yes | | Edge | Basic | No |
Recommended Setup:
π Tips & Best Practices
"ffmpeg not found"
# Install ffmpeg
brew install ffmpeg # macOS
sudo apt install ffmpeg # Linux
"API key not found"
Make sure your .env file is in the correct directory and API keys are set.
"Multi-speaker voices not available"
Google's multi-speaker TTS requires allowlisting. Use OpenAI or ElevenLabs instead.