AI Image & Speech
AI technology can do more than process text — it can generate high-quality images and process speech. This article covers how to install and use mainstream AI image generation and speech recognition tools on Ubuntu 26.04.
Stable Diffusion WebUI (AUTOMATIC1111)
AUTOMATIC1111’s Stable Diffusion WebUI is the most popular AI image generation tool, providing a feature-rich web interface that supports text-to-image (txt2img), image-to-image (img2img), inpainting, and more.
System Requirements
- NVIDIA GPU with 8GB+ VRAM recommended
- Python 3.10+
- At least 15GB disk space (base models are 4-7GB each)
Install Dependencies
# Install system dependencies
sudo apt update
sudo apt install -y git python3 python3-venv python3-pip \
libgl1 libglib2.0-0 libsm6 libxext6 libxrender1 \
wget curl
# Ensure NVIDIA drivers and CUDA are installed (see GPU Setup section)
nvidia-smiInstall Stable Diffusion WebUI
# Clone the repository
cd ~
git clone https://github.com/AUTOMATIC1111/stable-diffusion-webui.git
cd stable-diffusion-webui
# First launch (automatically installs Python dependencies and downloads base models)
# This may take 10-30 minutes depending on network speed
./webui.sh# Custom launch parameters
./webui.sh --listen --port 7860 --enable-insecure-extension-access
# Common launch parameters:
# --listen Allow remote access (default is local only)
# --port 7860 Specify port
# --xformers Enable xformers acceleration (recommended)
# --medvram Medium VRAM optimization (recommended for 8GB GPUs)
# --lowvram Low VRAM optimization (for 4-6GB GPUs)
# --no-half Disable half precision (needed for some AMD GPUs)Download Models
Stable Diffusion’s core is the checkpoint model file, which must be manually downloaded and placed in the specified directory.
# Model storage directory
mkdir -p ~/stable-diffusion-webui/models/Stable-diffusion
# Download Stable Diffusion XL base model
wget -O ~/stable-diffusion-webui/models/Stable-diffusion/sd_xl_base_1.0.safetensors \
"https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0/resolve/main/sd_xl_base_1.0.safetensors"
# Download SDXL Refiner (optional, for image refinement)
wget -O ~/stable-diffusion-webui/models/Stable-diffusion/sd_xl_refiner_1.0.safetensors \
"https://huggingface.co/stabilityai/stable-diffusion-xl-refiner-1.0/resolve/main/sd_xl_refiner_1.0.safetensors"# Download VAE model (improves image colors)
mkdir -p ~/stable-diffusion-webui/models/VAE
wget -O ~/stable-diffusion-webui/models/VAE/sdxl_vae.safetensors \
"https://huggingface.co/stabilityai/sdxl-vae/resolve/main/sdxl_vae.safetensors"Download Popular Community Models
The community provides many fine-tuned models on Civitai and Hugging Face:
# Place models in the corresponding directories
# Checkpoint models -> models/Stable-diffusion/
# LoRA models -> models/Lora/
# VAE models -> models/VAE/
# Embeddings -> embeddings/
mkdir -p ~/stable-diffusion-webui/models/Lora
mkdir -p ~/stable-diffusion-webui/embeddings
# Example: download a model from Hugging Face
# wget -O ~/stable-diffusion-webui/models/Stable-diffusion/model_name.safetensors "download-link"Using the WebUI
After launching, visit http://localhost:7860 in your browser:
- txt2img — Enter a prompt to generate images
- img2img — Upload a reference image and modify it based on a prompt
- Inpaint — Select an area of an image for local redrawing
- Extensions — Install plugins to extend functionality
Example prompts:
Positive prompt: a beautiful mountain landscape, sunset, golden hour,
photorealistic, 8k, masterpiece, best quality
Negative prompt: blurry, low quality, distorted, watermark, textCreate a Launch Script
cat > ~/stable-diffusion-webui/start.sh << 'EOF'
#!/bin/bash
cd ~/stable-diffusion-webui
./webui.sh --listen --xformers --enable-insecure-extension-access
EOF
chmod +x ~/stable-diffusion-webui/start.shIf your GPU has 8GB VRAM, we recommend launching with --medvram --xformers. For SDXL models, at least 10GB VRAM is recommended for the best experience. SD 1.5 series models run smoothly with just 4GB VRAM.
ComfyUI
ComfyUI is a node-based workflow AI image generation tool that builds image generation pipelines through visual node connections, offering extreme flexibility.
Install ComfyUI
# Install system dependencies
sudo apt install -y git python3 python3-venv python3-pip
# Clone ComfyUI
cd ~
git clone https://github.com/comfyanonymous/ComfyUI.git
cd ComfyUI
# Create a Python virtual environment
python3 -m venv venv
source venv/bin/activate
# Install PyTorch (CUDA version)
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu126
# Install ComfyUI dependencies
pip install -r requirements.txtLaunch ComfyUI
cd ~/ComfyUI
source venv/bin/activate
# Launch (default port 8188)
python main.py
# Allow remote access
python main.py --listen 0.0.0.0 --port 8188
# Low VRAM mode
python main.py --lowvramInstall Models
ComfyUI’s model directory structure:
# Model directories
ls ~/ComfyUI/models/
# checkpoints/ -- Checkpoint models
# clip/ -- CLIP models
# loras/ -- LoRA models
# vae/ -- VAE models
# controlnet/ -- ControlNet models
# upscale_models/ -- Upscaling models
# You can reuse AUTOMATIC1111 models via symlinks
ln -s ~/stable-diffusion-webui/models/Stable-diffusion/* ~/ComfyUI/models/checkpoints/
ln -s ~/stable-diffusion-webui/models/Lora/* ~/ComfyUI/models/loras/
ln -s ~/stable-diffusion-webui/models/VAE/* ~/ComfyUI/models/vae/Node Workflows
ComfyUI uses nodes and connections to build image generation pipelines:
- Open your browser to
http://localhost:8188 - The default workflow includes basic txt2img nodes
- Right-click the canvas to add new nodes
- Click “Queue Prompt” to start generating
Install ComfyUI Manager (recommended):
# ComfyUI Manager makes it easy to manage custom nodes
cd ~/ComfyUI/custom_nodes
git clone https://github.com/ltdrdata/ComfyUI-Manager.git
# After restarting ComfyUI, a Manager button will appear in the interfacePopular Custom Nodes
cd ~/ComfyUI/custom_nodes
# Advanced ControlNet nodes
git clone https://github.com/Fannovel16/comfyui_controlnet_aux.git
# IP-Adapter nodes (style transfer)
git clone https://github.com/cubiq/ComfyUI_IPAdapter_plus.git
# Image upscaling nodes
git clone https://github.com/ssitu/ComfyUI_UltimateSDUpscale.git
# Install dependencies for each node
cd ~/ComfyUI
source venv/bin/activate
for dir in custom_nodes/*/; do
if [ -f "$dir/requirements.txt" ]; then
pip install -r "$dir/requirements.txt"
fi
doneComfyUI has a steeper learning curve than AUTOMATIC1111 WebUI, but offers greater flexibility and reproducibility. Start with the default workflow and gradually add complex nodes. You can download community-shared workflows from https://comfyworkflows.com or https://openart.ai/workflows .
OpenAI Whisper
OpenAI Whisper is a powerful speech recognition model supporting speech-to-text in multiple languages with extremely high accuracy.
Install Whisper
# Install system dependencies
sudo apt install -y python3 python3-pip python3-venv ffmpeg
# Create a virtual environment (recommended)
python3 -m venv ~/whisper-env
source ~/whisper-env/bin/activate
# Install Whisper
pip install openai-whisper
# Or install with GPU acceleration (requires CUDA)
pip install openai-whisper torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu126Speech to Text
# Activate virtual environment
source ~/whisper-env/bin/activate
# Basic usage: transcribe an audio file
whisper audio.mp3 --language en --model medium
# Specify output format
whisper audio.mp3 --language en --model medium --output_format txt
whisper audio.mp3 --language en --model medium --output_format srt # Subtitle format
whisper audio.mp3 --language en --model medium --output_format vtt # WebVTT format
whisper audio.mp3 --language en --model medium --output_format json
# Auto-detect language
whisper audio.mp3 --model medium
# Specify output directory
whisper audio.mp3 --language en --model medium --output_dir ~/transcripts/Whisper Model Sizes
| Model | Parameters | VRAM Required | Relative Speed | Accuracy |
|---|---|---|---|---|
tiny | 39M | ~1 GB | Fastest | Fair |
base | 74M | ~1 GB | Very fast | Good |
small | 244M | ~2 GB | Fast | Better |
medium | 769M | ~5 GB | Medium | Very good |
large-v3 | 1.5B | ~10 GB | Slower | Best |
turbo | 809M | ~6 GB | Fast | Very good |
# Use different model sizes
whisper audio.mp3 --model tiny # Fast but lower accuracy
whisper audio.mp3 --model small # Balance of speed and accuracy
whisper audio.mp3 --model large-v3 # Highest accuracy
whisper audio.mp3 --model turbo # Balance of speed and large model accuracyPython Usage
import whisper
# Load model
model = whisper.load_model("medium")
# Transcribe audio
result = model.transcribe("audio.mp3", language="en")
# Output text
print(result["text"])
# Output segments with timestamps
for segment in result["segments"]:
print(f"[{segment['start']:.1f}s - {segment['end']:.1f}s] {segment['text']}")Using faster-whisper (Recommended)
faster-whisper uses the CTranslate2 engine and is 4x faster than the original Whisper with lower VRAM usage.
# Install faster-whisper
pip install faster-whisperfrom faster_whisper import WhisperModel
# Load model (auto-downloads)
model = WhisperModel("medium", device="cuda", compute_type="float16")
# Transcribe
segments, info = model.transcribe("audio.mp3", language="en")
print(f"Detected language: {info.language}, confidence: {info.language_probability:.2f}")
for segment in segments:
print(f"[{segment.start:.1f}s -> {segment.end:.1f}s] {segment.text}")# faster-whisper command line usage
pip install faster-whisper-cli
faster-whisper audio.mp3 --language en --model medium --output_format srtBatch Processing Audio
# Create a batch transcription script
cat > ~/transcribe.sh << 'SCRIPT'
#!/bin/bash
# Batch transcribe all audio files in a directory
INPUT_DIR="${1:-.}"
OUTPUT_DIR="${2:-./transcripts}"
MODEL="${3:-medium}"
source ~/whisper-env/bin/activate
mkdir -p "$OUTPUT_DIR"
for file in "$INPUT_DIR"/*.{mp3,wav,m4a,flac,ogg}; do
[ -f "$file" ] || continue
echo "Transcribing: $file"
whisper "$file" --language en --model "$MODEL" \
--output_format srt --output_dir "$OUTPUT_DIR"
done
echo "Transcription complete! Results saved to $OUTPUT_DIR"
SCRIPT
chmod +x ~/transcribe.sh
# Use the script
# ~/transcribe.sh ~/audio_files ~/output mediumFor general transcription, the medium model offers the best balance of accuracy and speed. For maximum accuracy, use large-v3. For real-time transcription or processing large volumes of files, use faster-whisper — it provides approximately 4x speed improvement at the same accuracy level.
Tool Comparison
| Tool | Type | GPU Required | Best For |
|---|---|---|---|
| AUTOMATIC1111 WebUI | Image generation | 4GB+ | Full-featured, beginner-friendly |
| ComfyUI | Image generation | 4GB+ | Advanced workflows, reproducibility |
| Whisper | Speech-to-text | Optional | Audio transcription, subtitles |
| faster-whisper | Speech-to-text | Recommended | High-volume audio processing |