Skip to Content
DocsAIAI Image & Speech

AI Image & Speech

AI technology can do more than process text — it can generate high-quality images and process speech. This article covers how to install and use mainstream AI image generation and speech recognition tools on Ubuntu 26.04.

Stable Diffusion WebUI (AUTOMATIC1111)

AUTOMATIC1111’s Stable Diffusion WebUI is the most popular AI image generation tool, providing a feature-rich web interface that supports text-to-image (txt2img), image-to-image (img2img), inpainting, and more.

System Requirements

  • NVIDIA GPU with 8GB+ VRAM recommended
  • Python 3.10+
  • At least 15GB disk space (base models are 4-7GB each)

Install Dependencies

# Install system dependencies sudo apt update sudo apt install -y git python3 python3-venv python3-pip \ libgl1 libglib2.0-0 libsm6 libxext6 libxrender1 \ wget curl # Ensure NVIDIA drivers and CUDA are installed (see GPU Setup section) nvidia-smi

Install Stable Diffusion WebUI

# Clone the repository cd ~ git clone https://github.com/AUTOMATIC1111/stable-diffusion-webui.git cd stable-diffusion-webui # First launch (automatically installs Python dependencies and downloads base models) # This may take 10-30 minutes depending on network speed ./webui.sh
# Custom launch parameters ./webui.sh --listen --port 7860 --enable-insecure-extension-access # Common launch parameters: # --listen Allow remote access (default is local only) # --port 7860 Specify port # --xformers Enable xformers acceleration (recommended) # --medvram Medium VRAM optimization (recommended for 8GB GPUs) # --lowvram Low VRAM optimization (for 4-6GB GPUs) # --no-half Disable half precision (needed for some AMD GPUs)

Download Models

Stable Diffusion’s core is the checkpoint model file, which must be manually downloaded and placed in the specified directory.

# Model storage directory mkdir -p ~/stable-diffusion-webui/models/Stable-diffusion # Download Stable Diffusion XL base model wget -O ~/stable-diffusion-webui/models/Stable-diffusion/sd_xl_base_1.0.safetensors \ "https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0/resolve/main/sd_xl_base_1.0.safetensors" # Download SDXL Refiner (optional, for image refinement) wget -O ~/stable-diffusion-webui/models/Stable-diffusion/sd_xl_refiner_1.0.safetensors \ "https://huggingface.co/stabilityai/stable-diffusion-xl-refiner-1.0/resolve/main/sd_xl_refiner_1.0.safetensors"
# Download VAE model (improves image colors) mkdir -p ~/stable-diffusion-webui/models/VAE wget -O ~/stable-diffusion-webui/models/VAE/sdxl_vae.safetensors \ "https://huggingface.co/stabilityai/sdxl-vae/resolve/main/sdxl_vae.safetensors"

The community provides many fine-tuned models on Civitai and Hugging Face:

# Place models in the corresponding directories # Checkpoint models -> models/Stable-diffusion/ # LoRA models -> models/Lora/ # VAE models -> models/VAE/ # Embeddings -> embeddings/ mkdir -p ~/stable-diffusion-webui/models/Lora mkdir -p ~/stable-diffusion-webui/embeddings # Example: download a model from Hugging Face # wget -O ~/stable-diffusion-webui/models/Stable-diffusion/model_name.safetensors "download-link"

Using the WebUI

After launching, visit http://localhost:7860 in your browser:

  1. txt2img — Enter a prompt to generate images
  2. img2img — Upload a reference image and modify it based on a prompt
  3. Inpaint — Select an area of an image for local redrawing
  4. Extensions — Install plugins to extend functionality

Example prompts:

Positive prompt: a beautiful mountain landscape, sunset, golden hour, photorealistic, 8k, masterpiece, best quality Negative prompt: blurry, low quality, distorted, watermark, text

Create a Launch Script

cat > ~/stable-diffusion-webui/start.sh << 'EOF' #!/bin/bash cd ~/stable-diffusion-webui ./webui.sh --listen --xformers --enable-insecure-extension-access EOF chmod +x ~/stable-diffusion-webui/start.sh
Tip

If your GPU has 8GB VRAM, we recommend launching with --medvram --xformers. For SDXL models, at least 10GB VRAM is recommended for the best experience. SD 1.5 series models run smoothly with just 4GB VRAM.


ComfyUI

ComfyUI is a node-based workflow AI image generation tool that builds image generation pipelines through visual node connections, offering extreme flexibility.

Install ComfyUI

# Install system dependencies sudo apt install -y git python3 python3-venv python3-pip # Clone ComfyUI cd ~ git clone https://github.com/comfyanonymous/ComfyUI.git cd ComfyUI # Create a Python virtual environment python3 -m venv venv source venv/bin/activate # Install PyTorch (CUDA version) pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu126 # Install ComfyUI dependencies pip install -r requirements.txt

Launch ComfyUI

cd ~/ComfyUI source venv/bin/activate # Launch (default port 8188) python main.py # Allow remote access python main.py --listen 0.0.0.0 --port 8188 # Low VRAM mode python main.py --lowvram

Install Models

ComfyUI’s model directory structure:

# Model directories ls ~/ComfyUI/models/ # checkpoints/ -- Checkpoint models # clip/ -- CLIP models # loras/ -- LoRA models # vae/ -- VAE models # controlnet/ -- ControlNet models # upscale_models/ -- Upscaling models # You can reuse AUTOMATIC1111 models via symlinks ln -s ~/stable-diffusion-webui/models/Stable-diffusion/* ~/ComfyUI/models/checkpoints/ ln -s ~/stable-diffusion-webui/models/Lora/* ~/ComfyUI/models/loras/ ln -s ~/stable-diffusion-webui/models/VAE/* ~/ComfyUI/models/vae/

Node Workflows

ComfyUI uses nodes and connections to build image generation pipelines:

  1. Open your browser to http://localhost:8188
  2. The default workflow includes basic txt2img nodes
  3. Right-click the canvas to add new nodes
  4. Click “Queue Prompt” to start generating

Install ComfyUI Manager (recommended):

# ComfyUI Manager makes it easy to manage custom nodes cd ~/ComfyUI/custom_nodes git clone https://github.com/ltdrdata/ComfyUI-Manager.git # After restarting ComfyUI, a Manager button will appear in the interface
cd ~/ComfyUI/custom_nodes # Advanced ControlNet nodes git clone https://github.com/Fannovel16/comfyui_controlnet_aux.git # IP-Adapter nodes (style transfer) git clone https://github.com/cubiq/ComfyUI_IPAdapter_plus.git # Image upscaling nodes git clone https://github.com/ssitu/ComfyUI_UltimateSDUpscale.git # Install dependencies for each node cd ~/ComfyUI source venv/bin/activate for dir in custom_nodes/*/; do if [ -f "$dir/requirements.txt" ]; then pip install -r "$dir/requirements.txt" fi done
Note

ComfyUI has a steeper learning curve than AUTOMATIC1111 WebUI, but offers greater flexibility and reproducibility. Start with the default workflow and gradually add complex nodes. You can download community-shared workflows from https://comfyworkflows.com  or https://openart.ai/workflows .


OpenAI Whisper

OpenAI Whisper is a powerful speech recognition model supporting speech-to-text in multiple languages with extremely high accuracy.

Install Whisper

# Install system dependencies sudo apt install -y python3 python3-pip python3-venv ffmpeg # Create a virtual environment (recommended) python3 -m venv ~/whisper-env source ~/whisper-env/bin/activate # Install Whisper pip install openai-whisper # Or install with GPU acceleration (requires CUDA) pip install openai-whisper torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu126

Speech to Text

# Activate virtual environment source ~/whisper-env/bin/activate # Basic usage: transcribe an audio file whisper audio.mp3 --language en --model medium # Specify output format whisper audio.mp3 --language en --model medium --output_format txt whisper audio.mp3 --language en --model medium --output_format srt # Subtitle format whisper audio.mp3 --language en --model medium --output_format vtt # WebVTT format whisper audio.mp3 --language en --model medium --output_format json # Auto-detect language whisper audio.mp3 --model medium # Specify output directory whisper audio.mp3 --language en --model medium --output_dir ~/transcripts/

Whisper Model Sizes

ModelParametersVRAM RequiredRelative SpeedAccuracy
tiny39M~1 GBFastestFair
base74M~1 GBVery fastGood
small244M~2 GBFastBetter
medium769M~5 GBMediumVery good
large-v31.5B~10 GBSlowerBest
turbo809M~6 GBFastVery good
# Use different model sizes whisper audio.mp3 --model tiny # Fast but lower accuracy whisper audio.mp3 --model small # Balance of speed and accuracy whisper audio.mp3 --model large-v3 # Highest accuracy whisper audio.mp3 --model turbo # Balance of speed and large model accuracy

Python Usage

import whisper # Load model model = whisper.load_model("medium") # Transcribe audio result = model.transcribe("audio.mp3", language="en") # Output text print(result["text"]) # Output segments with timestamps for segment in result["segments"]: print(f"[{segment['start']:.1f}s - {segment['end']:.1f}s] {segment['text']}")

faster-whisper uses the CTranslate2 engine and is 4x faster than the original Whisper with lower VRAM usage.

# Install faster-whisper pip install faster-whisper
from faster_whisper import WhisperModel # Load model (auto-downloads) model = WhisperModel("medium", device="cuda", compute_type="float16") # Transcribe segments, info = model.transcribe("audio.mp3", language="en") print(f"Detected language: {info.language}, confidence: {info.language_probability:.2f}") for segment in segments: print(f"[{segment.start:.1f}s -> {segment.end:.1f}s] {segment.text}")
# faster-whisper command line usage pip install faster-whisper-cli faster-whisper audio.mp3 --language en --model medium --output_format srt

Batch Processing Audio

# Create a batch transcription script cat > ~/transcribe.sh << 'SCRIPT' #!/bin/bash # Batch transcribe all audio files in a directory INPUT_DIR="${1:-.}" OUTPUT_DIR="${2:-./transcripts}" MODEL="${3:-medium}" source ~/whisper-env/bin/activate mkdir -p "$OUTPUT_DIR" for file in "$INPUT_DIR"/*.{mp3,wav,m4a,flac,ogg}; do [ -f "$file" ] || continue echo "Transcribing: $file" whisper "$file" --language en --model "$MODEL" \ --output_format srt --output_dir "$OUTPUT_DIR" done echo "Transcription complete! Results saved to $OUTPUT_DIR" SCRIPT chmod +x ~/transcribe.sh # Use the script # ~/transcribe.sh ~/audio_files ~/output medium
Tip

For general transcription, the medium model offers the best balance of accuracy and speed. For maximum accuracy, use large-v3. For real-time transcription or processing large volumes of files, use faster-whisper — it provides approximately 4x speed improvement at the same accuracy level.


Tool Comparison

ToolTypeGPU RequiredBest For
AUTOMATIC1111 WebUIImage generation4GB+Full-featured, beginner-friendly
ComfyUIImage generation4GB+Advanced workflows, reproducibility
WhisperSpeech-to-textOptionalAudio transcription, subtitles
faster-whisperSpeech-to-textRecommendedHigh-volume audio processing
Last updated on