What is Stable Diffusion Trained On?

Stable Diffusion is trained on LAION (Large-scale Artificial Intelligence Open Network) datasets, primarily LAION-5B containing 5.85 billion image-text pairs scraped from the internet. The Stable Diffusion 1.x checkpoints were trained first on the laion2B-en subset, then on laion-high-resolution, and the later checkpoints were fine-tuned on aesthetically filtered subsets — laion-improved-aesthetics and laion-aesthetics v2 5+. These datasets include images with associated alt-text, captions, and metadata from web sources like Common Crawl. Stable Diffusion does not train its own text encoder: it reuses a frozen, pre-trained CLIP text encoder (OpenAI CLIP ViT-L/14 in 1.x, OpenCLIP ViT-H/14 in 2.x), which was itself trained contrastively on image-text pairs beforehand. What is trained for Stable Diffusion is the autoencoder that maps images to and from latent space, and the diffusion model that learns to denoise in that space while conditioned on the frozen text embeddings. Additional fine-tuning uses curated datasets and safety filtering to reduce harmful content generation. For AI agents, understanding Stable Diffusion's training data helps predict model capabilities, limitations, and potential biases in generated outputs.

Work with us

Ready to put agentic AI to work?

Book a free 45-minute consultation. We'll map one real process worth automating with production-grade AI.