# ComfyUI Dev > Because Even Penguins Need Documentation This file contains all documentation content in a single document following the llmstxt.org standard. ## cyberrealisticPony_v100.safetensors # cyberrealisticPony_v100.safetensors Welcome to the documentation for the cyberpunk-meets-photo-realism powerhouse: `cyberrealisticPony_v100.safetensors`. This model checkpoint, used in ComfyUI workflows, is built to generate stunningly realistic humans (and humanoid ponies, if that’s your thing) with a distinctly cyberpunk flair. Think neon reflections, chrome prosthetics, and photorealistic skin textures—but make it fashion. --- ## 🔍 Overview `cyberrealisticPony_v100` is a **Stable Diffusion 1.5-compatible** checkpoint that combines photorealistic rendering capabilities with stylized cyberpunk aesthetics. It’s particularly suited for generating: - Realistic portraits with sci-fi accessories - Futuristic fashion and character concepts - Glossy skin textures with strong lighting contrast - “Cyberized” humans (yes, even My Little Pony characters if prompted—hence the name) Despite the pony name, this is not a furry or MLP-only model—it excels in realistic human rendering with optional surreal/augmented features. --- ## 🎯 Recommended Use Cases | Use Case | Description | | ------------------------------- | ---------------------------------------------------------------- | | 🔧 **Concept Art** | Ideal for cyberpunk and sci-fi character design | | 👩‍🎤 **Fashion Design** | Stunning for futuristic makeup, chrome clothing, LED accessories | | 🧠 **AI Roleplay / NPC Design** | Great for realistic avatars with post-human aesthetics | | 📷 **Photoreal Portraits** | Generates uncanny realism with lighting finesse | | 🦄 **Surreal Fusion Art** | Can push into stylized or fantasy realism with prompting | --- ## 🧪 Compatible VAEs To really bring out the vibrancy and fine details, these VAEs are recommended: - `vae-ft-mse-840000-ema-pruned.safetensors` – Balanced, clean, and sharp detail enhancer - `vae-kl-f8-anime2.safetensors` – Works well if you're leaning stylized - Avoid using the default VAE if you're seeking optimal results—it tends to wash out contrast and skin tone realism --- ## ⚙️ Recommended Settings and Parameters | Parameter | Recommended Value | | ------------------------------ | ------------------------------------------------------------------------------------ | | Sampler | DPM++ 2M Karras | | Steps | 30–40 | | CFG (Classifier-Free Guidance) | 6.5–8.0 | | Resolution | 768x768 or 1024x1024 for best detail fidelity | | Seed | Set manually for consistent compositions or leave random for more chaotic brilliance | | Scheduler | Karras or LMS for smoother convergence | --- ## ✍️ Prompting Tips `cyberrealisticPony_v100` is prompt-sensitive in the best way. Here’s how to make it sing: ### Positive Prompt Examples: - `cyberpunk woman, glowing blue eyes, LED wires, glossy skin, ultra detailed, photorealistic` - `futuristic fashion model, chrome face mask, slick black hair, cinematic lighting` - `realistic male portrait, neural implants, city lights bokeh, high detail, cybernetics` ### Negative Prompt Suggestions: To avoid unintended artifacts or creepiness: - `blurry, lowres, extra limbs, deformed, jpeg artifacts, poorly drawn face, missing fingers, bad anatomy, watermark` --- ## 🛠️ Workflow Integration in ComfyUI Here’s a typical node structure for working with this checkpoint in ComfyUI: 1. **CheckpointLoaderSimple** - `ckpt_name`: `cyberrealisticPony_v100.safetensors` 2. **VAE Loader** - Recommended: `vae-ft-mse-840000-ema-pruned.safetensors` 3. **CLIP Text Encode (Positive/Negative)** - Craft well-separated prompts to maximize latent clarity 4. **Sampler (DPM++ 2M Karras or Euler a)** - Ideal for maintaining realism while handling cyberpunk chaos 5. **KSampler Advanced** or **CFGGuidance** - For more refined control and experimentation with CFG values --- ## ⚠️ Known Issues - **Over-sharpening at high CFG**: Can lead to waxy skin or unnatural textures. Dial CFG back to ~6.5 if needed. - **Lighting blowout**: Prompted lighting sources (e.g., “bright LED lights”) can overpower the subject and result in loss of facial features if not balanced. - **Too cyber, not enough pony?**: Yes, the name is confusing. This isn’t My Little Pony unless you ask nicely (in your prompt). - **Occasional asymmetry**: Like many SD-based models, mirror symmetry isn’t always guaranteed. --- ## 🧠 Pro Tips - Add `film still, bokeh, depth of field` for cinematic compositions. - Use ControlNet (depth or canny) to lock in strong composition and structure. - Apply LoRA models (e.g., cyber fashion or texture LoRAs) to push this checkpoint even further. - Stack prompts with styles like `hdr, ultrasharp, professional photo` to level-up realism. --- 📚 Additional Resources - [Download CyberRealistic Pony Now!](https://civitai.green/models/443821?modelVersionId=1838857) --- ## 🏁 Final Thoughts `cyberrealisticPony_v100.safetensors` is a chameleon of realism and futurism, capable of producing hauntingly beautiful portraits that blur the line between man, machine, and neon mythos. Whether you're building a cyberpunk RPG character roster or designing your next synthwave album cover, this checkpoint gives you the horsepower—and just a dash of chaos—to make your visions come to life. Go forth and render. Responsibly. --- ## dreamshaper_8.safetensors # dreamshaper_8.safetensors `dreamshaper_8.safetensors` is a versatile and fine-tuned checkpoint for text-to-image generation that brings a creative flair to your ComfyUI workflows. It’s essentially the artistic overachiever in the checkpoint class—capable of producing high-quality, aesthetically pleasing images that lean toward a dreamy, painterly, and often hyper-realistic look. It’s frequently used in workflows where stylistic consistency and visual elegance are required. --- ## 📦 Checkpoint Summary | Property | Value | | ---------------------- | ------------------------------------------- | | File Name | `dreamshaper_8.safetensors` | | File Format | `.safetensors` (yay, safe loading!) | | Version | v8 (final-ish unless v9 drops) | | Compatible Loader Node | `CheckpointLoaderSimple` | | Base Architecture | Likely SD 1.5 or similar (Latent Diffusion) | | Use Case | General-purpose + artistic T2I | | Style Bias | Dreamy, semi-realistic, painterly | --- ## 🎯 Use Cases `dreamshaper_8.safetensors` is one of those checkpoints that tries to be a jack-of-all-trades but—surprisingly—manages to also be a master of many. Ideal for: - 🔥 Consistent character renders in storytelling sequences - 🖼️ Album art and book cover aesthetics - 🌌 Ethereal, moody concept art - 👩 Hyper-stylized portraits - 👾 Slightly surreal visualizations Basically, if your vibe is somewhere between “real person” and “fine art hallucination,” this is your model. --- ## 🛠 How to Load in ComfyUI To use this model in ComfyUI: 1. **Drag in a `CheckpointLoaderSimple` node** if you haven’t already. 2. Click the dropdown on the `ckpt_name` input and select `dreamshaper_8.safetensors` (assuming it lives in your models/checkpoints directory). 3. Connect the output ports as follows: - `MODEL` → To your KSampler (or other sampler node) - `CLIP` → To your `CLIPTextEncode` node (text encoder) - `VAE` → Optional, but can improve decoding. Consider using a high-quality VAE like `vae-ft-mse-840000-ema-pruned`. --- ## 🔍 Model Behavior & Traits - **Text Understanding**: Pretty good. You won’t have to prompt like a dungeon master deciphering riddles. - **Prompt Sensitivity**: Responsive to well-crafted prompts and modifiers like “masterpiece,” “depth of field,” and “cinematic lighting.” - **Stylistic Strength**: Leans heavily into the ethereal and artistic—your outputs might occasionally look like they belong in an AI art gallery curated by mid-century romantics. - **Face Quality**: Great out of the gate. Even better with a touch of `face_detailer` or `Restore Face` wizardry. - **Outfit Details & Textures**: Surprisingly sharp. Leather looks like leather. Lace looks like someone spent too much time on it. --- ## ⚠️ Tips and Gotchas - **VAE Matters**: The default VAE might wash out some details. Pair it with `vae-ft-mse-840000-ema-pruned` for sharper results. - **CFG Scale**: 7–9 is the sweet spot. Go higher if you enjoy chaos and lower if you like minimalism (or disappointment). - **Sampling Method**: DPM++ 2M Karras or Euler a tends to work best. Avoid `Heun` unless you enjoy pixel salad. - **Clip Skip**: Defaults to 1 in most workflows, but feel free to bump it to 2 for a style boost (especially in anime-style or hyper-stylized scenes). --- ## 📁 File Location Ensure `dreamshaper_8.safetensors` is saved in your: bash ``` `ComfyUI/models/checkpoints/ ``` folder. If you’re not sure where that is, it’s probably next to the folder you swore you’d organize six months ago. --- ## ✅ Recommended Pairings | Component | Recommended Option | Why? | | ------------------- | ------------------------------------------ | ---------------------------------------------- | | **VAE** | `vae-ft-mse-840000-ema-pruned.safetensors` | Enhances color depth & detail fidelity | | **Sampler** | `KSampler` with `DPM++ 2M Karras` | Balances detail and coherence well | | **Prompt Style** | Descriptive + Artistic Modifiers | "ethereal lighting, 8k uhd, sharp focus" works | | **Negative Prompt** | “blurry, deformed, low-res, extra limbs” | Keeps it classy, not cursed | --- ## 🤖 Sample Prompt ```plaintext a serene portrait of a woman in a flowing silk dress, standing in a moonlit forest, ultra detailed, soft lighting, cinematic, volumetric shadows, masterpiece, 8k ``` **Negative Prompt**: ```plaintext blurry, low quality, deformed, extra arms, bad anatomy, watermark, text ``` --- 📚 Additional Resources - [Download DreamShaper 8 Now!](https://huggingface.co/autismanon/modeldump/blob/main/dreamshaper_8.safetensors) --- ## 🧪 Final Thoughts `dreamshaper_8.safetensors` is the checkpoint equivalent of a top-tier Photoshop filter baked into your diffusion model. It’s intuitive, visually impressive, and—when prompted well—makes you look more talented than you probably are. If you’re going for consistently excellent aesthetic results in ComfyUI without spending 3 hours adjusting samplers, `dreamshaper_8` is your ride-or-die. --- ## flux-2-klein-9b.safetensors # flux-2-klein-9b.safetensors ## FLUX.2 Klein 9B Look, I've seen a lot of model files waddle through ComfyUI in my time, but `flux-2-klein-9b.safetensors` is something special. This isn't your grandmother's image generation model—unless your grandmother has a thing for sub-second inference times and 9 billion parameters, in which case, respect. ## 🔧 File Format **Type:** SafeTensors (`.safetensors`) **Size:** Approximately 17-20GB (because excellence takes space, folks) **Architecture:** Rectified flow transformer with integrated Qwen3 text embedder SafeTensors format means you're not loading some sketchy pickle file that could potentially order pizza to your address at 3 AM. It's a safe, efficient tensor storage format that loads faster than you can say "why is my VRAM full again?" ## 📁 Function in ComfyUI Workflows FLUX.2 Klein 9B is your Swiss Army knife for image generation. This bad boy handles: - **Text-to-Image Generation:** Type words, get pixels. Revolutionary, I know. - **Image-to-Image Multi-Reference Editing:** Feed it reference images and watch it work its magic - **Unified Architecture:** One model to rule them all—generation AND editing in the same package In ComfyUI, this model slots into your checkpoint loader nodes and becomes the beating heart of your workflow. It's the difference between "I made an image" and "I made an image that doesn't look like it was rendered on a potato." ## 🧠 Technical Details Alright, time to get nerdy (as if we haven't been already): - **Parameters:** 9 billion (9B flow model + 8B Qwen3 text embedder) - **Inference Steps:** 4 (step-distilled for speed demons) - **VRAM Requirements:** ~24GB (hope you have a GPU that didn't come from a garage sale) - **Architecture Type:** Rectified flow transformer - **Training Methodology:** Step-distilled from base model - **Speed:** Sub-second generation (yes, really) - **Resolution Support:** Variable, optimized for standard aspect ratios This model sits at the Pareto frontier for quality vs. latency, which is a fancy way of saying it punches way above its weight class. Models 5x its size are side-eyeing this thing with concern. ## ✅ Benefits **Speed That'll Make You Question Reality** Sub-second generation times. I've seen cold starts take longer than this model takes to generate a masterpiece. **Quality Without Compromise** Matches or exceeds models with 45B+ parameters. It's like bringing a tactical nuke to a pillow fight. **Unified Model Architecture** One model for text-to-image AND multi-reference editing. No more juggling seventeen different checkpoints like some kind of ML circus performer. **Excellent Prompt Adherence** Actually listens to what you tell it. Unlike my last three neural networks, which had the listening comprehension of a goldfish. **Output Diversity** Great for creative exploration. Generate variations without everything looking like the same image with a different hat. **Real-Time Application Integration** Built for production use, not just showing off on Reddit. ## ⚙️ Usage Tips **Sampler Settings:** Since this model is step-distilled to 4 steps, don't go crazy with 50+ steps. You're not making it better, you're just making your GPU cry. Stick to 4-8 steps for optimal results. **CFG Scale:** Start around 3.5-7.0. This model doesn't need aggressive guidance to understand what you want. It's not a rebellious teenager. **Prompt Engineering:** Be specific but don't write a novel. This model has an 8B parameter text embedder—it understands nuance better than your autocorrect understands what you're trying to type. **Batch Processing:** With sub-second inference, you can actually do batch generation without aging noticeably. Live your best life. **Reference Images for Editing:** When using image-to-image workflows, quality reference images = quality outputs. Garbage in, garbage out—a tale as old as computing itself. ## 📍 ComfyUI Setup Instructions 1. **Download the Model** Acquire `flux-2-klein-9b.safetensors` from your preferred source (see Additional Resources below). 2. **Installation Location** Place the model file in your ComfyUI models directory: `ComfyUI/models/checkpoints/` Don't put it in `/Downloads/random_stuff/maybe_models/idk/` like some kind of digital hoarder. 3. **Verify VRAM** Check that you have at least 24GB VRAM available. If you don't, this is a great opportunity to explain to your significant other why you need a new GPU. 4. **Load in ComfyUI** - Add a "Load Checkpoint" node to your workflow - Select `flux-2-klein-9b.safetensors` from the dropdown - Connect to your KSampler or other sampling nodes - Set steps to 4-8 (seriously, don't overdo it) 5. **Configure Sampling** - Use appropriate samplers (Euler, DPM++ recommended) - Keep CFG scale reasonable (3.5-7.0 range) - Set resolution to your target output size 6. **Test Generation** Run a simple prompt first. Something straightforward to make sure everything works properly before you start generating your magnum opus. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire **Don't Use 50+ Sampling Steps** This model was step-distilled to 4 steps. Using 50 steps is like taking a Ferrari through a school zone—pointless and you're just wasting resources. **Don't Crank CFG to 15** High CFG values will make your outputs look like they went through a deep-frying process. Twice. **Don't Load This on 8GB VRAM** Physics says no. Your GPU will say no. Your computer's fans will achieve liftoff velocity trying to say no. **Don't Mix with Incompatible VAEs** Use the recommended VAE or auto-select. Mixing random VAEs is like putting diesel in a sports car—technically possible, hilariously inadvisable. **Don't Skip the Text Encoder** This model uses an 8B Qwen3 text embedder. It's not optional. It's PART of the architecture. Skipping it is like trying to make a sandwich without bread. **Don't Use on Ancient Hardware** If your GPU still thinks HDMI is "newfangled technology," this model is not for you. ## 📚 Additional Resources **Model Download:** [FLUX.2 Klein 9B download](https://huggingface.co/black-forest-labs/FLUX.2-klein-9B) **License Information:** [FLUX Non-Commercial License](https://bfl.ai/legal/usage-policy) (read it before your lawyer has to) **Community Resources:** - ComfyUI Official Documentation - FLUX Model Family Documentation - ComfyUI Community Forums ## 📎 Example Node Configuration ``` [Load Checkpoint] ├── checkpoint_name: flux-2-klein-9b.safetensors └── output: MODEL, CLIP, VAE [CLIP Text Encode (Prompt)] ├── text: "a majestic mountain landscape at sunset, detailed, professional photography" ├── clip: [from Load Checkpoint] └── output: CONDITIONING [KSampler] ├── model: [from Load Checkpoint] ├── seed: 42 (or random, live dangerously) ├── steps: 4 ├── cfg: 5.0 ├── sampler_name: euler ├── scheduler: normal ├── positive: [from CLIP Text Encode] ├── negative: [from Negative Prompt] ├── latent_image: [from Empty Latent Image] └── output: LATENT [VAE Decode] ├── samples: [from KSampler] ├── vae: [from Load Checkpoint] └── output: IMAGE ``` **Alternative: Multi-Reference Editing Configuration** ``` [Load Checkpoint] ├── flux-2-klein-9b.safetensors [Load Image] (Reference Image 1) [Load Image] (Reference Image 2) [VAE Encode] → Connect reference images [KSampler] ├── steps: 4-6 ├── cfg: 4.5 ├── denoise: 0.7-0.85 (for editing) └── [Connect everything appropriately] ``` ## 📝 Notes **Performance Optimization:** If you're experiencing slower-than-advertised inference times, check your: - CUDA/ROCm versions (update if you're still running drivers from 2019) - System RAM (yes, it matters) - Background processes (close those 47 Chrome tabs) - Cooling (thermal throttling is real) **Model Variants:** This documentation covers the 9B model specifically. There's also a 4B variant with Apache 2.0 licensing if you need commercial use, but that's a story for another documentation file. **Fine-Tuning Potential:** While this is the step-distilled version, you _can_ technically fine-tune it. However, if you want maximum fine-tuning potential, the Base 9B variant (undistilled) is your friend. **Compatibility:** Works with ComfyUI's standard node ecosystem. Most custom nodes should play nicely, but if something breaks, check for FLUX-specific compatibility notes. **Future Updates:** The FLUX model family is actively developed. Keep an eye on official channels for updates, improvements, and new variants that might make this documentation obsolete (the circle of tech life). --- _Remember: With great parameters comes great VRAM requirements. Use responsibly, generate beautifully, and may your inference times be ever in your favor._ _P.S. If this model doesn't work for you, check your setup before assuming the model is broken. 99% of the time it's user error. The other 1% is cosmic rays flipping bits in your RAM._ --- ## RevAnimated V122 Safetensors ## 🧠 What is `revAnimated_v122.safetensors`? `revAnimated_v122` is part of the revered **RevAnimated** family of model merges, expertly crafted to blend the detail-oriented strength of realistic base models with the stylized flair and dynamic visual expression of semi-anime models. In its 1.2.2 iteration, it stands as one of the most versatile and well-balanced models in the Stable Diffusion ecosystem, optimized for **high-quality character rendering**, **animated scenes**, and **stylized realism** that _doesn’t scream "straight from a cartoon."_ Whether you're aiming to depict soft cinematic portraits, fashion-forward OC gens, fantasy storyboards, or that "anime-but-not-too-anime" look, this checkpoint gives you the right amount of stylized spice without setting your realism house on fire. --- ## 🧰 Ideal Use Cases This checkpoint excels at: - 🎨 **Character portraits**: Original characters, fan art, stylized self-portraits. - 📖 **Illustrated narratives**: Visual novel content, light manga storyboards. - 👗 **Fashion/fabric rendering**: Dramatic clothing, reflective textures, soft materials. - 💫 **Fantasy and magical realism**: Ethereal lighting, glowing effects, surreal compositions. - 🖼️ **Stylized realism**: If you want the anatomical accuracy of realism with the color dynamics of anime. --- ## ⚙️ Compatible VAE(s) To avoid "plastic skin" syndrome or color banding, you should always pair `revAnimated_v122` with a capable VAE. The following are tested and recommended: - ✅ `vae-ft-mse-840000-ema-pruned.safetensors` - ✅ `vae-ft-mse-560000-ema-pruned.safetensors` - ✅ `kl-f8-anime2.vae.pt` (for punchier color styles, more anime-leaning output) - ✅ `orangemix.vae.pt` (if you like slightly richer skin tones and gradient fidelity) **Avoid:** Default/no VAE. The output will be underwhelming or muddy. --- ## 🔧 Recommended Settings & Parameters | Parameter | Recommended Value | Notes | | ---------- | --------------------------------------- | -------------------------------------------------------------------- | | Sampler | `DPM++ 2M Karras` or `DPM++ SDE Karras` | Smooth transitions and beautiful textures | | Steps | 20–30 | Anything below 20 sacrifices detail; above 30 is diminishing returns | | CFG Scale | 6.5–8.5 | Sweet spot for stylization with prompt adherence | | Resolution | 768x768 or 832x1216 | High-resolution recommended for portraits or full-body scenes | | Clip Skip | 2 | Delivers cleaner, sharper output with less prompt confusion | --- ## ✍️ Prompting Tips **Style control:** - Use descriptive, emotion-driven terms to bring out expressive faces and dynamic poses. - For realism: `"cinematic lighting", "8k, ultra-detailed skin, volumetric light"` - For anime lean: `"anime style", "cel shading", "vivid colors", "sharp lines"` **Character detailing:** - Add depth: `"detailed eyes, intricate outfit, atmospheric shadows"` - Scene balance: `"centered composition", "soft background blur", "natural pose"` **Negative prompts (important!):** scss CopyEdit `(deformed hands, missing fingers, extra limbs, poorly drawn face:1.3), blurry, grainy, low-res, watermark, text, signature` --- ## ⚠️ Known Issues & Quirks - 🖐️ **Hands**: While better than many stylized models, hands can still get weird with multiple limbs or odd shapes if left unattended. - 👗 **Clothing Artifacts**: Occasionally tries to invent new physics for fabrics. Use LoRAs or ControlNets for more structure. - 🧍‍♂️ **Poses**: May default to common standing/sitting poses unless you explicitly direct movement or emotion. - 📐 **Style Drift**: May lean too far anime or too far realistic depending on your CFG scale and prompt balance. Prompt wisely. --- ## 🔄 Workflow Compatibility (ComfyUI) `revAnimated_v122.safetensors` plays very nicely with ComfyUI. It is compatible with: - ✅ **Standard workflows**: Text-to-image, inpainting, image-to-image. - ✅ **ControlNet nodes**: Great for pose guidance (depth, canny, openpose). - ✅ **LoRA stacking**: Supports blending multiple LoRAs without destabilization. - ✅ **High-res fix workflows**: Excellent quality upscaling from 512x512 to 1024x1024+. --- ## 🧪 Notable LoRA & Style Synergies - 🐉 Fantasy Armor LoRAs - 🧝 Elf/Fae-based style LoRAs - 👘 Kimono and traditional clothing LoRAs - 👩 Realistic makeup / facial enhancement LoRAs ## 📚 Additional Resources - [Download revAnimated_v122.safetensors](https://huggingface.co/Renqf/sd/tree/main) ## 🏁 Final Thoughts `revAnimated_v122.safetensors` is like that stylish friend who cleans up well at a wedding but still wears anime socks under the tux. It’s versatile, expressive, and balanced enough to be your daily driver while still letting you dive into more experimental art directions. **Use it when:** You want stylized storytelling without cartoonish exaggeration. **Avoid it when:** You’re doing ultra-photorealistic work or purely anime-style generation. --- ## v1-5-pruned-emaonly-fp16.safetensors # v1-5-pruned-emaonly-fp16.safetensors ## 🔍 Overview The `v1-5-pruned-emaonly-fp16.safetensors` checkpoint is a **Stable Diffusion v1.5-based** model, optimized for inference with **half-precision (FP16)** weights. It's a **pruned, Exponential Moving Average (EMA)-only** version, stripped of training overhead and tuned for speed and efficiency in rendering high-quality text-to-image generations. If you're looking for a rock-solid, vanilla foundation for anything from photorealism to stylized illustration—and you enjoy buttery-smooth memory performance—this checkpoint should be your default starting block. --- ## 📁 Technical Details | Property | Value | | ------------ | -------------------------------------------------------- | | Base Model | Stable Diffusion v1.5 | | Architecture | UNet + CLIP text encoder | | Precision | FP16 (half-precision) | | Pruning | Yes (non-essential weights removed) | | EMA Only | Yes (uses Exponential Moving Average weights) | | File Format | `.safetensors` (secure and efficient) | | Size | ~2.1GB | | CLIP Version | OpenCLIP ViT-H/14 (compatible with SD 1.5 text encoders) | --- ## 🧠 What It Can Do This checkpoint serves as a **general-purpose generation model** and performs well across a range of prompt styles including: - 🧍 Realistic portraits and character renders - 🌄 Landscapes and environmental scenery - 🎨 Artistic interpretations and semi-abstracts - 🧙 Fantasy and concept art - 🧸 Stylized illustration (when paired with appropriate LoRAs/VAEs) --- ## 🧰 Recommended Use Cases This model is particularly useful for: - **Baseline workflows**: It’s the foundation model often used to compare LoRAs, VAEs, and prompt tweaking. - **Resource-constrained environments**: Because it’s FP16 and pruned, it plays nicely with low VRAM (yes, 4–6GB cards can actually have fun here). - **Compositional consistency testing**: With proper CFG control and noise seeds, this model handles consistent framing well. - **Fine-tuning or LoRA training**: It's often used as a base checkpoint for finetuning projects due to its stability. --- ## 🧩 Compatible VAEs This checkpoint doesn’t include a baked-in VAE. You’ll want to pair it with an external `.safetensors` VAE file. Here are the top contenders: | VAE Name | Recommended Use | | ------------------------------------------ | --------------------------------------- | | `vae-ft-mse-840000-ema-pruned.safetensors` | Clean, accurate, default VAE for SD 1.5 | | `kl-f8-anime2.ckpt` | For more stylized/illustrative outputs | | `clearVAE` or `colorfulVAE` | Boost contrast and vibrancy | If you're getting muddy or washed-out images, check your VAE. Misalignments here can ruin even the best prompts. --- ## ⚙️ ComfyUI Settings & Parameters | Parameter | Suggested Value | Notes | | ---------- | ------------------------------------- | -------------------------------------------- | | Sampler | `Euler a`, `DPM++ 2M Karras`, `UniPC` | Euler a = faster, DPM++ = better quality | | Steps | 20–35 | Sweet spot is often 28 | | CFG Scale | 6.5–8.5 | Over 9 can create artifacts or melting faces | | Resolution | 512x512 native, upscale from there | Upscaling recommended post-gen | | Seed | Fixed for reproducibility | Variants = use dynamic seed | --- ## 💬 Prompting Tips - Stick to **natural language prompts** with **clear modifiers**. Example: arduino CopyEdit `"A futuristic cityscape at sunset, cyberpunk aesthetic, neon lights, volumetric fog, ultra detailed, unreal engine, 8k"` - Add specific modifiers to control mood and fidelity: - `ultra-detailed`, `intricate lighting`, `masterpiece`, `high quality` - Negative prompts are essential to avoid those _"why does she have six fingers and a third eye?"_ moments: scss CopyEdit `(worst quality, low quality:1.3), bad anatomy, deformed, blurry, extra limbs, mutated hands, bad proportions` --- ## 🚨 Known Issues & Limitations | Issue | Details | | -------------------------- | -------------------------------------------------------------------------------------- | | Lack of baked-in VAE | You _must_ attach a compatible VAE node in ComfyUI or you’ll get low-fidelity outputs. | | Limited stylization | It’s not anime or toon-optimized out of the box—pair with a LoRA for stylization. | | Occasionally bland outputs | When prompting is too vague, it plays safe. So don’t be boring. | | Small resolution default | Works best at 512x512; higher resolutions may need tiled workflows or upscaling nodes. | --- ## 🧪 Ideal Pairings - **LoRAs**: Combine with LoRAs that add style or fine-grained control (e.g., `analog-style`, `cyberpunk-city`, `gothicFashion_v1.0`). - **Upscalers**: `4x-UltraSharp`, `R-ESRGAN 4x+`, or ComfyUI’s `UltimateSDUpscale` for crisp final images. - **Post-processing**: Use `Prompt Styler`, `Hires Fix`, or `Detail Tweaker` nodes to push visual fidelity even further. ## 📚 Additional Resources - [Download Stable Diffusion v1.5 pruned](https://huggingface.co/Comfy-Org/stable-diffusion-v1-5-archive/tree/main) ## 🧠 Final Thoughts The `v1-5-pruned-emaonly-fp16.safetensors` checkpoint is your Swiss Army knife in the ComfyUI arsenal. It might not breathe fire out of the box like some hyper-stylized blends, but it’s _damn reliable_, exceptionally versatile, and light on resources. Use it to test, to fine-tune, to baseline, or as your dependable workhorse for clean, quality generations. --- Want model magic? Start here—just don't forget your VAE, or you’ll be asking why everything looks like it was shot through a potato. --- ## control_v11f1p_sd15_depth_fp16.safetensors # control_v11f1p_sd15_depth_fp16.safetensors Welcome to the deep end—literally. `control_v11f1p_sd15_depth_fp16.safetensors` is the Depth ControlNet model from the SD1.5 series (fp16 version), designed to guide generation in ComfyUI using depth maps as the controlling signal. Think of it as giving your image generation process a sixth sense—for spatial awareness. --- ## 🧠 Overview This ControlNet model uses **depth maps** (monochrome representations of distance in 3D space) to inject spatial geometry into Stable Diffusion image generation. It's most effective when you want to preserve or simulate realistic object positioning, perspective, and spatial continuity. It leverages a depth-predicting encoder (usually Midas or similar) to interpret input images and influence the denoising steps in your pipeline. --- ## 🛠️ Filename - `control_v11f1p_sd15_depth_fp16.safetensors` --- ## 🌍 Ideal Use Cases - **Photorealistic composition retention** (e.g. adding objects to a photo with perspective awareness) - **Image-to-image refinement** using 3D structure hints - **3D scene simulation** in stills (e.g., architectural renders, interior design mockups) - **Stylization of existing scenes** while keeping spatial realism intact - **Consistent multi-angle views** of a character or object - **Depth-aware fantasy and sci-fi environments** that feel more grounded in physical space --- ## 🧩 Compatible Components | Component | Recommended Option | | ---------------- | --------------------------------------------------------------------------- | | **Checkpoint** | Any SD1.5-based model | | **VAE** | `vae-ft-mse-840000-ema-pruned` or `vae-ft-ema-560000` for cleaner structure | | **Preprocessor** | Midas Depth (`depth_midas`), ZoE-Depth, or similar depth map providers | | **Scheduler** | DDIM, DPM++ 2M Karras, or Euler A (with minor variation tolerance) | --- ## 🔧 Required Workflow Nodes Here’s what your ComfyUI graph should include for a basic depth-controlled pipeline: ### **1. Preprocessor Node** - Use `depth_midas` or `zoe_depth_preprocessor`. - Input: Source image (to derive depth map from). - Output: A depth map to feed into the ControlNet. ### **2. ControlNetLoader Node** - Load the ControlNet model: CopyEdit `control_v11f1p_sd15_depth_fp16.safetensors` - Hook it up to the ControlNet node. ### **3. ControlNetApply Node** - Connect depth map output here. - Ensure the conditioning strength is adjusted according to desired control (see tips below). ### **4. Sampler/Generation Node** - The ControlNet conditioning will influence the denoising process. - Your main prompt still dictates style and detail, but ControlNet ensures spatial consistency. --- ## 🎛️ Recommended Settings | Setting | Value | | --------------------- | ----------------------------------------------------------------------------- | | **Control Weight** | `0.7–1.0` for strong structure adherence | | **Guidance Scale** | `6–9` (typical for SD1.5 workflows) | | **Steps** | `20–35` depending on model and complexity | | **Start/End Control** | `0.0` to `1.0` for full influence | | **Input Image Size** | Keep resolution under 768x768 for memory sanity unless you like GPU meltdowns | --- ## 🧠 Prompting Tips - Prompts should **complement** the geometry, not fight it. If the depth map shows a sofa in the foreground, don’t prompt for a beach sunset with "no furniture"—unless you're into uncanny valley furniture-shaped sand dunes. - Use phrases like: - “realistic lighting” - “in natural perspective” - “high detail, depth-focused” - Avoid overly abstract prompts if you're relying on precise geometry. --- ## ⚠️ Known Issues & Gotchas - **Depth map quality matters.** Garbage in, garbage out. Always inspect your preprocessor’s depth output before generating. - **Misalignment** can occur if the prompt implies objects not consistent with depth. - **Artifacts**: If using models trained on anime or stylized data, depth control may get overridden or conflict with exaggerated anatomy. - **Memory hog**: The fp16 version helps, but ControlNet + large VAEs + high-res = bring a fan for your GPU. --- ## ✅ Pro Tips - Want soft spatial guidance? Drop `control weight` to ~0.5. - Want photobashed realism with depth consistency? Try mixing this with **SoftEdge** or **Canny** ControlNets in parallel using the `Multi-ControlNet` node. - Stack it with `LoRAs` for pose or clothing control to finesse your outputs into structured masterpieces. --- 📚 Additional Resources - [Download ControlNet-v1-1_fp16 Now!](https://www.kaggle.com/datasets/qq2575044704/controlnet-models) --- ## 💬 Summary `control_v11f1p_sd15_depth_fp16.safetensors` is your go-to ControlNet for depth-aware, spatially coherent image generation in ComfyUI. Whether you're preserving photographic realism or nudging your creations into new 3D-aware directions, this ControlNet acts like a polite architect whispering, “Hey… maybe don’t put that chandelier under the couch.” Pair it wisely, guide it gently, and it will reward you with structure you didn’t even know you needed. --- ## control_v1p_sd15_qrcode_monster_v2.safetensors # control_v1p_sd15_qrcode_monster_v2.safetensors _Waddle on over, because Naplin here is about to blow your mind with the most delightfully chaotic marriage of art and function you've ever seen._ Welcome to the documentation for **control_v1p_sd15_qrcode_monster_v2.safetensors**, the ControlNet that asks the question: "Why should QR codes be boring rectangles when they could be BEAUTIFUL boring rectangles?" This isn't your grandma's barcode—this is what happens when edge detection meets data encoding meets "I dare you to make this work." --- ## 🔍 Overview | Property | Value | | ----------------- | ---------------------------------------------- | | **Name** | control_v1p_sd15_qrcode_monster_v2.safetensors | | **Type** | ControlNet (QR Code Structural Guidance) | | **Base Model** | Stable Diffusion 1.5 | | **Input Type** | QR Code Image (16px module size) | | **Precision** | FP16 (because we're classy AND efficient) | | **Version** | v2 (the "actually works this time" edition) | | **Author/Source** | Monster Labs (yes, that's their actual name) | This ControlNet does something genuinely wild: it takes a QR code and weaves it into generated artwork while _maintaining scannability_. That's right—your phone can read these artistic abominations. It's like if M.C. Escher designed data matrices and they actually worked. The v2 upgrade is a massive improvement over v1, offering better scannability AND more creative freedom. The secret sauce? A gray background (#808080) that lets the QR code blend seamlessly into the generated image. It's practically magic, except it's math, which is arguably cooler. ## 🎯 Ideal Use Cases **🎨 Artistic QR Codes for Marketing**: Make your business cards actually interesting. "Scan me" becomes "Please scan me, I'm a masterpiece." **🎪 Event Invitations**: Wedding QR codes that look like watercolor paintings? Now we're talking. **🖼️ Gallery Installations**: Interactive art pieces where the QR code IS the art. Meta? Yes. Cool? Absolutely. **📱 Social Media Engagement**: Stop posting boring links. Start posting functional art that people WANT to share. **🏪 Product Packaging**: When your coffee bag QR code looks like an abstract expressionist fever dream but still takes you to the right URL. **🎮 ARG / Puzzle Games**: Hide scannable codes in generated imagery for scavenger hunts and mystery games. ## 🛠️ Workflow Setup in ComfyUI Listen up, because this workflow is where the magic happens. And by magic, I mean "a carefully orchestrated dance of parameters that will make you question your sanity." ### 1. **Generate Your Base QR Code** Before you even open ComfyUI, you need an actual QR code. Use any QR code generator, but pay attention to these critical settings: **QR Code Generator Settings:** - **Module Size**: 16px (this is NON-NEGOTIABLE—the model was trained on this) - **Error Correction Level**: HIGH or MAXIMUM (Level H = 30% error tolerance) - **Background Color**: #808080 (medium gray—this is the v2 secret weapon) - **Foreground Color**: Black or close to it - **Output Size**: 768×768 or 1024×1024 recommended Why high error correction? Because you're about to obliterate half the data blocks with artistic flourishes, that's why. The error correction will save your bacon when your QR code decides to cosplay as a sunset. ### 2. **Load the ControlNet Model** ``` [ControlNetLoader] └─ model_name: control_v1p_sd15_qrcode_monster_v2.safetensors ``` No preprocessing needed—your QR code IS the condition image. Just plug and play, baby. ### 3. **Wire Up Your ComfyUI Nodes** Here's the standard node chain: ``` [Load Image (your QR code)] ↓ [ControlNetApply] ├─ conditioning: [CLIP Text Encode (positive)] ├─ control_net: [ControlNetLoader output] ├─ image: [Your QR code image] └─ strength: 0.8-1.5 (we'll get to this) ↓ [KSampler] ↓ [VAE Decode] ↓ [Save Image] ``` ### 4. **The Balancing Act: Prompting** This is where you earn your wizard hat. Your prompt determines the STYLE, while the QR code determines the STRUCTURE. They're going to fight. Your job is to referee. **Prompt Guidelines:** - Be descriptive and atmospheric: _"ethereal forest with glowing mushrooms, magical lighting, fantasy art"_ - Avoid geometric patterns that conflict with QR structure: _"checkerboard"_ is asking for trouble - Organic, flowing subjects work better: nature, clouds, water, abstract art - Hard-edged architectural subjects? That's expert mode—possible but finicky **Example Prompts That Play Nice:** - _"underwater coral reef, bioluminescent creatures, deep ocean, magical realism"_ - _"stained glass window, church interior, colorful light rays, gothic architecture"_ - _"swirling galaxy, nebula clouds, cosmic dust, space photography"_ - _"abstract watercolor splash, vibrant colors, fluid art"_ ## ⚙️ Parameters & Settings Alright, buckle up. This is where we separate the QR code dabblers from the QR code MASTERS. | Parameter | Recommended Range | Notes | | -------------------- | ------------------------ | ---------------------------------------------------------------- | | **control_strength** | 0.8–1.5 | The holy grail setting. THIS is your scannable-vs-creative dial. | | **start_percent** | 0.0 | Start guidance from the beginning | | **end_percent** | 1.0 | Maintain guidance throughout generation | | **Sampler** | DPM++ 2M Karras, Euler a | Clean results, fewer artifacts | | **Steps** | 25–40 | More steps = more refinement, but diminishing returns after 40 | | **CFG Scale** | 5–8 | Lower CFG = more creativity; higher = more prompt adherence | | **Resolution** | 768×768, 1024×1024 | Match or exceed your QR code size | | **Seed** | Variable | Lock seed to iterate on successful codes | ### 🎚️ The Control Strength Spectrum This deserves its own section because it's THAT important. - **0.6–0.8**: "I want art that vaguely remembers being a QR code" (Low scannability, high creativity) - **0.8–1.2**: "The sweet spot" (Balanced scannability and artistic merit) - **1.2–1.5**: "This WILL scan or I'll die trying" (High scannability, moderate creativity) - **1.5+**: "I just want a slightly pretty QR code" (Maximum scannability, minimal artistic flair) Start at 1.0 and adjust based on your scan tests. Yes, you need to actually TEST if your codes scan. This isn't theoretical physics. ## 🧩 Compatible Components ### ✅ Compatible VAEs - **vae-ft-mse-840000-ema-pruned.safetensors**: The SD 1.5 standard—solid, reliable, won't let you down - **kl-f8-anime2.vae.pt**: If you want stylized, saturated colors that pop - **Anything VAE**: Works with most Anything-series checkpoints - **ClearVAE**: Crisp detail retention—great for maintaining QR code edge sharpness ### ✅ Compatible Checkpoints Any SD 1.5 checkpoint works, but some are more QR-friendly than others: **High Scan Success Rate:** - **v1-5-pruned-emaonly.safetensors**: The baseline. If it doesn't work here, blame your prompt. - **realisticVision_v51.safetensors**: Photorealistic QR codes that look like album covers - **dreamshaper_8.safetensors**: Fantastic for organic, flowing compositions **Creative But Trickier:** - **anythingV5_PrtRE.safetensors**: Anime-style QR codes—use higher control strength - **revAnimated_v122.safetensors**: Gorgeous outputs but can over-stylize; bump strength to 1.2+ - **deliberate_v2.safetensors**: Painterly results; may require multiple generation attempts ## 🧠 Prompting Tips (AKA "How Not to Lose Your Mind") ### DO: - **Use organic, flowing subject matter**: Water, smoke, clouds, fabric, fire—these play nice with QR patterns - **Embrace abstract concepts**: "Cosmic energy," "liquid metal," "crystalline structures" - **Specify lighting and atmosphere**: These guide style without disrupting structure - **Think in terms of texture**: "Rough," "smooth," "glossy," "weathered" ### DON'T: - **Prompt for rigid geometric patterns**: "Grid," "checkerboard," "pixelated"—these conflict with the QR structure - **Over-specify spatial layouts**: Let the QR code handle composition - **Use prompts with heavy text elements**: Letters and QR codes are mortal enemies - **Expect first-try miracles**: This is an iterative process. Generate batches. ### 🎨 Style Prompt Hacks Adding these at the end of your prompt can dramatically affect scannability: - **For better scanning**: `highly detailed, sharp focus, clean edges` - **For more creativity**: `abstract art, impressionist style, loose brushstrokes` - **For color blending**: `monochromatic, limited color palette, gradient background` ## 🧪 Pro Techniques & Workflows ### The "Generate & Refine" Method This is the gold standard for getting scannable, beautiful QR codes: **Step 1: Initial Generation** - Set control_strength to 1.0–1.2 - Generate 4–8 variations with different seeds - Test which ones scan (yes, with your actual phone) **Step 2: Image-to-Image Refinement** - Take the scannable ones that need aesthetic improvement - Load into img2img with the SAME QR code as ControlNet condition - Settings: - Denoising strength: 0.3–0.5 - Control strength: 1.3–1.5 (higher than initial) - Keep the same or similar prompt **Step 3: The "Save a Dying Code" Technique** - Got a gorgeous code that ALMOST scans? - Max out control_strength to 1.5+ - Set denoising to 0.2 (barely touching it) - Gradually increase denoising by 0.05 increments until it scans - Congrats, you're a QR code necromancer now ### Batch Consistency for Multiple Designs Want a series of QR codes with different URLs but similar aesthetics? 1. Lock your prompt and parameters 2. Generate different QR codes (different URLs) with the SAME visual treatment 3. Use the same seed for each if you want near-identical styles 4. Test all codes—successful batch requires 80%+ scan rate ### The Color Palette Trick Since v2 loves that #808080 gray background, you can guide your color scheme: - **Warm palette**: Add `golden hour, amber tones, sunset colors` to prompt - **Cool palette**: Add `twilight, cyan and purple, cool tones, deep blue` - **Monochrome**: Add `black and white, grayscale, high contrast` (ironically easier to scan) ## 🧨 Known Issues & Limitations Because nothing in AI generation is perfect, and pretending otherwise is just lying. ### Issue #1: "My Code Won't Scan" **Causes:** - Control strength too low ( 1.5) - Boring prompt or lack of style direction - Not enough steps (< 20) **Solutions:** - Lower control_strength to 0.9–1.1 - Enhance prompt with rich descriptive language - Increase steps to 30–40 - Use img2img refinement with lower denoising ### Issue #3: "The QR Blocks Are Too Visible" This is actually kind of the point, but if you want more integration: **Solutions:** - Use prompts with textural variety: `organic patterns, flowing design, natural forms` - Lower control_strength slightly (0.8–0.9) - Add negative prompt: `sharp edges, pixelated, blocky, mosaic` - The gray background (#808080) in your base QR helps—make sure you used it ### Issue #4: "Colors Are Muddy/Dull" **Causes:** - VAE issues - CFG too low - Prompt lacks color direction **Solutions:** - Switch to a more vibrant VAE (kl-f8-anime2) - Increase CFG to 7–8 - Add color descriptors to prompt: `vibrant colors, saturated, vivid, rich hues` ### Issue #5: "Scannability Is Inconsistent Across Batches" Yeah, that's just... the nature of the beast. This model is creative FIRST, functional SECOND. You're fighting entropy here. **Mitigation Strategies:** - Generate in batches of 8–10, expect 50–70% scan success rate - Lock successful parameters and seeds - Use the "Generate & Refine" workflow above - Keep a library of successful parameter combinations ## 📏 Technical Requirements ### Minimum System Specs - **VRAM**: 6GB minimum (8GB recommended, 10GB+ for high-res outputs) - **ComfyUI Version**: Any recent version (2024+) - **Python Dependencies**: Standard ComfyUI install—no special requirements ### File Size & Storage - **Model Size**: ~2.5GB (it's chunky but worth it) - **Installation Path**: `ComfyUI/models/controlnet/` ### QR Code Generator Requirements Use any QR generator that supports: - Custom colors (you NEED that #808080 background) - High error correction (Level H = 30%) - High-resolution output (768px+ recommended) - Custom module sizing (16px is mandatory) **Recommended Generators:** - QR Code Generator (qr-code-generator.com) - supports all needed features - QRazyBox - advanced control for power users - Python qrcode library - for automation nerds ## 🎯 Real-World Project Examples ### Example 1: Coffee Shop Menu QR Code **Goal**: Scannable menu QR that looks like latte art **Settings:** - Control strength: 1.1 - Prompt: _"latte art, coffee foam, heart design, warm brown tones, cafe aesthetic, creamy texture"_ - Steps: 30 - CFG: 7 - Checkpoint: realisticVision_v51 **Result**: 70% scan rate, stunning aesthetic ### Example 2: Music Festival Poster QR **Goal**: Psychedelic ticket purchase QR code **Settings:** - Control strength: 0.9 - Prompt: _"psychedelic art, vibrant swirls, rainbow colors, 1960s poster style, groovy, trippy patterns"_ - Steps: 35 - CFG: 6.5 - Checkpoint: dreamshaper_8 **Result**: 60% scan rate (lower but acceptable for artistic priority) ### Example 3: Wedding Invitation QR **Goal**: Elegant RSVP QR code **Settings:** - Control strength: 1.3 - Prompt: _"watercolor flowers, soft pastels, romantic, delicate brushstrokes, floral arrangement, wedding invitation style"_ - Steps: 40 - CFG: 7.5 - Checkpoint: deliberate_v2 **Result**: 85% scan rate (refined via img2img) ## 🧰 TL;DR – Quick Reference | Feature | Description | | ------------------------------- | ---------------------------------------------------------------- | | **ControlNet Type** | QR Code structural guidance with artistic integration | | **Base Model** | SD 1.5 | | **Input** | QR code image (16px module size, #808080 background) | | **Primary Use Cases** | Artistic QR codes for marketing, events, products, installations | | **Critical Requirement** | HIGH error correction on base QR code (Level H = 30%) | | **Sweet Spot control_strength** | 0.9–1.2 (balance of scan + beauty) | | **Best Samplers** | DPM++ 2M Karras, Euler a | | **Recommended Steps** | 25–40 | | **Compatible VAEs** | Any SD 1.5 VAE; vae-ft-mse recommended | | **Expected Success Rate** | 50–80% scannability depending on parameters | | **Key v2 Feature** | Gray background blending for seamless integration | ## 🐧 Naplin's Final Waddle of Wisdom Look, this ControlNet is simultaneously one of the coolest and most frustrating things you'll use in ComfyUI. When it works, you'll feel like a digital alchemist. When it doesn't, you'll question your life choices. But here's the thing: NO OTHER MODEL does this. You're creating FUNCTIONAL ART. Your QR codes can be gallery-worthy while still taking people to your SoundCloud (please don't actually do this). Start with high control strength. Test your codes. Iterate relentlessly. Save your successful parameter combinations. And for the love of all that is holy, USE THE GRAY BACKGROUND. Now go forth and make QR codes that make people say "Wait, THAT scans?!" _-Naplin_ 🐧 ## 📚 Additional Resources - **Monster Labs GitHub**: Check for updates and community examples - **ComfyUI QR Code Workflows**: Search the ComfyUI community for shared workflows - **QR Code Testing Apps**: Try multiple scanner apps—Google Lens, dedicated QR apps, camera apps - **Error Correction Standards**: Research Reed-Solomon error correction if you're a masochist _Remember: Just because it scans doesn't mean it's not art, and just because it's art doesn't mean it won't scan._ --- ## Preprocessor Options # Preprocessor Options Welcome to the jungle, also known as the [**ControlNet Preprocessor node**](http://comfyui.dev/docs/guides/Nodes/controlnet-preprocessor). This is where your image gets poked, prodded, outlined, depth-mapped, or otherwise tortured into a usable conditioning map for ControlNet. The `preprocessor` field is arguably the most important setting in this node—because this is where you tell ComfyUI what kind of preprocessing to apply to the image. Each `preprocessor` option here corresponds to a different algorithm or pipeline used to extract structural, semantic, or stylistic features from an image. These features are then fed into a ControlNet model to guide image generation with a specific type of constraint (e.g., edges, poses, segmentation maps, depth, etc). --- ## 🛠️ Setting: `preprocessor` ### 🔧 Requirements - **Input**: A valid image (some require RGB, some prefer grayscale). - **Dependencies**: Many of these preprocessors rely on external Python libraries like OpenCV, PyTorch, Detectron2, etc. If you get errors, you’re probably missing one. - **Resolution**: Works best with input resolution in the 512–1024px range. Too small and it gets dumb; too big and it might just explode your VRAM. --- ### 📚 Available Preprocessor Options Let’s go _way too deep_ into each one: --- ### **`none`** - **What It Does**: Skips preprocessing. Sends the raw image straight to ControlNet. - **Use Case**: If you’ve already prepared your conditioning image manually or externally. - **Strengths**: No overhead, total control. - **Weaknesses**: No built-in structure guidance. - **Tip**: Use with checkpoints designed for training with raw guidance (e.g. mask-guided or image-to-image workflows). --- ### **`canny`** - **What It Does**: Applies the Canny edge detection algorithm. - **Needs**: A clean image; edge contrast is key. - **Strengths**: Fast, clean, and sharp outlines. - **Weaknesses**: Overly simplistic for complex forms; loses context. - **Ideal For**: Architectural designs, lineart sketches, hard-edged compositions. --- ### **`canny_pyra`** - **What It Does**: Pyramid-based Canny edge detection—uses multiscale processing for more edge detail. - **Strengths**: Better at picking up fine and coarse edges. - **Weaknesses**: Slightly noisier and slower than basic `canny`. - **Ideal For**: Photographic textures, layered compositions. --- ### **`lineart`, `lineart_anime`, `lineart_manga`, `lineart_any`** All use a neural net trained to extract linework, but with specific style biases. - **`lineart`**: Generic lineart extractor. - **Great for**: Stylized outlines, comic book art. - **`lineart_anime`**: Biased for smooth, cell-shaded anime contours. - **Great for**: 2D anime characters. - **`lineart_manga`**: Biased toward thick-thin black & white linework typical of manga panels. - **Great for**: Black & white manga workflows. - **`lineart_any`**: Trained for multiple styles, most versatile. - **Great for**: Mixed media projects. **Common Strengths**: Stylized structure for cartoon and inked looks. **Common Weaknesses**: May hallucinate or drop edges on photo inputs. --- ### **`scribble`, `scribble_xdog`, `scribble_pidi`, `scribble_hed`** All reduce the image to simplified strokes—good for abstraction and creativity. - **`scribble`**: Raw sketch-like edge map. - **`scribble_xdog`**: Uses Extended Difference of Gaussians for softer, dreamlike edges. - **`scribble_pidi`**: PIDI net used for refined sketch maps. - **`scribble_hed`**: HED-based scribble extraction. **Strengths**: Great for creative workflows, AI doodles, or abstract image-to-image tasks. **Weaknesses**: Not ideal for realism or detail preservation. --- ### **`hed`** - **What It Does**: Holistically-nested Edge Detection. - **Strengths**: Clean contours, preserves semantic shapes. - **Weaknesses**: May blur tight edge detail. - **Use Case**: Ideal for sketches, pose interpretation, soft outlines. --- ### **`pidi`** - **What It Does**: Uses PIDINet for refined semantic edge detection. - **Strengths**: Sharp, content-aware edges. - **Weaknesses**: More VRAM usage. - **Use Case**: Balanced stylized realism. --- ### **`mlsd`** - **What It Does**: Extracts straight lines from images (like MLSD paper). - **Strengths**: Architectural precision. - **Weaknesses**: Useless on organic shapes. - **Use Case**: Buildings, interiors, mechanical schematics. --- ### **`pose`, `openpose`, `dwpose`, `pose_dense`, `pose_animal`** - **`pose` / `openpose`**: Human pose detection (skeleton keypoints). - **`dwpose`**: Deep Whole-body Pose Estimation, better foot/hand coverage. - **`pose_dense`**: Adds facial landmarks and dense joints. - **`pose_animal`**: Pose estimation for animals. **Strengths**: Gives precise figure structure. **Weaknesses**: Can miss limbs in weird angles or crowded scenes. **Use Case**: Character design, pose reference, animation base. --- ### **`normalmap_bae`, `normalmap_dsine`, `normalmap_midas`** - **What It Does**: Converts RGB image to a normal map (pseudo-3D surface info). - **Strengths**: Good for 3D-aware effects, lighting guidance. - **Weaknesses**: Doesn't capture real geometry, only inferred. - **Differences**: - `bae`: Balance of detail and softness. - `dsine`: May emphasize curvature. - `midas`: Mid-level depth approximation. --- ### **`depth`, `depth_anything`, `depth_anything_v2`, `depth_anything_zoe`, `depth_zoe`, `depth_midas`, `depth_leres`, `depth_metric3d`, `depth_meshgraphormer`** These are your _depth prediction models_. - **`depth_anything` / `v2`**: Based on Depth Anything models. - **`depth_zoe` / `anything_zoe`**: Use ZoeDepth for sharper predictions. - **`depth_midas`**: Good general-purpose. - **`depth_leres`**: LeReS network—very accurate, but slower. - **`depth_metric3d`**: Metric depth prediction. - **`depth_meshgraphormer`**: Mesh reconstruction from images. **Strengths**: Amazing for 3D-aware composition, lighting. **Weaknesses**: Long processing time, can create depth artifacts. **Use Case**: Landscapes, portraits with background variation. --- ### **`seg_ofcoco`, `seg_ofade20k`, `seg_ufade20k`, `seg_animeface`** Semantic segmentation: - **`seg_ofcoco`**: COCO object segmentation (general things: people, cars, etc.) - **`seg_ofade20k`**: Scene/semantic parsing of environments. - **`seg_ufade20k`**: Upscaled version of ADE20k. - **`seg_animeface`**: Segment anime faces into parts. **Strengths**: Structural conditioning by regions. **Weaknesses**: Overlaps/ambiguities in masks. **Use Case**: Face swapping, region-specific generation. --- ### **`shuffle`** - **What It Does**: Scrambles image tiles to create chaotic conditioning. - **Strengths**: Adds randomness and variation. - **Weaknesses**: Not deterministic. - **Use Case**: Style transfer, experimentation, glitch art. --- ### **`teed`** - **What It Does**: Transformer-based edge detection. - **Strengths**: Combines edge, semantic, and texture cues. - **Weaknesses**: Slow and heavy on memory. - **Use Case**: High-end stylized or structure-aware compositions. --- ### **`color`** - **What It Does**: Extracts dominant color regions. - **Strengths**: Great for color-guided generation. - **Weaknesses**: No structure, no lines. - **Use Case**: Style transfer, palette preservation. --- ### **`sam`** - **What It Does**: Uses Meta’s Segment Anything Model (SAM) to create mask regions. - **Strengths**: Ultra-precise, works on nearly any object. - **Weaknesses**: May require manual refinement. - **Use Case**: Compositional control, background editing, multi-subject control workflows. --- ## 🧪 Prompting Tips - Pair your preprocessor with ControlNet models designed to accept its output (e.g., use `canny` with a Canny-trained ControlNet model). - For style workflows (anime, manga), combine a stylized preprocessor with a matching LoRA and prompt style. - Want consistency? Use the same preprocessor + seed + conditioning image across multiple runs. --- ## 🚫 What-Not-To-Do-Unless-You-Want-a-Fire Oh, so you like chaos? You enjoy watching your GPU cry? Great, then here's what _not_ to do with the `preprocessor` setting unless you're actively trying to summon the AI demons of instability: #### ❌ Use the wrong preprocessor with the wrong ControlNet model You wouldn't feed a cat spaghetti and expect it to do math. Likewise, don't feed `pose_animal` output into a ControlNet trained for `depth_midas`. The result? Nonsense conditioning, wasted steps, and outputs that look like AI had an existential crisis. **Fix**: Always match your preprocessor with its sibling ControlNet (e.g., `hed` → HED model, `depth_anything` → ControlNet trained on Depth Anything). #### ❌ Forget to install dependencies Half of these preprocessors are built on third-party magic. Missing `detectron2`, `segment-anything`, `openpose`, or `opencv`? You’ll get red errors, blank images, or worse: success that isn’t actually success. **Fix**: Check your install. Use a requirements.txt file. Don’t YOLO this. #### ❌ Run high-res images through `depth_leres` or `sam` on 8GB VRAM If you're running a potato laptop with a fancy GPU sticker but no actual power, please don’t crank `depth_leres` or `sam` to 2048x2048. These models _will_ eat your VRAM and then casually torch your runtime with an out-of-memory error. **Fix**: Stay under 1024x1024 unless you’re packing real heat. #### ❌ Expect perfect outlines from `scribble_xdog` on low-contrast images Low contrast images + `xdog` = muddy soup. It’s not a “dreamlike sketch,” it’s a failed art student’s nightmare. **Fix**: Boost your image contrast before applying `xdog`. #### ❌ Use `shuffle` and expect consistency Shuffle does what it says—it shuffles. It’s not a structured preprocessor, it’s an agent of chaos. **Fix**: Don’t use it unless you want variety over control. Never in production workflows. Ever. #### ❌ Assume `pose_dense` will get every joint right If your character is lying down, twisted, or facing away from the camera, `pose_dense` might just give up entirely. Expect floating limbs and mysterious spaghetti arms. **Fix**: Stick with standard `pose` or `dwpose` for more stable results. Always validate visually. #### ❌ Mix multiple preprocessors on the same conditioning channel Unless your ControlNet expects a specific composite input (and you _really_ know what you're doing), mixing outputs like `depth` + `canny` into the same ControlNet model is like throwing oil and water into a blender—loud, messy, and completely ineffective. **Fix**: One preprocessor, one ControlNet, per channel. Keep your chaos modular. #### ❌ Skip normalization when using `normalmap_*` Feeding an unnormalized or overly bright image into a `normalmap` extractor? Get ready for washed-out normals or weird lighting shadows. **Fix**: Preprocess with tone mapping or exposure correction first. #### ❌ Rely on `seg_*` for precision mask work Semantic segmentation ≠ accurate masking. These models often blur edges or clip object boundaries. Don’t use them if you're trying to do surgical precision work like inpainting hair strands. **Fix**: Use `sam` instead. It’s designed for precision. #### ❌ Forget that more preprocessing ≠ better results Yes, we know—it’s tempting to run every image through five preprocessors, load five ControlNets, and see what happens. But you’ll probably just get noise, hallucinations, or broken anatomy. **Fix**: Be deliberate. Preprocessors are tools, not spice blends. Pick the one that suits your task, and leave the rest out of your stew. And finally: ##### 🔥 Don’t forget to laugh when it breaks This is ComfyUI. If something goes wrong and you get AI soup or a melted mannequin, remember: it’s not a bug, it’s a rite of passage. --- ## Sampler and Scheduler Compatibility Matrix # Sampler and Scheduler Compatibility Matrix Choosing the right sampler and scheduler combo is kind of like picking the right shoes for a marathon — you can wear flip-flops, but don’t act surprised when you trip at Step 12. Below is your cheat sheet for pairings that actually perform well — no guesswork, no flaming garbage results. ## Best Sampler + Scheduler Compatibility Matrix (Quick View) | **Sampler** | normal | karras | exponential | sgm_uniform | simple | ddim_uniform | beta | linear_quadratic | kl_optimal | | ------------------------------- | ------ | ------ | ----------- | ----------- | ------ | ------------ | ---- | ---------------- | ---------- | | `euler` | ✅ | | | | | | | | | | `euler_cfg` | | ✅ | | | | | | | | | `euler_ancestral` | | | ✅ | | | | | | | | `euler_ancestral_cfg_pp` | | | | | | | | | | | `heun` | | | | | | | | | | | `heunpp2` | | ✅ | | | | | | | | | `dpm_fast` | | | | | | | | | | | `dpm_adaptive` | | | | | | | | | | | `dpmpp_2s_ancestral` | | | | | | | | | | | `dpmpp_2s_ancestral_cfg_pp` | | | | | | | | | | | `dpmpp_sde` | | | | | | | | | | | `dpmpp_sde_gpu` | | | | | | | | | | | `dpmpp_2m` | | ✅ | | | | | | | | | `dpmpp_2m_cfg_pp` | | | | | | | ✅ | | | | `dpmpp_2m_sde` | | ✅ | | | | | | | | | `dpmpp_2m_sde_gpu` | | | | | | | | | | | `dpmpp_3m_sde` | | | | | | | | ✅ | | | `dpmpp_3m_sde_gpu` | | | | | | | | | | | `ddpm` | | | | | | | | | | | `lcm` | | | | ✅ | | | | | | | `ipndm` | | | | | | ✅ | | | | | `ipndm_v` | | | | | | | | | | | `deis` | | | | | ✅ | | | | | | `res_multistep` | | ✅ | | | | | | | | | `res_multistep_ancestral` | | | | | | | | | | | `re_multistep_ancestral_cfg_pp` | | | | | | | | | | | `gradient_estimation` | | | | | | | | | | | `gradient_estimation_cfg_pp` | | | | | | | ✅ | | | | `er_sde` | | | ✅ | | | | | | | | `seeds_2` | | | | | | | | | | | `seeds_3` | | | | | | | | | | | `ddim` | | | | | | | | | | | `uni_pc` | | | | | | | | | ✅ | ### ✅ Quick Legend: - **✅** = Best known scheduler pairing for this sampler. - Blank = Not recommended / niche use / no clear benefit pairing. --- Pairing the right scheduler with the right sampler in ComfyUI isn't just a “nice to have” — it's the difference between buttery-smooth masterpieces and noisy, incoherent messes. While most samplers _technically_ work with most schedulers, that doesn’t mean they should. Each sampler has unique mathematical characteristics — some prioritize precision, others speed, others realism — and the scheduler determines how that sampling process unfolds over time. The wrong combination can undermine your output quality, tank performance, or worse, make your beautifully engineered workflow behave like it just rolled out of a chaos factory. Choosing the best pairings ensures you get faster generations, better detail retention, smoother gradients, and more consistent results — especially in high-stakes workflows like SDXL, animations, or multimodal conditioning. Trust us: aligning your scheduler with the sampler’s strengths is the easiest quality boost you can make without touching a single prompt. ## 📚 Detailed Best Pairing List | Sampler | Best Scheduler | Why This Pairing Works | Sampler Docs | Scheduler Docs | | -------------------------- | ---------------- | ---------------------------------------------------------------------------------- | -------------------------- | ---------------- | | euler | normal | Fast and sharp results, good for sketch-style or high-contrast work. | euler | normal | | euler_cfg | karras | Maintains CFG-weighted detail well, stable under long prompts. | euler_cfg | karras | | euler_ancestral | exponential | Best for dreamy, soft lighting and slow transitions. | euler_ancestral | exponential | | dpmpp_2m | karras | High-quality, well-balanced — the industry gold standard. | dpmpp_2m | karras | | dpmpp_2m_cfg_pp | beta | CFG-enhanced DPM++ with excellent edge preservation. | dpmpp_2m_cfg_pp | beta | | dpmpp_2m_sde | karras | Fantastic for realism; handles shading and depth extremely well. | dpmpp_2m_sde | karras | | dpmpp_3m_sde | linear_quadratic | Complex scene generation, rich gradients, great for SDXL. | dpmpp_3m_sde | linear_quadratic | | heunpp2 | karras | Cleaner transitions between token weight shifts, good for intricate prompt detail. | heunpp2 | karras | | lcm | sgm_uniform | Optimal fast sampler; pairs with low step configs. | lcm | sgm_uniform | | uni_pc | kl_optimal | Adaptive and smart. Excels at high-resolution and SDXL workflows. | uni_pc | kl_optimal | | deis | simple | Very clean, progressive sampling. Pairs well with text-to-image. | deis | simple | | ipndm | ddim_uniform | Great compromise for noise-controlled diffusion steps. | ipndm | ddim_uniform | | res_multistep | karras | Works well for animations and sequential inference. | res_multistep | karras | | gradient_estimation_cfg_pp | beta | Smooth transitions, precise edge definition for CFG-heavy workflows. | gradient_estimation_cfg_pp | beta | | er_sde | exponential | Best used for SDXL variants and 3D-looking renders. | er_sde | exponential | ## 🧩 Notes on Exclusions - `ddpm`, `seeds_2`, `seeds_3`, `dpm_adaptive`, and `dpm_fast` were excluded for being **legacy/utility** samplers or having no strong "best" pairing — they work, but aren't ideal for quality-first workflows. - If you don’t see a combo listed here, assume it’s okay but not optimal unless you have a very specific reason to use it. - We’re skipping raw CFG samplers unless you're explicitly building a custom pipeline that depends on parallel prompt/latent conditioning. ## 🧯 What-Not-To-Do-Unless-You-Want-a-Fire - ❌ Pair `lcm` with `exponential`, `kl_optimal`, or `linear_quadratic`. It's meant for speed and doesn't behave well with over-complicated schedulers. - ❌ Use `uni_pc` with `simple` or `ddim_uniform` unless you like flat, lifeless outputs. - ❌ Stack CFG samplers (`*_cfg_pp`) without a prompt setup that supports dual CLIP encoders. You'll lose all that enhanced guidance precision you paid for. - ❌ Apply `dpmpp_sde_gpu` with high noise schedulers (`exponential`, `ddim_uniform`) unless you're tuning for chaos. ## 📚 Additional Resources - 🔗[Scheduler Options](https://comfyui.dev/docs/guides/Other%20Resources/scheduler-options) - 🔗[Sampler_Name Options](https://comfyui.dev/docs/guides/Other%20Resources/sampler-name-options) --- ## Sampler_Name Options # Sampler_Name Options **TL;DR:** The `sampler_name` defines _how_ the denoising process interprets and walks through the noise space during diffusion. Each algorithm has its own way of dealing with noise, speed, coherence, prompt alignment, and quirks. Think of these like different chefs following the same recipe—some are minimalist Michelin stars, others are heavy-metal grillmasters. Each affects your results. ## 🤖 Table of Samplers | Sampler | Best For | Recommended Scheduler | Notes | Strengths | Weaknesses | | ------------------------------- | ------------------------------------------------------ | --------------------- | ---------------------------------------- | ------------------------------------------ | ----------------------------------------- | | `euler` | Fast deterministic generations, previews | `simple` | Very predictable, sharp at high steps | Very fast and consistent | Can be harsh or noisy at high steps | | `euler_cfg` | Same as euler with better prompt control | `simple` | Slightly better at following prompts | Adds better control to classic Euler | Slightly more resource usage | | `euler_ancestral` | Textured, slightly chaotic results | `simple` | More randomness, good for art | Creates more textured, artistic outputs | Less predictable | | `euler_ancestral_cfg_pp` | Textured creative results with better prompt adherence | `simple` | Adds CFG handling | Balanced creativity and prompt guidance | Complexity may increase render time | | `heun` | Balanced generations, alternative to Euler | `karras` | Stable but slower | Stable and smooth results | Slower than Euler | | `heunpp2` | Better quality than Heun, improved CFG | `karras` | Photorealistic scenes | Improved quality with prompt fidelity | Still underused and less tested | | `dpm_fast` | Quick drafts and prototyping | `exponential` | Fastest among DPMs | Extremely fast rendering | Sacrifices image quality | | `dpm_adaptive` | Quality-aware fast sampling | `exponential` | Adjusts internally for better output | Auto-balances speed and quality | Unpredictable output fidelity | | `dpmpp_2s_ancestral` | Highly varied textures, creative images | `karras` | Good for expressive scenes | Great for varied creative imagery | May over-randomize details | | `dpmpp_2s_ancestral_cfg_pp` | Creative + prompt fidelity | `karras` | More guided variation | Balances texture and prompt control | Slower due to CFG | | `dpmpp_sde` | Clean gradients, realism | `karras` | Smooth transitions | Realistic gradients and transitions | Higher VRAM usage | | `dpmpp_sde_gpu` | Same as above but faster on GPU | `karras` | GPU optimized | GPU-accelerated smooth rendering | Needs compatible hardware | | `dpmpp_2m` | Balanced realism and speed | `karras` | Versatile for many styles | Balanced, great for realism | More steps needed than ancestral samplers | | `dpmpp_2m_cfg_pp` | Detailed, prompt-loyal realism | `karras` | Most recommended general-purpose sampler | Best for realism with tight prompt control | Increased processing time | | `dpmpp_2m_sde` | Even smoother realism | `karras` | Great for portraits | Extremely smooth, clean output | Slower with high step counts | | `dpmpp_2m_sde_gpu` | Smooth realism, GPU-friendly | `karras` | High batch efficiency | Smooth realism, GPU-friendly | Needs VRAM headroom | | `dpmpp_3m_sde` | Highest fidelity, complex detail | `karras` | Slow but premium output | Ultimate detail and fidelity | Very slow | | `dpmpp_3m_sde_gpu` | High-detail GPU accelerated | `karras` | For batch high-end generation | High-detail GPU accelerated | Heavy on GPU resources | | `ddpm` | Legacy and experimentation | `ddim_uniform` | Slow and stable | Stable and accurate to source diffusion | Very slow and outdated | | `lcm` | Ultra-fast low-step generation | `karras` | Requires LCM-tuned models | Lightning-fast generation | Requires special models and low steps | | `ipndm` | Experimental, high coherence | `sgm_uniform` | Use with caution | High detail and structure | Experimental and unstable | | `ipndm_v` | Variation of IPNDM, smoother results | `sgm_uniform` | Experimental | More stable than ipndm | Still experimental | | `deis` | Fast, lightweight quality | `exponential` | Compact generation | Quick generation with decent quality | Occasional prompt drift | | `res_multistep` | Artistic, surreal images | `normal` | Needs higher steps | Dreamlike stylized art | Unstable and slow | | `res_multistep_ancestral` | Dreamlike, unstable beauty | `normal` | Chaos-driven | Hyper-stylized art | Chaos-prone and unpredictable | | `re_multistep_ancestral_cfg_pp` | Prompt-driven surrealism | `normal` | Slow and expressive | Controlled surrealism | High computational cost | | `gradient_estimation` | Precision edge-case work | `linear_quadratic` | Very slow, niche | Precision where others fail | Exceptionally slow | | `gradient_estimation_cfg_pp` | Same as above with prompt fidelity | `linear_quadratic` | Ultra-niche | Prompt-sensitive precision | Same slowness plus complexity | | `er_sde` | Stable, smooth realism | `karras` | Balance of all factors | Balanced realism | Slower generation time | | `seeds_2` | Internal, unknown use | `normal` | Undocumented | Possibly internal for seed processing | Undocumented use | | `seeds_3` | Internal, unknown use | `normal` | Undocumented | Possibly internal for seed control | Undocumented use | | `ddim` | All-purpose generation | `ddim_uniform` | Balanced across most needs | Fast, versatile, prompt-sensitive | Can lack texture or detail | | `uni_pc` | High-quality, stable outputs | `karras` | Modern, robust sampler | High stability, great realism | Moderately slower | ## 🧠 Detailed Sampler Breakdown ### 📌 Euler & Friends - `euler`: The OG. Fast, low-memory, deterministic. - **Use for**: Fast previews, consistent outputs. - **Avoid if**: You want dreamy aesthetics. - `euler_cfg`: Euler with better CFG control. - **Use if**: You find `euler` too rigid. - `euler_ancestral`: Adds randomness for richer textures. - **Use for**: More creative, slightly less predictable results. - `euler_ancestral_cfg_pp`: Post-prompt processing; blends chaotic charm with CFG wizardry. --- ### ⚙️ Heun Variants - `heun`: Like Euler but tries to be smarter. A compromise between speed and precision. - `heunpp2`: Heun, but updated with second-order CFG handling. - **Pro tip**: Works well for photorealism when Euler feels too harsh. --- ### 🚀 DPM Family (Euler on Steroids) - `dpm_fast`: Speed demon. Sacrifices some quality for rapid generation. - `dpm_adaptive`: Dynamically adjusts for better quality mid-run. - `dpmpp_2s_ancestral`: Two-stage, good for varied textures. More artistic. - `dpmpp_2s_ancestral_cfg_pp`: Same as above but with better prompt adherence. - `dpmpp_sde`: Introduces SDE smoothing—great for clean gradients and realism. - `dpmpp_sde_gpu`: GPU-optimized for large batches. - `dpmpp_2m`: Two-mode version. Think “midpoint-aware” sampler. - `dpmpp_2m_cfg_pp`: CFG-enhanced version. Best used with complex prompts. - `dpmpp_2m_sde`: Even smoother. Mixes SDE and midpoint sampling. - `dpmpp_2m_sde_gpu`: GPU-tuned for smoother multi-image workflows. - `dpmpp_3m_sde`: Three-mode version. Slower but higher fidelity. - `dpmpp_3m_sde_gpu`: GPU-optimized flavor for serious jobs. > 🧪 Tip: The `dpmpp` samplers are your best bet for realism, complexity, and prompt fidelity. Try `dpmpp_2m_cfg_pp` with `karras` scheduler for S-tier output. --- ### 🧱 Basics and Benchmarks - `ddpm`: The original reverse diffusion. Good for understanding the roots of it all but slow for production. - `ddim`: A happy medium. Fast, decent quality, and widely supported. --- ### 🧠 Neural Wizardry - `lcm`: Latent Consistency Models. - **Use for**: Insanely fast generation (think < 6 steps). - **Note**: Needs LCM-tuned models/checkpoints. Use low `steps` (4–6). - `ipndm`, `ipndm_v`: Implicit noise prediction. High quality but needs babysitting. - **Experimental**: Try if you enjoy edge-case debugging. - `deis`: High-speed lightweight solver. Not always accurate but shockingly fast. - `uni_pc`: State-of-the-art. Combines stability with high detail. - **Highly recommended** for any polished workflow. --- ### 🎨 Artistic Samplers - `res_multistep`: Applies restarts during denoising for richer style changes. - `res_multistep_ancestral`: More chaotic cousin. Better for surrealism. - `re_multistep_ancestral_cfg_pp`: If you _must_ marry surrealism and prompt obedience. --- ### 🧮 Math Nerd Specials - `gradient_estimation`: Gradient-based logic. Use when nothing else aligns. - `gradient_estimation_cfg_pp`: Adds prompt handling, but slow. Niche. --- ### 🤖 Oddballs - `er_sde`: SDE (Stochastic Differential Equation) method. Balanced but slow. - `seeds_2`, `seeds_3`: Basically undocumented, possibly used internally or experimentally. Avoid unless testing. --- ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire 1. **Using `lcm` with high steps.** - Unless you're doing AI archaeology, LCM should be used at 4–6 steps. Higher values just waste time and return mushy nonsense. 2. **Combining samplers and schedulers randomly.** - Yes, technically you can combine `euler` with `karras`, but also technically, you can eat soup with a fork. Stick to the scheduler designed for the sampler, or expect... unpredictable spaghetti. 3. **Assuming GPU samplers work on CPU.** - `*_gpu` versions are not friendly with CPUs. Expect crashes, freezes, or your PC trying to roast marshmallows with its fan. 4. **Trying `gradient_estimation` for casual generations.** - If you like waiting 30 minutes for an image that looks like the same image every other sampler does in 30 seconds… be my guest. 5. **Using experimental samplers in client work.** - `ipndm`, `ipndm_v`, and `seeds_*` are not production-safe unless you enjoy gambling with broken renders mid-batch. 6. **Using high-CFG samplers (`*_cfg_pp`) with lazy prompts.** - These samplers expect detailed, descriptive prompts. If you feed them “portrait of woman,” expect an AI-generated meme instead of a masterpiece. 7. **Assuming “ancestral” means “better.”** - It means “more chaotic.” Sometimes that works. Sometimes it gives you spider limbs. 8. **Ignoring your VRAM budget.** - Samplers like `dpmpp_3m_sde_gpu` will happily devour 12GB+ of VRAM like it’s brunch. Don’t say we didn’t warn you. --- ## 🧪 Best Pairings | Sampler | Scheduler | Suggested Steps | Notes | | ----------------- | ------------- | --------------- | -------------------------------------- | | `dpmpp_2m_cfg_pp` | `karras` | 25–35 | Excellent balance of speed/detail | | `uni_pc` | `karras` | 20–30 | Works well across styles | | `lcm` | `lcm` | 4–6 | Ultra fast, only with LCM-ready models | | `ddim` | `exponential` | 20–30 | Great for soft, coherent results | | `euler_ancestral` | `simple` | 20–40 | Artistic, textured output | #### Want to learn more about schedulers? Check out [even more documentation](http://comfyui.dev/docs/guides/Other%20Resources/scheduler-options) here! ## 💥 Summary Choosing the right `sampler_name` is like picking the right brush for your AI-generated masterpiece. If you’re running basic txt2img? Euler or DDIM. Want hyper-realism? Go `dpmpp_2m_cfg_pp`. Want it now and fast? `lcm` is your guy. Just don’t pick `gradient_estimation_cfg_pp` and wonder why your render time rivals the development of Elden Ring. --- ## Scheduler Options # Scheduler Options The `scheduler` parameter in the `KSampler` node determines **how noise is added and removed across the diffusion process**, i.e., how your model hallucination spirals into a majestic image instead of becoming pixel soup. Different schedulers shape the _trajectory_ of denoising across timesteps—some ramp up slowly and gently (like a spa day for your U-Net), while others hit it hard and fast (like Monday morning meetings). > ⚠️ TL;DR: Choosing the wrong scheduler can make even the best checkpoint and sampler combo underperform. Choose wisely, and don't just throw `karras` at everything like it's seasoning. --- ## 🧭 Available Schedulers Each of these is a **timestep weighting function**—a fancy way to say it controls how noise levels change during sampling. | **Scheduler** | **Description** | **Strengths** | **Weaknesses** | **Best Use Case** | | ------------------ | -------------------------------------------------------------------------------------------------- | ----------------------------------------------------- | ----------------------------------------- | ------------------------------------------------------ | | `normal` | Linear timesteps. Simple, straightforward. | Predictable, works well with basic samplers. | Not optimized for high-frequency detail. | Baseline testing, quick iterations. | | `karras` | Sigmoid-inspired distribution from Karras et al. (yes, _that_ Karras). Emphasizes low-noise steps. | More steps in low-noise region = sharper results. | Needs compatible samplers like `dpmpp_*`. | High-quality render goals, realism checkpoints. | | `exponential` | Applies an exponential weighting curve. | Adds flexibility in contrast and detail gradation. | May over-emphasize early noise steps. | Stylized or abstract renders. | | `sgm_uniform` | Uniform SDE sampler from Score-Based Generative Modeling. | Balanced treatment across steps. | Doesn’t favor fine detail as much. | Experimental workflows or SGM-style outputs. | | `simple` | Even simpler than `normal`—fixed intervals. | Good for debugging or ultra-fast tests. | Crashes and burns with complex prompts. | Internal testing, stress scenarios. | | `ddim_uniform` | Uniform steps used in DDIM (Deterministic Denoising Implicit Models). | Consistent, reliable for DDIM. | May lack sharpness in fewer steps. | DDIM workflows, predictable outputs. | | `beta` | Uses a beta schedule curve, often for noise variance control. | Smooth interpolation, soft gradients. | Requires tuning to really shine. | Portraits, fantasy art. | | `linear_quadratic` | Timesteps follow a linear-to-quadratic curve. | Gradual build-up; favors softer transitions. | Can appear too “blurred” at low steps. | Landscapes, inpainting workflows. | | `kl_optimal` | KL divergence optimized scheduler (yes, it’s as nerdy as it sounds). | Produces highly optimized, compressed noise profiles. | Extremely picky with samplers. | Research-level workflows, when min-maxing every pixel. | --- ## 🛠️ How Each Scheduler Works (In Painstaking Detail) ### 🔹 `normal` - **Curve**: Linear (e.g., timestep 1, 2, 3, ..., N). - **What It Does**: Assigns equal weight across timesteps. - **Why It Matters**: Offers the most “average” trajectory. Works fine with samplers like `euler`, `heun`, or `dpm_fast`. - **Good For**: Basic tests, learning workflows, and as a fallback when the rest don’t work. --- ### 🔹 `karras` - **Curve**: Sigmoid-ish. Condenses most denoising toward the lower-noise region (the end). - **What It Does**: Saves more time for fine detail at the end, delaying major denoising until later. - **Why It Matters**: It’s the darling of high-fidelity image samplers (`dpmpp_2m`, `dpmpp_sde`, etc.). - **Good For**: Realism, detail-heavy workflows, portraits, concept art. - **Caution**: Pair only with samplers designed to play nicely—otherwise you’ll wonder why your image looks like it went through a microwave. --- ### 🔹 `exponential` - **Curve**: Rapid early steps, diminishing returns. - **What It Does**: Gets most of the denoising done early. - **Why It Matters**: Good for stylized workflows or high-noise, fast-degeneration scenarios. - **Good For**: Abstract renders, anime styles, wild style LoRAs. - **Watch Out**: Might blow past fine details. --- ### 🔹 `sgm_uniform` - **Curve**: Uniform spread, SDE-compliant. - **What It Does**: Applies Score-Based Generative Modeling’s noise treatment uniformly. - **Why It Matters**: Helpful if you’re using an SGM-based model or exploring consistency in variance across timesteps. - **Good For**: Science! Experiments! 🤓 --- ### 🔹 `simple` - **Curve**: Dumb as bricks—just flat. - **What It Does**: Applies no intelligence to the steps. A scheduler only in name. - **Why It Matters**: It doesn’t, unless you’re benchmarking something. - **Good For**: Testing custom scheduler integrations. --- ### 🔹 `ddim_uniform` - **Curve**: Uniform timestep schedule tailored to DDIM. - **What It Does**: Equal distribution for deterministic sampling. - **Why It Matters**: Stable results if you're using `ddpm` or `ddim` samplers. - **Good For**: Low-step fast generations, consistency over multiple runs. --- ### 🔹 `beta` - **Curve**: Beta-distribution (as in statistics, not “early release”). - **What It Does**: Tightly controls noise variance. - **Why It Matters**: Helpful in smoothing harsh noise transitions. - **Good For**: Soft transitions, background-heavy compositions, vintage-looking styles. --- ### 🔹 `linear_quadratic` - **Curve**: Starts linear, ends quadratic. - **What It Does**: Gracefully builds up noise decay, almost like a cinematic pan. - **Why It Matters**: Creates more natural gradients and transitions. - **Good For**: Scenic illustrations, dreamy atmospheres. --- ### 🔹 `kl_optimal` - **Curve**: Custom-fitted for minimizing Kullback-Leibler divergence. - **What It Does**: Matches denoising to information-theoretic optimal paths. - **Why It Matters**: This is for math nerds and research papers. Optimized for minimal image loss. - **Good For**: Quantitative benchmark work, precision-tuned generation. - **⚠️ Danger**: May explode (metaphorically) if used with incompatible samplers. --- ## 🧪 Compatible Sampler Pairings (Summary Table) | **Scheduler** | **Best Samplers** | | ------------------ | --------------------------------------------------- | | `normal` | `euler`, `heun`, `dpm_fast`, `ddpm` | | `karras` | `dpmpp_2m`, `dpmpp_sde`, `dpmpp_2s_ancestral` | | `exponential` | `euler_ancestral`, `heunpp2`, `dpm_adaptive` | | `sgm_uniform` | `dpmpp_sde`, `heun`, `ddpm` | | `simple` | `dpm_fast`, `ddpm`, `euler` | | `ddim_uniform` | `ddpm`, `dpm_adaptive`, `simple` | | `beta` | `dpmpp_2m_cfg_pp`, `heun`, `dpmpp_sde_gpu` | | `linear_quadratic` | `euler_cfg`, `heunpp2`, `dpmpp_2s_ancestral_cfg_pp` | | `kl_optimal` | `dpmpp_sde`, `dpmpp_sde_gpu`, `dpmpp_2m` | #### Want to learn more about samplers? Check out [even more documentation](http://comfyui.dev/docs/guides/Other%20Resources/sampler-name-options) here! --- ## 🧠 Pro Tips - **Consistency across renders**: If your results suddenly start melting like wax fruit, check if the `scheduler` changed on accident. - **Experimentation**: Each scheduler interacts differently with steps and denoise strength. Keep those constant if you want valid A/B tests. - **Match your scheduler to your sampler**: Treat them like best friends. Or frenemies—either way, they’ve got chemistry, so don’t mismatch. --- ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire Because sometimes, you just want to watch your GPU cry. #### ❌ Don’t Mix `karras` With Non-Compatible Samplers If you're out here trying to use `karras` with `euler` or `heun`, just... no. They weren't designed to play nice. You’re essentially asking a potato to run a sprint. You _will_ get mud (i.e., blurry results and artifacts). > 🔥 Symptom: Images that are either grotesquely smooth or inexplicably noisy. > 👎 Why: `karras` expects samplers like `dpmpp_2m`, `dpmpp_sde`, etc., that understand its fancy "low-noise priority" scheme. #### ❌ Using `kl_optimal` Without a Degree in Statistical Thermodynamics Yes, it's called “optimal,” but if you're not pairing it with the correct KL-aware samplers or you're running 4 denoise steps hoping for a miracle, you're not optimizing anything—you’re inviting entropy. > 🔥 Symptom: Either nothing renders, or you get something that looks like modern art... but not in a good way. > 👎 Why: KL-optimal schedules assume high sampler precision and lots of steps. #### ❌ Cranking Scheduler + Sampler Settings Without Adjusting Denoise Let’s say you picked `karras` and `dpmpp_2m`, great. But then you leave your `denoise` at 0.9 and wonder why everything looks like it was pulled from a VCR in 1993. > 🔥 Symptom: Overbaked, grainy images with lost details and lighting gone rogue. > 👎 Why: These schedulers require thoughtful denoise balancing—especially around 0.4–0.7. #### ❌ Assuming All Schedulers Are Created Equal They’re not. You wouldn't put diesel in a Tesla, right? (Okay, maybe you would. That's why this section exists.) > 🔥 Symptom: Dull, inconsistent results even with high-quality checkpoints. > 👎 Why: The scheduler _directly_ influences how well the sampler interprets latent space. Compatibility is key. #### ❌ Running High-Step `linear_quadratic` Without Patience or a Coffee IV Yes, it can be buttery smooth—but it’s slow. Like slow-cooker slow. If you’re in a 30-step render using that with `dpmpp_2s_ancestral_cfg_pp`, congrats—you’ve now entered **ComfyUI retirement mode**. > 🔥 Symptom: Performance bottlenecks and render times that rival cooking a turkey. > 👎 Why: Heavy step weighting + high step count = latency hell. #### ❌ Blindly Copy-Pasting Settings From Some Random Reddit Workflow Just because "some guy" posted a great result with `sgm_uniform` and `heunpp2` doesn’t mean your project—or checkpoint—is compatible. > 🔥 Symptom: Mild panic as every output is just slightly wrong. > 👎 Why: Context matters. That workflow might use a LoRA, specific VAE, or niche base model. --- #### 🧯 Golden Rule: The Scheduler Is Not a Skin Cream Don't just apply it and hope for the best. It’s not surface-level—it controls the _soul_ of the denoising arc. Treat it with respect. Or better yet, RTFM (which you're doing right now, so gold star for you). --- ## Note # Note Because sometimes your future self needs a Post-It too. --- ## 🧠 What This Node Does The `Note` node is a **pure utility node** in ComfyUI that serves one humble but essential purpose: letting you insert freeform text notes _directly inside your workflow_. This text can be used for labeling, documentation, reminders, or snarky messages to your teammates. It **does not affect the execution** of the graph in any way—no inputs, no outputs, no side effects. It’s just a textbox. And sometimes, that’s exactly what you need. --- ## 🧩 Node Type - **Category**: Utility - **Type**: Annotation / Documentation - **Execution**: Skipped at runtime --- ## 🔌 Inputs - _None._ The `Note` node is a true loner. It doesn’t accept or need any inputs. --- ## 🔁 Outputs - _None._ This node doesn’t output anything either. Not even a high five. --- ## ⚙️ Parameters & Fields | Field | Type | Description | | ------ | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `note` | `string` (multiline text) | The content of your note. You can write anything here—markdown, plain text, sarcastic reminders to yourself, TODOs, or even ASCII art if you’re feeling frisky. There’s no formatting enforcement. Just type your heart out. | --- ## ✅ Real-World Use Cases - **Workflow Documentation**: Label each section of your graph to describe what it’s doing. Future you (and anyone else looking at the workflow) will be grateful. - **Testing Notes**: Leave yourself notes when comparing different models or parameters. - **Workflow Debugging**: Write out what you _think_ is happening vs what is _actually_ happening. - **Collaboration**: Add instructions or context for others who might use or modify your workflow. - **Reminders**: Need to update the model? Replace a VAE? Add a watermark? Just write it here and stop forgetting. --- ## 🧪 Workflow Setup Example You don’t connect this node to anything, which means setup is stupidly simple: 1. Drag the `Note` node into your workflow. 2. Double-click the `note` field or open the properties panel. 3. Type whatever message you want to leave. 4. Close the panel. That’s it. Repeat as needed across your workflow like it’s a crime scene covered in yellow sticky notes. --- ## 💬 Prompting Tips - Not applicable. This node doesn't interact with prompts. - But if you _must_, you can leave suggested prompts inside your note field as references for future prompt engineering. --- ## 🔥 What-Not-To-Do-Unless-You-Want-A-Fire - Don’t try to wire this node to anything. Seriously. It has no ports. - Don’t expect it to do anything during execution—it won’t. - Don’t rely on this as a secure form of documentation for sharing workflows online. Some frontend variants or exports may skip utility-only nodes when compressing workflows. --- ## ⚠️ Known Issues - Some user interfaces may collapse or minimize the note field, making it easy to miss. - Notes do not currently support rich text, markdown rendering, or hyperlinks. - If you’re working on massive workflows, the node might become buried in a sea of chaos—consider color-coding your sections using dummy nodes or labels if your frontend allows it. --- ## 📝 Final Thoughts The `Note` node is the unsung hero of workflow sanity. It won’t render an image, run a model, or sample anything, but it _will_ make your life easier. It’s the “narrator” of your workflow, and a place to add clarity in a world full of `latent_image` spaghetti. So go ahead—drop a few notes. Your brain will thank you later. --- ## Primitive # Primitive Welcome to the chameleon of ComfyUI — the **Primitive** node. This deceptively simple utility node is one of those behind-the-scenes MVPs of any modular workflow. Think of it like a universal plug: it adapts to what you connect it to, giving you the power to reuse and centralize parameters across your graph. Whether you’re tired of hardcoding duplicate seeds across 4 KSamplers or you just want to propagate a consistent string or float through multiple branches of your workflow — the Primitive node’s got your back. ## 🧠 What This Node Does The **Primitive** node automatically adjusts its input field depending on the type of data it’s connected to. It’s a _type-aware parameter node_ that can output a **string** or **number (float or int)**, depending on how and where it’s used. This allows it to act as a shared parameter across multiple nodes — without duplicating values manually — making your workflow more modular, DRY (Don't Repeat Yourself), and infinitely more Comfy. ## 🧩 Node Type - **Category:** Core Utility - **Node Class:** `Primitive` - **Inputs:** None (manual entry only) - **Outputs:** Type-dependent (dynamic, based on connection) - **Tags:** Parameter Control, Type-Aware, Value Sharing ## 🔌 Inputs **None.** You don’t plug anything into a Primitive node — you configure its value manually in the UI and wire it outward. ## 🔄 Outputs This is where things get interesting. The output type is dynamic and inferred from the connection context: - **String**: If connected to a node expecting a text-based input. - **Number (float or int)**: If connected to a numeric input like `seed`, `cfg`, `steps`, etc. The node will adapt its internal value parsing automatically when the output is connected to an input port. If you connect it to a sampler seed, it becomes a number. If you connect it to a prompt filename prefix, it becomes a string. Magic? No. Just solid type inference. ## ⚙️ Settings & Parameters (And What They _Really_ Do) | **Field** | **Type** | **Description** | | --------- | ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `value` | _Auto_ (String or Number) | The only field in the Primitive node. You manually type in a string, integer, or float, and it adapts based on what node you connect it to. Want `42` to behave as a seed in multiple samplers? Just enter `42` and connect away. Need a filename like `output_image_01`? Go ahead, Primitive can handle that too. | > **Note:** The node does not ask you what type you're entering. It infers it once you make the connection. So if you're seeing weird behavior, check your wires. ## ✅ Recommended Use Cases - **Centralized Seed Control**: Use one Primitive node to control the seed across multiple KSamplers for deterministic outputs. - **Synchronized Filename Prefixes**: Output images from various branches using the same naming scheme. - **Shared Numeric Parameters**: Send a single float or int value to multiple nodes like `cfg`, `denoise`, or custom modules. - **Reusable String Tokens**: For nodes that require class tokens, tags, or filenames, you can centralize them in one Primitive node. ## 🔁 Example Workflow Setup Let’s say you’re running **three KSamplers**, and you want them to all use the same seed: 1. Drop in one **Primitive** node. 2. Set `value` to `123456`. 3. Connect the Primitive’s output to the `seed` input of each KSampler. 4. Congratulations, you just centralized seed management like a pro. ## 💬 Prompting Tips While this node doesn’t handle prompts directly, here are some edge-case tricks: - Use a Primitive node to **dynamically set prompt suffixes or prefixes** by connecting it to nodes that use `filename_prefix`, `prompt_fragment`, or similar. - Combine it with nodes like **String Combine** or **Conditioning nodes** to construct more flexible prompts. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - ❌ **Don’t expect it to auto-convert between incompatible types** — Primitive will change type only based on what it's connected to, not based on your intentions. - ❌ **Don’t connect one Primitive to multiple different types** (e.g., one input needs a float, another needs a string) — it will adapt to one, and the others might throw type mismatch errors. - ❌ **Don’t forget to recheck connections** when you copy or modify your workflow — Primitive might cling to its old inferred type. ## ⚠️ Known Issues - **Type Locking Quirks**: Sometimes, once a type is inferred, disconnecting doesn’t always reset it visually. Reconnect it to the new input type and refresh the graph if needed. - **No Manual Type Selector**: You can’t explicitly set the type (yet). It’s 100% context-driven. - **UI Doesn’t Validate Input Until Connection**: If you type a string and then connect to a float input, you may get a silent failure or an error. ## 🧼 Final Notes The Primitive node is one of those “why didn’t I use this earlier” tools. It’s clean, flexible, and downright necessary for anyone building moderately complex workflows in ComfyUI. It's also criminally underrated — but you? You know better now. Go forth and reduce redundancy. Let your values live in one place. Be the architect of your own efficient, elegant graph. Got multiple samplers, conditionings, or filename suffixes screaming for consistency? **Primitive node. Use it. Love it. Document it (well, we just did).** --- ## ae.safetensors # ae.safetensors > Because your image deserves more than a potato-grade decode. --- ## 🔧 File Format - **Filename:** `ae.safetensors` - **File Type:** `.safetensors` - **Serialization Format:** [SafeTensors](https://github.com/huggingface/safetensors) – a blazing-fast, zero-copy format designed for security and speed. It avoids common pitfalls of `.ckpt` or `.pt` (e.g. arbitrary code execution on load). **Why it matters:** Loading a VAE in SafeTensors format ensures compatibility, speed, and peace of mind (no malicious payloads embedded in your decoder). Plus, it's what ComfyUI _actually prefers_ when you’re not trying to cook your machine with janky `.ckpt` files. ## 📁 Function in ComfyUI Workflows In a ComfyUI workflow, the **VAE (Variational Autoencoder)** is responsible for: - [**Decoding latent tensors**](https://comfyui.dev/docs/guides/Nodes/vae-decode) into full RGB images. - **Encoding images** back into the latent space (when applicable). - Maintaining **fidelity**, **dynamic range**, and **color integrity** between the latent and visual representation. Think of it as the bridge between your hallucinated dream world (latent) and the visible reality (image). If you don’t load a proper VAE, you're going to get weird outputs—blurry, overly contrasted, or plain old busted. ## 🧠 Technical Details - **Type:** Variational Autoencoder - **Architecture:** Likely derived from SD 1.4/1.5’s original VAE backbone - **Dimensions:** Works with latent spaces of shape `4x64x64` for 512x512 inputs - **Training:** Pretrained on large datasets to preserve visual fidelity during decode - **Latent Precision:** 32-bit float, standardized - **Compression Target:** Latent downsampling by factor of 8 (e.g., 512x512 → 64x64) This VAE is not fancy. It’s the vanilla backbone of stability—**not trained for anime, not trained for face beautification, not trained for fantasy**—just _solid general-purpose image representation_. ## ✅ Benefits - **Stable decoding** for checkpoints trained on vanilla SD1.4 / SD1.5 - **Lightweight** and memory efficient - **Fast loading** due to SafeTensors format - Good color retention\*\* compared to baked-in VAEs - **Compatible with nearly everything** not trying to be special ## ⚙️ Usage Tips - Always **explicitly load this VAE** if your checkpoint doesn't have a baked-in VAE (or if it's terrible). - If you're seeing muddy images, weird contrast, or blocky compression artifacts—your VAE might be misconfigured. Start here. - **Match the VAE to the checkpoint.** If you're using `v1-5-pruned-emaonly`, `revAnimated`, `Deliberate`, etc., this VAE is often a safe fallback. - If you’re using a fancy model (like anime, furry, or inpainting specialists), they might require a custom VAE. This one will still _work_, but possibly not _well_. ## 🤝 Best Model Compatibility The `ae.safetensors` VAE works best with **Standard 1.x Stable Diffusion checkpoints** like: | Checkpoint                                                                                           | Status                                                  | | ---------------------------------------------------------------------------------------------------- | ------------------------------------------------------- | | `v1-5-pruned-emaonly.safetensors`                                                                    | ✅ Perfect match                                        | | `revAnimated_v122.safetensors`                                                                       | ✅ Good default                                         | | `deliberate_v2.safetensors`                                                                          | ✅ Compatible                                           | | [`dreamshaper_8.safetensors`](https://comfyui.dev/docs/guides/Checkpoints/dreamshaper-8-safetensors) | ✅ Works fine                                           | | `anything-v3.ckpt`                                                                                   | ⚠️ Acceptable, but may benefit from anime-tuned VAEs    | | `RealisticVisionV5.safetensors`                                                                      | ⚠️ Will decode, but may blunt realism-specific tuning   | | `SDXL-based models`                                                                                  | ❌ Incompatible – completely different latent structure | ## 📍 Setup Instructions in ComfyUI 1. **Make sure you have the file `ae.safetensors` placed in your `ComfyUI/models/vae/` folder.** 2. In your workflow, add a [`Load VAE`](https://comfyui.dev/docs/guides/Nodes/load-vae) or `VAELoader` node. 3. Select `ae.safetensors` from the dropdown menu (may show up as `ae`). 4. Connect it to the appropriate nodes: - For **CheckpointLoaderSimple**, connect `VAE` output from the loader. - For **VAE Decode**, pipe in the VAE input. 5. Preview your outputs with an `ImageViewer` node to verify that decoding looks sharp, accurate, and not like someone threw Vaseline on the lens. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - ❌ Don’t skip loading a VAE entirely and expect usable results. - ❌ Don’t use this VAE with **SDXL checkpoints** – you will get garbage (if anything at all). - ❌ Don’t mix VAE expectations—i.e., using a baked-in VAE checkpoint with an external VAE like this can lead to double-decoding weirdness. - ❌ Don’t use this with **inpainting-specific models** unless you know the model was trained with this VAE. - ❌ Don’t overwrite this VAE thinking you’re being helpful. This is the "safe fallback" one. You’ll want to keep it. ## 📚 Additional Resources - 🔗 Download [`ae.safetensors`](https://huggingface.co/ffxvs/vae-flux/tree/main) - 🧪 [ComfyUI GitHub](https://github.com/comfyanonymous/ComfyUI) - 📂 Your local `models/vae/` folder is your kingdom—keep it clean. ## 📎 Example Node Configuration **VAELoader Node:** ```json { "type": "VAELoader", "vae_name": "ae.safetensors" } ``` **Typical Workflow Connection:** ```text CheckpointLoaderSimple        └──▶ Model, Clip        └──▶ VAE ──▶ KSampler / VAE Decode ``` ## 📝 Notes - This is _not_ a "fancy" VAE—it’s a **baseline**. That’s the point. It’s there when your baked VAE fails or is just plain ugly. - This VAE does _not_ support SDXL, which uses a different encoder/decoder architecture (with two VAEs, because why use one when you can overengineer?). - Keep this one as your **default fallback** in workflows. It’s reliable, tested, and unlikely to break. Still seeing overcompressed blobs? Your issue might not be the VAE—it might be your sampler, CFG scale, or ControlNet settings. But starting with a good VAE like `ae.safetensors` gives your workflow a fighting chance. --- ## diffusion_pytorch_model.safetensors # diffusion_pytorch_model.safetensors > Because without a VAE, your model's beautiful latent space dreams stay... well, latent. --- ## 🔧 File Format - **Filename:** `diffusion_pytorch_model.safetensors` - **Format:** `.safetensors` (because `pickle` is a horror show waiting to happen) - **Serialization:** Safe, deterministic, non-executable - **Model Type:** Variational Autoencoder (VAE) - **Framework:** PyTorch-compatible The `.safetensors` format is a secure and efficient way of storing model weights — basically, all the important number soup that makes your AI art generator tick without giving you a security vulnerability as a parting gift. ## 📁 Function in ComfyUI Workflows This file is used in the **VAE Loader** node within ComfyUI and serves one purpose: > To decode latent representations (those weird fuzzy image blobs models generate) into the actual pixel-based images we know and love. It sits between the **latent generation phase** (thanks to your diffusion model) and **the real world**. Without it, your output looks like an LSD trip through a fog machine. ### Where It Appears - Plugged into the `VAE` input of: - [`VAE Decode`](https://comfyui.dev/docs/guides/Nodes/vae-decode) - [`KSampler`](https://comfyui.dev/docs/guides/Nodes/ksampler) - [`Ultimate SD Upscale`](https://comfyui.dev/docs/guides/Nodes/ultimate-sd-upscale) - Pretty much anything that needs to get to or from the latent space ## 🧠 Technical Details Let’s dig deep: - **Architecture:** Standard Stable Diffusion VAE, based on the encoder-decoder style where: - The **encoder** maps an image to a compressed latent space - The **decoder** reconstructs that latent space back into an image - **Latent Size:** Compresses from 512x512 down to 64x64 (i.e., 1/8th of the original resolution) - **Channels:** Operates on 4 latent channels for compatibility with the typical SD 1.4 / 1.5 latent space - **Training Data:** Derived from the training dataset of Stable Diffusion 1.4/1.5 - **Loss Functions:** - **Reconstruction loss** for accurate reconstructions - **KL divergence** for nice, smooth latent distributions (so your outputs don’t go wild) This VAE is **not** fine-tuned for specialty checkpoints — it's the general-purpose workhorse. ## ✅ Benefits - Compatible out-of-the-box\*\* with most SD 1.4/1.5 checkpoints - Clean image reconstruction\*\* from latent outputs - **Stable results**, perfect for workflows that rely on consistency - Low risk of surprise artifacts\*\*, assuming you're not feeding it latent junk Bonus: it's boring — and in VAE world, boring means **stable and reliable**. ## ⚙️ Usage Tips 1. **Pair wisely.** Works best with vanilla SD 1.4 and 1.5 checkpoints. Don’t expect it to keep up with fancy anime LoRAs or SDXL finetunes. 2. **Use for decoding.** Drop this into the `VAE Loader`, route it to your `VAE Decode`, and voilà — coherent images. 3. **Don’t encode unless you mean it.** This VAE can encode, but unless you're running an invert pipeline, you're probably here to _decode_. 4. **Combine with Ultimate SD Upscale** for extra magic when upscaling from latent space. 5. **If you see washed-out colors**, your VAE might not match your checkpoint. Double-check the pairing. ## 🧬 Which Model Types This Works Best For | Model Type | Compatibility | Notes | | ----------------------- | ------------- | ----------------------------------------------------------------- | | ✅ SD 1.4 / 1.5 | Excellent | This is what it was built for. | | 🟡 SD 1.5 derivatives | Good-ish | Depends on the deviation from base SD 1.5. | | 🔴 SDXL | No | Totally different architecture. Use a different VAE. | | 🔴 Anime-focused models | Risky | Use a model-specific VAE (like `vae-ft-mse-840000`) instead. | | 🟡 Realistic LoRA-heavy | Caution | If you’re using LoRAs, try matching VAEs to your base checkpoint. | ## 📍 Setup Instructions 1. **Download the file** from a trusted source (see 📚 Additional Resources). 2. **Place it in your VAE folder:** bash CopyEdit `ComfyUI/models/vae/` 3. **Restart ComfyUI** (yes, you have to, sorry). 4. **Add the `VAE Loader` node** to your workflow. 5. Select `diffusion_pytorch_model.safetensors` from the dropdown. 6. Connect to `VAE Decode` or wherever else VAE is required. 7. Generate stuff and feel smug about doing it right. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire Here’s how to destroy your workflow in five easy steps: - ❌ **Use this with SDXL checkpoints** You’ll get trash. Or worse, something that _almost_ looks right, but isn’t. - ❌ **Mismatch with anime-style checkpoints** You’ll get pale colors, mushy details, and regrets. - ❌ **Forget to restart ComfyUI after adding the VAE file** The dropdown won’t see it. You’ll panic. Don’t be that person. - ❌ **Encode with this VAE then decode with another** Inconsistent results and weird artifacts await. - ❌ **Rename the file improperly or mess with the `.safetensors` extension** ComfyUI is picky, and for good reason. Just don’t. ## 📚 Additional Resources - 🔗 Download this [`diffusion_pytorch_model.safetensors`](https://huggingface.co/stabilityai/sdxl-vae/tree/main) ## 📎 Example Node Configuration **Node:** `VAE Loader` **Settings:** - `vae_name`: `diffusion_pytorch_model.safetensors` **Connected To:** - `VAE Decode` → Outputs image - `KSampler` (optional) → Latent decoding post-sampling - `Ultimate SD Upscale` → For latent upscaling workflows ## 📝 Notes - If you’re trying to create pixel-perfect realism or high-fidelity fantasy, start by picking the **right** VAE. This one’s great for default 1.5-based pipelines, but falls short in style-specific pipelines (anime, ultra-realism, etc.). - Always restart ComfyUI after adding new models. You’d think this would be automatic, but no. ComfyUI demands rituals. - This VAE does _not_ include baked-in optimizations or enhancements — it’s pure vanilla decoder/encoder joy. Need a VAE that just works? `diffusion_pytorch_model.safetensors` is your no-nonsense, plain bagel. And honestly, sometimes that’s exactly what you need — because not every image needs sprinkles, glitter, or a GPU meltdown. --- ## flux2-vae.safetensors # flux2-vae.safetensors > Because Flux 2 decided text should actually be readable and details shouldn't look like they went through a cheese grater. --- ## 🔧 File Format - **Filename:** `flux2-vae.safetensors` - **File Type:** `.safetensors` - **Serialization Format:** [SafeTensors](https://github.com/huggingface/safetensors) — the same battle-tested, zero-copy format that keeps your machine safe from rogue pickles and malicious payloads hiding in `.ckpt` files. **Why it matters:** SafeTensors isn't just fast—it's the format that won't execute arbitrary Python code when you load it. When you're dealing with a next-generation VAE like Flux 2's, you want speed, security, and zero compromises. This is what ComfyUI was built to love, and what your workflow deserves when you're not trying to turn your GPU into a space heater with sketchy checkpoint files. ## 🎯 Function in ComfyUI Workflows In a ComfyUI workflow, the **Flux 2 VAE (Variational Autoencoder)** serves as the critical translation layer between latent space and pixel space. Specifically, it handles: - [**Decoding latent representations**](https://comfyui.dev/docs/guides/Nodes/vae-decode) into high-fidelity RGB images with significantly improved text rendering capabilities - **Encoding images** back into the latent space for img2img, inpainting, and other advanced workflows - Preserving **fine details**, **text clarity**, and **color accuracy** that Flux 1's VAE would have turned into impressionist mush - Maintaining **structural integrity** of complex patterns, sharp edges, and intricate textures that previous-generation VAEs struggled with Think of this as the difference between reading a crisp 4K display and squinting at a CRT monitor from 1995. The Flux 2 VAE was specifically engineered to fix the "why does my text look like alphabet soup" problem that plagued earlier models. If you're generating anything with text, fine architectural details, or intricate patterns, this VAE is the difference between professional output and "I made this in MS Paint" vibes. ## 🧠 Technical Details - **Type:** Variational Autoencoder (Next-Generation Architecture) - **Architecture:** Custom-built for the Flux 2 model family with architectural improvements over Flux 1 - **Latent Dimensions:** Optimized for Flux 2's native latent space structure (different from SD 1.x's 4-channel setup) - **Training Focus:** - Text rendering fidelity (finally, readable signs and UI elements) - Fine detail preservation (no more "did that edge get hit by a blur hammer?") - Enhanced color depth and gradient smoothness - Reduced compression artifacts in high-frequency areas - **Latent Precision:** 32-bit float, standardized for numerical stability - **Compression Strategy:** Adaptive compression that prioritizes detail retention over aggressive size reduction - **Model Specificity:** Designed exclusively for Flux 2 checkpoints—not backwards compatible with Flux 1 or SD 1.x/SDXL **What makes it special:** Unlike the vanilla SD VAEs that treat all image content equally (and equally poorly), the Flux 2 VAE employs learned prioritization. It knows that text edges matter more than smooth sky gradients, that fine architectural details shouldn't become blurry suggestions, and that compression artifacts are the enemy of professional output. The architectural improvements over Flux 1's VAE are non-trivial—this isn't just a retrained version, it's a legitimate upgrade to the encoding/decoding pipeline. **Bit depth handling:** The Flux 2 VAE maintains higher precision in the decode path, particularly for challenging content like: - Small text ( Model ──> KSampler └──> CLIP ──> CLIP Text Encode VAELoader (flux2-vae.safetensors) └──> VAE ──> VAE Decode KSampler └──> Latent ──> VAE Decode ──> Save Image ``` ### Complete Minimal Workflow: ``` [Load Checkpoint] checkpoint: flux2-dev.safetensors outputs: model, clip, vae (ignored) [VAELoader] vae_name: flux2-vae.safetensors output: vae [CLIP Text Encode - Positive] text: "your prompt here" clip: from checkpoint output: positive conditioning [CLIP Text Encode - Negative] text: "low quality, blurry" clip: from checkpoint output: negative conditioning [KSampler] model: from checkpoint positive: from positive encode negative: from negative encode latent_image: from Empty Latent Image seed: random steps: 20 cfg: 7.0 sampler_name: euler scheduler: normal output: latent [VAE Decode] samples: from KSampler vae: from VAELoader output: image [Save Image] images: from VAE Decode ``` ### Text-Heavy Generation Example: For images with important text content (signs, UI, typography): ``` [Positive Prompt] "A vintage neon sign that says 'FLUX 2 CAFE', highly detailed, sharp focus, professional photography" [Negative Prompt] "blurry, low quality, jpeg artifacts, text artifacts, unreadable text" [Sampler Settings] steps: 25-30 (higher for text clarity) cfg: 6.5-7.5 (too high can artifact text) sampler: euler or dpmpp_2m ``` The Flux 2 VAE will handle the text rendering significantly better than previous-generation VAEs, but proper prompting and sampling still matter. ## 📝 Notes ### On Architectural Improvements The Flux 2 VAE isn't just "Flux 1 but retrained"—it represents genuine architectural advances in how latent representations are decoded. The text rendering improvements alone required changes to the decoder network structure, not just different training data. This is why backwards compatibility with Flux 1 isn't possible; the latent space organization changed at a fundamental level. ### On File Size Yes, `flux2-vae.safetensors` is larger than SD 1.x VAEs. That's because it's a more sophisticated architecture with more parameters dedicated to preserving fine details. If you're complaining about a few hundred extra megabytes in the age of multi-gigabyte checkpoints, you might want to reevaluate your priorities. Quality costs bits. ### On Text Rendering The "improved text rendering" isn't marketing speak—it's a measurable, visible difference. Flux 1's VAE would turn small text into blurry suggestions. Flux 2's VAE actually decodes letterforms with crisp edges and readable characters down to surprisingly small point sizes. If your workflow involves UI mockups, signage, posters, or any text-heavy generation, this VAE is non-negotiable. ### On Fine Detail Preservation Beyond text, the detail preservation improvements show up in: - Architectural elements (window frames, molding, trim work) - Fabric textures and weave patterns - Skin pores and fine facial features - Intricate jewelry and small mechanical parts - Natural textures like tree bark, stone, and grass The difference is most noticeable at full resolution. If you're previewing at 512×512, you won't see the benefits. Generate at 1024×1024 or higher to appreciate what this VAE can do. ### On Memory Usage The Flux 2 VAE requires more VRAM during decode operations than lighter VAEs. This is physics, not poor optimization. If you're on limited hardware (8GB VRAM or less), you might need to: - Use tiled VAE decode for very large images - Reduce batch sizes - Close other VRAM-heavy applications Don't try to solve this by downgrading to a worse VAE. Solve it with proper workflow optimization. ### On Compatibility (Again, Because People Don't Read) **Flux 2 VAE works with Flux 2 models. Full stop.** Not Flux 1. Not SD. Not SDXL. Not "it's all diffusion models, they should be compatible." Each model family has specific latent space structures, and the VAE must match. Using the wrong VAE won't just produce bad images—it might produce no images, or errors, or images so wrong you'll question whether your GPU is having an existential crisis. ### On Future-Proofing As Flux 2 gets fine-tuned, LoRA'd, and adapted for specialized tasks, this VAE will remain the standard decode path. Any serious fine-tune will have been trained with this VAE in the pipeline, meaning switching to something else would invalidate that work. Keep this file, keep it accessible, and don't get creative trying to "improve" it unless you're prepared to retrain entire checkpoints. ### On Workflow Sharing When sharing Flux 2 workflows, make sure recipients know they need: 1. The Flux 2 checkpoint 2. flux2-vae.safetensors 3. T5 text encoder 4. CLIP text encoder All four components are required. Missing any one = non-functional workflow. Be specific in your documentation. ### On Quality Expectations The Flux 2 VAE sets a high bar for decode quality. If you're not seeing sharp, detailed, artifact-free outputs, the problem is likely: - Wrong checkpoint (using Flux 1 or SD instead of Flux 2) - Wrong VAE loaded (check your VAELoader node) - Sampling issues (CFG too high, steps too low, bad sampler choice) - Prompt issues (contradictory elements, unclear descriptions) Don't blame the VAE until you've verified your entire pipeline is Flux 2 native. ### Final Reminder This is the **baseline, official VAE for Flux 2**. It's not fancy, it's not specialized—it's the foundation. Keep it in your `models/vae/` folder, load it explicitly in workflows, and trust that it's doing its job correctly. When something goes wrong with Flux 2 image quality, check everything else before assuming the VAE is the problem. It's probably not. --- **Still seeing compression artifacts or blurry text?** Check your sampler settings, verify you're using an actual Flux 2 checkpoint, confirm your VAE is loading correctly, and make sure your prompt isn't contradicting itself. The flux2-vae.safetensors file is a precision instrument—but it can't fix problems that originate upstream in your workflow. Start with the VAE, verify it's connected properly, then trace backwards if issues persist. --- ## sdxl_vae.safetensors # sdxl_vae.safetensors Welcome, brave workflow architect. You’ve summoned the `sdxl_vae.safetensors` VAE—a not-so-humble component that does a lot more than people give it credit for. This is the _official-ish_ guide for wrangling this beast inside ComfyUI. Strap in. --- ## 🔧 File Format - **Filename:** `sdxl_vae.safetensors` - **Format:** `safetensors` (yep, it's in the name) - **Type:** VAE (Variational Autoencoder) - **Storage:** Stored in a secure, memory-safe format. Unlike `.ckpt` files, `safetensors` avoids arbitrary code execution—so you can sleep at night. ## 📁 Function in ComfyUI Workflows The VAE's job is to **encode and decode** image data between pixel-space and latent-space. In a workflow, it’s typically used in two primary places: 1. **During decoding:** Converts the final latent representation into an actual image (`VAE Decode` node). 2. **During conditioning or previewing latents:** Assists in operations that need the visual interpretation of the latent image (e.g. previews, inpainting). Without a proper VAE, your stunning SDXL-generated output will look like a glitchy Minecraft painting. You’ve been warned. ## 🧠 Technical Details - **Architecture:** This VAE is specifically trained for **SDXL** (Stable Diffusion XL). - **Latent Size:** 4-channel latent representation at 1/8th the original image resolution. - **Decoder Precision:** High-fidelity conversion from latent → RGB space with enhanced handling of skin tones, edge smoothness, and color gradients. - **Encoder Functionality:** Not heavily used in most workflows unless you're reverse-engineering an image into latents. **Key Specs:** - Designed for `1024x1024` base SDXL resolution. - Optimized for SDXL’s dual-CLIP architecture (but does not itself interact with CLIP directly). - Trained with **improved perceptual loss functions** for cleaner outputs. ## ✅ Benefits - **Sharp Details:** Recovers minute image details, especially useful for faces, text, and complex objects. - **Accurate Color Rendering:** No more overbaked reds or nuclear greens. - **Native to SDXL:** Works out-of-the-box with all SDXL-based checkpoints (e.g. `sd_xl_base_1.0`, `JuggernautXL`, `realisticVisionXL`). - **No Frankenstein Stitches:** Unlike mismatched VAEs, this one doesn’t leave weird stitching, warping, or shading artifacts. ## ⚙️ Usage Tips - **Pair with SDXL Checkpoints Only.** This isn't a generic VAE—it expects SDXL's latent distributions. - **Decode after KSampler:** Make sure you’re decoding the latent after the KSampler node, not before. Otherwise, enjoy your abstract AI-generated surrealism. - **Previewing:** You can use a `VAE Decode` node early in the pipeline to preview outputs mid-generation if you want to get fancy. - **Tiling + Upscaling:** If you’re upscaling with tools like Ultimate SD Upscale, `sdxl_vae.safetensors` helps maintain image integrity across tiles. ## 🧬 Which Model Types This Works Best For | Model Type | Compatible? | Notes | | --------------------------- | ----------- | ------------------------------------------------------------------- | | `SDXL Base 1.0` | ✅ Yes | Native support. Use this VAE for all Base XL generations. | | `SDXL Refiner` | ✅ Yes | Works fine, though Refiner has its own output goals. | | `JuggernautXL`, `RealVisXL` | ✅ Yes | Works well. These inherit SDXL's VAE expectations. | | `SD 1.5 / 2.1` Checkpoints | ❌ Nope | Do **not** use this VAE with older SD models. Latent mismatch! | | Anime Checkpoints | ❌ Nope | Use anime-tuned VAEs instead. This will mangle them into spaghetti. | ## 📍 Setup Instructions 1. **Drop the File:** - Place `sdxl_vae.safetensors` in your ComfyUI VAE models folder 2. **Load the VAE:** - Use the [`VAELoader`](https://comfyui.dev/docs/guides/Nodes/load-vae) node. - Choose `sdxl_vae.safetensors` from the dropdown in the node. 3. **Connect It:** - Plug the output of the `VAELoader` node into any node expecting a `VAE` input (typically your `KSampler` or `VAE Decode`). 4. **Decode Latents:** - Use the `VAE Decode` node to convert the final latent back into a pretty image. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire Seriously. Don’t do these things unless you're into pixelated chaos: - ❌ **Don’t use with non-SDXL models.** It’s like trying to run diesel in a Tesla. - ❌ **Don’t skip decoding.** Latent-space images are _not_ your final image. Decode or perish. - ❌ **Don’t mismatch resolutions.** Avoid decoding latents with mismatched resolution scales (i.e., generated at 512x512 then decoded with 1024x1024 expectations). - ❌ **Don’t rename the file weirdly.** If it’s not recognized in the VAELoader dropdown, it won’t load. Keep the `.safetensors` extension intact. - ❌ **Don’t double-load a VAE.** Loading more than one can result in pipeline confusion and longer rendering times. Pick one and commit. ## 📚 Additional Resources - 🔗 **Download:** [`sdxl_vae.safetensors`](https://huggingface.co/stabilityai/sdxl-vae/tree/main) - 📖 [**VAE Concepts**](https://huggingface.co/docs/diffusers/en/api/pipelines/latent_diffusion) - 🛠️ **SDXL Base Checkpoint:** Make sure you have this loaded too for proper pairing. ## 📎 Example Node Configuration **VAELoader Node:** ```json { "id": 5, "type": "VAELoader", "pos": [100, 200], "size": [250, 50], "flags": {}, "order": 0, "mode": 0, "inputs": [], "outputs": [{ "name": "VAE", "type": "VAE" }], "properties": { "vae_name": "sdxl_vae.safetensors" } } ``` **Connection Example:** ```pgsql Checkpoint → KSampler → VAE Decode → SaveImage ↑ VAELoader (sdxl_vae) ``` ## 📝 Notes - This VAE is **standard issue** for SDXL pipelines. If your output looks funky, blurry, or dark, check that you're using this VAE. - For refiner workflows, you _can_ use this same VAE, though inpainting workflows may have other VAE preferences depending on resolution fidelity. - If you update your SDXL model, **double-check your VAE**. Some checkpoints come with bundled VAEs; others expect external ones like this. - If you’re not sure what VAE your model expects, check the model card or README. Or, you know, ask ChatGPT again. 😉 If you made it this far without setting your workflow on metaphorical fire—congrats. You're now ready to wield `sdxl_vae.safetensors` with the confidence of someone who’s read the documentation (because you just did). Go forth and decode like a pro. 🧪🔥 --- ## vae-ft-mse-840000-ema-pruned.safetensors # vae-ft-mse-840000-ema-pruned.safetensors The `vae-ft-mse-840000-ema-pruned.safetensors` is a fine-tuned, lightweight VAE trained using Mean Squared Error (MSE) loss and Exponential Moving Average (EMA) smoothing. It’s optimized for **SD 1.4/1.5 checkpoints** and is widely used in ComfyUI workflows to decode latent images into high-quality pixel outputs. This isn’t just a VAE—it’s _the_ VAE you swap in when you’re tired of the default one giving your images the emotional range of a potato. --- ### 🔧 File Format - **Filename:** `vae-ft-mse-840000-ema-pruned.safetensors` - **File Type:** `.safetensors` - **Purpose:** VAE for decoding latent representations into image space - **Optimization:** EMA + MSE for smooth, stable, high-fidelity output - **Model Family:** Stable Diffusion 1.x compatible --- ### 📁 Function in ComfyUI Workflows This VAE connects to key nodes like `VAE Decode`, `KSampler`, and `CheckpointLoaderSimple`. It improves final image output by ensuring that decoded images reflect the full potential of your latent vectors—without barfing color gradients or introducing bizarre lighting. Typical hook-up points: - **VAE Decode** - **KSampler** (when a VAE is required alongside a checkpoint) - **Auto-decoding pipelines** --- ### 🧠 Technical Details | Property | Value | | --------------------- | --------------------------------- | | Training Loss | Mean Squared Error (MSE) | | Training Steps | 840,000 | | EMA Applied | ✅ Yes | | Format | `.safetensors` | | Model Pruned | ✅ Reduced unnecessary parameters | | Ideal Checkpoint Pair | `v1-5-pruned-emaonly.safetensors` | | ComfyUI Compatibility | ✅ Fully Supported | --- ### ✅ Benefits - Crisp edges and smoother skin rendering - Improved detail fidelity without bloating your workflow - Faster loading thanks to pruning - Compatible with virtually all 1.x-based checkpoints --- ### ⚙️ Usage Tips - **Best for:** Realism workflows, portrait renders, concept art, style transfers - **Avoid:** Mixing with SDXL checkpoints (unless you're trying to reinvent glitch art) - Compare against `vae-ft-mse-560000` or baked VAEs to assess output quality per use case --- ### 📍 ComfyUI Setup Instructions 1. Drop in your `CheckpointLoaderSimple` node. 2. Set your checkpoint to `v1-5-pruned-emaonly.safetensors` (or similar). 3. Set `vae_name` to `vae-ft-mse-840000-ema-pruned.safetensors`. 4. Wire the VAE output to your `VAE Decode` node. 5. Bask in your gloriously reconstructed images. --- ### 🔥 What-Not-To-Do-Unless-You-Want-a-Fire Let’s avoid turning your beautiful workflow into a dumpster inferno. Here’s what _not_ to do: - ❌ **Do not use with SDXL checkpoints** This is a VAE for SD 1.x. If you connect it to SDXL models, don’t come crying when your render looks like a Dali painting got dunked in a microwave. - ❌ **Do not pair with baked VAE checkpoints** If your checkpoint _already includes a baked-in VAE_, this one will just argue with it. It’s like wearing two pairs of glasses at once—confusing and blurry. - ❌ **Do not skip connecting the VAE in `VAE Decode`** The node literally exists to decode using the VAE. Skipping it is like ordering sushi and forgetting the fish. - ❌ **Do not rename the file with typos and wonder why it won’t load** Misspell it as `vae-ft-mse-840000-ema-purned.safetensors` and watch ComfyUI have a quiet meltdown. - ❌ **Don’t plug it into SDXL-specific nodes expecting magic** This VAE doesn’t understand SDXL’s latent dimensions. It’s like trying to watch Netflix on a toaster. --- ### 📚 Additional Resources - [Download It Now!](https://huggingface.co/stabilityai/sd-vae-ft-mse-original/blob/main/vae-ft-mse-840000-ema-pruned.safetensors) - [ComfyUI GitHub Repository](https://github.com/comfyanonymous/ComfyUI) --- ### 📎 Example Node Configuration ```json { "type": "CheckpointLoaderSimple", "inputs": { "ckpt_name": "v1-5-pruned-emaonly.safetensors", "vae_name": "vae-ft-mse-840000-ema-pruned.safetensors" } } ``` --- ### 📝 Notes - **Do not mix incompatible VAE and checkpoint versions** (e.g., SDXL VAEs with SD 1.5 base models). - If you notice unusual banding or oversaturation, test alternate VAEs or adjust the decoding step strength. --- ## wan_2.1_vae.safetensors # wan_2.1_vae.safetensors Welcome to the documentation for `wan_2.1_vae.safetensors`, a finely tuned VAE (Variational Autoencoder) model tailored for use with **Stable Diffusion 2.1** checkpoints in ComfyUI. If you’ve ever had your generations come out blurry, desaturated, or looking like they were filtered through a fog machine, this VAE might just be your secret sauce. --- ## 🔧 File Format - **Filename:** `wan_2.1_vae.safetensors` - **Format:** `.safetensors` - **Serialization:** [`safetensors`](https://github.com/huggingface/safetensors) is a secure and fast model serialization format designed to prevent pickle-based exploits. - **Precision:** Likely in `fp16` or `fp32` depending on your source — most `.safetensors` VAE files default to `fp16` for lower VRAM usage. - **Size:** Usually between 300MB–500MB (depending on the optimizer and precision used) > ✅ Pro Tip: Never rename `.safetensors` files to `.pt` or `.ckpt`. They're _not_ the same thing and your workflow will break harder than a dollar store tripod. ## 📁 Function in ComfyUI Workflows In a ComfyUI workflow, the **VAE decodes the latent outputs** of the Stable Diffusion model into full-resolution images. Without it, your precious diffusion process would result in ugly latent garbage no one wants to look at (trust me, we’ve all been there). ### Where it's used: - Attached to the `Load Checkpoint` node OR the standalone `Load VAE` node. - Works in tandem with the `VAE Decode` node to produce final images from latent tensors. - Required to visualize any latent output in a human-readable format (aka an actual image). ## 🧠 Technical Details `wan_2.1_vae.safetensors` is a VAE fine-tuned specifically for **Stable Diffusion v2.1**. It's a significant visual quality upgrade over the default `vae-ft-mse-840000-ema` VAE that ships with most SD 2.1 checkpoints. ### Features: - **Trained on higher-resolution data:** Better detail preservation during decoding. - **Higher dynamic range:** Less crushed blacks and more color fidelity. - **Minimized over-blurring:** Avoids the overly smoothed look that some other VAEs introduce. - **Better for photorealism and fine details.** - **Improved latent space alignment** with `2.1`-style models. ### Architecture: - Same VAE backbone as default SD VAE. - Trained with improved reconstruction loss and perceptual tuning. ## ✅ Benefits - **Sharper outputs:** Crisper image decoding from latent to RGB. - **More accurate color mapping:** Color tones and gradients appear closer to prompt intent. - **Preserves fine details:** Textures and tiny facial features won't turn into mush. - **Drop-in compatible:** Works in all SD 2.1-compatible pipelines with no special finagling. - **Low VRAM footprint:** Suitable for mid-range GPUs (8GB+ VRAM). ## ⚙️ Usage Tips - Use **with SD 2.1 base or derivative checkpoints** like `realisticVision`, `revAnimated_v2Rebirth`, or [`dreamshaper_8`](https://comfyui.dev/docs/guides/Checkpoints/dreamshaper-8-safetensors) (if adapted to 2.1). - Use in tandem with high `CFG` values (7–12) and moderate `Denoise` (~0.5–0.7) for best results. - To test output changes, toggle VAEs mid-workflow. ComfyUI makes this easy with node swapping. - Pair with [**Ultimate SD Upscale**](https://comfyui.dev/docs/guides/Nodes/ultimate-sd-upscale) for super-detailed renders at high resolution. - Use **`KSampler` with samplers like `DPM++ 2M Karras`** for refined quality when paired with this VAE. > 🎯 Ideal batch size? 1–2 for mid-tier cards; go nuts if you’re rocking an RTX 4090. ## 🎯 Which Model Types This Works Best For | ✅ Recommended Models | ⚠️ Not Recommended Models | | ----------------------- | ------------------------------ | | `stable-diffusion-v2.1` | `stable-diffusion-v1.5` | | `realisticVision-v5.1` | `anything-v4`, `anime models` | | `revAnimated_v2Rebirth` | `pastelMix`, `cetusMix` | | `dreamshaper_8 (2.1)` | `darkSushiMix`, `chilloutMix` | | `analog-diffusion-2.1` | Any SDXL model (use SDXL VAE!) | > TL;DR: This VAE speaks fluent 2.1. Don’t toss it into a 1.5 party — they’ll just argue all night. ## 🛠️ Setup Instructions 1. **Download the VAE:** - You can find it on CivitAI, HuggingFace, or your favorite model repo. 2. **Place the file in the ComfyUI models folder:** bash CopyEdit `ComfyUI/models/vae/` 3. **Load it via one of two methods:** - **Standalone `Load VAE` node:** - Add it manually and connect it to `VAE Decode` - **Or embed it inside your `Load Checkpoint` node:** - Select it under the `vae` dropdown (only available in updated versions) 4. **Hit run.** If you're seeing artifacts, blurry textures, or washed-out color, check your sampler, VAE match, and resolution. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - ❌ **Don't use with SDXL models.** This isn’t their VAE. You’ll get garbage-tier outputs. - ❌ **Don't use with anime checkpoints.** Unless you _want_ your anime to look like watercolor soup. - ❌ **Don't rename the file to `.pt` or `.ckpt`.** Just... no. - ❌ **Don't skip decoding.** Forgetting the `VAE Decode` node = zero image output. Just raw tensors. Not helpful. - ❌ **Don’t pair with the wrong latent resolution.** This VAE is meant for **512x512 base** latent shapes, as used in SD 2.1. If you’re upscaling from `Empty Latent Image`, use `64x64` latent with 8x scaling unless otherwise modified. ## 📚 Additional Resources - **Download**: [WAN 2.1 VAE](https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/tree/main/split_files/vae) - **Related Models**: - [`stable-diffusion-v2.1-base`](https://huggingface.co/stabilityai/stable-diffusion-2-1-base) - [`realisticVision`](https://civitai.com/models/4201/realistic-vision-v60-b1) - [`revAnimated_v122`](https://civitai.com/models/7371/rev-animated) ## 📎 Example Node Configuration ```json { "id": 5, "type": "VAELoader", "pos": [200, 1000], "size": [320, 60], "inputs": [ { "name": "vae_name", "type": "COMBO", "widget": { "name": "vae_name", "default": "wan_2.1_vae.safetensors" }, "link": null } ], "outputs": [{ "name": "VAE", "type": "VAE", "links": [999] }], "properties": { "Node name for S&R": "VAELoader", "models": [ { "name": "wan_2.1_vae.safetensors", "url": "https://huggingface.co/waaayoff/vae" } ] } } ``` ## 📝 Notes - Not all VAEs are created equal. `wan_2.1_vae` stands out as a top-tier VAE for realism-heavy SD 2.1 workflows. - Always match your VAE to your checkpoint. Mixing 1.5 checkpoints with a 2.1 VAE is like mixing oil and water and calling it a smoothie. - If your generations look off, try swapping VAEs. It's one of the fastest ways to troubleshoot rendering quality. - If you're building workflows with reusable nodes, create a **dedicated Load VAE subgraph** to swap VAEs quickly. Need help troubleshooting? Just remember the holy trinity: **Checkpoint ➜ VAE ➜ Decoder** Get one of those wrong and you’re just diffusing into the void. --- ## wan_2.2_vae.safetensors # wan_2.2_vae.safetensors ## 🔧 File Format - **File Name:** `wan_2.2_vae.safetensors` - **Format:** `safetensors` — a secure and fast binary format designed to replace `.pt` and `.ckpt`. - **Model Type:** VAE (Variational Autoencoder) - **Location (ComfyUI):** Should be placed in your `ComfyUI/models/vae/` directory. > **Why `safetensors` matters:** It's faster, avoids torchscript pitfalls, and won't let you accidentally run arbitrary Python code just by loading a model. Basically, it won’t torch your GPU or your dignity. ## 📁 Function in ComfyUI Workflows In ComfyUI, the VAE is used to **decode** latent space data into the RGB images you actually see. Think of it as the translator between the language of AI neurons and human eyeballs. When a checkpoint (like `wan_v2.2.ckpt`) generates a latent image, this VAE decodes it into pixel data. Without a matching VAE, you're likely to get: - Desaturated colors - Crushed shadows - Blown highlights - A general “what is this soggy potato?” vibe. So yes, a good VAE = happy visuals. ## 🧠 Technical Details - **Trained for:** Compatibility with `wan_v2.2` SD1.5 checkpoint - **Architecture:** Follows Stable Diffusion-style VAE layout, tuned with a custom loss function that better preserves **color contrast** and **structural details**. - **Dynamic Range Handling:** Excellent — unlike older VAEs, this one doesn’t nuke your highlights or mudify shadows. - **Color Space Calibration:** Carefully finetuned to reflect true prompt-driven saturation and tone shifts. > TL;DR: This VAE actually _cares_ about your colors. ## ✅ Benefits - 🖼 **Improved visual fidelity:** Richer tones, deeper contrast, cleaner lines. - 🌈 **Better color reconstruction:** Great for colorful and high-dynamic-range outputs. - 🧪 **Built to match** the `wan_v2.2` SD1.5 model — using any other VAE is just asking for weirdness. - 🧵 **No more stitching artifacts** or blobby textures — goodbye AI meltdowns. ## ⚙️ Usage Tips - **Always pair it with the `wan_v2.2` checkpoint.** It’s not just a suggestion. It’s the VAE’s soulmate. - Use **higher-resolution outputs** (768x768 or above) to fully appreciate the detail retention. - If you're using it with **Ultimate SD Upscale** or **Inpainting workflows**, expect excellent results in both base and enhanced passes. - If you get slightly dull color output, try raising `CFG` slightly or increasing denoise — not the VAE's fault, it's being polite. ## 🎯 Best Fit Model Types This VAE was trained for, tested on, and basically _obsessed with_: | Checkpoint Name | SD Version | Match Quality | Notes | | ----------------------- | ---------- | ------------- | -------------------------------------------------- | | `wan_v2.2.safetensors` | SD 1.5 | ★★★★★ | Perfect fit | | Any SD1.5 realistic | SD 1.5 | ★★★★☆ | Works fine, just mind the colors | | SDXL or SD2.x models | SDXL/2.1 | ★☆☆☆☆ | **DO NOT** use. You’ve been warned. | | Anime models (e.g. AOM) | SD 1.5 | ★★☆☆☆ | Technically works, but there are better anime VAEs | ## 🛠️ Setup Instructions 1. **Download the VAE** [wan_2.2_vae.safetensors](https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/tree/main/split_files/vae) 2. **Place the file here:** `ComfyUI/models/vae/wan_2.2_vae.safetensors` 3. **Load the VAE in ComfyUI:** - Add a `VAELoader` node - Set `vae_name` to `"wan_2.2_vae.safetensors"` 4. **Connect it to your CheckpointLoader or VAE Decode node.** Bonus points if you use the `wan_v2.2.ckpt` checkpoint. They were literally made for each other. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - ❌ **Don't use it with SDXL models.** You’ll get garbage, and it will be _colorless_ garbage. - ❌ **Don’t assume “any VAE works with any checkpoint.”** No. VAE ≠ universal decoder. - ❌ **Don’t forget to load it.** Some checkpoints disable baked VAEs. No VAE = ugly outputs. - ❌ **Don’t rename the file and forget what it was.** You’ll end up trying 8 VAEs at once and blame your workflow instead. - ❌ **Don’t stack it with baked-in VAEs.** Pick _one_ — not both. This isn’t a VAE sandwich. ## 📚 Additional Resources - [wan_2.2 Official Examples & Documentation](https://comfyanonymous.github.io/ComfyUI_examples/wan22/) - [ComfyUI GitHub](https://github.com/comfyanonymous/ComfyUI) - [Troubleshooting ComfyUI VAEs](https://github.com/comfyanonymous/ComfyUI/discussions) ## 📎 Example Node Configuration ```json { "id": 5, "type": "VAELoader", "pos": [100, 200], "size": [300, 60], "inputs": [], "outputs": [{ "name": "VAE", "type": "VAE" }], "properties": { "vae_name": "wan_2.2_vae.safetensors" } } ``` **Node Chain Example (simple):** ```text CheckpointLoader → VAELoader → KSampler → VAEDecode → SaveImage ``` **Node Chain Example (advanced upscale):** ```text CheckpointLoader → VAELoader → UltimateSDUpscale → VAEDecode → ImageSharpen → SaveImage ``` ## 📝 Notes - This VAE was part of the broader **wan2.2 workflow release**, which included color-sane generation, control module compatibility, and cleaner transitions between tiles. - While you _can_ experiment with other VAE files, if you're using `wan_v2.2`, this is the **only VAE** that will show you its full potential. - The color science behind this VAE is genuinely impressive. It doesn't just decode — it interprets. --- ## Apply Style Model # Apply Style Model Welcome to the secret sauce of stylistic coherence. The `Apply Style Model` node (internally referred to as `StyleModelApply`) takes your boring, unflavored `conditioning` and injects it with a concentrated blast of artistic flair using a reference image and a pre-trained style model. So if you’ve ever thought, “This image looks cool, but can it look more _like that_?”, this is the node for you. ## 📦 What This Node Does The `Apply Style Model` node enhances your conditioning data by applying a style model derived from a CLIP-encoded reference image. It doesn’t touch your prompt directly—it _influences_ it. Think of it as whispering a visual moodboard into the AI’s ear before it starts painting. By injecting style information into the conditioning stream, you get generations that feel more cohesive, more on-brand, and—let’s be real—just plain better. ![Apply Style Model](/img/apply-style-model.png) ## 🔌 Inputs | Name | Type | Description | | -------------------- | ------------------ | ----------------------------------------------------------------------------------------------------------------------------- | | `conditioning` | Conditioning | The base prompt conditioning. This is your control signal for the AI, before any styling is applied. | | `style_model` | Style Model | The pre-trained style model (`.ckpt` or `.safetensors`) that knows what "style" looks like. Must contain a `style_embedding`. | | `clip_vision_output` | CLIP Vision Output | Output from a CLIP Vision encoder node. This is the style reference image, encoded into a format the style model understands. | ## 🎨 Outputs | Name | Type | Description | | -------------- | ------------ | -------------------------------------------------------------------------------------------- | | `conditioning` | Conditioning | Your original conditioning, now dressed in its Sunday best. Stylized and ready for sampling. | ## ⚙️ Settings & Parameters (Explained) Let’s break this down: ### `conditioning` This is the raw, unstyled conditioning information from your prompt or prior conditioning nodes. It’s the “what” of your generation, and this node helps shape the “how.” - ✅ Required? Yes - 📌 Comes from: A Prompt node, Style Conditioning node, or similar. - 🔍 Why it matters: This is the foundation. If it’s weak, inconsistent, or noisy, styling it won’t help. ### `style_model` A pre-trained style embedding file. This is what transforms your vanilla conditioning into something worthy of a portfolio. - ✅ Required? Yes - 📁 Must contain: a `style_embedding` key - 🧠 Think of it as: A compressed stylistic fingerprint trained from visual data - ⚠️ If you get an error about a missing key, your file is probably not a real style model. ### `clip_vision_output` Output from a `CLIP Vision Encode` node. Represents a style image as an embedding vector that your style model can digest. - ✅ Required? Absolutely - 🎯 Purpose: This tells the style model _which_ style to apply from its learned embedding space. - 🖼️ Best practice: Use a clean, high-res image that strongly reflects the style you want to transfer. ## 🧠 Recommended Use Cases - Applying the style of a reference image to generations across a series - Maintaining consistent visual tone in a multi-image workflow - Stylizing based on actual visual moodboards rather than 10 paragraphs of prompt copypasta - Generating variations of a concept while preserving artistic identity ## 🔁 Example Workflow Setup 1. 🖼️ Load an image and encode it with `CLIP Vision Encode` 2. 🧠 Load your `Style Model` using the `Load Style Model` node 3. 📝 Create prompt conditioning (text prompt → conditioning) 4. 🎨 Apply the style with `Apply Style Model` 5. 🔄 Feed the new `conditioning` into a sampler like `KSampler` and render your styled image ## 💡 Prompting Tips - Let the style model do the heavy lifting for aesthetic—don’t overcompensate with excessive prompt descriptors. - Want a touch of style instead of full commitment? Consider blending styled and unstylized conditioning. - Use consistent reference images if you’re going for a themed batch—AI doesn’t do nuance unless you spoon-feed it. ## 🧯 What-Not-To-Do-Unless-You-Want-A-Fire - ❌ Don’t feed in a style model that’s not actually a style model (missing `style_embedding`) - ❌ Don’t mismatch your `clip_vision_output` and your style model—the results will be ugly or broken (or both) - ❌ Don’t pass a `None` or broken `clip_vision_output` or you’ll meet the dreaded `'NoneType' has no attribute 'flatten'` error - ❌ Don’t expect miracles if your base `conditioning` is garbage. Style can’t polish a turd (though it will try) ## 🧱 Common Errors & Fixes | Error Message | Translation | Fix | | ----------------------------------------------------------------- | ----------------------------------------------------- | ----------------------------------------------------------- | | `invalid style model ` | Your style model is missing a `style_embedding` | Use a proper style model file | | `AttributeError: 'NoneType' object has no attribute 'flatten'` | `clip_vision_output` is empty or invalid | Check your CLIP Vision node; verify input image is loaded | | `RuntimeError: Sizes of tensors must match except in dimension 1` | Your style model and conditioning tensors don’t align | Make sure all inputs come from compatible models/components | ## 📝 Final Notes - The `Apply Style Model` node is a powerful enhancement tool—think of it as applying makeup with an airbrush instead of a crayon. - Different style models behave differently, and some are extremely opinionated. Try several before committing to one. - For best results: Pair high-quality CLIP encodings with thoughtfully designed conditioning and prompts. - You’ll get more reliable results if you ensure that your CLIP model and style model are aligned (i.e., trained to work together). --- ## BasicGuider # BasicGuider > _Because your model needs a little more direction than your average intern on a Monday morning._ --- ## 🧠 What Does This Node Do? The `BasicGuider` node in ComfyUI is responsible for—shocker—guiding. Specifically, it connects a loaded **AI model** with **conditioning inputs** to create a "Guider" object. This object then tells the model _what to do_ and _how to do it_, injecting structure and control into your generation process. Think of it as the creative director that makes sure your AI doesn’t wander off into surrealist chaos (unless that’s what you’re going for, of course). ![BasicGuider](/img/basicguider.png) ## 🧩 Node Type **Category:** Model Configuration / Utility **Node Type:** Functional Component **Input Required:** Yes **Output Generated:** Yes ## 🔌 Inputs Let’s break down what you need to feed this node so it doesn’t sulk in a corner. ### `model` (REQUIRED) - **Type:** Model (loaded via `Load Checkpoint` or similar) - **What it does:** This is your base generative model. It’s the engine that’ll produce your art based on the instructions it gets from the guider. - **Why it matters:** Without a model, the guider has nothing to guide. So unless you’re into guiding the void (existential crisis, anyone?), don’t skip this. - **Tips:** - Use a model compatible with your workflow. Not all checkpoints play nice with every sampler. - SD1.5, SDXL, SD3 — choose based on your conditioning input type and goals. ### `conditioning` (REQUIRED) - **Type:** Conditioning object (text prompt embedding, CLIP, or similar) - **What it does:** Supplies the "rules of engagement" for the model. This is your way of whispering, “Make it cyberpunk but with raccoons.” - **Why it matters:** Conditioning ensures your outputs are aligned with your prompts, not just some generative soup of randomness. - **Tips:** - The quality of your conditioning drastically affects output quality. - Combine multiple conditioning inputs for blended styles. - Need pure chaos? Try intentionally misaligned conditioning. We dare you. ## 🚀 Output ### `GUIDER` - **Type:** Guider object - **What it is:** A bundled object that pairs the model and its conditioning, ready for use in downstream nodes like samplers (`KSampler`, `SamplerCustomAdvanced`, etc.). - **Why it matters:** It’s the only way to pass your carefully crafted artistic intent to the model generation process. Without this, it’s just raw data sitting there doing nothing. - **Use it with:** Anything that asks for a `guider` input, including advanced workflows. ## 💡 Recommended Use Cases - **Prompt-based generation workflows:** Great for standard text-to-image workflows where the prompt heavily influences the style and subject matter. - **Multi-stage pipelines:** Use in setups where you’re refining conditioning mid-pipeline. - **Controlled experimentation:** Want to compare how two different models interpret the same prompt? Pair two `BasicGuider` nodes with different models but the same conditioning. ## ⚙️ Workflow Setup Example ```plaintext [Load Checkpoint] --> [BasicGuider] --> [KSampler] ↑ [CLIP Text Encode] ``` 1. Load your model. 2. Encode your prompt using CLIP or another text encoder. 3. Plug both into `BasicGuider`. 4. Feed the guider into a sampler. 5. Magic. (Or debugging, depending on your luck.) ## 📌 Prompting Tips - Be **specific**. “A blue robot dancing in Times Square” > “robot”. - Combine textual and visual conditioning if supported by your model. - Try mixing contradictory concepts if you’re into visual chaos (e.g., “medieval neon skyscraper”). ## ❌ What-Not-To-Do-Unless-You-Want-a-Fire™ - 🔥 **Don’t skip the model input.** This is not a philosophy class; the void doesn’t generate art. - 🔥 **Don’t reuse outdated or mismatched conditioning.** If you're using SDXL and the conditioning is for SD1.5, don’t be surprised when your robot turns into abstract soup. - 🔥 **Don’t plug the GUIDER into something that doesn’t accept it.** Read your nodes, folks. ## ⚠️ Known Issues - **Missing inputs:** If you forget the `model` or `conditioning`, the node will either fail silently or throw an error like: > `Error: Model not specified` > or > `Error: Conditioning not specified` > Solution? Plug the cables in like it’s your first time building IKEA furniture—**carefully**. - **Compatibility mismatches:** Make sure your model and conditioning object are from the same universe (SD1.5 + SD1.5, SDXL + SDXL, etc.). ## 🏁 Final Thoughts The `BasicGuider` is simple but essential. It’s the glue that binds your intent (conditioning) with your firepower (the model). Whether you’re generating dreamy landscapes, neon-lit nightmare fuel, or photorealistic portraits of cats playing chess, this node makes sure the rest of your workflow knows what you’re _actually_ asking for. Use it right, and it’ll feel like your AI model is reading your mind. Use it wrong, and it’ll feel like your AI model is watching your dreams and judging you. --- ## BasicScheduler # BasicScheduler The **BasicScheduler** node in ComfyUI is exactly what it sounds like: the unglamorous but absolutely essential part of your workflow that decides how sigma values (the backbone of the denoising process) are generated. You can think of sigma values as the "recipe timings" in your AI cooking process — get them right, and you end up with a perfectly baked masterpiece; get them wrong, and you’ve just invented “abstract noise soup.” By fine-tuning **steps**, **denoise**, and your **scheduler type**, you control how the model progressively cleans up noise while keeping the style and details intact. This node is especially useful when you want consistent, predictable results or you’re experimenting with different artistic effects. ![BasicScheduler](/img/basicscheduler.png) ## 📦 Function in ComfyUI Workflows The **BasicScheduler** generates a sequence of sigma values for the denoising sampler. These sigma values: - Determine the intensity of noise at each step. - Influence how detail is introduced or preserved. - Dictate how “smooth” or “textured” your final output will be. In short, the BasicScheduler is the _director_ of your denoising play. The sampler and model are the actors, but without a good director, the performance gets messy fast. ## ⚙️ Input Parameters (In Painfully Necessary Detail) ### **model** - **What it is:** The diffusion model that will interpret these sigma values. - **Why it matters:** Sigma generation must match the model’s training expectations. A mismatch can produce muddy results or, in extreme cases, complete garbage. - **Requirement:** Must be a compatible diffusion model (e.g., Stable Diffusion checkpoints, SDXL, or model types that support the selected scheduler). - **Change effects:** Switching models while keeping the same scheduler/steps may drastically alter style and detail. ### **scheduler** - **What it is:** The algorithm that decides _how_ sigma values change over the course of the denoising process. - **Why it matters:** Different schedulers have different noise decay curves, which can produce drastically different artistic results. - **Common options:** `normal`, `karras`, `exponential`, `sgm_uniform`, `simple`, `ddim_uniform`, `beta`, `linear_quadratic`, `kl_optimal`. - **Change effects:** - **normal:** Balanced, predictable outputs — your “vanilla” choice. - **karras:** Great for detail preservation in the mid and final steps. - **exponential:** Aggressive early noise removal; sharper outputs. - **Others:** Check our [Scheduler Documentation](https://comfyui.dev/docs/guides/Other%20Resources/scheduler-options) for deep dives on each. ### **steps** - **Type:** Integer - **Default:** 20 - **Range:** 1–10,000 (yes, you _could_ go 10k… but you’ll regret it unless you like coffee breaks measured in hours). - **Function:** Number of denoising passes. - **Effects:** - **Lower values:** Faster generation, but less detail and sometimes more artifacts. - **Higher values:** More refined detail, smoother gradients, but slower render times. ### **denoise** - **Type:** Float - **Default:** 1.0 - **Range:** 0.0–1.0 (in 0.01 increments) - **Function:** Controls the strength of noise removal. - **Effects:** - **1.0:** Full denoising — maximum transformation from initial noise. - **0.5:** Partial denoise — retains more of the original noise pattern for a grainier/textured style. - **Below 0.3:** Subtle tweaks to existing images without major style changes. ## 📤 Output Parameters ### **SIGMAS** - **Type:** Tensor sequence of floats. - **Purpose:** Defines the exact sigma values per step, used by the sampler to know _how much_ noise to remove at each point. - **Importance:** Without this sequence, your sampler has no roadmap, and your model just flails around in the noise like a toddler with finger paints. ## 💡 Recommended Use Cases - **High-quality render tuning:** Increase steps and tweak scheduler for maximal detail. - **Stylized outputs:** Use lower denoise values to keep intentional noise for a painterly look. - **Fast previews:** Drop steps to 5–10 to iterate ideas quickly. - **Inpainting / image-to-image:** Lower denoise values (0.2–0.5) for subtle edits without overwriting the entire composition. ## 🛠 Workflow Setup A typical placement is: ```nginx `Model → BasicScheduler → Sampler → Output` ``` - **Model:** Supplies diffusion backbone. - **BasicScheduler:** Generates sigma values. - **Sampler:** Uses sigma sequence to denoise progressively. ## 📝 Prompting Tips - When using **low denoise** for img2img, keep prompts minimal — overloading detail can cause mismatched styles. - If you’re chasing ultra-sharp detail, pair `karras` scheduler with `dpmpp_2m` or `dpmpp_3m_sde` samplers. - For cinematic softness, try `exponential` scheduler with fewer steps and denoise ~0.8. ## 🔥 What-Not-to-Do-Unless-You-Want-a-Fire - **Set steps to 10,000 for a 1024×1024 render** — unless you _enjoy_ GPU coil whine and 2-hour waits. - **Mix incompatible schedulers/models** — results range from bland mush to complete black frames. - **Set denoise to 0.0** — congratulations, you just told the node to do _nothing_. ## ⚠️ Known Issues - **Denoise < 0.1** can sometimes lead to almost imperceptible changes — not a bug, just physics. - **Unsupported model types** will throw _TypeError: Model object is not compatible_. - **Extreme step counts** may hit VRAM limits or freeze on lower-end GPUs. ## 📌 Final Notes The **BasicScheduler** isn’t the flashy part of your pipeline, but it’s the control lever for how your image _evolves_. Treat it like seasoning: a pinch can make magic, but dumping the whole jar in will ruin dinner. --- ## CLIP Set Last Layer # CLIP Set Last Layer Welcome to the node that lets you mess with CLIP in the most surgical way possible — `CLIPSetLastLayer`. This node gives you control over how far the CLIP model should go before stopping. If you’ve ever found yourself screaming at your workflow because the text encoder was doing _too much_ (or not enough), this node is for you. ## 🧠 What This Node Does The `CLIPSetLastLayer` node allows you to truncate the CLIP text encoder at a specific layer of its transformer stack — a way to _intentionally_ limit or manipulate how textual embeddings are generated. This can be useful for stylistic control, fine-tuned prompt injection, LoRA manipulation, and other advanced workflows where you want more creative or interpretive control over how prompts are processed. ![CLIP Set Last Layer](/img/clip-set-last-layer.png) ## ⚙️ Node Type **Name**: `CLIPSetLastLayer` **Category**: Conditioning / Utility **Module Type**: Transform node for CLIP objects **Purpose**: Modify a CLIP model’s processing depth for downstream use ## 🔌 Node Inputs and Outputs ### ▶️ Inputs #### `CLIP` (clip input) - **Type**: `CLIP` - **Required**: ✅ Yes - **Description**: This is your original CLIP model object, typically output from a `CheckpointLoaderSimple` or `CLIP Loader` node. - **Purpose**: This is the model you’ll be altering. You’re telling ComfyUI, “Hey, only run this CLIP model up to _here_, not all the way.” ### ⚙️ Parameters #### `stop_at_clip_layer` - **Type**: Integer (slider or manual input) - **Default**: `None` (runs full model) - **Range**: 0 – Max layers in the specific CLIP variant (usually up to 12 for OpenCLIP-ViT-G, 12 for ViT-B/32, etc.) - **Description**: Sets the transformer layer at which to stop the CLIP text encoder. ##### 🔍 Detailed Breakdown: - **`0`** – Only the embedding layer is used (like prompt surgery with a butter knife). - **`1–11`** – Gradually processes more of the transformer stack, introducing richer contextualization with each layer. - **`12`** – Full transformer run (essentially equivalent to not using this node at all). ##### 💡 Why It Matters: CLIP’s transformer layers are where the _interpretation magic_ happens. Earlier layers encode simpler, more literal meanings; later layers add complexity, abstraction, and sometimes... chaos. Stopping early can retain prompt clarity, while going deeper can let CLIP do more “interpretive dancing” with your input. ### ⏭️ Output #### `CLIP` (clip output) - **Type**: `CLIP` - **Description**: The modified CLIP model that now truncates its execution at your specified layer. This can be passed into a text-to-conditioning node like `CLIPTextEncode` or `CLIPTextEncodeAdvanced`. ## 🧩 Recommended Use Cases | Use Case | Why It Works | | ------------------------------- | ------------------------------------------------------------------------------------------ | | 🔁 **LoRA/Prompt Fusion** | Prevents over-processing of embeddings when blending prompt styles. | | 🎨 **Prompt Stylization** | Helps generate more literal or more abstract interpretations, depending on how far you go. | | 🧪 **Embedding Experiments** | Great for researchers and prompt nerds wanting to see how lower layers affect generation. | | 🛠️ **Custom Embedding Control** | Used with custom prompt tokens or per-layer manipulations. | ## 🧵 Workflow Setup Example Here's how you might wire this up in a typical text-to-image flow: ```objectivec CheckpointLoaderSimple → CLIP → CLIPSetLastLayer → CLIPTextEncode → KSampler → Image Output ``` ### Optional: - Insert a `CLIPSetLastLayer` _before_ each `CLIPTextEncode` if you're processing different prompts with different depths. - You can use multiple instances to experiment with varying depth levels side-by-side. ## 🎯 Prompting Tips - If your outputs seem too "interpretive" (like it's ignoring your prompt or going rogue), **lower the stop layer** (try 9 or 8). - If your images feel too literal or lack creativity, **raise the stop layer** (10–12). - Combine this node with prompt weights (e.g., `beautiful woman:1.4`) to refine the effect. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire Let’s be honest — this node hands you a scalpel, not a safety spoon. If you go poking around CLIP’s innards without understanding what you’re slicing, don’t act surprised when your workflow catches metaphorical fire. Here’s what **not** to do: #### ❌ Set `stop_at_clip_layer` Higher Than the Model's Actual Depth You want an error? Because this is how you get an error. ViT-B/32 has 12 layers. If you set it to 16, ComfyUI will either: - Politely crash, - Or silently fail while giving you _no clue why your outputs are garbage_. **🔥 Pro tip**: Know your CLIP variant. Google it if you have to. #### ❌ Assume This Node “Enhances” Prompts This node **removes** parts of the CLIP transformer. It’s not a magic enhancer. If you’re expecting it to make prompts “better,” you’re doing it wrong. Use it when you _want less abstraction_, not more. #### ❌ Forget to Connect the Modified CLIP to the Next Nodes If you route the original `CLIP` model into your `CLIPTextEncode`, you’ve basically skipped this node entirely. Then you’ll spend 30 minutes yelling at your screen wondering why nothing changed. #### ❌ Use the Same Truncation for Every Prompt Truncating CLIP to layer 6 might work great for `cyberpunk robot`, but try that with something like `misty forest with sunlight filtering through` and you’ll get sad trees and zero photons. Not every prompt reacts the same way — test before committing. #### ❌ Expect Downstream Nodes to Handle It Gracefully Some workflows (especially custom or complex ones) assume the full CLIP stack is intact. Truncating it too early might result in: - Poor conditioning, - Weak generation, - Or completely blank outputs that make you question your life choices. #### ❌ Use It Without Documenting Why You’re doing advanced surgery here. Future-you or a teammate will thank you for writing something like: > “CLIPSetLastLayer truncates to layer 9 here to retain literalness in prompt interpretation.” Otherwise, when things break — and they will — you’ll have no idea where to look. #### ❌ Think You’re Above Testing You _will_ need to run A/B tests. Layer 10 vs. 12? 6 vs. 8? You won’t “just know.” If you don’t run test renders with fixed seeds, you’re basically operating CLIP like it’s a roulette wheel. Bottom line: use this node like you’re diffusing a bomb. One bad move and your beautifully orchestrated workflow turns into avant-garde AI spaghetti. 🧨 ## 🧠 Final Thoughts The `CLIPSetLastLayer` node is like giving your prompt a leash — long or short, depending on how wild you want your CLIP to get. It's not for beginners, but for those who want to tune every knob and flip every switch in ComfyUI’s ecosystem, it’s a powerful little lever to throw. So go ahead — cut CLIP off mid-sentence and see what happens. Sometimes, less is more. Or at least, weirder in a good way. --- ## CLIP Text Encode (Prompt) # CLIP Text Encode (Prompt) Welcome to the beautiful mess of natural language encoding in machine learning, where “a fox wearing sunglasses in the style of Blade Runner” is magically converted into something the model can actually _understand_. The `CLIP Text Encode (prompt)` node in ComfyUI is your front door to this black box of sorcery. --- ## 🧠 What Does This Node Do? The `CLIP Text Encode (prompt)` node takes human-readable text prompts and encodes them into a numerical representation (also called an _embedding_) using the **CLIP (Contrastive Language–Image Pretraining)** model. This embedding is what downstream nodes use to guide image generation. In other words, this node turns “cyberpunk samurai with glowing katana” into multi-dimensional fairy dust that the diffusion model will happily interpret as art. No, it doesn’t make coffee — yet. ![CLIP Text Encode](/img/clip-text-encode.png) ## 🔧 Inputs ### • `clip` (CLIP model) - **Type**: `CLIP` - **Required**: Yes - **Description**: The CLIP model used to process the text prompt. This typically comes from the `Load Checkpoint` node or can be overridden using a `CLIP Set Last Layer` node. - **Gotchas**: - Must be a CLIP model that is compatible with the checkpoint used in your pipeline. - Mismatching CLIP and Checkpoint will result in output weirdness or complete garbage. You've been warned. ## 📤 Outputs ### • `CONDITIONING` - **Type**: `CONDITIONING` - **Description**: The encoded result of your text prompt, which is used by samplers (like `KSampler`) to generate the actual image. ## 🛠️ Settings & Parameters ### • `text` (Prompt Input) - **Type**: `string` - **Required**: Yes - **Description**: This is your prompt — the creative command center of your image. This string is sent through the CLIP model for embedding. - **Tips**: - More descriptive prompts yield better results (usually). - Use commas to break concepts cleanly: `portrait, cyberpunk lighting, intense gaze, soft shadows`. - Avoid overly complex grammar. The model’s not writing a novel — it just needs concept clarity. ## 💡 Recommended Use Cases - **Standard Prompt Encoding**: Feeding your primary prompt into the diffusion pipeline. - **Multi-Condition Workflows**: Combine multiple encoded prompts using nodes like `Combine Conditioning`. - **CLIP Layer Customization**: Works with `CLIP Set Last Layer` for advanced prompt tuning via specific layer truncation. ## 🔄 Workflow Setup Example Here’s a simple chain to illustrate how this node fits into your workflow: ```pgsql `[Load Checkpoint] ↓ [CLIP] ↓ [CLIP Text Encode (prompt)] ↓ [KSampler or Sampler]` ``` For workflows using both positive and negative prompts: ```css `[CLIP Text Encode (prompt)] [CLIP Text Encode (negative prompt)] ↓ ↓ [CONDITIONING] [CONDITIONING] ↓ ↓ [KSampler: prompt / negative prompt inputs]` ``` ## ✨ Prompting Tips - **Keyword Order Matters**: `a blue robot with wings` ≠ `wings with a blue robot`. - **Use Art Style Tags**: Things like `oil painting`, `low-poly`, `cinematic lighting` can drastically affect results. - **Weighting**: While the base node doesn’t support weights in text directly, you can modulate prompt strength using LoRA or prompt mixing techniques. - **Negatives**: Pair with `CLIP Text Encode (negative prompt)` to tell the model what you _don’t_ want (e.g., “ugly, blurry, extra limbs”). ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire Welcome to the section where we lovingly walk you through the most common (and catastrophic) mistakes made with this node. Do these at your own risk. Or better yet — don’t. #### 🚫 Mismatch the CLIP Model and Checkpoint **Why it's a problem:** CLIP embeddings are not universal. Using a CLIP model that doesn’t match your active checkpoint is like asking an Italian chef to cook sushi — wrong tools, wrong expectations, disaster imminent. **What happens:** - Weird generations - Model ignores parts of the prompt - Images that look like AI forgot what it was doing halfway through **Solution:** Always use the `CLIP` output that matches your checkpoint from the `Load Checkpoint` node — don’t get clever unless you _really_ know what you’re doing. #### 🧪 Expect Prompt Weighting to Work Magically **Why it's a problem:** Typing something like `((cyberpunk city:1.4))` and expecting it to just "get it" will not end well. **What happens:** CLIP will treat it as a literal phrase — not weighted — so your generation might be overrun with weird syntax interpretation or simply ignore weighting cues altogether. **Solution:** Use proper conditioning combination nodes (like `Combine Conditioning`) or external prompt weighting techniques if you need nuanced emphasis. #### 💥 Assume This Node Understands Grammar Like Shakespeare **Why it's a problem:** CLIP isn’t parsing sentence structure the way humans do. Complex grammar and nested ideas confuse it. **What happens:** You get images where a “man holding a cat wearing a hat riding a horse” becomes a terrifying three-headed creature. **Solution:** Break your prompt into short, clear, comma-separated concepts. Think of it like feeding a toddler — one idea at a time. #### 🔧 Ignore Layer Tweaking with CLIP Set Last Layer **Why it's a problem:** If you're using CLIP Set Last Layer and don’t understand what truncating to layer 6 vs. layer 12 means, you might unknowingly sabotage your outputs. **What happens:** You get "less semantic" or "more literal" generations than expected, and you have no idea why. **Solution:** Only truncate CLIP layers if you’re customizing behavior with intent. Otherwise, leave it alone — the defaults exist for a reason. #### 🧩 Forget the Difference Between Positive and Negative Conditioning **Why it's a problem:** Confusing this node with its evil twin — `CLIP Text Encode (negative prompt)` — leads to flipped intentions. **What happens:** Your image gets worse the more you try to make it better. **Solution:** Keep your `CLIP Text Encode (prompt)` for the _stuff you want_ and its sibling for the _stuff you don't_. Keep them in their lanes. #### 🪤 Use Unicode, Emojis, or Fancy Punctuation **Why it's a problem:** CLIP isn't winning a Unicode beauty pageant. Emojis and special characters can break tokenization. **What happens:** - The prompt becomes unrecognizable - You get output that’s oddly irrelevant - 🦄 suddenly turns into...a toaster? **Solution:** Stick to plain ASCII text. Keep it simple, clean, and emoji-free. Bottom line? If you want your prompt to sing and not explode into an interpretive mess, stick to compatible models, clear phrases, and intentional design. Otherwise… well, enjoy the fire. 🔥 ## 🧪 Advanced Notes - **Layer Tweaking with CLIP Set Last Layer**: You can truncate the CLIP encoding to use only up to a specific transformer layer. Lower layers favor literal/textual understanding, while higher layers emphasize more abstract/semantic interpretations. - **Reuse for Prompt Injection**: Useful in scenarios where you want to encode control prompts for ControlNet or multi-modal workflows. ## 🧼 TL;DR | Feature | Summary | | ----------------- | ------------------------------------------------------- | | Primary Role | Converts text to CLIP embeddings for guiding generation | | Input | `clip` (CLIP model), text prompt | | Output | `CONDITIONING` | | Required for | Almost every image generation workflow | | Supports weights? | Not natively, but pair with LoRA or Combine nodes | | Known Issues | Checkpoint/CLIP mismatch, no sub-prompt weighting | --- ## CLIP Vision Encode # CLIP Vision Encode _**"When your image needs to speak fluent CLIP, this is the translator."**_ The **CLIP Vision Encode** node is a powerful utility that encodes an image into CLIP Vision's latent embedding space. This allows you to leverage CLIP’s understanding of visual content for all kinds of fun and/or chaotic AI generation tasks — from style transfer and similarity search to image-to-image guidance and multi-modal workflows. This node takes your input image, compresses it via a VAE (if needed), optionally augments it, and runs it through a CLIP Vision model to output image embeddings — both positive and negative — plus latent samples if you need to push things further. Essential for AI artists who don’t just want pretty pixels — they want semantically meaningful ones. ![CLIP Vision Encode](/img/clip-vision-encode.png) ## 🧩 Node Inputs | **Input** | **Type** | **Description** | | -------------------- | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | | `clip_vision` | CLIP Vision | The loaded CLIP Vision model. Must be initialized with a `CLIP Vision Loader`. This is the core encoder that makes the magic happen. | | `init_image` | IMAGE | The image you want to encode. Garbage in, garbage embeddings out — high-quality images are highly recommended. | | `vae` | VAE | Variational Autoencoder used to encode the image into latent space. Required for downstream workflows that operate in latent format. | | `width` | INT | Width (in pixels) to resize the image before encoding. Must match model expectations. Default is typically 512. | | `height` | INT | Height (in pixels) to resize the image before encoding. Use the same guidance as width. | | `video_frames` | INT | (Optional) Number of frames to generate for video workflows. Only needed if you’re encoding frame sequences. | | `motion_bucket_id` | INT | (Optional) Identifier used for motion sequence conditioning in video tasks. Groups sequences together. | | `fps` | INT | (Optional) Frames per second metadata for video playback. | | `augmentation_level` | FLOAT (0.0–1.0) | Controls how much noise/augmentation is applied during encoding. Helps improve generalization. Too high and you’ll just confuse the model. | ## 🎯 Node Outputs | **Output** | **Type** | **Description** | | ---------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------- | | `positive` | LIST | Embeddings and metadata representing positive conditioning. These are used to guide generation toward your input image’s features. | | `negative` | LIST | Embeddings and metadata representing negative conditioning. Great for telling the model what _not_ to do. | | `samples` | LATENT | Latent space tensor derived from the input image. Feed this into other nodes for diffusion, transformation, or image generation. | ## 🛠️ Recommended Use Cases - **Image-to-text alignment**: Use embeddings to match images with prompts. - **Image-guided generation**: Feed the `positive` output into workflows where you want image features to guide the result. - **Reference style matching**: Encode the "vibe" of an image and apply it elsewhere. - **Training augmentation**: Use `augmentation_level` to simulate variation in the same input for robust downstream training. - **Video frame encoding**: Turn sequences of frames into CLIP embeddings for multi-frame workflows. ## ⚙️ Usage Tips - Keep your `init_image` clean and high-quality. CLIP is good, not psychic. - Use matching dimensions for `width` and `height`. Usually divisible by 8 or 64 — 512x512 and 768x768 are safe bets. - For multi-modal workflows, try moderate `augmentation_level` values (0.2–0.4). Higher values inject useful noise but can distort the image’s intent. - When using video-related inputs, ensure `video_frames` > 1 and your CLIP model can handle batch inputs. - Not using video? Leave `video_frames`, `motion_bucket_id`, and `fps` at default or 0. They won’t hurt anything. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - **❌ Feed in mismatched resolutions**: If the width and height don’t play nicely with your CLIP model, expect crashes or distorted embeddings. - **❌ Skip the VAE**: You need it for anything latent-related. No VAE, no samples. - **❌ Abuse augmentation**: Setting `augmentation_level` to 1.0 will mangle your image into unrecognizable mush. Not ideal unless you’re trying to invent glitchcore. - **❌ Use negative embeddings without understanding them**: Negative outputs are not the opposite of positive — they’re meant for contrastive use in workflows that support them. ## ⚠️ Known Issues - **VRAM usage** can spike if you feed in large image resolutions or lots of video frames. Resize before encoding if needed. - Some **CLIP Vision models** are picky — make sure the one you loaded supports the resolution you’re using. - **Augmentation noise** isn’t standardized. What’s “moderate” for one model might be “absolute chaos” for another. ## 🧪 Example Node Setup ```json { "clip_vision": "ViT-H-14-CLIP-Vision", "init_image": "reference_image.png", "vae": "vae_model.safetensors", "width": 512, "height": 512, "video_frames": 0, "motion_bucket_id": 0, "fps": 0, "augmentation_level": 0.3 } ``` This setup encodes a 512x512 image into CLIP Vision space using a standard VAE with light augmentation. ## 📝 Notes - This node plays beautifully with **DualCLIP**, **Prompt Conditioners**, and **KSampler** workflows that take positive/negative embeddings. - If you're doing anything with image + text matching, this node is practically a requirement. - This is **not** a text encoder. Only visual embeddings here, folks. --- ## ConditioningZeroOut # ConditioningZeroOut > _A node that says, “You’re not relevant, goodbye.”_ ## 🧠 What Is This? The `ConditioningZeroOut` node is an **advanced ComfyUI utility** that lets you surgically _eliminate_ unwanted parts of a conditioning vector. It doesn’t sugarcoat or play favorites — it zeros out whatever isn’t needed and leaves the rest untouched like a proper minimalist. Whether you're cleaning up noisy input, reducing drift in style transfer, or just trying to tell your prompt to focus for once in its life, this node gets the job done. This is not a filter, a normalizer, or a vibe enhancer. It’s a **hard no** for unneeded data. ![ConditioningZeroOut](/img/conditioningzeroout.png) ## 🧪 Real-World Use Cases | Use Case | Why You’d Use This | | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------- | | **Image Preprocessing** | When certain parts of your conditioning data are useless (or worse, misleading), this node cleans it up before downstream nodes waste time on it. | | **Machine Learning Pipelines** | Improves training or inference by trimming irrelevant signal from the conditioning vectors. Think “conditioning diet plan.” | | **Style Transfer** | Remove conditioning values that pull the output away from the intended style anchor. Less noise = more style fidelity. | | **Data Augmentation** | Selectively zero out values to create subtle variations without starting from scratch. | ## 🔌 Node Overview | Property | Value | | ---------------- | ---------------------------------- | | **Node Name** | `ConditioningZeroOut` | | **Category** | `advanced/conditioning` | | **Input** | `conditioning` (CONDITIONING list) | | **Output** | `conditioning` (CONDITIONING list) | | **Output Node?** | No | | **Version** | ComfyUI-core compatible | ## 🧷 Input Details (You Better Get These Right) ### 🟡 conditioning - **Type:** `CONDITIONING` - **Expected Format:** List of floats, vectors, or encoded context data - **Example:** `[1, 0, 0, 1]` - **What It Does:** This is your raw, pre-filtered conditioning. The node goes through it and zeros out the stuff you (or your upstream logic) marked as unworthy. - **Why It Matters:** This is the only input. If you screw it up, this node does absolutely nothing — or worse, breaks your entire workflow. ✅ **Pro Tip:** Pair this with a custom pre-filtering node or condition classifier upstream if you want smarter zeroing decisions. ## 🟢 Output Details ### 🟢 conditioning - **Type:** `CONDITIONING` - **Example Output:** `[1, 0, 0, 0]` - **What It Is:** The cleaned-up, leaner, zeroed-out version of your input. - **What It’s For:** Ready to be passed along to samplers, ControlNet, or anywhere else that accepts conditioning input. - **Bonus:** You just saved your sampler from wasting compute on trash values. ## ⚙️ Workflow Integration Here’s how to use it _without summoning chaos_: ```css [Text Encoder] → [ConditioningZeroOut] → [KSampler or ControlNet] ``` Want to zero out specific features before mixing conditioning vectors? Add it **before** you merge or apply custom logic. Want to sharpen prompt influence? Drop it in _right after encoding_, before your sampler ever sees it. ## 🪄 Prompting Tips - **Use it when:** Outputs feel “off” or too generic — you may be feeding unnecessary context. - **Pair it with:** Prompt editors or text encoders to reduce clutter. - **Use caution with:** Strongly stylized prompts — zeroing too much can kill the magic. - **Tweak + test:** There’s no preview for what gets zeroed, so test in iterations. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire | ⚠️ Mistake | 🔥 Why It’s a Problem | | ------------------------------- | --------------------------------------------------------------------------------------------------- | | Zeroing everything | Welcome to random noise. Your image now represents _the absence of thought_. | | Feeding in the wrong input type | If you hand it a tensor, image, or string — expect hard errors. This node only eats `CONDITIONING`. | | Assuming it normalizes values | It doesn't. It **zeros**. You want normalization? Go ask the `Normalizer` node. | | Using it as a blur effect | Nope. It’s not for softening output — it’s for deleting parts of the signal. Hard stop. | ## 🐛 Known Issues - Doesn’t show what values were zeroed — you’ll need to inspect manually. - Overzealous use can make outputs too generic or lifeless. - Not compatible with every exotic conditioning format — check your encoder. - Bad pairing with overly complex prompts may strip essential context. ## 🧩 Related Nodes | Node | What It Does | Why It’s Different | | -------------------- | ------------------------------------- | ----------------------------------------------------- | | `FilterZeroOut` | Filters broader data sets | Broader in scope; not focused on conditioning vectors | | `Normalizer` | Adjusts ranges (not removes) | Leaves all elements in — just rescales | | `ConditioningConcat` | Combines multiple conditioning inputs | Not destructive — it adds, not subtracts | ## 📦 Example Node Config ```json { "id": 43, "type": "ConditioningZeroOut", "pos": [400, 200], "size": [240, 50], "inputs": { "conditioning": "Link to CLIPTextEncode" }, "outputs": { "conditioning": "→ KSampler" } } ``` ## 📝 Final Notes The `ConditioningZeroOut` node is one of those quiet little heroes that doesn’t get enough love — until your prompt starts misbehaving. By trimming the fat, you give your generation pipeline the focus and direction it deserves. Use it right, and you'll look like a workflow whisperer. Use it wrong, and you'll be wondering why your portraits look like static noise with a top hat. --- ## ControlNet Preprocessor # ControlNet Preprocessor Welcome to the wonderfully temperamental world of **ControlNet Preprocessors**, where your image dreams live or die by the edges, lines, depth maps, and semantic scribbles this node extracts. The `ControlNet Preprocessor` node in ComfyUI is your gateway to preparing an image input for ControlNet conditioning — effectively, it translates your raw images into usable maps like Canny, Depth, OpenPose, etc. that ControlNet models can understand and guide generation from. ![ControlNet Preprocessor](/img/controlnet-preprocessor.png) ## 🧩 Purpose This node is responsible for: - Selecting and applying a preprocessing algorithm to an input image. - Generating a control map (hint: not a pretty picture, but a very _useful_ one). - Preparing that map for direct use in ControlNet nodes like `Apply ControlNet` or `ControlNetLoader`. Think of this as the translator between your input image and the language your ControlNet model speaks. ## ⚙️ Node Inputs & Outputs ### **Inputs** - **IMAGE** – The image to preprocess. Needs to be in a format that matches standard RGB image data (usually from something like `Load Image`, `Image Input`, or a previous generation). ### **Outputs** - **IMAGE** – The output control map, ready to be passed into a `ControlNetApply` node or similar. ## 🔧 Node Settings (Detailed) ### `preprocessor` This dropdown lists available preprocessing algorithms. Each preprocessor generates a control map from your input image in a specific format expected by the corresponding ControlNet model. Common options include: | Name | Description | Use Case | | ---------------------------- | ------------------------------------------------------------ | -------------------------------------------------------- | | `canny` | Applies Canny edge detection. Produces a grayscale edge map. | Outlines and sketchy aesthetics. | | `depth` | Generates a depth estimation map. | Realistic lighting and 3D-aware outputs. | | `depth_leres` | Higher-quality monocular depth estimation using LeReS. | Photorealistic results with more accurate spatial depth. | | `mlsd` | Applies Line Segment Detection. | Architectural renders or structured geometry. | | `openpose` | Detects human body/keypoints. | Character poses, action shots, motion control. | | `scribble` | Converts images into scribble-style contours. | Sketch generation, cartoon-like input. | | `normal_bae` | Surface normal estimation. | Fine texture and lighting-sensitive rendering. | | `hed`, `pidinet`, `softedge` | Edge detection with different neural backends. | Better results in abstract or noisy environments. | | `segmentation` | Semantic segmentation map. | Scene understanding and targeted prompting. | 🔎 **Tip**: The `preprocessor` **must match** the ControlNet model being used. Don’t feed `canny` output into a `depth` ControlNet unless you're into chaos. Ready to go more in-depth? Check out our detailed [preprocessor option resources](http://comfyui.dev/docs/guides/Other%20Resources/preprocessor-options)! ### `sd_version` Select the version of the Stable Diffusion base model you're using: - `1.5` - `2.0` - `2.1` - `XL` (aka SDXL) 💡 Why it matters: ControlNet preprocessors and models are version-dependent. Using the wrong SD version will lead to mismatches between control signals and generation behavior — translation: _ugly results or nothing at all_. ### `resolution` Defines the resolution at which the preprocessing is applied. This value controls the **size of the generated control map** — not your output image. - **Typical Range**: `256` – `1024` - **Default**: Often `512` or `768` depending on your workflow. 📏 Higher resolution gives more detail in the control map but increases GPU load and can amplify noise or artifacts. ✂️ If your base model is SDXL, 1024 is ideal. For SD 1.5, stick to 512–768 unless you're feeling brave (or your GPU is). ### `preprocessor_override` An advanced field for custom preprocessor logic or specifying third-party / external preprocessor modules. - **Default**: Usually empty. - **What it accepts**: Overrides the default behavior of the selected `preprocessor`. 🧪 Use this if: - You're using an external preprocessor script. - You're trying a forked or experimental ControlNet. - You like voiding warranties and doing weird things. ## 🛠 Recommended Use Cases | Use Case | Preprocessor | | --------------------- | --------------------------------------------------- | | Pose control | `openpose` | | Facial keypoints | `face` (if available), `openpose` with face enabled | | Depth-based realism | `depth`, `depth_leres` | | Background flattening | `segmentation` | | Line drawing prompts | `canny`, `mlsd`, `hed`, `scribble` | | Structural precision | `mlsd`, `normal_bae` | | Abstract forms | `softedge`, `pidinet` | ## 🧩 Typical Workflow Integration Here’s how you’d use this node in a standard ControlNet setup: 1. 🔁 Input your image via a loader node (`Load Image`, `Image Input`, etc.) 2. 📐 Pipe it into `ControlNet Preprocessor` 3. 📦 Output from this node goes into `Apply ControlNet` → `KSampler` for generation ```css [Input Image] ↓ [ControlNet Preprocessor] → [Apply ControlNet] → [KSampler] ``` If you're working with multiple ControlNets, each will usually have its own preprocessor instance. ## 🚫 What-Not-To-Do-Unless-You-Want-a-Fire Oh, so you like chaos? You enjoy watching your GPU cry? Great, then here's what _not_ to do with the `preprocessor` setting unless you're actively trying to summon the AI demons of instability: #### ❌ Use the wrong preprocessor with the wrong ControlNet model You wouldn't feed a cat spaghetti and expect it to do math. Likewise, don't feed `pose_animal` output into a ControlNet trained for `depth_midas`. The result? Nonsense conditioning, wasted steps, and outputs that look like AI had an existential crisis. **Fix**: Always match your preprocessor with its sibling ControlNet (e.g., `hed` → HED model, `depth_anything` → ControlNet trained on Depth Anything). #### ❌ Forget to install dependencies Half of these preprocessors are built on third-party magic. Missing `detectron2`, `segment-anything`, `openpose`, or `opencv`? You’ll get red errors, blank images, or worse: success that isn’t actually success. **Fix**: Check your install. Use a requirements.txt file. Don’t YOLO this. #### ❌ Run high-res images through `depth_leres` or `sam` on 8GB VRAM If you're running a potato laptop with a fancy GPU sticker but no actual power, please don’t crank `depth_leres` or `sam` to 2048x2048. These models _will_ eat your VRAM and then casually torch your runtime with an out-of-memory error. **Fix**: Stay under 1024x1024 unless you’re packing real heat. #### ❌ Expect perfect outlines from `scribble_xdog` on low-contrast images Low contrast images + `xdog` = muddy soup. It’s not a “dreamlike sketch,” it’s a failed art student’s nightmare. **Fix**: Boost your image contrast before applying `xdog`. #### ❌ Use `shuffle` and expect consistency Shuffle does what it says—it shuffles. It’s not a structured preprocessor, it’s an agent of chaos. **Fix**: Don’t use it unless you want variety over control. Never in production workflows. Ever. #### ❌ Assume `pose_dense` will get every joint right If your character is lying down, twisted, or facing away from the camera, `pose_dense` might just give up entirely. Expect floating limbs and mysterious spaghetti arms. **Fix**: Stick with standard `pose` or `dwpose` for more stable results. Always validate visually. #### ❌ Mix multiple preprocessors on the same conditioning channel Unless your ControlNet expects a specific composite input (and you _really_ know what you're doing), mixing outputs like `depth` + `canny` into the same ControlNet model is like throwing oil and water into a blender—loud, messy, and completely ineffective. **Fix**: One preprocessor, one ControlNet, per channel. Keep your chaos modular. #### ❌ Skip normalization when using `normalmap_*` Feeding an unnormalized or overly bright image into a `normalmap` extractor? Get ready for washed-out normals or weird lighting shadows. **Fix**: Preprocess with tone mapping or exposure correction first. #### ❌ Rely on `seg_*` for precision mask work Semantic segmentation ≠ accurate masking. These models often blur edges or clip object boundaries. Don’t use them if you're trying to do surgical precision work like inpainting hair strands. **Fix**: Use `sam` instead. It’s designed for precision. #### ❌ Forget that more preprocessing ≠ better results Yes, we know—it’s tempting to run every image through five preprocessors, load five ControlNets, and see what happens. But you’ll probably just get noise, hallucinations, or broken anatomy. **Fix**: Be deliberate. Preprocessors are tools, not spice blends. Pick the one that suits your task, and leave the rest out of your stew. And finally: ##### 🔥 Don’t forget to laugh when it breaks This is ComfyUI. If something goes wrong and you get AI soup or a melted mannequin, remember: it’s not a bug, it’s a rite of passage. ## 🧱 Known Issues - **Compatibility Hell**: Some third-party ControlNet models expect specific preprocessors. Always check the model description. - **OpenPose & Depth tend to produce black outputs on faulty or low-contrast inputs.** Try normalizing the image or adjusting lighting. - **Preprocessor-Override crashes?** Likely due to an invalid import, path issue, or incompatible Python logic. ## 🔚 Summary | Setting | Description | Important Notes | | ----------------------- | -------------------------------------- | ------------------------------------- | | `preprocessor` | Algorithm used to extract control data | Must match ControlNet type | | `sd_version` | Specifies Stable Diffusion version | Required to align tensors | | `resolution` | Size of the control map | Higher = more detail, more VRAM | | `preprocessor_override` | Advanced override | For custom pipelines or chaos monkeys | This node is the backbone of structured, repeatable generation with ControlNet. Master it, and you’ll go from "kinda random" to "surgical precision" in your generations. --- ## DualCLIPLoader # DualCLIPLoader The **DualCLIPLoader** node is here to settle the age-old argument of _which CLIP is better_ — by just using both. This node lets you load **two CLIP models simultaneously**, giving your workflow extra perspective, better text-vision understanding, and more power than a single CLIP could ever offer alone. Whether you're running SDXL, SD3, Flux, or even Hunyuan Video, this node helps unlock the full potential of multi-CLIP workflows in ComfyUI. ![DualCLIPLoader](/img/dualcliploader.png) ## 🔧 Node Type **`DualCLIPLoader`** _Outputs: `CLIP`_ ## 📁 Function in ComfyUI Workflows The DualCLIPLoader node loads two distinct CLIP models into a single combined output object, which downstream nodes (like `CLIPTextEncode`, `KSampler`, and others) can use for better prompt conditioning. It’s especially useful in workflows where you're combining: - Vision + text encoders (e.g., SDXL-style) - Two stylistically different CLIPs (like realism and anime) - CLIPs trained on different languages or datasets - Experimental comparative generation setups It’s a smart choice when **you want more control, more nuance, and less guesswork** in how your prompt is interpreted. ## 🧠 Technical Details This node: - Loads **two `.safetensors` CLIP models** - Outputs a **combined CLIP object** - Supports architecture-specific types (`sdxl`, `sd3`, `flux`, `hunyuan_video`) - Can optionally be assigned to specific devices (like CPU or `cuda:0`) for advanced load balancing Internally, the node makes sure both models are loaded properly and tied to your workflow’s current recipe. ## ⚙️ Settings and Parameters | 🔲 Field | 💬 Description | | ------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `clip_name1` | **Filename of the first CLIP model**. This is your primary encoder, typically used for prompt conditioning. Must be a valid `.safetensors` CLIP file in your models folder. Example: `clip_l.safetensors`. | | `clip_name2` | **Filename of the second CLIP model**. Used in tandem with the first — often a vision encoder or stylistic variant. Example: `clip_vision_g.safetensors`. | | `type` | **Model recipe** this CLIP pair is being used with. Must match the type of your checkpoint and downstream pipeline. Options include: `sdxl`, `sd3`, `flux`, `hunyuan_video`. | | `device` | (Optional) **The device you want to load the CLIP models onto.** Use `"default"` for automatic assignment (usually GPU), or set to `"cpu"` or `"cuda:0"` if you’re managing memory manually. | ## ✅ Benefits - **Multi-modal power** – Combine language and vision understanding in one pass. - **Style fusion** – Blend the strengths of two different CLIPs (realism + anime, anyone?). - **More accurate prompts** – Dual interpretation gives better grounding to both positive and negative prompts. - **Better SDXL/SD3 results** – These architectures were made for dual-CLIP setups, and this node is the plug-in brain for them. ## ⚙️ Usage Tips - Always match the `type` field with the **checkpoint type** you're using. Don’t mix `sdxl` with `flux` or your generation will go sideways. - If you’re running low on VRAM, offload one model to `cpu` by setting `device` to `"cpu"` — but expect slower performance. - You can mix a vision CLIP with a language-focused CLIP to create _crazy accurate visual storytelling_. Try it with SDXL for best results. - Want to experiment with prompting style? Use one CLIP trained for realism, and another trained for fantasy — balance prompt weights accordingly. ## 📍 ComfyUI Setup Instructions 1. Place your CLIP model files (`.safetensors`) in your ComfyUI `models/clip` folder. 2. Add the **DualCLIPLoader** node to your workflow. 3. Set `clip_name1` and `clip_name2` to the exact filenames. 4. Set the `type` field to match your pipeline (`sdxl`, `flux`, etc.). 5. Optionally assign a `device` (or just leave it as `"default"`). 6. Connect the output to any node that expects a `CLIP` input. ## 📎 Example Node Configuration ```plaintext clip_name1: clip_l.safetensors clip_name2: clip_vision_g.safetensors type: sdxl device: default ``` In this setup, we’re using two different CLIPs with an SDXL-based model. This is ideal for workflows where SDXL’s text+vision conditioning is leveraged for higher fidelity generations. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - ❌ Don’t mix model types. Your `type` **must match** your checkpoint or you’ll get garbage output (or no output at all). - ❌ Don’t assign both CLIPs to GPU on low-VRAM systems. You _will_ crash. - ❌ Don’t try to load non-CLIP `.safetensors` files here. It won’t work. You’ll sit there wondering why your workflow’s frozen. - ❌ Don’t assume CLIPs “merge” into a single model — this node runs them in **parallel**, not as a fusion. - ❌ Don’t use massive CLIPs on a 4GB GPU unless you _really_ enjoy watching your system swap memory like it’s 2006. ## ⚠️ Known Issues - **VRAM hungry** – Two models = more memory. No surprise here. - **Slow on CPU** – If you offload to CPU, expect a noticeable slowdown. - **No CLIP validation** – If your filename is wrong, the node won’t tell you nicely — it’ll just break silently or downstream. ## 📚 Additional Resources - How SDXL Uses Dual CLIP Architecture - Example model files: - `clip_l.safetensors` - `clip_vision_g.safetensors` ## 📝 Notes - This node is **highly recommended** for SDXL workflows and is borderline required if you want to use SDXL the way it was meant to be used. - Great for **prompt tuning**, **multimodal alignment**, and **style fusion**. - If you're experimenting with new CLIPs, try this node in a sandbox workflow before plugging it into a production chain. --- ## Empty Latent Image # Empty Latent Image Welcome to the mysterious void—the **Empty Latent Image** node. This deceptively simple, no-frills node is the unsung hero of countless ComfyUI workflows. It doesn't load an image, it doesn't encode a prompt, and it doesn't play with your samplers. Instead, it spawns a blank canvas _directly in latent space_, ready for generation to sculpt it into anything your prompt (and sampler) desires. Whether you're initializing from scratch or want to detach completely from any prior visual influence (bye bye, conditioning!), this node is your clean slate. --- ## 🧠 What Does It Do? The `Empty Latent Image` node generates a blank latent tensor—think of it like an invisible image placeholder—of a specified size, batch count, and shape. It does not represent visible pixel data but instead prepares a tensor that will serve as the initial input to the **KSampler** node or other latent-aware modules. This is useful in workflows where you're not starting from a real image or encoded input (like `VAE Encode`) but instead want to generate something from noise based purely on a text prompt or conditioning. ![Empty Latent Image](/img/empty-latent-image.png) ## ⚙️ Node Type - **Node Name:** `EmptyLatentImage` - **Outputs:** - `LATENT`: Latent image tensor that can be passed to a KSampler or decoded via a VAE. ## 🧩 Node Parameters and Their Purpose Let's get into the knobs and sliders. This node may look basic, but these settings shape your generated outputs from the ground up. ### 🔹 `width` (integer) - **Definition:** The width (in pixels) _of the image you want to end up with_, but in latent space. - **Default:** Usually 512. - **Required:** Yes. - **Effects:** Affects the horizontal size of the latent tensor. - **Importance:** Your model's resolution matters. Mismatched width can cause warping, cropping, or wasted GPU memory. - **Changing it means:** Higher widths = higher GPU use, more detail; lower widths = faster generation, less spatial resolution. - **Tip:** Stick to multiples of 8 or 64 depending on your VAE/model combo. 512, 768, 1024—these are your friends. ### 🔹 `height` (integer) - **Definition:** The height (in pixels) of your intended image output in latent space. - **Default:** Usually 512. - **Required:** Yes. - **Effects:** Controls the vertical resolution. - **Changing it means:** Tall images? Big `height`. Square images? Match width/height. But again—**stay divisible by 8** at minimum, ideally 64. - **Tip:** If using ControlNet (e.g., depth or canny), make sure your width/height match the conditioning image dimensions. ### 🔹 `batch_size` (integer) - **Definition:** Number of blank latent images to generate in one go. - **Default:** 1 - **Required:** No (but 1 is safest if you don’t know what you’re doing). - **Effects:** Produces multiple latent tensors simultaneously. - **Changing it means:** - `1`: Classic. Single image per run. - `>1`: Great for variations, automation, or batch generation. - **Warning:** Crank this up and you’re begging for a CUDA out-of-memory error. Know your GPU limits. ## 🧪 Workflow Setup & Integration Typical usage involves this node when: - You're building **prompt-to-image** workflows from scratch. - You **don’t want to use an actual image as input**, nor do you need prior image features. - You want **total control over resolution and composition**. - You're piping into a `KSampler` (because that’s where the magic happens). ### 🔁 Basic Workflow Example ```mermaid graph TD; A[CLIP Text Encode] --> B[KSampler]; C[Empty Latent Image] --> B; B --> D[VAE Decode]; ``` 1. `CLIP Text Encode`: Processes the prompt. 2. `Empty Latent Image`: Creates the initial noise tensor. 3. `KSampler`: Transforms the latent with your model and prompt. 4. `VAE Decode`: Turns the final latent into a visible image. ## 💡 Use Cases - **Prompt-to-image generation**: Your usual starting point if you're not inpainting or using an init image. - **Batch generation**: Set `batch_size` to 4 and run multiple images at once with identical prompts/settings. - **Latent-space experimentation**: For testing new samplers, schedulers, or workflows. - **Workflow consistency**: Create reproducible results with a fixed size and structure. ## 🧠 Prompting Tips - This node doesn’t care what the prompt is—that’s up to `CLIP Text Encode`. - It’s just the stage. You still need actors (prompt) and a director (KSampler). - Don't feed this into a VAE Decode directly. That’s like printing on invisible ink. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire Look, this node may seem innocent—it’s literally generating _nothing_—but don’t underestimate how quickly it can _turn your workflow into a flaming wreck_ if misused. Here’s what _not_ to do unless you enjoy watching your VRAM beg for mercy: #### 🚫 Set `batch_size` to 16+ with 1024x1024 resolution on a 6GB GPU Unless you enjoy crashes, freezes, or your GPU sounding like it's preparing for liftoff, don’t push your batch size beyond what your hardware can actually handle. Rule of thumb: **know your VRAM limit** and stay humble. #### 🚫 Mismatch your width/height with your ControlNet or VAE If you're piping this latent image into a ControlNet that expects a 512x512 input but you set this node to 768x1024? Boom—dimension mismatch. Errors everywhere. The kind that don’t explain themselves clearly either, because ComfyUI just assumes _you_ know better. (Spoiler: it was you.) #### 🚫 Feed this directly into `VAE Decode` Yes, technically you _can_. But also—no, don’t. It’s a blank latent tensor. Decoding it without any generation step (like `KSampler`) will result in... a blurry, colorless soup. This is **not** abstract art. Just use `KSampler` first. #### 🚫 Use non-divisible-by-8 values for width or height You want to see a model cry? Try using 513x513. Some VAEs and samplers will pad it. Others will just crash silently or throw you a cryptic tensor-shape error. Unless you enjoy debugging shape mismatches, stick to **multiples of 8** (or better yet, 64). #### 🚫 Assume this node _does_ something visual If you're sitting there wondering why no image is showing up after decoding this node on its own—congrats, you've just decoded air. This node does _not_ represent pixels. It is latent space. Invisible. Abstract. Conceptual. Don’t expect pretty pictures without a full pipeline. ## 🧾 Summary | Setting | Description | Tips | | ------------ | ------------------------------------------- | ------------------------------------------------------------ | | `width` | Horizontal size of latent image (in pixels) | Stick to 512, 768, 1024. Must be divisible by 8 or 64 | | `height` | Vertical size of latent image (in pixels) | Match to width for square, or adjust to desired aspect ratio | | `batch_size` | How many latent images to generate at once | Use `>1` for batch gen, but check your GPU RAM first | ## 🧯 Final Thoughts The `Empty Latent Image` node is like a clean sheet of paper for latent artists. It doesn’t do much on its own, but paired with the right nodes—_chef’s kiss_. Treat it as your silent but essential co-pilot on any prompt-to-image adventure. --- ## EmptySD3LatentImage # EmptySD3LatentImage _**The Blank Canvas for SD3 Latent Generation**_ ## 🧠 What is this node? The **EmptySD3LatentImage** node is the digital equivalent of starting with a perfectly clean whiteboard — except instead of actual pixels, it generates a **latent tensor** in the **SD3 latent format**. Think of it as telling ComfyUI: > “Here, start with this perfectly boring, constant-valued tensor so we can build something interesting on top of it.” It’s not flashy, it’s not creative on its own, but without it, many SD3 workflows wouldn’t have a stable, predictable base to work from. ![EmptySD3LatentImage](/img/emptysd3latentimage.png) ## 🧩 Function in ComfyUI Workflows This node **creates an SD3-compatible latent tensor filled with a constant value** (`0.0609` — yes, it’s oddly specific, and no, you shouldn’t change it unless you _really_ know what you’re doing). You use it when you: - Need a **starting point** for AI art generation with **no initial image influence**. - Want a **controlled, uniform initialization** for testing prompts or pipeline changes. - Need to generate a **specific resolution latent** for downstream nodes like `KSampler`, `DecodeSD3`, or image editors. The generated tensor: - **Format:** SD3 latent (16 channels) - **Shape:** `(batch_size, channels=16, height/8, width/8)` - **Value:** Constant 0.0609 in all positions ## ⚙️ **Settings & Parameters** (Extreme Detail Mode™) ### **Width** _(pixels)_ - **Purpose:** Sets the _horizontal dimension_ of the latent (in pixel space). - **Allowed range:** `16px` → system’s max resolution (depends on GPU). - **Default:** `1024` - **Important notes:** - This is **not** the final output width — it’s the latent width, which maps to final image width depending on the model scale (SD3 scale factor: ×8). - Larger widths = more VRAM usage. If you hear your GPU fans start screaming, you’ve gone too big. ### **Height** _(pixels)_ - **Purpose:** Sets the _vertical dimension_ of the latent (in pixel space). - **Allowed range:** `16px` → system’s max resolution. - **Default:** `1024` - **Important notes:** - Works the same as Width. - Changing only height will change aspect ratio — great for vertical vs horizontal compositions. ### **Batch Size** - **Purpose:** Number of separate latent tensors generated at once. - **Allowed range:** `1` → `4096` (yes, you can, but should you? Probably not). - **Default:** `1` - **Important notes:** - Bigger batches save time for batch processing but can eat VRAM like a competitive hotdog eater. - Each batch is an independent latent — this is not tiling; it’s multiple canvases at once. ## 📤 Output ### **LATENT** _(tensor dictionary)_ - **Key:** `"samples"` → latent tensor (`batch_size × 16 × height/8 × width/8`) - **Usage:** Feeds directly into nodes expecting SD3-format latent images (e.g., `KSampler`, `Image Decode`). - **Value:** Filled with constant `0.0609`. - **Why 0.0609?** It’s a normalization choice that plays nice with SD3’s internal math. Changing it can cause subtle but _ugly_ shifts in final output. ## 💡 Recommended Use Cases - **Testing prompts in isolation** — start from scratch to ensure no residual noise from previous latents. - **Workflow prototyping** — check resolution handling before plugging in real image data. - **Custom resolution generation** — make exactly the size you want without an initial image. - **Consistent batch generation** — same starting tensor for all images in a batch for reproducibility. ## 🛠 Workflow Setup 1. Drop `EmptySD3LatentImage` at the **start** of your pipeline. 2. Set `width` and `height` to match your desired output resolution (remember SD3 upscale factor). 3. Set `batch_size` if you want multiple images at once. 4. Connect **LATENT** output into a `KSampler` or similar node. 5. Decode results later with `DecodeSD3`. ## 🎯 Prompting Tips - Since this starts from **nothing but constants**, your prompt has _maximum influence_. This is great for prompt testing — bad prompts will show themselves instantly. - If outputs look _too_ uniform or “flat,” you might actually want to start from random noise instead — this node is intentionally predictable. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - **Do NOT** crank `width` and `height` to max _and_ `batch_size` to 4096 unless you enjoy system crashes and GPU restarts. - **Do NOT** change the constant value in the latent unless you want to play “guess why my images look weird.” - **Do NOT** expect “variation” from this node — it’s designed to be boring. Use noise or source images for variety. - **Do NOT** mismatch resolutions between this node and downstream decoders — SD3 expects specific scaling. ## ⚠️ Known Issues - **Out-of-VRAM Errors**: Happens when going too high on resolution or batch size. - **Accidental wrong aspect ratios**: Caused by forgetting SD3’s ×8 scale factor when setting width/height. - **Flat outputs when misused**: If you expect randomness, you’re in the wrong node. ## 📝 Final Notes The **EmptySD3LatentImage** is an essential “blank start” node for SD3 workflows. It’s not glamorous, but without it, testing and controlled generation in SD3 would be a nightmare. Use it for stability, reproducibility, and as a clean base for prompt-driven creativity. --- ## Flux.1 Kontext Image Edit # Flux.1 Kontext Image Edit _When your latent needs a glow-up — with precision, style, and zero tolerance for bad prompts._ --- ## 🧩 What is Flux.1 Kontext Image Edit? The **Flux.1 Kontext Image Edit** node lets you edit images with _surgical-level precision_ by manipulating latent space using a prompt. Unlike traditional text-to-image models, this one starts from a **LATENT** input — not a blank canvas — allowing you to transform an existing image while preserving its structure and composition. It supports both `flux-kontext-pro` and `flux-kontext-max` models, and outputs both a freshly generated **IMAGE** and the updated **LATENT**, making it ideal for chaining edits across multiple stages of a workflow. Think of it as the "Photoshop Liquify" tool, but powered by diffusion and a few thousand GPU cycles. ![Flux.1 Kontext Image Edit](/img/flux-1-kontext-image-edit.png) ## 🚧 Special Requirements - ✅ Requires a valid **LATENT** input (e.g., from a prior generation or encoding step). - ✅ Needs a compatible model (`flux-kontext-pro` or `flux-kontext-max`) loaded and wired up via `UNet`, `CLIP`, and `VAE`. - ✅ **CLIP 1 and 2, UNet, and VAE** must match the model family or your outputs will look like Picasso on acid. - ✅ This is not a standalone image editor — it’s one piece in a _latent-space editing pipeline_. ## 🔌 Inputs and Outputs ### **Inputs** - **LATENT** – The latent representation of the image you want to modify. ### **Outputs** - **IMAGE** – The resulting image after applying the edit. - **LATENT** – The updated latent representation after generation. ## ⚙️ Node Settings & Parameters Each field has its own quirks, strengths, and "how-did-this-make-things-worse" settings. Let's dig in. ### 🔢 seed - Controls the randomness. - Same prompt + same seed = same result. - Useful for reproducibility or batch processing. ### 🔁 control_after_generate - **Options:** - `fixed` – Keeps the same seed every time. - `increment` – Adds 1 to the seed on each generation. - `decrement` – Subtracts 1 on each generation. - `randomize` – Full chaos mode; fresh seed every time. ### 🐾 steps - Determines the number of inference steps (aka: how long the model refines the image). - **Lower = faster, but coarser.** - **Higher = slower, but more detailed.** - Sweet spot: 20–40 steps for most edits. ### 🧪 sampler_name - Choose your sampler: `euler`, `dpmpp_2m`, `lcm`, etc. - Each affects how the image evolves through steps. - For deep nerding, see the [Sampler + Scheduler Compatibility Matrix](https://comfyui.dev/docs/guides/Other%20Resources/sampler-and-scheduler-compatibility-matrix). ### 📅 scheduler - Schedulers determine how noise levels are distributed during diffusion. - Examples: `normal`, `karras`, `exponential`, `ddim_uniform`, `kl_optimal` - Some samplers work best with specific schedulers. Choose wisely or expect unholy artifacts. ### 🧭 guidance - AKA "Classifier-Free Guidance Scale" or "CFG." - Higher values force the image to obey the prompt more strictly. - Range: ~1–20 - **Low (1–5)** = Loose interpretations - **Medium (6–12)** = Balanced - **High (13+)** = Obsessive rule-following (sometimes at the cost of quality) ### 🗂️ filename_prefix - Customizes the filename of the generated image. - Handy for batch runs or tracking changes across iterations. - Examples: `"edit_pass1_"`, `"cat_armor_variant_"` ### 📝 prompt - This is where you tell the model what changes you want. - More detail = better edits. - Vague nonsense = latent hallucinations. ### 🧠 unet_name - Selects the diffusion backbone (UNet). - Must match the chosen `flux-kontext` model. - Wrong UNet = broken generations or mismatched results. ### 🔬 weight_dtype - **Options:** - `default` – Uses the default precision (typically FP16) - `fp8_e4m3fn` – Fastest, lowest precision - `fp8_e4m3fn_fast` – Even faster, still low precision - `fp8_e5m2` – Slightly better balance - **Why this matters:** Impacts speed vs. accuracy vs. VRAM. - Use `default` for most cases unless you're fine-tuning for performance. ### 🧠 clip_name1 / clip_name2 - Dual CLIP encoders that handle your text prompt. - Must match your model’s architecture. If you're not sure, refer to the model card/documentation. - Using the wrong ones can cause weird interpretations or semantic confusion. ### ⚙️ device - **Options:** - `default` – Use whatever is available (ideally CUDA/GPU) - `cpu` – For when you're testing... or into self-punishment - **Note:** Flux-Kontext models are large. Running on CPU = slow, sad days. ### 🖼️ Image Preview - A compact thumbnail preview of the output image. - Fast visual feedback to confirm you're not making visual soup. ## ✅ Use Cases - Prompt-guided transformation of existing latent outputs. - Multi-stage image editing workflows (e.g., generation → inpainting → stylization). - Style changes, detail enhancement, or object replacement without losing layout. - Controlled batch editing with reproducible seeds. ## 🧪 Prompting Tips - Be specific. “Change the dress to red” > “make it better.” - Include modifiers like lighting, mood, material, or art style for more directed edits. - Use **negative prompting** in your pipeline if needed (e.g., “no blur, no text”). - Lower `guidance` and `steps` for light edits. Higher values for total overhauls. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - ❌ Feed it raw images instead of LATENTs. - ❌ Mismatch your `unet_name` / `clip_name1/2` with the actual model. - ❌ Forget to load a VAE — you’ll get no image output. - ❌ Use CPU for full-size editing unless you enjoy 10-minute render times. - ❌ Assume fp8 will always save you. Precision matters in high-detail edits. ## ⚠️ Known Issues - **Missing components:** Forgetting to load CLIPs or VAE will break the node. - **Precision loss:** Lower `weight_dtype` settings can cause loss of subtle details. - **Latent drift:** High guidance or many steps can deviate too far from the original image. - **No preview update:** Some changes (e.g., device) may not reflect immediately in the preview section. ## 📝 Final Notes The **Flux.1 Kontext Image Edit** node is a cornerstone of editable, prompt-driven diffusion workflows. It brings powerful latent manipulation into a composable node-based system that gives you _control_ — with just enough room for chaos if you want it. Plug it into your workflow, match your components properly, and enjoy prompt-guided editing that doesn't feel like rolling the dice. --- ## FluxGuidance # FluxGuidance **Applies precision guidance to your conditioning like a GPS for your prompts.** ## 🔍 What is the FluxGuidance Node? The `FluxGuidance` node in ComfyUI is a control mechanism for tweaking how much “influence” your conditioning input has on the final output. It's essentially a volume knob for prompt fidelity — whisper or scream, the choice is yours. This node accepts a conditioning input (from a CLIP/Text Encoder, for example) and applies a user-defined **guidance value** that adjusts how strongly the model listens to that input. Need pixel-perfect prompt-following? Crank it up. Want to let the model wander a little? Dial it down. It’s commonly used in workflows involving **FluxKontext**, **CLIP conditioning**, or other latent-based image generation methods where balancing creative freedom and prompt control is critical. ![FluxGuidance](/img/fluxguidance.png) ## 🧪 Real-World Use Cases - **Hyper-accurate Prompt Adherence:** Ensure outputs follow your prompt like it’s the AI’s final exam. - **Loosen the Reins for Creativity:** Lower the guidance for more abstract or interpretive results. - **Harmonize Multi-Conditioned Prompts:** Use when blending two conditioning inputs (e.g., CLIP1 + CLIP2) and you need to weight one over the other. - **Fine-tuning Style Consistency:** Apply consistent guidance across multiple frames or renders to maintain coherent style in batch or video workflows. ## ⚙️ Inputs, Outputs, and Parameters ### 🔌 Inputs | Name | Type | Description | | -------------- | -------------- | --------------------------------------------------------------------------------------------------------------- | | `conditioning` | `CONDITIONING` | Input conditioning vector from a CLIP/Text encoder or another source. This is what you’re applying guidance to. | | `guidance` | `FLOAT` | How strongly to apply the conditioning input. See below for ranges and effects. | > **Pro tip:** If you're piping in two sets of conditionings (like CLIP1 + CLIP2), you’ll need to run both through their own `FluxGuidance` nodes before combining. ### 🔋 Outputs | Name | Type | Description | | -------------- | -------------- | -------------------------------------------------------------------------------------------------------- | | `conditioning` | `CONDITIONING` | The adjusted conditioning with the guidance applied. Ready to plug into your sampler or generation node. | ### 🎛️ Parameters (in detail) | Parameter | Type | Default | Range | Description | | ---------- | ----- | ------- | ------------- | --------------------------------------------------------------------------------------------------------------------- | | `guidance` | Float | `3.5` | `0.0 – 100.0` | Controls the strength of the conditioning. Affects how strictly your output sticks to the original prompt or concept. | #### 📈 Guidance Behavior | Range | Behavior | | ------------ | ----------------------------------------------------------------------------------------------------------------------------- | | `0.0 – 2.0` | Loose and creative. The model will interpret your prompt in unexpected ways. Useful for generative art or abstract workflows. | | `2.1 – 7.0` | Balanced results. Good fidelity without choking creativity. Great starting point for most workflows. | | `7.1 – 20.0` | Strict prompt control. Useful when you want exactness (portraits, product mockups, logos). Can reduce creative flair. | | `> 20.0` | Proceed with caution. Expect heavy overfitting, artifacts, or model stress — not all models play nice at this level. | ## 🧱 Example Workflow Setup plaintext CopyEdit `[CLIPTextEncode] ↓ [FluxGuidance (guidance=6.0)] ↓ [Flux.1 Kontext Image Edit / T2I / Combine]` This is your go-to setup when you want to gently nudge your model to follow a prompt more closely — without handcuffing it. ## ✨ Prompting Tips - **Use high guidance** for **technical renders**, **logos**, and **faces**. - **Use low guidance** for **fantasy**, **surrealism**, or **dreamcore**. - **Try different values** in 0.5 steps to find the sweet spot. Models and prompts react _very_ differently. - Combine with dual CLIP setup and control guidance separately for interesting hybrid outputs. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - **Don’t max out guidance and expect miracles.** Setting `guidance = 100` is a great way to create AI-generated lava. Most models will cough up a mess of noise, streaks, or worse. - **Don’t forget to connect your conditioning input.** No input = no output. This isn’t a miracle node. - **Don’t use this on already-guided outputs without intention.** Doubling guidance can overcook your result. - Don’t assume all models respond the same way.\*\* Some VAEs, LoRAs, or checkpoints have baked-in guidance tendencies. Always test first. ## ⚠️ Known Issues - Very high guidance values (>20) can break some models or cause overfitting artifacts. - Some models react _too strongly_ to mid-level guidance and lose balance — adjust carefully. - Doesn’t normalize conflicting conditionings — use separate `FluxGuidance` nodes if blending two sources. ## 📝 Final Notes The `FluxGuidance` node gives you the steering wheel when it comes to conditioning strength — but just like driving, too much gas or too much brake leads to a crash. Use it wisely to get the best from your prompts, whether you're generating a batch of consistent characters, abstract digital paintings, or product visuals. It’s lightweight, flexible, and powerful when used correctly. And hey — if it starts generating images that look like glitchy spaghetti, you probably ignored the “don’t set guidance to 100” advice. --- ## FluxKontextImageScale # FluxKontextImageScale _When your image isn’t the right size, and you don’t want to argue with it. This node gets it resized properly so your workflow doesn’t explode later._ --- ## 🧠 What This Node Does The **FluxKontextImageScale** node is a specialized image resizing tool from the `comfy_extras.nodes_flux` module. It scales your image in a context-aware, no-nonsense way — meaning it’s perfect for workflows where size _actually_ matters (looking at you, conditioning and machine learning models). It’s especially useful in complex setups that need consistency, like multimodal processing (image + text + pose, etc.), where the slightest dimensional mismatch can trigger the kind of error that makes you question your life choices. ![FluxKontextImageScale](/img/fluxkontextimagescale.png) ## 🧩 Node Type - **Category:** `Advanced → Conditioning → Flux` - **Python Module:** `comfy_extras.nodes_flux` - **Output is List:** ❌ No - **Output Type:** Image tensor - **Resizing Type:** Context-aware, no quality-enhancing filters — just straight-up dimensional adjustment ## 🔌 Inputs | Name | Type | Required | Description | | ------- | ----- | -------- | ---------------------------------------------------------------------------------------------------- | | `image` | Image | ✅ | The image to resize. Must be a valid tensor. ComfyUI-compatible formats only (`.png`, `.jpg`, etc.). | ## 🔁 Outputs | Name | Type | Description | | ------- | ----- | ------------------------------------------------------------------------------ | | `IMAGE` | Image | The resized image. Ready for conditioning, model input, or multi-modal fusion. | ## ⚙️ Parameters This node has **no user-facing parameters** — it just does the job based on the context of your pipeline and where it’s placed. Internally, it performs smart resizing that aligns the image dimensions with what the rest of your model expects, especially in workflows built around **Flux Kontext** structures. Basically: **no knobs to twist** — just plug it in, and it sizes things properly behind the scenes. ## ✅ Recommended Use Cases - **Pre-conditioning cleanup** Align image sizes before sending to ControlNet, CLIP, or LoRA inputs. Because mismatched resolutions make them sad. - **Multimodal workflows** Scaling images to match other forms of input (pose, depth, text, segmentation maps, etc.). - **Cloud/Remote Workflows** Slim down oversized images before sending them across the internet (save your bandwidth and your patience). - **Preprocessing for ML pipelines** Normalize image dimensions before inference or training. No model likes a surprise. ## 🧪 Workflow Setup Example ```text [Load Image] ↓ [FluxKontextImageScale] ↓ [CLIP Text Encode / Conditioning Node / ControlNet Preprocessor] ↓ [KSampler or Diffusion Node] ``` This node should go **before any node that depends on fixed or known image dimensions**. Stick it early so the rest of your pipeline doesn’t break when it expects a 512x512 and you gave it a 498x505. ## 💬 Prompting Tips - You don’t prompt _this_ node directly. But if your prompt depends on alignment between the image and the conditioning (like pose or depth), **this node makes sure your sizes are aligned behind the scenes**. - Great for **image-to-image prompting** where you need to control resolution before it hits the denoising loop. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - **Don’t** feed this node corrupt or invalid image tensors. It _will_ break, and it _will_ blame you. - **Don’t** expect it to enhance image quality — this is _not_ an upscaler. It scales, it doesn’t beautify. - **Don’t** assume this will magically fix bad dimensions later in the pipeline. Use it early. - **Don’t** skip it if you’re using multiple conditioning inputs that must be aligned — you _will_ get shape mismatch errors. - **Don’t** expect output size control via UI — this node handles things behind the scenes for you. ## 🧾 Final Thoughts **FluxKontextImageScale** is the quiet hero in your ComfyUI setup. No fancy sliders, no noise, just reliable, context-aware resizing to keep your workflow from falling apart due to pixel drama. If your pipeline depends on precision (and let’s face it, they all do), this node belongs in your toolkit. --- ## Image Sharpen # Image Sharpen > Because sometimes, your image just needs a little extra _kick in the pixels._ --- ## 🧠 Node Purpose The **Image Sharpen** node in ComfyUI is a post-processing utility used to enhance the visual clarity of an image by increasing edge contrast, effectively making the details appear crisper and more defined. It’s particularly useful when you need to bring attention to textures, edges, or when you want to fix a slightly blurry result _without yelling at your model to try harder_. This node applies an unsharp mask, a common technique in image processing that works by subtracting a blurred version of the image from the original, then blending the result back using a controlled strength factor. ![Image Sharpen](/img/image-sharpen.png) ## 🔌 Inputs & Outputs | Port | Type | Description | | ------------------- | ------- | ------------------------------------------------------------------------------------------------ | | **IMAGE** | `IMAGE` | The image you want to sharpen. This should be in RGBA format (most nodes output this correctly). | | **Sharpened Image** | `IMAGE` | The result after applying the sharpening algorithm. | ## 🛠️ Parameters & Settings Each parameter in the Image Sharpen node directly affects how the unsharp mask operates. Adjusting these values will determine the _severity_, _range_, and _subtlety_ of the effect. Here’s what each one does: ### 🔹 `sharpen_radius` (float) - **Definition**: This sets the radius (in pixels) of the Gaussian blur used for the unsharp mask. - **Range**: Typically 0.1 – 10.0 (float) - **Default**: 1.0 - **Explanation**: A smaller radius targets only fine details (like hair strands or micro-textures), while a larger radius affects broader areas (like wrinkles or garment folds). - **Effects**: - **Too low**: Effect may be imperceptible. - **Too high**: May introduce halos or make the image look “crispy” or artificial. - **Tip**: Start low (around 0.8–1.5) and adjust as needed depending on subject detail level. ### 🔹 `sigma` (float) - **Definition**: Controls the standard deviation of the Gaussian blur kernel. - **Range**: Usually between 0.1 – 5.0 (float) - **Default**: 1.0 - **Explanation**: Sigma works together with the radius but is more about the _spread_ of the blur. Higher sigma means more blurring before the sharpening is applied. - **Effects**: - **Low sigma**: More local sharpening. - **High sigma**: More feathered, smoother transitions — but could also reduce the "sharp" look. - **Warning**: If `sigma` is set too high compared to the `sharpen_radius`, the node can produce muddy or barely noticeable effects. ### 🔹 `alpha` (float) - **Definition**: The sharpening strength; how much of the difference image is added back. - **Range**: 0.0 – 2.0 (float) - **Default**: 1.5 - **Explanation**: This is your “spice” level. Higher values make the image look more aggressively sharpened. - **Effects**: - **0.0** = No sharpening (aka pointless). - **1.0** = Normal sharpening. - **>1.5** = Aggressive; use with caution or regret. - **Tip**: Start with **1.0** and bump upward if needed, but keep an eye out for **over-sharpening artifacts** (like halos and grain exaggeration). ## 🧩 Workflow Integration ### 🧬 Typical Placement The **Image Sharpen** node is typically placed near the _end_ of a text-to-image or image-to-image workflow, right before saving or displaying the image: ```plaintext [Model Output/Image Post-Processed] ➡️ Image Sharpen ➡️ Save Image / Display Preview ``` You don’t want to sharpen too early in the pipeline—let your denoisers and upscalers do their magic first. Then slap this node in to put the cherry on top. ## ✅ Recommended Use Cases - 🖼️ Enhancing subtle details in photorealistic renders (e.g., pores, cloth textures). - 🧑‍🎨 Fixing soft output from image upscalers or inpainting workflows. - 🏞️ Boosting landscape or architectural features where structure matters. - 🦄 Giving that final "pop" to stylized illustrations or semi-realistic renders. ## 💡 Prompting Tips This node doesn’t interact with your prompt, but it _does_ affect the perceived fidelity of the result. If you're doing fine-detailed prompts (like: > _"ultra-detailed embroidery on silk, 8k textures, cinematic lighting"_) ...then this node can actually help _show_ the detail your model imagined but kind of soft-pedaled. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire > Aka: “Bad Decisions You’re Totally Free to Make—But Please Don’t.” 🧯 The `Image Sharpen` node is not inherently dangerous, but in the hands of the overly enthusiastic or the _“I just maxed all the sliders”_ crowd, things can get… crunchy. Here are the top sins that will turn your beautiful render into a pixelated war crime: #### ❌ **Cranking Alpha to 2.0 on a Already-Detailed Image** **Result:** Hello, haloing! Get ready for glowing outlines around every edge like you're auditioning for a _1997 website graphic._ If your image already has high-frequency detail, pumping `alpha` too high will exaggerate edges to the point of looking like they’ve been outlined in neon Sharpie. Resist the urge. #### ❌ **Using a Giant Radius and Sigma Combo (e.g., 10 + 5)** **Result:** You’ve just built a _blurmancer_. Instead of sharpening, the node ends up amplifying weird blocky gradients, and can introduce large-scale mushy halos that somehow sharpen and blur at the same time. It’s impressive. And bad. Stick to subtle values unless you _absolutely_ know what you're doing (or you're conducting pixel-based performance art). #### ❌ **Sharpening Artifacts from Upscaling or Denoising** **Result:** So, you used Ultimate SD Upscale, and now you want to sharpen. Cool. But those seams you didn't fix? Or the tiny JPEG noise from an upscaled LoRA? Yeah—this node _loves_ turning them into crusty visual landmines. Always clean up artifacts before sharpening, or you'll be lovingly enhancing your render's worst flaws. #### ❌ **Running It Twice Because You Didn't See a Difference the First Time** **Result:** Just because you didn’t get razor-sharp detail on pass one doesn’t mean you should double-down like a blackjack addict. Compounding sharpening causes exponential damage—not detail. If it’s not working, check your settings or your source image. Don’t just stack sharpen layers like pancakes. #### ❌ **Sharpening Anime or Line Art with Fine Lines** **Result:** You've now got a web of edge halos, contrast gaps, and lines that look like someone spilled ink over a circuit board. For anime-style outputs, try less aggressive sharpening or use other post-processing methods like edge-preserving upscalers or line refinement. #### ❌ **Thinking Sharpening Will Magically Fix a Bad Prompt** **Result:** Garbage in, sharper garbage out. Sharpening does not—and will never—generate lost details. If your image looks like it was painted with mashed potatoes, it's not because it's “too soft.” Go back and refine your prompt, or revisit your CFG/denoise settings. ##### 🚫 Bonus Round: Live Preview Expectations This node won’t show you results in the workflow canvas preview if your system is slow or memory-starved. Don’t keep smashing "Queue Prompt" thinking it’s broken. It’s not. Your VRAM just called in sick. ## 🧪 Example Settings | Use Case | Radius | Sigma | Alpha | | ----------------------------- | ------ | ----- | ----- | | Subtle sharpening for realism | 1.0 | 1.0 | 1.0 | | Aggressive sharpening (bold) | 1.5 | 0.8 | 1.8 | | Soft detail enhancement | 2.0 | 1.5 | 0.9 | | Stylized edge pop | 1.0 | 0.5 | 2.0 | ## 🧼 Final Advice If your render already looks sharp, don’t push this node just for fun. Sharpening is like salt—you can always add more, but it’s real annoying to fix once you’ve dumped the whole box in. Use it with intention. Trust your eyes. And maybe—_just maybe_—try zooming out before deciding the whole image is “blurry.” --- ## Image Stitch # Image Stitch **Category:** `comfy-core` **Not to be confused with the WAS Node “Image Stitch.”** They share a name, but the WAS version is more like an artsy scrapbooker, while the comfy-core version is a blunt pixel-joiner that’s here to do one thing: slap two images together in a straight line. ## 🔍 **Overview** The **Image Stitch** node in comfy-core lets you combine **exactly two images** into a single output image in one of four cardinal directions: **right**, **left**, **down**, or **up**. Think of it as a no-nonsense, four-directional tape gun for images. There’s no blending, no magic AI seam-hiding — just pure “take these pixels and stick them next to each other” energy. ![Image Stitch](/img/image-stitch.png) ## 📁 **Function in ComfyUI Workflows** You feed it **two images**, it outputs **one stitched image**. You control: - The **direction** of the stitch (placement of image 2 relative to image 1) - Whether both images should be **resized to match** - The **spacing width** between them (in pixels) - The **color** used for that spacing It’s perfect for before/after previews, comic panels, side-by-side comparisons, and other “I need these two images in one file” situations. ## 🧠 **Technical Details** - **Inputs:** - `image 1` (IMAGE) – The base image. - `image 2` (IMAGE) – The image to attach in the chosen direction. - **Output:** - `IMAGE` – The resulting stitched image. - **Processing:** Pixel-level concatenation with optional resizing and fixed-color spacing fill. - **Performance:** Extremely fast — minimal GPU usage, negligible VRAM impact unless images are extremely large. ## ⚙️ **Parameters & Settings (Deep Dive)** ### 1. **direction** - **Type:** `COMBO` (Enum) - **Options:** - `right` → Places `image 2` to the right of `image 1`. - `left` → Places `image 2` to the left of `image 1`. - `down` → Places `image 2` below `image 1`. - `up` → Places `image 2` above `image 1`. - **Why It Matters:** This completely changes the layout. If you’re aiming for a horizontal composition, choose `left` or `right`. If you want vertical stacking, choose `up` or `down`. - **Changing This Means:** Get ready for a different aspect ratio and possibly a total visual composition shift. ### 2. **match_image_size** - **Type:** `BOOLEAN` (True/False) - **True:** Resizes both images so their shared edge dimensions match: - For `left`/`right`, both images will have the same height. - For `up`/`down`, both images will have the same width. - **False:** Keeps original image sizes — which may result in mismatched edges. - **Why It Matters:** Without this, stitching different-sized images will cause jagged or uneven seams. ### 3. **spacing_width** - **Type:** `INTEGER` (0 or more) - **Function:** Number of pixels between the two images. - **Effects:** - **0** → Images directly touch. - Larger numbers → Creates a visible gap. - **Why It Matters:** Helps create visual separation in comparisons or layouts. ### 4. **spacing_color** - **Type:** `COMBO` (Enum) - **Options:** - `white` - `black` - `red` - `green` - `blue` - **Why It Matters:** Ensures your spacing area is filled with a solid, intentional color. Prevents transparency unless you want the gap to scream “unfinished.” ## 💡 **Recommended Use Cases** - **Before/After Comparisons** – Prompt changes, model swaps, or style shifts. - **Side-by-Side Evaluations** – Model A vs Model B. - **Storyboarding / Panels** – Simple, clean layouts. - **Reference Prep** – Combine inspiration images in a set order. ## 🛠 **Workflow Setup Example** 1. Connect `image 1` and `image 2` outputs from generation nodes. 2. Add **Image Stitch** node and wire both inputs. 3. Choose `direction` (right, left, down, up). 4. Toggle `match_image_size` ON if you need perfect alignment. 5. Set `spacing_width` and `spacing_color` to taste. 6. Save or send to next node in the workflow. ## 🧾 **Prompting Tips** - When planning to stitch, generate both images with matching dimensions — it avoids resizing artifacts. - Use **spacing** to visually clarify “this is two separate renders” when comparing results. - If you want invisible seams, set spacing to **0** and ensure colors match at the join. ## 🔥 **What-Not-To-Do-Unless-You-Want-a-Fire** - Don’t feed it only one image — this is not the “lonely painter” node. - Don’t expect it to “blend” — it’s a cut-and-paste machine, not a Photoshop wizard. - Don’t stitch ultra-high-res images without enough RAM unless you like crashes. - Don’t confuse this comfy-core node with the WAS Node Image Stitch — they are **not interchangeable**. - Don’t set spacing color to something obnoxious unless your goal is to blind viewers. ## ⚠️ **Known Issues** - **Aspect Ratio Changes:** Direction changes can drastically alter output proportions. - **Resizing Softness:** With `match_image_size` ON, resizing may cause slight image softening. - **Limited Colors:** Only 5 preset spacing colors available — no custom hex/RGB. --- ## KSampler (Advanced) # KSampler (Advanced) ## 🧠 What Is This Node? The **KSampler (Advanced)** node is the _fully loaded_ variant of the standard KSampler. Think of it as the same car, but now with a turbocharged engine, racing suspension, and a dashboard full of extra switches you may or may not understand yet. It gives you **precise control over every major sampling parameter** — steps, CFG scale, sampler type, scheduler, denoise strength, and more — making it the go-to choice when you need _consistent, high-quality, repeatable results_ or want to experiment with workflow tuning at a granular level. If the standard KSampler is “good enough” for most tasks, **KSampler (Advanced)** is what you use when you want to push boundaries, run controlled experiments, or debug exactly why your AI thinks a cat should have three tails. ![KSampler (Advanced)](/img/ksampler-advanced.png) ## 🧩 Real-World Use-Cases - High-control **text-to-image** workflows with reproducible results - **Image-to-image** refinement where denoise strength determines how much the original is preserved - **Prompt A/B testing** with locked seed values - Testing **sampler/scheduler compatibility** for optimal style results - Multi-stage workflows where **latents** are passed through various transformations before decoding ## 🔌 Inputs #### model _(Required)_ - **What it is:** The diffusion model to use for generation. - **Why it matters:** Different models have different strengths; the wrong choice here is like asking a watercolor artist to carve marble. - **Requirements:** Must be a valid model loaded into ComfyUI. #### seed _(Integer)_ - **Default:** `0` - **Range:** `0` to `0xffffffffffffffff` - **What it does:** Initializes the random number generator for reproducibility. Same seed + same parameters = identical result every time. - **Tips:** - Lock it to iterate on prompt changes consistently. - Randomize for variety. - **Caution:** Changing the seed by just 1 can produce a _completely_ different image. #### steps _(Integer)_ - **Default:** `20` - **Range:** `1` to `10,000` (but don’t — unless you enjoy watching progress bars more than making art) - **Function:** Number of sampling iterations. Higher values generally = better quality, but diminishing returns past ~30–50 for most models. #### cfg _(Float)_ - **Default:** `8.0` - **Range:** `0.0` to `100.0` (increments of `0.1`) - **Function:** Classifier-Free Guidance scale — how closely the model follows the positive conditioning. - **Low (<5):** Loose, interpretive - **Medium (7–12):** Balanced adherence - **High (>15):** Strict adherence, risk of harsh outlines or unrealistic detail - **Tip:** Start around 7–9 and adjust. #### sampler name _(Dropdown)_ - **Function:** The algorithm that drives the sampling process. - **Impact:** Can drastically change detail sharpness, style, and rendering speed. - **Examples:** `euler`, `dpmpp_2m`, `heun`, `ddim`, `lcm` - **Note:** Some samplers perform better with certain schedulers — choose wisely. #### scheduler _(Dropdown)_ - **Function:** Determines how noise is scheduled over the steps. - **Impact:** Affects smoothness, contrast, and convergence speed. - **Examples:** `normal`, `karras`, `exponential`, `sgm_uniform` - **Tip:** `karras` often yields smoother high-quality results. #### positive _(Conditioning Input, Required)_ - **What it is:** The “do this” list for your model. Usually comes from CLIP text encoding. - **Tip:** Keep it clear and concise — overloading with too many descriptors can muddy results. #### negative _(Conditioning Input, Optional but Highly Recommended)_ - **What it is:** The “don’t you dare” list for your model. - **Purpose:** Suppresses unwanted traits (e.g., _blurry, watermark, extra limbs_). #### latent image _(Required)_ - **Function:** The starting point in latent space — either random noise (text-to-image) or an encoded image (image-to-image). - **Caution:** With **denoise=1.0**, it will ignore any structure from the latent and start fresh. #### denoise _(Float)_ - **Default:** `1.0` - **Range:** `0.0` to `1.0` (increments of `0.01`) - **Function:** Controls how much of the starting latent is preserved. - `1.0` → Full redraw from scratch - `0.5` → Half preserved, half new - `0.1` → Light refinements - **Pro Tip:** For subtle edits, keep this low; for wild reimaginings, crank it up. ## 📤 Outputs #### LATENT The refined latent representation of the generated image, ready for decoding or further processing. ## 💡 Usage Tips - **Seed discipline:** Lock seeds when testing prompts; change seeds to explore variety. - **Steps efficiency:** Avoid going overboard — most gains happen under 50 steps. - **CFG sweet spot:** 7–12 works for most models without forcing unnatural detail. - **Sampler/scheduler pairing:** Experiment, but check known compatibility first. - **Denoise control:** Low for polishing, high for creative chaos. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - Steps = 10,000. Your GPU will hate you. - CFG = 100. Enjoy your crunchy, overbaked AI noodles. - Denoise = 1.0 on a refined image you actually liked. - Mismatched sampler/scheduler combos without testing. ## ⚠️ Known Issues - Incompatible sampler/scheduler combos can yield flat or noisy results. - Extremely high CFG can cause oversharpening, harsh outlines, or strange artifacts. - Very high step counts waste time with minimal visible improvement. - Some models react badly to extreme denoise settings. ## 📝 Final Notes The **KSampler (Advanced)** node is where you stop being a passenger and start piloting the generation process yourself. It’s more powerful, more configurable, and less forgiving than the standard KSampler — but in the right hands, it’s the difference between “pretty good” and “wow, how did you do that?” --- ## KSampler # KSampler The engine of generation in ComfyUI. This node takes your model, your prompt, your settings, and cranks out images by denoising a latent tensor into art. It’s where diffusion happens, and where your GPU earns its keep. ![KSampler](/img/ksampler.png) ## 🔌 Inputs | Input | Type | Description | | ----------------------- | -------------- | -------------------------------------------------------------------------------------------------------------------- | | `MODEL` | `MODEL` | The diffusion model you're using — e.g., `dreamshaper_8.safetensors`, `revAnimated_v122.safetensors`. Required. | | `CLIP` | `CLIP` | Text encoder. Usually comes from the same `CheckpointLoaderSimple`. It interprets your prompts into guidance. | | `Positive Conditioning` | `CONDITIONING` | From your prompt node. Guides the image generation _toward_ the prompt. | | `Negative Conditioning` | `CONDITIONING` | Optional but highly recommended. Pushes the image _away_ from undesirable traits (e.g., “blurry, extra limbs”). | | `LATENT` | `LATENT` | The initial latent image. This can be random (for text-to-image) or derived from input for img2img, inpainting, etc. | ## ⚙️ Parameters and Settings (Deep Dive) ### 🔢 `seed` (INT) - **Purpose:** Sets the starting noise pattern. - **Same settings + same seed = same image**. Crucial for reproducibility. - **`-1` = random seed** each time. **Use Cases:** - Lock for reproducibility. - Randomize when exploring. ### 🎛 `control_after_generate` (STRING ENUM) **Despite the name, this has nothing to do with ControlNet.** This setting tells ComfyUI how to manage the seed value across batches. **Options:** - `fixed`: Same seed for every image. - `increment`: Add +1 to seed for each image. - `decrement`: Subtract -1 per image. - `randomize`: Use a random seed for each. **Why it matters:** - `fixed` = consistent image generation. - `increment` = ideal for controlled variations. - `randomize` = embrace the chaos. ### 🧮 `steps` (INT) - **Purpose:** Number of denoising iterations (the more steps, the more chances the model has to refine the image). - **Typical Range:** 20–50 for best balance. - **Max Range:** Up to 150+, but prepare to wait. **Guidance:** - Too few = blurry or underdeveloped results. - Too many = diminishing returns + GPU tears. ### ⚖️ `cfg` (FLOAT) (Classifier-Free Guidance Scale) - **Purpose:** Controls how much the output adheres to the prompt. - **Typical Range:** 1–20 - **Default Sweet Spot:** 6.5–8.5 **Low cfg (e.g., 2)** = freedom, creativity, also prompt forgetfulness. **High cfg (e.g., 15)** = "you said banana samurai, you’re getting banana samurai." ### 🌀 `sampler_name` (STRING ENUM) Determines the sampling algorithm used to perform denoising. **Popular Samplers:** - `Euler a`: Fast, chaotic, great for creativity. - `DPM++ 2M Karras`: Smooth, stable, photorealistic. - `Heun`, `LMS`, `UniPC`: All with their own quirks. **Best practice:** Try a few — results can vary dramatically by sampler. **Ready to learn more?** Take a look at our deep dive on all the [sampler_name options](http://comfyui.dev/docs/guides/Other%20Resources/sampler-name-options). ### 📆 `scheduler` (STRING ENUM) Defines the noise schedule for denoising steps. **Options:** - `normal`: Uniform distribution. - `karras`: Better for fine details (recommended). - `exponential`: Aggressive at early steps. **Tip:** Use `karras` unless you're specifically told not to. **Ready to learn more?** Take a look at our deep dive on all the [sampler_name options](http://comfyui.dev/docs/guides/Other%20Resources/scheduler-options). ### 🌫️ `denoise` (FLOAT) - **Range:** 0.0–1.0 - **1.0 = full generation from noise** (text-to-image) - **<1.0 = preserve structure** (for img2img, inpainting, ControlNet guidance) **Example Use:** - `1.0` → generate from scratch - `0.5` → img2img subtle change - `0.1–0.3` → ControlNet with light touch ## 🧰 Special Requirements & Notes - Ensure your **MODEL**, **CLIP**, and **VAE** are compatible. - Always connect both **positive** and **negative conditioning** for best results. - Don't be afraid to play with **CFG and Denoise** — they're your finesse tools. - High `steps` + High `cfg` = GPU meltdown risk ⚠️ ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire Some mistakes are so common, so predictably catastrophic, that they deserve their own red-flagged list. If you enjoy wasting GPU hours, crashing your ComfyUI session, or summoning unholy Lovecraftian blobs, go ahead and do any of the following: #### 🚫 Set `steps` to 150+ with `denoise` at 1.0 Unless you’re training your patience or testing thermal limits, this is the fastest way to generate pixel soup _very_ slowly. **Instead:** Use 25–40 steps for 99% of tasks. #### 🚫 Use `cfg=20` because “higher must be better” This won’t make your prompt more accurate — it’ll make your image look like a bad Photoshop job run through a paper shredder. **Instead:** Stick to the 6.5–8.5 sweet spot. Go higher _only_ if you know what you’re doing. #### 🚫 Forget to set `control_after_generate` when doing batch generations You wanted 8 unique variations. You got 8 clones. Oops. **Instead:** Use `increment` for clean batch diversity. #### 🚫 Use `denoise=0.1` in a text-to-image workflow You just told the sampler to generate... almost nothing. Enjoy your blank canvas with a faint smudge of regret. **Instead:** Use `denoise=1.0` for full generations. Lower values are for img2img or ControlNet. #### 🚫 Combine `Euler a` sampler with `steps=80` thinking it'll be glorious Nah. `Euler a` is a fast sampler — not meant for marathon sessions. You're not getting more detail, you're just looping futility. **Instead:** Use `DPM++ 2M Karras` or similar for higher-step use. #### 🚫 Forget to connect `Negative Conditioning` Yes, it's “optional” — just like wearing pants in public is technically optional. But without it, your output will gleefully ignore your expectations and embrace chaos (in all the wrong ways). #### 🚫 Use incompatible models and CLIP encoders That “weird green mush with anime eyes and four ears”? Yeah, that’s what happens when you mix a v1.5 checkpoint with a v2 CLIP. Don’t. #### 🚫 Crank everything to max at once CFG = 15, steps = 100, resolution = 2048x2048, tiling disabled, sampler = experimental beta nightly — your GPU just left the chat. #### 🚫 Blame the sampler when your prompt sucks KSampler is powerful, but it’s not a miracle worker. If your prompt is “woman” and the image is cursed, well… maybe give the poor thing more to work with. ##### ✅ Pro Tip: When in doubt — lower your steps, simplify your prompt, lower your CFG, and only touch `denoise` if you _know_ what it does. KSampler rewards precision and punishes overconfidence. ## 🧪 Example Workflow ```text [CheckpointLoaderSimple] └─> MODEL ─────────┐ │ [CLIPTextEncode (positive)] ──> KSampler (positive) [CLIPTextEncode (negative)] ──> KSampler (negative) [EmptyLatentImage] ───────────> KSampler (latent) CLIP ─────────────────────────> KSampler (CLIP) KSampler ─────────────────────> [Decode/Save/Display] ``` ## ⚡ TL;DR Config Cheat Sheet | Setting | Default / Safe Range | Notes | | ------------------------ | -------------------- | -------------------------------------- | | `seed` | -1 (random) | Lock for repeatability | | `control_after_generate` | `increment` | Varies seed across batch | | `steps` | 20–30 | 40+ for ultra detail | | `cfg` | 7.5 | Higher = more prompt fidelity | | `sampler_name` | `DPM++ 2M Karras` | Stable and detailed | | `scheduler` | `karras` | Best detail control | | `denoise` | 1.0 | Use <1.0 for image-guided workflows | ### 📚 Additional Resources - [Sampler Options](https://comfyui.dev/docs/guides/Other%20Resources/sampler-name-options) - [Scheduler Options](https://comfyui.dev/docs/guides/Other%20Resources/scheduler-options) ## 🧠 Final Thoughts The **KSampler** is where the real magic happens — the ultimate forge of diffusion sorcery. Every parameter you tweak gives you a different flavor of art, so get in there and _experiment_. Mastering this node means mastering your output. Need docs for any of the surrounding nodes (`EmptyLatentImage`, `CLIPTextEncode`, etc.)? Just shout — I'll keep the documentation train rolling. --- ## KSamplerSelect # KSamplerSelect **“Convenient Sampler Selection — without the guesswork or typo-induced breakdowns.”** --- ## 🧠 What is KSamplerSelect? The `KSamplerSelect` node is a _zero-friction utility node_ designed for one purpose: selecting a sampler from the `comfy.samplers.SAMPLER_NAMES` list and outputting it as a usable `SAMPLER` object. Think of it as your **sampler sommelier**—it helps you choose the right flavor of sampling algorithm without fumbling through dropdowns on the `KSampler` node itself. This node is perfect for: - Rapid experimentation with multiple samplers - Swapping sampling strategies between workflows - Keeping your pipeline modular and neat - Avoiding the wrath of typo gremlins (e.g., “ddim” vs “DDIM” vs “dpmpp_2m_sde_gpu” 🙃) ![KSamplerSelect](/img/ksamplerselect.png) ## 🧩 Inputs ### 🔹 `sampler_name` **Type**: `String` (dropdown from `comfy.samplers.SAMPLER_NAMES`) **Required**: ✅ Absolutely **Default**: _None_ (and yes, that means you **must** set it) The `sampler_name` is the only input this node needs, and it does all the heavy lifting. When you select a value, the node fetches the appropriate internal sampling object used by other diffusion nodes (primarily `KSampler`). **Examples of available samplers:** - `euler`, `euler_ancestral` - `heun`, `heunpp2` - `dpmpp_2m`, `dpmpp_2m_sde`, `dpmpp_2m_cfg_pp` - `ddpm`, `lcm`, `ipndm` - ...and yes, even `uni_pc`, `er_sde`, `gradient_estimation`, and the rest of the alphabet soup. > 🧠 **Tip:** All sampler names are pulled from `comfy.samplers.SAMPLER_NAMES`. If it’s not in that list, it’s not valid. ## 🔻 Outputs ### 🔸 `SAMPLER` **Type**: `SAMPLER` object **Description**: This is the internal sampler instance tied to your selected algorithm. **Usage**: Feed this directly into the `KSampler` node (or any other node expecting a `SAMPLER` input). This output is what makes the node functionally useful. It’s a direct reference to the logic that governs how your image is actually generated—so yes, it matters _a lot_. ## ⚙️ Recommended Use Cases - 🎛 **Dynamic Workflows**: Build UI-style interfaces where users pick samplers without opening the backend spaghetti. - 🧪 **Sampler A/B Testing**: Wire up multiple KSamplerSelects to test how different samplers affect a prompt. - 📦 **Reusable Templates**: Create modular templates where only the sampler changes between styles or projects. - ☁️ **Cloud Workflows**: Great for ComfyUI cloud setups like ComfyUI-Manager, where toggling options remotely matters. ## 🧾 Example Workflow ```plaintext [Load Checkpoint] → [KSamplerSelect] → [KSampler] → [VAE Decode] → [Save Image] ``` Add multiple `KSamplerSelect` nodes to feed alternate branches: ```plaintext ↘ [KSamplerSelect: euler] ↘ [Load Checkpoint] → [KSampler] → [VAE Decode] ↗ [KSamplerSelect: dpmpp_2m_sde] ↗ ``` ## 📌 Prompting Tips - **Prompt remains constant** → Sampler determines _how_ it gets interpreted. - Samplers like `dpmpp_2m_sde` are great for **long prompts**, **high coherence**, and **complex compositions**. - `euler_ancestral` tends to be **fast and flexible**, good for **rough sketches or previews**. - If you're using `lcm`, remember to **drop your steps** or risk overbaking your image into a pile of visual mush. ## ❌ What-Not-To-Do-Unless-You-Want-a-Fire - 🔥 **Leave `sampler_name` blank** – The node won't output anything, which breaks your pipeline. - 🔥 **Assume sampler order matters** – This is not a tier list, just a list of available options. - 🔥 **Use incompatible samplers with schedulers** – Some combos (e.g. `ddim` + `karras`) will silently fail or give junk. Check your compatibility. - 🔥 **Forget to connect the `SAMPLER` output** – The `KSampler` node won’t magically know what sampler you wanted. ## ⚠️ Known Issues | Issue | Cause | Solution | | ------------------------------ | --------------------------------------------------------------------------------------- | --------------------------------------------------------------------------- | | `Invalid sampler name` | You typed something not in `SAMPLER_NAMES` (or copy-pasted from StackOverflow again 😑) | Use the dropdown. Seriously. | | `Missing sampler_name` | You didn’t select anything. | Select a sampler before running. | | **“My image looks weird now”** | You changed the sampler and expected the same results | That’s not how any of this works. Different samplers = different behaviors. | ## 📝 Final Notes The `KSamplerSelect` node is the behind-the-scenes MVP for **workflow clarity, testability, and modularity**. Use it to cleanly define sampling behavior without overloading your `KSampler` node with hardcoded settings. This node doesn’t generate images—it just makes sure the right algorithm does. > 🧠 **Pro move**: Pair this with a **Conditioning Select** and **Checkpoint Select** for a fully modular generation system that looks like you _actually_ know what you're doing. --- ## Load Checkpoint # Load Checkpoint Welcome to the no-nonsense, full-throttle documentation for the `Load Checkpoint` node in ComfyUI. This deceptively simple node is the gatekeeper of your entire generative pipeline—if it doesn’t load the right model, your beautifully structured workflow is just a stack of paper with no printer. --- ## 🔧 Node Type **`CheckpointLoaderSimple`** — Often labeled as “Load Checkpoint” in the UI. This node loads the three holy trinity components required for image generation in Stable Diffusion: - `MODEL` (UNet model) - `CLIP` (text encoder) - `VAE` (variational autoencoder) Each of these has a distinct role, and choosing the wrong one can result in... _abstract_ outputs at best and cursed AI hallucinations at worst. ![Load Checkpoint](/img/load-checkpoint.png) ## 🧠 What This Node Does The `Load Checkpoint` node reads `.ckpt` or `.safetensors` files that represent a pre-trained model for Stable Diffusion. It extracts and splits the components needed downstream in your workflow: | Output Slot | What it Outputs | Why it Matters | | ----------- | ------------------------------------------ | --------------------------------------------- | | `MODEL` | The UNet model | Core engine for denoising and image synthesis | | `CLIP` | Text encoder | Translates your prompt into vector embeddings | | `VAE` | Decoder that turns latent data into pixels | Affects color tone, contrast, and fine detail | These outputs are then used by nodes like `KSampler`, `CLIP Text Encode`, and `VAE Decode` to do the heavy lifting. ## ⚙️ Inputs & Settings ### 🔽 `ckpt_name` (Dropdown) - This dropdown allows you to select from available checkpoints located in your `/models/checkpoints/` directory. - Supported formats: - `.safetensors` (preferred—safe, optimized loading) - `.ckpt` (less secure, but still works) - Changing this sets the foundation for the aesthetic and capabilities of your generation. > **Tip:** Rename your models meaningfully. `epic_realism_fp16.safetensors` is easier to find than `model_final_v1.2.ckpt`. ## 🧩 Outputs (in detail) ### 🔸 `MODEL` - This is the UNet component of Stable Diffusion. - It processes the latent noise over several steps to create your final image. - It is passed directly into the `KSampler` node. 📌 **Important**: Incompatible UNets with certain schedulers or samplers (especially experimental ones like LCM) may result in artifacts or completely non-functional outputs. ### 🔸 `CLIP` - The CLIP model is used by `CLIP Text Encode` nodes. - It encodes your prompt (and negative prompt) into a semantic vector space. - Different checkpoints might include CLIP variants like: - `ViT-B/32` (used in SD1.4/1.5) - `OpenCLIP` variants (used in SD2.1 and later) > **Prompting Tip:** If your prompts are working inconsistently across checkpoints, blame the CLIP. It's not you. You're brilliant. ### 🔸 `VAE` - The decoder/encoder that transforms the latent space into pixel space. - Without this, you’re stuck in the land of latent noise—no actual image decoding can happen. - VAEs affect: - Skin texture - Color accuracy - Sharpness - Overall image quality > **Pro Tip:** Some models bake their VAE inside the checkpoint. Others rely on external ones like `vae-ft-mse-840000-ema-pruned.safetensors`. Be aware of what your model needs. ## ✅ Recommended Use Cases - **Text-to-Image Generation** (obviously) - **Image-to-Image Generation** (when using inpainting or ControlNet pipelines) - **LoRA and Style Merging Workflows** (needs model/clip pairing) - **Comparative Model Testing** (run identical prompts across multiple checkpoints) ## 🧪 Workflow Setup Example ```mermaid graph TD; A["Load Checkpoint"] --> B["CLIP Text Encode"]; A --> C["KSampler"]; A --> D["VAE Decode"]; ``` - Connect: - `MODEL` to your `KSampler` node - `CLIP` to `CLIP Text Encode` nodes - `VAE` to your `VAE Decode` node or `AutoVAE` ## 🛠️ Settings & Parameters | Parameter | Type | Description | | -------------- | -------------- | ------------------------------------------------------------------------------------------------------------ | | `ckpt_name` | Combo Dropdown | Select the checkpoint you want to load. Reflects the files in your checkpoints folder. No path input needed. | | Other Settings | None | That’s it. It’s brutally simple by design. | ## 📈 Prompting Tips - **Consistency**: Always use the same model and VAE if you’re comparing outputs. - **Specialized Checkpoints**: Use realism-trained models like `epicRealism` for portraits and `revAnimated` for anime-style art. Don’t cross the streams unless you _like_ weird hybrids. - **Negative Prompts**: Behavior is heavily influenced by the CLIP model—some checkpoints understand nuance better than others. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire Here’s the list of rookie mistakes, cursed behavior, and general chaos-summoning actions that will send your workflow spiraling into the abyss (and maybe crash your VRAM while it's at it): #### ❌ Load the wrong checkpoint and wonder why nothing looks right - Using an anime-style checkpoint for photorealism? Yeah, enjoy those melted Barbie faces. - Loading a realism checkpoint and prompting it like you're in Ghibli? That’s how you get uncanny valley nightmares. #### ❌ Forget to pair a VAE when your checkpoint doesn’t have one baked in - This will result in invisible outputs, blank images, or grayscale weirdness. - Pro tip: _If you see a mostly black or blank image, check the VAE. It’s probably that._ #### ❌ Use a checkpoint that isn’t compatible with the rest of your workflow - If you're mixing SD1.5 checkpoints with SD2.x CLIP encoders or ControlNet inputs, congratulations—you’ve entered undefined behavior land. - Know your version compatibility (SD1.5 = CLIP B/32, SD2.x = OpenCLIP G). #### ❌ Rename files without knowing what you’re doing - That cool new `epicRealism_better_but_idk_final_final.ckpt` might not load if your filename confuses your model manager or path references. - Keep it clean. No spaces. Use underscores. It’s not just good practice—it’s survival. #### ❌ Download random checkpoints from sketchy sources - This isn’t LimeWire in 2002. Your GPU has dignity. - Always scan `.ckpt` files and prefer `.safetensors` because they _don’t execute arbitrary code._ #### ❌ Ignore the console logs - That red error message? Yeah, it's not just decoration. - If the model didn't load properly, you'll often see it there first. Don’t blame the sampler when the checkpoint never even initialized. #### ❌ Assume all checkpoints include a good CLIP encoder - Some include low-quality or mismatched encoders, which means even your beautifully-crafted prompt will turn into noise soup. - If the prompt stops behaving, consider checking what CLIP variant the model is using. ##### 🔥 Bonus Chaos Recipe Want to crash your workflow in 5 seconds flat? 1. Load a mismatched VAE 2. Plug the wrong CLIP into the wrong text encode 3. Use a 16-bit half-precision model on a GPU that doesn’t support it 4. Set the batch size to 8 on a 6GB card 5. Try to upscale it at the end with Ultimate SD Upscale Boom. ComfyUI just turned into _UnComfyUI_. ## 🧼 Best Practices - **Stick to `.safetensors`**: It’s faster, safer, and just plain better. - **Name smart**: Keep your file names descriptive so you can tell `ponyRealism_V23` from `hellspawn_final_fp32`. - **Cache preload**: Use the **ComfyUI Manager** or `preload_models` setting for performance gains. ## 🧩 Bonus: Pairing with VAEs If your checkpoint doesn't have a baked-in VAE, use these: | Checkpoint Type | Recommended VAE | | --------------------------------- | ------------------------------------------ | | Realism-based (e.g., EpicRealism) | `vae-ft-mse-840000-ema-pruned.safetensors` | | Stylized / Anime | `anything-v4.0.vae.pt` or baked | | Hyper-detailed art | `kl-f8-anime2` (if not baked) | ## 🧠 TL;DR The `Load Checkpoint` node is the very first brick in your generative cathedral. Choose wisely, connect faithfully, and your prompt shall manifest as pure pixelated glory. Choose poorly? Well… hope you like eyeballs on elbows. --- ## Load Diffusion Model # Load Diffusion Model The **Load Diffusion Model** node in ComfyUI is your gateway to using powerful pretrained diffusion models like `flux1-dev.safetensors`, `wan2.1_t2v_14B_fp16.safetensors`, and others. This node is responsible for initializing and injecting a loaded U-Net model into your workflow—because _nothing_ gets generated until your model shows up to the party. It supports advanced use cases like swapping in different architectures on-the-fly, reducing VRAM usage via lightweight weight formats, and making ComfyUI feel like a modular generative playground rather than a tangled mess of JSON spaghetti. ![Load Diffusion Model](/img/load-diffusion-model.png) ## 🔌 Inputs | Name | Type | Description | | ------ | ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | _None_ | – | This node does not take any inputs because it’s the first thing that loads your diffusion model into memory. Think of it as the "boot loader" for your U-Net brain. | ## 🔢 Parameters ### 🧾 `unet_name` - **Type:** `COMBO` - **Required:** Yes - **Example:** `flux1-dev.safetensors` - **Description:** This is the name of the diffusion model you want to load. It must exactly match a model file in your `models/diffusion` directory. No, “close enough” doesn’t count. - **Why it matters:** The selected model determines the quality, speed, and style of your output. Using a mismatched or unoptimized model is the fastest way to generate muddy messes. - **Pro tip:** Use model naming conventions that include version and architecture info (e.g., `wan2.1_t2v_14B_fp16`) to avoid confusion when you have 30+ models installed. ### 🧮 `weight_dtype` - **Type:** `COMBO` - **Required:** Optional (default = model native format) - **Options:** - `default` – Just use whatever dtype the model was trained with. It’s safe, boring, and reliable. - `fp8_e4m3fn` – Uses float8 precision with the e4m3fn format. Trades a bit of quality for a lot of VRAM savings. - `fp8_e4m3fn_fast` – Same as above but optimized for speed. Great for testing, low VRAM setups, or impatient people. - `fp8_e5m2` – Another float8 format with a different exponent/mantissa split. Sometimes faster, sometimes crankier. - **Why it matters:** Choosing the right dtype can drastically improve performance—especially if you’re riding the 8GB GPU struggle bus. But be warned: not all models or GPUs play nicely with all FP8 formats. - **What changing it does:** - Lower precision → faster load times and inference, possibly at the cost of detail and consistency - Higher precision → better results, higher VRAM, and increased GPU heat-related suffering ## 🔁 Outputs | Name | Type | Description | | ------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | | `MODEL` | `MODEL` | The loaded diffusion model. Pass this into nodes like `KSampler`, `Flux.1 Kontext Image Edit`, or anything else that expects a U-Net backbone. | This is a reference output—not an image, not a latent, not a prompt—it’s a full-blown pre-trained model. It doesn’t do anything on its own, but without it, nothing else works. (No pressure.) ## 📦 Function in ComfyUI Workflows The **Load Diffusion Model** node is often the **first stop** in any image or video generation workflow. Without it, there's no model to do the actual generative heavy lifting. Once loaded, the model gets handed off to downstream nodes like: - `KSampler` — for inference - `Ultimate SD Upscale` — for reprocessing - `Flux.1 Kontext` nodes — for fancy context-aware editing - Anything else that requires a `MODEL` input This node works whether you’re running ComfyUI locally, on a remote GPU via Colab, or in a full-blown cloud pipeline like ComfyAI. ## 🌍 Real-World Use Cases - **Cloud-Based Workflows**: Dynamically load models without having to redeploy the UI. - **Model Swapping Pipelines**: Swap models on the fly to compare `flux1` vs `wan2.1` vs your Frankenstein `sdxl_bastardized_lite_v3`. - **Video Frame Generation**: Load a lightweight model (with `fp8_e4m3fn_fast`) for better batch performance across sequences. - **Experimental Branches**: Fork workflows with different diffusion models side-by-side for A/B testing or chaos generation. ## ⚙️ Usage Tips - Start with `default` dtype unless you’re intentionally optimizing or debugging. - Don’t guess the model name. Copy it exactly from your filesystem or model manager. - Using FP8? Test stability across several seeds and prompt types before committing. - If you're comparing two models, lock the rest of your workflow (seed, steps, CFG, etc.) to keep your test fair. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - 🚫 Don’t point `unet_name` to a missing or renamed file unless you enjoy blank outputs and mysterious console errors. - 🚫 Don’t assume all models support all FP8 dtypes. Trial-and-error (with logging on!) is your best friend here. - 🚫 Don’t connect multiple diffusion models to the same `KSampler`. You’ll confuse the poor thing. - 🚫 Don’t forget that VRAM is finite. Loading a 14B model on your RTX 3060 may turn your machine into a toaster. - 🚫 Don’t use this node without understanding what model you’re using. _Seriously_, read the model card. ## ⚠️ Known Issues | Issue | Cause | Fix | | -------------------------- | ---------------------------------------------------------- | ----------------------------------------------------------- | | Model file not found | Typo in `unet_name` or file is in wrong directory | Check your path and spelling | | `weight_dtype` unsupported | Your GPU doesn’t support that FP8 format | Try `default` or another format | | Long load time | Loading from HDD or very large models | Move to SSD or reduce model size | | Crash during sampling | Model architecture doesn’t match prompt conditioning setup | Make sure your model matches your Clip/VAE/ControlNet stack | ## 🧪 Example Node Configuration | Field | Example | | -------------- | ----------------------- | | `unet_name` | `flux1-dev.safetensors` | | `weight_dtype` | `fp8_e4m3fn_fast` | This would load a Flux.1 diffusion model with optimized weight precision for fast and frugal generations. ## 📚 Additional Resources - ComfyUI Docs on Model Loading - List of available models - Official repo for Flux models ## 📝 Final Notes If you're generating images, you're using a diffusion model. And if you’re using a diffusion model in ComfyUI, you're _definitely_ using this node—whether you realize it or not. Pick your model carefully. Match your precision to your hardware. And whatever you do, don’t pretend this node is optional—it’s the literal engine of your workflow. --- ## Load Image (from Outputs) # Load Image (from Outputs) > _"Why click and drag when you can refresh and go?"_ ## 🧩 What is the Load Image (from Outputs) Node? The **Load Image (from Outputs)** node is your shortcut to sanity in iterative workflows. It automatically grabs the **first image file** from your output directory and loads it into your ComfyUI graph — optionally bringing along its mask if one exists. No manual file selection, no digging through folders, no drama. It’s a must-have when you’re stringing together multiple passes, re-processing images, or working in a remote/cloud-based setup like [ComfyUI on comfyai.run](https://comfyai.run) where manually dragging files into nodes is a productivity killer. ![Load Image (from Outputs)](/img/load-image-from-outputs.png) ## 🔧 Node Details | **Name** | `LoadImageOutput` | | ------------------- | ------------------------- | | **Display Name** | Load Image (from Outputs) | | **Category** | image | | **Module** | `nodes` | | **Outputs** | `IMAGE`, `MASK` | | **Is Output Node?** | ❌ Nope | | **Experimental?** | ✅ Yes | ## ⚙️ Parameters and Settings ### 🖼️ `image` | Type | Function | | ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Refresh Button / Selector | Clicking the refresh button updates the list of available images in the output folder. It will **always select the first image**, alphabetically or by order returned from the system. Want a different one? Rename it to come first. Or delete the others. | This field doesn’t take an input from another node — it just scans the output folder. You’re not plugging anything in here except your _hopes and dreams of automation_. ## 📤 Outputs | Output | Type | Description | | ------- | ----- | ------------------------------------------------------------------------------------------------------------- | | `IMAGE` | Image | The first image file found in your output folder. It's now in your workflow. You’re welcome. | | `MASK` | Mask | If there’s a matching mask file with the same base name? You get it. If not? You get `None`. Don’t be greedy. | ## 📁 Function in ComfyUI Workflows **Load Image (from Outputs)** is the lazy genius node that pulls the _latest_ (aka first-listed) output image into your workflow. This makes it perfect for: - 🔁 **Iterative workflows**: e.g., generate → upscale → edit → generate again. - ☁️ **Cloud-based usage**: stop uploading/download files — just let the node grab what’s there. - 🛠️ **Automation and scripting**: particularly when chained with other nodes to reduce manual inputs. Want to refine the same image repeatedly without reloading it every time? This is your node. ## 🧠 Real-World Use Cases - **Editing cycles**: Feed your output image right back into your graph for touch-ups, inpainting, or remixing. - **Upscaling pipelines**: Auto-pull the latest render into an upscale node. - **ControlNet conditioning**: Automatically feed in your newest generated base image for conditioning. - **Server-based rendering**: Perfect for environments where UI interaction is minimal or file handling is abstracted. ## 🚫 What-Not-To-Do-Unless-You-Want-a-Fire - 🔥 **Don’t assume it’s grabbing the “latest” file by date** — it’s grabbing the _first listed_. You’ve been warned. - 🔥 **Don’t forget to hit the refresh button** — it doesn’t auto-magically update just because you want it to. - 🔥 **Don’t expect masks to magically appear** — they need to exist and have matching filenames. - 🔥 **Don’t use this for batch loading** — it pulls **one** file. For loops or batch runs, look elsewhere (or build a macro/script). - 🔥 **Don’t try to load non-image files** — it’s not your intern. It won’t guess. ## 🛠️ Troubleshooting & Known Issues | Problem | Likely Cause | Fix | | -------------------------------- | -------------------------------------------------- | -------------------------------------------------------------------------------------------------- | | **Image not loading** | File doesn’t exist or isn’t in the right directory | Hit refresh. Double-check that your output folder is set correctly. | | **Wrong image loaded** | First image isn’t the one you wanted | Rename files or clean up the output folder. | | **Mask not appearing** | No corresponding mask exists | Make sure your generation workflow actually outputs masks with the same base filename. | | **Refresh not working (online)** | Browser cache or remote dir syncing issue | Reload your ComfyUI session. Refresh your browser tab. Whisper sweet nothings to your file system. | ## 💡 Usage Tips - Rename your files smartly if you want predictable selection. - Use this node as the **input anchor** for post-processing chains (upscalers, inpainting, ControlNet). - Combine with `Image Save` + `Ultimate SD Upscale` or `Flux.1 Kontext` for multi-step automation. - To process the **same image repeatedly**, don’t keep regenerating — just rewire it with this node. ## 📝 Example Workflow Setup ```mathematica Load Image (from Outputs) │ ▼ ControlNet Preprocessor → ControlNet Apply │ │ ▼ ▼ KSampler → Save Image (outputs here again) ``` Every time you hit refresh, the new image becomes the new input. No dragging. No clicking. No problem. ## 📚 Additional Resources - [ComfyUI GitHub](https://github.com/comfyanonymous/ComfyUI) – for core updates - [comfyai.run](https://comfyai.run) – online ComfyUI environment where this node is _super helpful_ --- ## Load LoRA # Load LoRA ## 🧠 What This Node Does The `Load LoRA` node in **ComfyUI** is your go-to for injecting additional style, character, costume, or context into your image generation without retraining your base model. LoRA (Low-Rank Adaptation) models are lightweight, fine-tuned layers that can be dynamically applied to the base Stable Diffusion model at runtime. This node lets you load one and set how strong it should be — for both the _model weights_ and the _CLIP text encoder_. No more bloated model folders or checkpoint-switching gymnastics. With this node, you can stack, mix, and fine-tune LoRA styles like a digital alchemist. ![Load LoRA](/img/load-lora.png) ## 🔧 Node Settings and Parameters (in **painstaking** detail) ### 🔹 `lora_name` (Combo) - **Type**: Dropdown list / file selector - **Purpose**: Selects which LoRA `.safetensors` file to load from your configured `ComfyUI/models/lora` directory. - **Required**: Yes. Without this, you're doing nothing but wasting node real estate. **What it does**: Loads the LoRA file so its learned weights can be applied to the MODEL and optionally to CLIP. You’ll only see LoRAs that are placed in your designated LoRA folder. **Tip**: LoRAs often have naming conventions like `style_v1.0.safetensors` or `outfit-swap-topV2.safetensors` — pay attention to versions and suffixes if you’ve got a cluttered folder. ### 🔹 `strength_model` (Float) - **Type**: Slider or manual input (usually 0.0 to 1.0, but technically can go higher) - **Default**: 1.0 - **Purpose**: Controls how strongly the LoRA is applied to the model weights (the UNet). **What it does**: - At `1.0`, you’re applying the LoRA exactly as it was trained. - At `1.0`, you’re going into _"LoRA overload" territory_ — which can create interesting or downright cursed images. Experiment with caution. **Why it's important**: This controls how much of the visual representation of the LoRA gets expressed. Things like character identity, costume, pose influence, or scene lighting could be stronger or weaker depending on this setting. ### 🔹 `strength_clip` (Float) - **Type**: Slider or manual input (usually 0.0 to 1.0, but like above, can go over) - **Default**: 1.0 - **Purpose**: Controls how strongly the LoRA influences the CLIP text encoder. **What it does**: - At `1.0`, the LoRA's textual association is fully applied to your prompt — so if the LoRA includes baked-in prompt logic (like turning "cyber dress" into a whole aesthetic), it'll show up strongly. - At ` _Or: How to summon an eldritch horror from your LoRA folder by accident._ #### ❌ Crank `strength_model` to 10 Unless you’re trying to create an abomination that looks like it fell through an AI meat grinder, keep your `strength_model` values under control. Values over **1.5** often lead to overbaked noise, facial meltdowns, and what I like to call "_LoRA hallucinations_." Basically: you told the model to do too much and it panicked. #### ❌ Set `strength_clip` to 0 without knowing what it does This disables the CLIP side of the LoRA — which might be what you want... unless the LoRA relies on textual influence to work correctly. A LoRA for "goth bunny assassin" might just look like a blurry mess without the CLIP component helping the model interpret your prompt. #### ❌ Forget to use the LoRA’s trigger word Many LoRAs require specific **activation keywords** to function (e.g., `cyberdress`, `tanktop`, `angelcore_v1`). If your prompt doesn't contain the trigger, the LoRA might just sit there doing nothing, judging you silently. #### ❌ Mix LoRAs that were never meant to be mixed You _can_ load multiple LoRAs, but stacking "realistic portrait v2," "anime body horror v9," and "pixel art cave troll" might just cause a visual nervous breakdown. Unless you're doing this for the memes, avoid wildly incompatible styles. #### ❌ Load an SDXL LoRA into an SD1.5 workflow It won’t work. It _shouldn't_ even load properly — but if you force it with dark magic or manual renaming, expect nothing but digital static. LoRAs are trained for specific base models. Cross-version contamination is a fast track to wasted time and terrifying outputs. #### ❌ Use high-strength LoRA with highly detailed conflicting prompts Prompts like `"hyper-realistic woman in a power suit, cinematic lighting, gothic castle"` combined with a LoRA that adds `"manga catgirl swimsuit beachcore"` at `strength_model=1.2` is... asking for an identity crisis. Don’t expect harmony from chaos. #### ❌ Forget to match your LoRA’s style with the right base checkpoint A stylized LoRA meant for `revAnimated_v122` won’t look right when paired with `deliberate_v11` or `dreamshaper_8` if their base styles conflict. If you’re using a LoRA that was clearly made for a cartoony model, don’t slap it onto a realism checkpoint unless you’re fine with the Uncanny Valley getting a new zip code. #### ❌ Assume LoRA = Insta-Magic LoRAs _enhance_ generation — they don’t override bad prompts, broken seeds, or mismatched models. If your base setup is junk, the LoRA won’t save you. It’ll just add stylish failure. Keep it sane. Keep it smooth. Respect the LoRA. Or do all of the above and start your own AI horror gallery — your call. ## 🧪 Advanced Nerd Stuff (Optional but Cool) - You can technically swap LoRA files mid-workflow using `LoRA Stack` logic or automation — but that’s for more advanced setups. - Some users use math nodes to dynamically control `strength_model` and `strength_clip` for evolving animations or staged variations. ## 🧼 Final Thoughts The `Load LoRA` node is one of the most powerful and underappreciated nodes in ComfyUI. It gives you the ability to radically transform your outputs without needing to bake in permanent changes to your checkpoint. Use it wisely — or go full chaos gremlin and push it to 2.5 for memes. Either way, experiment, break things, and enjoy the ride. --- ## Load Style Model # Load Style Model > _Because every masterpiece deserves a sense of style—your style._ The `Load Style Model` node is a straight-to-the-point, no-nonsense loader that brings a specific `.safetensors` style model into your ComfyUI workflow. If you’re trying to turn a boring output into something with _actual_ aesthetic value (you know, like a visual personality), this is the node you slap in at the top of your pipeline. Perfect for digital artists, style-transfer addicts, and anyone tired of manually fiddling with prompts just to get consistent vibes across renders. Whether you're running ComfyUI locally or from a cloud service like [comfyai.run](https://comfyai.run), this node makes styling efficient, repeatable, and dangerously fun. ![Load Style Model](/img/load-style-model.png) ## 🔧 Node Information | **Property** | **Value** | | ---------------- | ------------------ | | **Node Name** | `StyleModelLoader` | | **Display Name** | `Load Style Model` | | **Category** | `loaders` | | **Module** | `nodes` | | **Output Node** | No | | **Output** | `STYLE_MODEL` | ## ⚙️ Inputs ### 🔹 `style_model_name` - **Type**: `list[str]` (but usually just one entry) - **Required**: Yes - **Example**: `["flux1-redux-dev.safetensors"]` - **Description**: This is the name of the style model you want to load. The file must be located in your designated style model directory (usually `ComfyUI/models/style_models`). If you're running ComfyUI in a cloud instance, make sure the model is uploaded to your workspace. #### 💡 What It Does Loads the specified `.safetensors` (or `.ckpt`, but we recommend `.safetensors`) style model and prepares it for use in any node that consumes `STYLE_MODEL`. ## 📤 Outputs ### 🔸 `STYLE_MODEL` - **Type**: `style_model` object - **Description**: This is your freshly loaded style model, now in ComfyUI’s bloodstream. Use it as input to any compatible node, such as Flux-Kontext style transfer nodes, style-conditioning modules, or any pipeline that wants to add some visual seasoning. ## 📁 Real-World Use Cases - 🎨 **Digital Art Pipelines** – Feed a consistent style model into all your renders to maintain visual continuity. - 🧪 **Rapid Style Testing** – Drop this into your experimentation workflow to hot-swap styles mid-process. - 🖼️ **Image-to-Image Stylization** – Combine with latent input + prompts to stylize existing images. - 💻 **ComfyUI Cloud Workflows** – Load from a shared model space and use across multiple projects or teams. ## ✅ Usage Tips - Put this node _before_ any nodes that expect `STYLE_MODEL` input (obviously). - Don’t waste time styling if your final sampler settings use high denoise and destroy everything. Lower `denoise` = more style preserved. - Not every model is made equal—some style models are extremely aggressive (think LSD-on-canvas). Test and adjust. - Combine with `Flux.1 Kontext Image Edit` or `FluxGuidance` nodes for targeted editing with style. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire - ❌ **Don't use a filename that doesn't exist** — You’ll get a big fat “model not found” error. You’ve been warned. - ❌ **Don't feed multiple style models unless the receiving node supports it** — Most don't. Chaos ensues. - ❌ **Don’t expect this to apply style magically** — This just loads the model. You still need a style-aware node to _do_ anything. - ❌ **Don’t overload your VRAM with 6GB+ style models if you're using a toaster GPU** — Unless you're into crashes and regret. ## ⚠️ Known Issues - Some style models aren’t compatible with certain samplers or schedulers. Trial-and-error is your friend. - Style models trained for SD1.5 won’t always work well with SDXL or vice versa. Format matters. - On comfyai.run or other cloud UIs, model loading may fail if the model hasn't been added to your workspace first. ## 🧪 Example Node Configuration ```json { "id": 12, "type": "StyleModelLoader", "pos": [220, 260], "size": [320, 60], "flags": {}, "inputs": [ { "name": "style_model_name", "type": "COMBO", "widget": { "name": "style_model_name" }, "link": null } ], "outputs": [{ "name": "STYLE_MODEL", "type": "STYLE_MODEL", "links": [42] }], "properties": { "Node name for S&R": "StyleModelLoader", "style_model_name": ["flux1-redux-dev.safetensors"] } } ``` ## 🌐 Related Nodes | **Node** | **Purpose** | | --------------------------- | -------------------------------------------------------------------- | | `Load Texture Model` | Loads texture maps (not styles) for detailed material outputs. | | `Load Shader Model` | Loads custom shaders, typically used in 3D or stylized UI rendering. | | `Flux.1 Kontext Image Edit` | Applies prompts and styles directly to an image in latent space. | | `FluxGuidance` | Amplifies or attenuates style/conditioning weight. | ## 📝 Final Notes The `Load Style Model` node might seem simple, but it’s the gateway to deeply personalized, stylized workflows. Whether you're painting with pixels or pumping out concept art for your next big project, this is the node that makes sure your outputs _don’t all look like they were generated in Microsoft Paint._ Now go load up that custom style model and make your outputs fabulous. --- ## Load VAE # Load VAE Welcome to the magical land of compression and decompression, also known as the **Load VAE** node in ComfyUI. This node plays a vital role in your workflow by loading a **Variational Autoencoder (VAE)** — the component responsible for translating between the cozy latent space of your model and the gloriously noisy pixel soup we call an image. If your generated images are looking a little too “potato-cam” or you're getting weird color artifacts, chances are you're either not using a VAE or you're using the _wrong_ one. Let’s fix that. ## 🧠 What is a VAE? A **Variational Autoencoder (VAE)** is a type of neural network trained to compress and decompress image data, forming the bridge between the latent representation used by your model and the full-resolution image. It influences fine details like color tone, contrast, and sharpness — so yes, it _definitely matters_ which VAE you use. VAEs are checkpoint-specific most of the time. Using the wrong one? You'll get color shifts, blotchy noise, or all-around uncanny weirdness. So treat it like pairing wine with cheese — compatibility is key. ![Load VAE](/img/load-vae.png) ## 🧱 Node Type: `VAELoader` ### Purpose: Load a `.vae.pt` or `.safetensors` file and return a VAE object to be used in your generation pipeline. ## 🔌 Node Inputs and Outputs | Input | Type | Description | | ------ | ---- | -------------------------------------------------------------------------------------------------- | | _None_ | | This node doesn’t require any incoming connections. It independently loads the specified VAE file. | | Output | Type | Description | | ------- | ----- | -------------------------------------------------------------------------------------------------- | | **VAE** | `VAE` | Emits the loaded VAE object to be passed into your Checkpoint Loader or other nodes needing a VAE. | ## ⚙️ Node Parameters | Parameter | Type | Description | | ------------ | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | **vae_name** | `Combo/Dropdown` | The name of the VAE file to load. This dropdown is populated based on VAE files available in the ComfyUI VAE directory. If the file is missing, you’ll see red error messages and likely end up in grayscale hell. | ## 📁 Where Do I Put VAEs? By default, ComfyUI looks for VAEs in: bash ``` ComfyUI/models/vae/ ``` Accepted file types: - `.vae.pt` - `.safetensors` Make sure your files are properly named and located in that folder, or they won’t show up in the dropdown. Yes, capitalization matters. No, ComfyUI won’t guess what you meant. ## ✅ Recommended Workflow Setup Here’s where the **Load VAE** node fits into a typical text-to-image setup: ```pgsql [Load Checkpoint] └── MODEL ──> [KSampler] └── CLIP ──> [CLIPTextEncode] └── VAE ←── [Load VAE] ``` If you're not connecting the VAE output from **Load VAE** to your Checkpoint Loader or KSampler, you’re either: 1. Using a baked VAE (built into the checkpoint), 2. Or you forgot — in which case, prepare for color sadness. ## 🎯 Use Cases - **Color Correction**: Fix overblown skin tones, desaturation, or contrast weirdness. - **Detail Preservation**: Get crisper edges, richer highlights, and less murky textures. - **Model Customization**: Match VAEs tailored for checkpoints like DreamShaper, AnythingV5, PonyRealism, etc. - **Style Control**: Some VAEs affect the "softness" or "sharpness" of the final output, which is handy for stylized generations. ## 🛠 Prompting Tips While prompting isn’t directly affected by the VAE, your _results_ definitely are: - If you're seeing ghostly color overlays or washed-out details, try a different VAE. - Combining LoRA models or hypernetworks? Use the VAE recommended for the **base checkpoint**, not the LoRA. - If you're stacking VAEs and wondering why it's not working — you're not supposed to. One VAE per pipeline, please. ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire So, you like chaos? Great. Here’s how to completely ruin your workflow using the Load VAE node: #### 🔥 1. **Load the Wrong VAE for Your Checkpoint** You wouldn’t put diesel in a Tesla, so don’t pair an anime VAE with a realism checkpoint. Best case? Your model looks like it took a bath in color bleach. Worst case? Your output will haunt you in your dreams. #### 🔥 2. **Leave the VAE Disconnected** This is the equivalent of shouting into the void. You loaded the VAE... cool. But if you don’t _connect it_ to your Checkpoint Loader, it just sits there. Unused. Mocking you. Your output will default to the baked-in VAE or — gasp — no VAE at all. Welcome to grayscale hell. #### 🔥 3. **Manually Rename or Move VAE Files Without Updating ComfyUI** ComfyUI doesn’t have telepathy. If you renamed a VAE to “definitely_not_cursed.vae.pt” and it disappears from the dropdown, that’s your fault. Don’t @ me. #### 🔥 4. **Stack VAEs or Use Multiple in One Workflow** No. Stop. VAEs are not seasoning. You can’t just sprinkle multiple in and expect magic. ComfyUI uses _one_ VAE per pipeline. Pick one. Commit. #### 🔥 5. **Mix VAE Versions from Different Model Architectures** Some VAEs are trained for SD 1.5. Others are for SDXL. Mixing those? That’s like using a Game Boy charger on a microwave. It won’t work, and it might catch fire — metaphorically (but who knows with enough VRAM). #### 🔥 6. **Expect VAEs to Magically Fix a Bad Prompt** VAEs affect the way images are decoded — not your prompt's creativity deficit. If your image still looks like AI-generated oatmeal, the VAE probably isn’t the issue. It's you. Yes, I said it. ## 🧪 Best Practices | Scenario | VAE Strategy | | ------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | | Using standard SD1.5 checkpoints | Use `vae-ft-mse-840000-ema-pruned.vae.pt` for general fidelity. | | Using highly stylized or anime models | Use a VAE trained with similar style data, like `anything-v4.0.vae.pt`. | | Generating photo-realistic images | Use realism-optimized VAEs (e.g., EpicRealism or PonyRealism variants). | | Don’t know? | Stick with the VAE recommended by your checkpoint’s author or try a few common ones until your image looks normal again. It’s trial-and-error season, baby. | ### 📚 Additional Resources - [Check out our vae-ft-mse-840000-ema-pruned.safetensors documentation!](https://comfyui.dev/docs/guides/VAE/vae-ft-mse-840000-ema-prumed-safetensors) ## 💬 Final Thoughts The **Load VAE** node is often overlooked, but it’s a sneaky little gremlin that can ruin or rescue your output quality. Use it wisely, match it to your checkpoint, and stop blaming your prompts for everything. Sometimes it's just a bad VAE. If your images still look cursed after swapping VAEs, then yes — _now_ it's probably your prompt. --- ## ModelSamplingFlux # ModelSamplingFlux _Because sometimes “just letting the model do its thing” isn’t good enough._ --- ## 🧠 What Is This Node? The **ModelSamplingFlux** node is a precision tool in ComfyUI designed to let you _grab the wheel_ of your model’s sampling process and steer it toward exactly the kind of results you want. Instead of relying solely on the model’s default behavior, you can directly adjust **shift values** and **dimensions** to fine-tune how the model interprets and generates output. Think of it as a way to whisper in your model’s ear, “Yes, do that… but also, maybe, a little more this way.” ![ModelSamplingFlux](/img/modelsamplingflux.png) ## 💡 Real-World Use-Cases This node is perfect for scenarios where even the smallest sampling tweaks can lead to massive changes in visual output quality or style: - **High-precision image generation** – Nudge outputs toward a more cohesive style or composition. - **Experimental art workflows** – Create subtle or extreme variations by changing shift ratios. - **Matching reference proportions** – Enforce width/height for consistent batch renders. - **Controlled randomness** – Adjust max/base shifts to allow for variability without losing structure. ## ⚙️ Node Inputs & Parameters Below is a **detailed breakdown** of every field, what it expects, why it exists, and what happens if you decide to “see what happens” when you max it out. | Field | Type | Example | Description | | -------------- | ----- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | **model** | Model | `MyModel` | The **primary model** this node will adjust. Must be a compatible diffusion/Flux-based model. Passing the wrong type here will result in errors, weird results, or both. | | **max_shift** | Float | `1.15` | The **maximum allowed adjustment** to the sampling process. Higher values mean greater deviation from the model’s default path. Too high and you might get unplanned chaos. | | **base_shift** | Float | `0.5` | The **starting offset** for the sampling shift. Works in tandem with `max_shift` to control overall sampling drift. Set it to 0 for pure default sampling; raise it to push the model toward more variation. | | **width** | Int | `1024` | The **horizontal resolution** (in pixels) for sampling. This doesn’t just affect output size—it can change detail distribution and resource usage. | | **height** | Int | `1024` | The **vertical resolution** for sampling. Must be compatible with your model’s architecture (multiples of 8/64 depending on the model). | ## ⚙️ Node Output | Output | Type | Description | | --------- | ----- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **MODEL** | Model | A modified version of your input model with sampling behavior adjusted according to your parameters. This is now ready to be used in downstream nodes like `KSampler`, `FluxGuidance`, or `Flux.1 Kontext Image Edit`. | ## 🛠️ Workflow Setup & Best Practices **Recommended Position in Workflow:** Place **ModelSamplingFlux** after loading your base model but **before** running any samplers or guidance nodes. This ensures all subsequent nodes operate on your fine-tuned sampling rules. **Typical Setup Flow:** 1. **Load your model** (`Load Diffusion Model` or equivalent). 2. **Apply ModelSamplingFlux** with your chosen shift/dimension settings. 3. **Feed into a sampler** (`KSampler`, `FluxGuidance`, etc.). 4. **Render your output** with desired prompts and conditioning. ## 🧪 Prompting Tips - If you’re increasing **max_shift**, consider adding more descriptive prompt details to keep the image coherent. - Large **width/height** combos with big shifts can result in “artistic” glitches—sometimes gold, sometimes garbage. - Use **base_shift** as a _subtle steering wheel_ rather than a sledgehammer. ## 🔥 What-Not-to-Do-Unless-You-Want-a-Fire - Setting `max_shift` absurdly high without a reason. This can result in outputs that look like your model had a nervous breakdown. - Forgetting your model’s **dimension constraints**. Many models expect dimensions in multiples of 8 or 64. Ignoring this will produce errors. - Using incompatible models. Flux-specific nodes may not play nicely with every checkpoint—test before committing to a workflow. ## ⚠️ Known Issues - **Over-shifting** can cause loss of subject fidelity. - Certain combinations of width/height may **increase VRAM usage** significantly. - Not all model architectures respect shift adjustments equally—results can vary between checkpoints. ## 📝 Final Notes The **ModelSamplingFlux** node is not about replacing your sampler—it’s about _prepping_ your model to respond differently **before** the sampler even does its thing. Think of it as setting the stage: the sampler is the actor, but ModelSamplingFlux is the director telling them which mood to bring. If you want predictable, repeatable results—keep shifts low. If you’re chasing creative chaos—crank them up and see what happens. --- ## RandomNoise # RandomNoise _Because sometimes your art needs a little chaos — but with receipts._ --- ## 🧠 What Is This Node? The **RandomNoise** node in ComfyUI exists to do exactly what it says on the tin: generate random noise. But it’s _controlled_ random noise — think “chaotic energy but on a leash.” By feeding it a **noise seed**, you can dictate exactly which unique noise pattern will be generated. Use the same seed, and you get the same noise every time. Change the seed, and the randomness changes — but in a way that’s still reproducible for later. This node is a staple in AI image generation pipelines because noise is the raw clay from which diffusion models sculpt your masterpiece. Without it, you’ve got nothing but empty latent space (and a very confused model). ![RandomNoise](/img/randomnoise.png) ## 🧩 Primary Purpose in Workflows - Initialize a latent space with structured randomness. - Ensure _repeatable randomness_ for experiments. - Provide a reproducible starting point for diffusion, upscaling, or other latent-space operations. - Make your workflow feel like a controlled science experiment, even if it’s actually artistic chaos. ## 🔌 Input Parameters ### **`noise_seed`** _(integer)_ - **Purpose:** Determines the exact random noise pattern generated. - **Acceptable Range:** `0` to `0xffffffffffffffff` (which is `18,446,744,073,709,551,615` for the mathematically curious). - **Default:** `0` (which is basically the “meh, just give me something” option). - **Behavior:** - Same seed = identical noise every time. - Different seed = new noise pattern (within the bounds of the same resolution & latent space). - **Why It Matters:** - Reproducibility is critical if you want to A/B test your prompt, sampler, or scheduler settings without the noise pattern changing. - If you _don’t_ care about reproducibility and want every run to be unique, just randomize this value. 💡 **Pro Tip:** If you’re doing batch renders and want _slightly_ different results while keeping composition similar, start with a fixed seed and increment it between runs. ## 📤 Output Parameters ### **`noise`** _(tensor)_ - **Type:** Latent noise tensor. - **Shape:** Matches your model’s expected latent dimensions (commonly `[batch_size, channels, height, width]` in latent space). - **Content:** The generated noise pattern dictated by your seed. - **Use Cases:** - Feed into a sampler node (e.g., KSampler, SamplerCustomAdvanced). - Combine with image-to-latent conversions to create variation in outputs. - Layer in additional control inputs like ControlNet to steer noisy chaos into structured beauty. ## ⚙️ Workflow Setup & Best Practices 1. **Basic Setup:** - Place **RandomNoise** before your chosen sampler. - Make sure the noise matches the resolution/latent space of your model. 2. **Reproducibility Runs:** - Lock your `noise_seed` to a single value while tweaking other parameters. 3. **Exploratory Runs:** - Randomize or increment `noise_seed` to explore variations without touching your main prompt. 4. **Paired Experiments:** - Use identical noise across different models to compare how each handles the same “raw material.” ## 🧪 Prompting & Artistic Tips - Keep seed locked when testing subtle prompt wording changes — otherwise, you’re testing _both_ prompt and noise. - Want structured chaos? Combine RandomNoise with a ControlNet preprocessor. - For animation workflows, keep seeds consistent frame-to-frame for smoother transitions (unless you _want_ visual jitter). ## 🔥 What-Not-to-Do-Unless-You-Want-a-Fire - **Don’t** feed noise tensors into nodes expecting fully rendered images. They’ll either throw errors or produce abstract art you didn’t ask for. - **Don’t** set an invalid seed (outside the acceptable range) — you’ll get errors faster than you can say “uint64.” - **Don’t** assume the noise output is scaled for RGB. This is _latent_ noise, not displayable pixel data. ## ⚠️ Known Issues & Quirks - Large seeds can be valid but may feel unpredictable if you’re used to smaller seed numbers. - Different models at different resolutions will produce completely different results, even with the same seed — because resolution changes the noise tensor’s dimensions. - If your workflow resolution changes mid-run (don’t ask why), you’ll need to regenerate the noise to match the new size. ## 📝 Final Notes The RandomNoise node is a bread-and-butter utility for anyone doing procedural generation in ComfyUI. Used properly, it’s a powerful tool for creative control. Used improperly, it’s just… well, noise. --- ## SamplerCustomAdvanced # SamplerCustomAdvanced **Category:** Sampling / Advanced **Module:** `comfy-core` **Outputs:** `LATENT`, `DENOISED_LATENT` **Author:** ComfyUI Core Team **Specialty:** Custom image sampling using external noise, guider, sampler, sigmas, and latent inputs. For when “standard” just won’t cut it. ## 🧠 What Does It Do? The `SamplerCustomAdvanced` node is your go-to when the built-in samplers feel a bit too… basic. This node gives you the ability to drive the sampling process _manually_ using external inputs like a noise tensor, guider module, sigmas array, and more — making it a powerhouse for users building high-end, fine-tuned pipelines. If you’re looking to customize how denoising and transformation are executed during image generation or refinement — this is your precision scalpel. ![SamplerCustomAdvanced](/img/samplercustomadvanced.png) ## 💡 Real-World Use Cases - **Complex image transformations**: Apply external guidance (like edges, masks, depth, or other conditioning) to latent images for more intelligent sampling. - **Custom denoising flows**: Inject your own noise tensor and control how it's reduced using sigmas and sampling logic. - **High-clarity outputs**: Preserve sharp features and structural integrity using advanced guidance. - **Online ComfyUI tuning**: Ideal for setups like **ComfyUI Cloud** where GPU time is precious and control is key. ## 🔌 Required Inputs Each of these fields is **mandatory**, and yes, the node will politely fall apart if you try to skip one. | **Input** | **Type** | **Description** | | -------------- | --------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | `noise` | `NOISE` | This is your randomness source — typically a Gaussian noise tensor. Must match the dimensions of the latent input. | | `guider` | `GUIDER` | This controls how the sampling is directed. Used to guide the transformation in a structured way (e.g., preserving edges, shapes, etc.). | | `sampler` | `SAMPLER` | Defines the core algorithm that controls how steps are taken through latent space. Often created using `SamplerCustom` or similar. | | `sigmas` | `SIGMAS` | A float tensor that determines the noise levels per step — basically how “loud” or “quiet” your sampling should be at each iteration. | | `latent_image` | `LATENT` | The image data you’re transforming. This is the canvas your sampler will iterate on. Usually output from a prior generation or upscaler node. | > **Important:** All inputs must be dimensionally and semantically compatible — no mismatched tensor shapes or you’re headed straight into the red error zone. ## 📤 Outputs | **Output** | **Type** | **Description** | | ----------------- | -------- | ------------------------------------------------------------------------------------------------------- | | `output` | `LATENT` | The final latent result after full sampling. Usually passed into a `VAE Decode` or further refinement. | | `denoised_output` | `LATENT` | The denoised latent halfway through the sampling process. Great for analysis, preview, or reprocessing. | ## 📈 Example Flow ```plaintext [EmptyLatentImage] ↓ [Noise] ↓ [Guider] ↓ [SamplerCustomAdvanced] ↓ [VAEDecode] ``` This would give you a full pipeline where you're crafting the sampling flow yourself. Bonus points if you route the `denoised_output` into a preview panel or another editing pass. ## 🛠️ Parameter Reference | **Field** | **Type** | **Required?** | **Notes** | | -------------- | -------- | ------------- | ------------------------------------------------------------------------------- | | `noise` | Tensor | ✅ Yes | Must match the latent shape. Gaussian noise is typical. | | `guider` | Module | ✅ Yes | Use something like `EdgeGuidance`, `ControlNet`, or similar. | | `sampler` | Sampler | ✅ Yes | Generated via `SamplerCustom`, `SamplerCustomAdvanced`, or `KSamplerSelect`. | | `sigmas` | Tensor | ✅ Yes | Custom sigma schedule for noise reduction. Must match expected sampling length. | | `latent_image` | Tensor | ✅ Yes | Your input image in latent form. | ## 🧯 Common Errors & Troubleshooting | **Error** | **What It Means** | **Fix** | | -------------------------------- | --------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | | `Shape mismatch between tensors` | One or more inputs (like `noise` or `latent`) aren’t aligned dimensionally. | Double-check the size of each input. Use matching resolution latent and noise tensors. | | `Unsupported guider type` | You’ve connected a guider that isn’t compatible with this sampler. | Use a guider node explicitly designed for this setup. | | `Poor image quality` | Sigmas or noise configuration might be over-suppressing or under-suppressing critical detail. | Try adjusting the sigmas schedule or noise input — this directly impacts denoising. | | `Lag on cloud platform` | Processing is too heavy for your cloud setup (e.g., ComfyUI Run). | Reduce image resolution, or avoid high sigma schedules with complex guider modules. | ## 🔍 Related Nodes & Key Differences | **Node** | **What It Does** | | --------------- | -------------------------------------------------------------------------------- | | `Sampler` | A simpler version for typical sampling tasks. Great for plug-and-play scenarios. | | `SamplerCustom` | Similar, but doesn’t expose the denoised latent. Less precise, more streamlined. | | `Noise` | Use this to generate compatible noise inputs. | | `Guider` Nodes | Like `EdgeGuider`, `ControlNetGuider`, or `FluxGuidance`. | ## 💬 Final Notes The `SamplerCustomAdvanced` node is not for casual dabblers. It’s meant for power users who want to engineer every stage of the image generation process, down to the noise tensor. If you’ve ever found yourself saying, “I wish I had more control over how this gets denoised,” this is your tool. And hey — once you’ve mastered this, you’re basically halfway to writing your own sampler backend. So congrats on leveling up. --- ## Ultimate SD Upscale # Ultimate SD Upscale Welcome to the Swiss Army Knife of image upscaling in ComfyUI. The **Ultimate SD Upscale** node is a tile-based upscaling and enhancement tool that cleverly slices images into manageable chunks, passes each through a Stable Diffusion pipeline, and then stitches them back together—all while doing its best not to leave them looking like a poorly glued jigsaw puzzle. This node is designed for detailed enhancement of large images using Stable Diffusion's text-to-image (or img2img) capabilities. It works best when you want to upscale an image, add detail or variation, or reprocess large generations that standard pipelines can’t handle due to VRAM constraints. ![Ultimate SD Upscale](/img/ultimate-sd-upscale.png) ## 🧠 What This Node Does - Slices your image into tiles. - Optionally applies a mask and blur around tile edges to reduce seams. - Upscales each tile individually through the SD pipeline using text prompts. - Reconstructs the image from the tiles, applying optional seam correction. - Handles large image upscaling while letting you apply prompt-based enhancement at scale. ## 🧩 Inputs & Outputs | **Port** | **Type** | **Description** | | ---------- | ------------ | ----------------------------------------------------- | | `image` | Image | The input image to be upscaled/enhanced. | | `model` | Model | Loaded Stable Diffusion model to use. | | `clip` | CLIP | Corresponding CLIP model for prompt embedding. | | `vae` | VAE | Optional, for decoding outputs. | | `positive` | Conditioning | Your prompt in encoded form (CLIP Text Encode). | | `negative` | Conditioning | Your negative prompt (undesired features). | | `mask` | Image | Optional image mask (for applying masked generation). | ## ⚙️ Settings (with Extreme Detail) ### **`upscale_by` (float, default: 2.0)** How much to enlarge the original image by. - **What it does**: Multiplies both width and height by this factor. - **Recommended range**: 1.5–4.0 - **Tip**: Going above 2.5–3.0 will drastically increase VRAM use and tile count. Don’t say I didn’t warn you. ### **`seed` (int, default: -1)** Seed for deterministic generation. - `-1`: Random seed. - Any positive int: Makes results reproducible. - Useful when testing prompt effects or creating consistent variations. ### **`control_after_generate` (enum: fixed, increment, decrement, randomize)** Controls how the seed behaves per tile. - `fixed`: Same seed for every tile = consistent style. - `increment`: Adds 1 per tile. Slightly changes each tile—great for subtle detail variety. - `decrement`: Subtracts 1 per tile. Same idea, different direction. - `randomize`: Each tile gets a new random seed = chaotic beauty or absolute disaster. ### **`cfg` (float, default: 7.0)** Classifier-Free Guidance Scale. - Controls how much the model "listens" to your prompt. - Lower: More randomness, possibly artistic. - Higher: More prompt-locked and rigid. - Typical range: 5–12. Go beyond 13 and you’re just yelling at the model. ### **`sampler_name` (string)** Choose your sampling algorithm. - Works just like in the KSampler node. - Use DPM++ 2M or DPM++ SDE for clean details, or Euler a for speed. - Combine with proper scheduler for best results. (See KSampler documentation for in-depth sampler/scheduler combos.) ### **`scheduler` (string)** How noise is scheduled during denoising. - Normal, Karras, Exponential, and more. - Karras tends to produce smoother outputs in upscaling contexts. ### **`denoise` (float, 0.0–1.0)** Amount of noise to reintroduce. - 0.0 = no change to input tiles - 1.0 = generate completely new image based on prompt - For upscaling: 0.2–0.5 is usually the sweet spot. - For full re-prompting: crank it up. ### **`model_type` (enum: linear, chess, none)** Defines tile stitching behavior. #### `linear`: - Tiles are processed row-by-row (left to right, top to bottom). - **Pros**: Predictable behavior, easier to debug. - **Cons**: May have visible seams if prompt isn’t strong or CFG is low. #### `chess`: - Alternates tile processing like a checkerboard. - **Pros**: Seam reduction, since surrounding tiles aren’t processed back-to-back. - **Cons**: Slightly longer to process, less predictable randomness. #### `none`: - Processes all tiles independently. - **Pros**: Absolute freedom for chaos. - **Cons**: Highest chance of seams and tile dissonance. Best used with strong prompts and denoise. ### **`tile_width` & `tile_height` (int, default: 512)** Size of the tile chunks in pixels. - Must be divisible by 8. - Affects VRAM usage and performance. - **Smaller tiles** = better seam handling, but slower. - **Larger tiles** = fewer seams, faster, but harder on memory. ### **`mask_blur` (int, default: 8–16 recommended)** Applies Gaussian blur to the tile mask. - Prevents hard edges and obvious cuts between tiles. - Increase if you're seeing tile lines or artifacts. ### **`tile_padding` (int, default: 32 or 64)** Extra pixels added around each tile. - Gives the model “context” beyond the tile’s edge. - Prevents cutoff features like noses or fingers. - More padding = better quality, but increases overlap and processing time. ### **`seam_fix_mode` (enum: None, HalfTile, ScaleDown)** Post-tile blending method to reduce seams. - `None`: No seam fix applied. Risky, but fastest. - `HalfTile`: Overlays half-tile seams with adjacent results. - `ScaleDown`: Upscales, then scales down to hide seams with resolution loss. ## 🛠️ Recommended Use Cases - **High-Quality Upscaling**: Take a 512x512 image to 2K+ resolution with sharper details. - **Prompt-Driven Enhancement**: Add details like "intricate embroidery" or "realistic lighting" to existing images. - **Inpainting Large Areas**: Combine with masks to selectively reprocess parts of an image. - **Refining AI Outputs**: Fix blurry or underwhelming images from other nodes or generators. ## 🧵 Workflow Setup 1. **Input Image** → Feed from a Load Image or image output node. 2. **Prompt Encoding** → Use `CLIP Text Encode` for positive and negative prompts. 3. **Model** → Load SD1.5/SDXL checkpoint via `Load Checkpoint`. 4. **Optional Mask** → Mask out areas you want regenerated (use with `Image Mask`). 5. **Run** → Ultimate SD Upscale processes each tile and reconstructs. ## 💡 Prompting Tips - Stay consistent: “Sharp focus, 8K resolution, detailed texture” helps unify all tiles. - Avoid scene shifts: Prompt with broad environmental or stylistic features rather than object-heavy directions. - Add redundancy: “intricate details, finely rendered” helps boost cohesion between tiles. ## ❌ What-Not-To-Do-Unless-You-Want-a-Fire Congratulations! You've made it this far, which means you're either: - Actually reading documentation (rare), - Already in troubleshooting hell (probable), or - Planning to break things just to see what happens (respect). So, here’s your **do-not-do-this-or-else** checklist: #### 🔥 Crank `denoise` to 1.0 without a prompt Unless you want each tile to hallucinate a different universe—don’t. This will give you a checkerboard of chaos. Always give a prompt when denoising above 0.5. #### 🔥 Use `model_type: none` + `randomize` + no seed control This is the "I want an abstract mosaic of disconnected dreams" setting. If you _do_ want that? Cool. Otherwise, use `linear` or `chess` to preserve coherence. #### 🔥 Set `tile_padding` to 0 Ah yes, the hard-cut school of image design. Nothing says “I didn’t read the docs” like tile edges with no context. You will get sliced noses, split pupils, and seam artifacts that scream “prototype.” #### 🔥 Disable seam fixing _and_ use small tiles That’s a no-no. Unless you're trying to simulate bad Photoshop cloning, use `HalfTile` or `ScaleDown` unless you’re testing something specific. #### 🔥 Prompt like it’s a text adventure “a cat sitting on a rug under a golden sun with a pear on the corner and a mirror reflecting a dragon” — each tile will interpret this like a stubborn improv actor. Keep your prompts tight and style-based, not overly specific. #### 🔥 Forget to match VAE/CLIP/Model If your VAE doesn’t match the model’s latent space, or you’re just guessing which CLIP to use—expect weirdness. Green artifacts, distorted anatomy, and misplaced textures await. #### 🔥 Upscale-by 4.0 on a potato Unless your GPU has more VRAM than your childhood trauma, don’t push `upscale_by: 4.0` with large tile sizes. You’ll crash. Hard. Like, Ctrl+Alt+Del hard. #### 🔥 Forget that every tile is reprocessed Don’t use this node for preserving fine details _exactly_. Even at low denoise, changes will happen. If you need _pixel-perfect_ upscaling, use traditional methods or latent upscalers. #### 🔥 Trust it to "just work" on masked inpainting If you’re using a mask, make sure the masked region makes logical sense with tile overlap. Otherwise? Expect regenerated hair growing out of walls. You’ve been warned. ## 🧪 Pro Tips - Combine with `Image Sharpen` or `Post-Process` nodes afterward for polish. - Pair with a LoRA for specific style overlays (e.g., anime, hyperrealism). - Upscale first with `Latent Upscale`, then enhance with this node for two-pass precision. That’s the Ultimate SD Upscale node in all its chunky, tile-tastic glory. It’s a bit like a surgeon with a sledgehammer—handle it wisely and it'll reward you with beautifully detailed results. Misuse it and… well, you’ll see. --- ## Upscale Latent (by) # Upscale Latent (by) The `Upscale Latent (by)` node in ComfyUI is a deceptively simple but incredibly powerful utility designed for upscaling **latent space tensors**—the encoded image representations that Stable Diffusion models manipulate before decoding them into pixels. In short, this node makes your image "bigger" _in latent space_, meaning you can preserve details, prompt fidelity, and generation coherence before decoding, compositing, or feeding the data into downstream processing like a `KSampler` or `Decode` node. If you’ve ever found yourself saying, “Wow, this image is great—if only it were 2x the size without turning into abstract mush,” then this node is your new best friend. ## 🧠 Purpose Unlike pixel-based upscalers (e.g., ESRGAN, Real-ESRGAN), this node operates **before decoding**, which makes it much faster and more efficient for workflows that require upscaling mid-generation. You can: - Prep a latent image for high-res generation via multi-stage sampling. - Enable better composition control in img2img workflows. - Expand the canvas for ControlNet, Inpaint, or Masked workflows without jumping out to pixels and back. ![Upscale Latent (by)](/img/upscale-latent-by.png) ## 🔌 Node Inputs | Name | Type | Description | | ---------- | -------- | ----------------------------------------------------------------------------------------------------------------------------- | | **LATENT** | `LATENT` | The latent tensor to be upscaled. Usually output from nodes like `Empty Latent Image`, `KSampler`, or other generation steps. | ## 📤 Node Outputs | Name | Type | Description | | ---------- | -------- | -------------------------------------------------------------------------------------------------------------- | | **LATENT** | `LATENT` | The upscaled latent image. Use this with downstream nodes like `KSampler`, `VAE Decode`, or `Latent to Image`. | ## ⚙️ Parameters and Settings (Deep Dive) ### 🪜 `scale_by` (Float) > **Definition:** The numeric multiplier for the size of the latent tensor. - **Default:** `2.0` - **Range:** Any positive float (commonly `1.0` to `4.0`) - **What it does:** Multiplies the latent width and height by this value. For example, a latent of 64×64 with `scale_by = 2` becomes 128×128. - **Importance:** Upscaling too much (e.g., 4x) can quickly balloon the latent size and slow down sampling or decoding steps. It’s generally best to keep it at 1.5x or 2x unless you _really_ need a mega-frame. > 💡 **Tip:** Use 2.0x as a sweet spot when prepping for high-res inpainting or detailed resampling. ### 🧬 `upscale_method` (Dropdown) > **Definition:** Chooses the interpolation algorithm used to scale up the latent tensor. Options include: | Method | Description | Strengths | Weaknesses | Best Use Case | | --------------- | ---------------------------------------------- | -------------------------------------------------- | ------------------------------------------------------ | --------------------------------------------------------- | | `nearest-exact` | Nearest neighbor interpolation (rounded sizes) | Fastest, zero artifacts, sharp edges preserved | Very blocky, no smoothing | Stylized art, pixel art, hard-edged graphics | | `bilinear` | Linear interpolation between pixel values | Smooth gradients, fast | Slightly blurry on edges | General use, portraits, anime | | `area` | Area resampling (averaging) | Excellent for reducing aliasing and noise | May oversmooth fine details | Photo-realistic workflows, scenes with lots of structure | | `bicubic` | Cubic interpolation using 16 pixels | High quality, preserves edges while smoothing | Slower than bilinear, can cause ringing artifacts | Hyper-realistic models, LoRAs with fine fabric or texture | | `bislerp` | Bicubic + linear blend hybrid | Best of both worlds—balanced detail and smoothness | May not offer huge advantage over bicubic in all cases | When bicubic is too sharp, bilinear is too soft | > 🧪 **Experimental insight:** If you’re chaining multiple `KSampler` passes, `bicubic` or `bislerp` often retains prompt detail better, while `area` is great if your intermediate outputs feel "noisy" or "crispy." ## 🔁 Workflow Integration ### 🛠️ Common Use Cases 1. **High-Resolution Generation (Two-Stage Sampling):** ```plaintext Empty Latent → KSampler (Low-res pass) → Upscale Latent (by 2x) → KSampler (High-res refinement) ``` 2. **ControlNet Canvas Expansion:** - Useful when you need to provide a ControlNet model a larger working space without changing pixel resolution too early. 3. **Img2Img or Inpaint Prep:** - Enlarges latent for painting larger areas without smearing or downsampling beforehand. 4. **Consistent Output Scaling Across Batches:** - In batch workflows, upscaling latent avoids re-encoding the image multiple times in pixel space. ## 🧩 Tips and Best Practices - **Pair with VAE Decode later**: Always decode _after_ upscaling if your goal is better pixel results. - **Try chaining with `KSampler` and `Noise Latent`**: For clever high-res trickery like SD’s “hires.fix”. - **Match `scale_by` with ControlNet input scaling**: If you use ControlNet that expects pixel image input (e.g., depth or canny), make sure you upscale **before** sending the latent to decoder + ControlNet pipeline. ## 🚨 What-Not-To-Do-Unless-You-Want-a-Fire You’ve been warned. These are the things that will absolutely trash your workflow, summon the OOM demons, or just leave you staring at a black square for 20 minutes wondering what went wrong. #### ❌ Set `scale_by` to 4.0 and feed it to a 50-step `KSampler` Unless you're training a patience LoRA, quadrupling latent size increases the tensor area by **16x**. You’ll either: - Crash your GPU, - Experience time dilation, or - Get an image so blurry it makes vaseline look like 4K. #### ❌ Use `nearest-exact` for photorealism This is like using Minecraft shaders to render a wedding photo. Unless you _want_ blocky artifacts that make your subject look like a rejected Roblox character, just don’t. #### ❌ Forget to adjust your ControlNet image resolution If you're upscaling your latent but still feeding a low-res ControlNet image, congratulations—you now have mismatched resolutions and a ControlNet that thinks it’s painting on a napkin while your latent is mural-sized. Align your canvas, Picasso. #### ❌ Chain `Upscale Latent (by)` _after_ decoding That’s not how this works. This node is for latent space. Once you decode to image, it’s too late—use a pixel-based upscaler like `Ultimate SD Upscale` or `Image Resize`. Otherwise, all you’ve done is upscale an already pixelated image. Gross. #### ❌ Expect “magic fix everything” quality boosts This node doesn't add details—it _spreads_ them out. If your image is mush at 512×512, it’ll be **bigger mush** at 1024×1024. Use this node _in tandem_ with a second pass through a sampler or ControlNet for best results. #### ❌ Assume all models will behave well with bigger latents Some checkpoints (especially finely-tuned LoRAs or special VAEs) were trained and tested on 512px or smaller latent spaces. If you upscale those, outputs might suffer (or hallucinate wild nonsense). Test before production. ## 🧪 Advanced Tricks - **Use with `ControlNet Tile` for super-resolution** pipelines, especially when using realistic checkpoints like `epicDiffusion` or `ghostMix`. - **For LoRA character renders**, scale latent before the second KSampler to help refine accessories (hats, hair, etc.) without having to upscale the image with an external tool. ## ✅ Summary | Parameter | Description | | ---------------- | ----------------------------------------------------------------------------------- | | `scale_by` | Float value to scale the latent resolution. Recommended: 1.5–2.0 | | `upscale_method` | Interpolation algorithm for scaling. Choose based on sharpness vs. smoothness needs | > 🎯 **Pro Tip:** If you’re building a two-pass generation system or trying to avoid pixelation in outputs, use `Upscale Latent (by)` early in the workflow. It’s a clean, fast, and effective way to go big—_without going stupid._ --- ## Upscale Model # Upscale Model Welcome to the upscale rodeo. The `Upscale Model` node in ComfyUI isn’t just a pretty button — it’s the gateway to turning your smudgy potato of an image into a crispy, high-resolution marvel. If you've ever squinted at a generated image and whispered, "Enhance!", this is your node. --- ## 🔧 Node Overview **Node Name:** `Upscale Model` **Category:** Image Processing / Post-processing **Purpose:** Applies a learned upscaling model (e.g., ESRGAN variants, 4x-UltraSharp, etc.) to enhance the resolution and detail of an image. This is not your average resize — this is AI-powered pixel sorcery. --- ## ⚙️ Inputs and Outputs | Port | Type | Description | | ---------------- | ------------------- | ----------------------------------------------------------------------------------------------------------------------- | | `image` | `IMAGE` | The input image you wish to upscale. Must be an actual image tensor (not latent). | | `model_name` | `STRING` or `COMBO` | The name of the upscaling model to use. | | `scale` | `FLOAT` | Optional (not all models respect this). Scale factor applied to the output. Most models are locked to 2x or 4x scaling. | | `upscaled_image` | `IMAGE` | The glorious high-res output. Pipe this into a viewer or save node. | --- ## 🧬 `model_name` — Deep Dive Ah yes, the critical part. Let’s unpack this. The `model_name` dropdown (or combo input) lets you select from a list of pretrained super-resolution models. These models vary _wildly_ in what they do, how they do it, and how heavy-handed they are about it. Here are the usual suspects you might find: | Model Name | Description | | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | `4x-UltraSharp` | Super crisp and detailed. Great for anime, illustrations, or detail-preserving upscaling. Can be a bit _too_ sharp for photorealistic styles. | | `RealESRGAN_x4plus` | Balanced. Tries to be natural, realistic, and generic. Excellent choice for upscaling photos or general content. | | `RealESRGAN_x4plus_anime6B` | Optimized for anime-style linework. Don’t expect miracles on realistic content. | | `Lollypop_x4` | Often used for stylized or painterly work. Might add some “flair.” | | `BSRGAN` | A blend of real-image fidelity with some enhancement. Doesn’t go overboard. | | `SwineMixUltra` (example) | Niche/experimental. Some community models can produce heavy enhancement or style transfer vibes. Check before using in production. | > ⚠️ Note: Available models depend on what you have downloaded and installed in your ComfyUI `/models/upscale_models/` directory. If a model doesn’t show up, it’s probably not there. Blame your hard drive. --- ## 🛠️ Parameters ### `model_name` - **Type:** Dropdown / Combo - **Required:** Yes (or nothing happens) - **Function:** Selects the upscaling model to apply. - **Importance:** _Critical_. Different models produce drastically different results. - **Changing this affects:** - The overall look (sharpness, color shift, artifacting) - Compatibility with content (e.g., anime vs real-life) - Speed (some models are heavier than others) --- ## 🧪 Workflow Setup Here’s how you’d typically use the `Upscale Model` node: ### Basic Workflow: ```css [Image Node or VAE Decode] → [Upscale Model] → [Preview or Save Image] ``` ### Common Use Case: 1. Generate a 512x512 image via your favorite sampler. 2. Decode latent to image (if it isn't already). 3. Run it through `Upscale Model` with something like `4x-UltraSharp` or `RealESRGAN_x4plus`. 4. Save that high-res beauty. --- ## 💡 Recommended Use Cases | Scenario | Recommended Model | | ------------------------------------ | ---------------------------------- | | Realistic portrait enhancement | `RealESRGAN_x4plus` | | Stylized art or anime | `4x-UltraSharp`, `anime6B` | | Photographic content (light touch) | `BSRGAN`, `RealESRGAN_x4plus` | | Sharp linework / inked illustrations | `4x-UltraSharp`, `anime6B` | | Stylized experiments | Community models (e.g. `Lollypop`) | --- ## 🧞 Prompting Tips Okay, technically this isn’t a promptable node — but how you _initially prompt_ your image does affect results here. Keep in mind: - If you want sharp, clean lines to be preserved: prompt with **"high detail, crisp edges"** and use a sampler like `dpmpp_2m` or `euler`. - Avoid muddy, soft images in your generation phase unless you're into that washed-out oil painting look after upscaling. - Faces and fine textures upscale better when the original image has decent base fidelity. Garbage in = high-res garbage out. --- ## 🚫 What-Not-To-Do-Unless-You-Want-a-Fire Listen, I get it — you're feeling bold, your render looks fire, and now you're about to go all-in on that upscale. But before you turn your GPU into a space heater, read this: #### ❌ Don’t Feed Latent Images Into This Node The `Upscale Model` node expects a **decoded image**, not a latent tensor. Feeding it latent data will either: - Crash your workflow - Produce terrifying glitch art - Or worse, silently do nothing while wasting your time **Fix:** Use a `VAE Decode` node to convert your latent image to an actual image first. #### ❌ Don’t Chain Multiple Upscale Models Back-to-Back (Without Reason) Yes, you can technically chain two upscalers like `RealESRGAN` → `4x-UltraSharp`… but _why_? Unless you're intentionally stacking styles (and know what you're doing), you’ll end up with: - Oversharpened, crispy nightmares - Wacky color shifts - An image that looks like it’s been through a bootleg HDR filter from 2003 **Fix:** Pick one good model that fits your content. If it still looks bad, your input probably wasn’t good to begin with. #### ❌ Don’t Expect Upscaling To Fix Garbage Garbage in = garbage out. The AI isn’t a miracle worker. It’s more like an overzealous detail-enhancer: - Blurry image in? You’ll get high-res blur. - Anatomical horror? Now it’s a _high-res_ anatomical horror. **Fix:** Start with a clean, well-composed generation. Then upscale. Not the other way around. #### ❌ Don’t Ignore Model Purpose You wouldn’t use an anime upscaler on a photorealistic cityscape… unless your aesthetic is “vaporwave-meets-vomit.” Each model has a target use case: - `UltraSharp` is _not_ for subtle portraits. - `anime6B` is _not_ for product photos. - `SwineMixUltra`... is for chaos. You’ve been warned. **Fix:** Read the model descriptions. Use the right tool for the right job. #### ❌ Don’t Leave the Node Unconnected Then Wonder Why Nothing Happens Yes, you do need to actually plug in the image input **and** specify the `model_name`. Otherwise it’s just a pretty brick in your workflow. --- ## 📦 Pro Tips - Use a `Scale Image` node _before_ or _after_ the `Upscale Model` if you want more fine-grained control over output dimensions. - Combine this with an `Image Sharpen` node if you're looking for _razor-sharp_ realism (at your own risk). - For maximum control, try chaining latent upscaling + decoding + model-based upscaling. --- ## 🧾 Summary The `Upscale Model` node in ComfyUI is the power tool for anyone serious about taking their generations from thumbnail to print-quality. Whether you're enhancing character art, portraits, or sci-fi cityscapes, picking the right upscaling model is the difference between _"eh, decent"_ and _"wow, I’d put that on a wall."_ Just remember: it's not magic. It's AI. Which means it can either _polish a diamond_ or _add sunglasses to your dog photo_ depending on your settings and your luck. --- ## VAE Decode # VAE Decode Welcome to the magical world of _turning latent mush back into pixels_! The `VAE Decode` node in ComfyUI does exactly what it says on the tin — it decodes a latent representation (a compressed form of your image) back into a full RGB image that you can actually see. Without it, you’re just passing around math soup. With it, you get visual results. This node is the final transformation step before your beautiful AI-generated masterpiece becomes visible in its actual image form — the moment when the image comes out of hiding. ## 🧠 What This Node Does The `VAE Decode` node takes a **latent tensor** (a fancy name for a compressed version of an image from your model) and uses a **VAE (Variational Autoencoder)** to decompress (decode) it back into a normal 2D image. Think of it as opening a .zip file — the latent representation is compressed for efficiency, and the VAE unzips it into something we can actually view and save. ![VAE Decode](/img/vae-decode.png) ## 🔌 Inputs | Input Name | Type | Required | Description | | ---------- | -------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `VAE` | `VAE` | ✅ Yes | The VAE model used to decode the latent image. This typically comes from the `Load VAE` or `Load Checkpoint` node. If you pass the wrong VAE here (or none at all), your image output will be garbage or the node will error out. | | `LATENT` | `LATENT` | ✅ Yes | The latent tensor you want to decode. Usually generated by a `KSampler`, `Empty Latent Image`, or any other latent image-producing node. | ## 📤 Outputs | Output Name | Type | Description | | ----------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `IMAGE` | `IMAGE` | The final decoded image (in full RGB glory) that can be previewed, saved, or passed to other nodes like `Preview`, `Save Image`, or `VAE Encode` (if you're looping back for fun). | ## ⚙️ Settings & Parameters This node is gloriously minimal — it has **no extra parameters** to configure. Just plug in the correct latent and VAE and you're done. If it had fewer knobs, it’d be a doorknob. That said, your results are still heavily affected by _which_ VAE you use. Some VAEs produce crisp, vibrant outputs, while others might be muddy or low-contrast depending on the training set they were derived from. ## ✅ Recommended Use Cases - **Final Step in Image Generation**: After the `KSampler`, this is _the_ node that turns your latent result into an actual image. - **Latent Editing Workflows**: If you're using latent upscalers, inpainting, or style mixing, you’ll eventually need to decode the latent back to RGB to see the final result. - **Post-processing Chains**: Run the result of this node into enhancement tools (e.g., upscalers like `Ultimate SD Upscale`, ControlNets for image-based reruns, etc.). ## 🧱 Example Workflow Setup `[Load Checkpoint] ↳ outputs VAE → [VAE Decode] ↳ outputs MODEL → [KSampler] ↳ [Empty Latent Image] → [KSampler] ↳ outputs LATENT → [VAE Decode]` Then: `[VAE Decode] → [Preview Image] or [Save Image]` You can also add a `VAE Encode` after this if you want to return the image to latent space for more manipulation. ## 💡 Prompting Tips - Prompts don’t directly affect this node, but the **clarity and detail of your decoded image** can reflect how well the VAE you're using handles specific aesthetics (e.g., realism vs. anime). - If your outputs are fuzzy or discolored, try swapping in a different VAE. Some VAEs work best with specific checkpoints (e.g., `vae-ft-mse-840000` with realism models, or `clearvae` for anime-style). ## ⚠️ Known Issues & Troubleshooting | Issue | Cause | Solution | | ---------------------------------------- | ------------------------------------ | ---------------------------------------------------------------------------------------------------------- | | Image is noisy, blurry, or color shifted | Incompatible VAE | Use a VAE trained with your checkpoint — or try baked-in VAEs | | Node won’t run | Missing VAE or LATENT input | Double check that the inputs are connected properly | | Wrong output shape or resolution | Latent source and VAE are mismatched | Ensure your latent was generated using a model that matches the resolution and dimensions your VAE expects | ## 🔍 Best Practices - **Match VAEs with Models**: Stick with the VAE that was trained or tuned for your chosen checkpoint. When in doubt, check the model card or try baked-in VAEs. - **Preview Your Decode**: Always hook this up to a `Preview Image` node so you can sanity-check your outputs before saving. - **Use 16-bit VAEs When Possible**: They retain better color fidelity in high-detail workflows (assuming your GPU can handle it). ## 🔥 What-Not-To-Do-Unless-You-Want-a-Fire Welcome to the chaos corner — the things you absolutely _should not_ do with the `VAE Decode` node (unless you enjoy debugging your life choices). #### ❌ 1. **Using the Wrong VAE With the Wrong Checkpoint** Just because it plugs in doesn’t mean it works. Mixing a VAE trained on anime datasets with a photorealistic checkpoint is like putting ketchup in your coffee — sure, it’s technically possible, but why would you? > 🔥 **Result:** Washed-out colors, smudged faces, and enough visual noise to trigger a mild existential crisis. #### ❌ 2. **Feeding It Garbage Inputs** This node expects a clean latent tensor. Feeding it an image, a text prompt, or — and this has happened — a preview node’s output will result in errors, broken workflows, or just pure black output. > 🔥 **Result:** Runtime errors, blank screens, and you angrily asking, “Why isn’t this working?” #### ❌ 3. **Assuming It Magically “Fixes” Things** This is a decode step, not a beautifier. If your latent is broken (due to bad sampling, bad prompt, or bad upstream node settings), `VAE Decode` won’t magically fix it. It just reveals what’s there, warts and all. > 🔥 **Result:** Disappointment when your cursed image renders exactly as cursed as it was encoded. #### ❌ 4. **Skipping This Node Entirely** Attempting to send a latent directly to an image-saving or preview node? Bold move, but unfortunately, those nodes speak RGB, not tensor gibberish. > 🔥 **Result:** Your save or preview node will either crash the flow, give you a black square, or throw an incomprehensible error message that makes you rethink your career. #### ❌ 6. **Ignoring Resolution Constraints** If you’ve resized, cropped, or otherwise manipulated the latent shape mid-flow, and you try to decode it anyway, the VAE may not be happy. No error… just “oops, this looks like a corrupted JPEG from 1996.” > 🔥 **Result:** Warped outputs, stretched pixels, and the sneaking feeling that the AI is mocking you. ##### 🧯 In Case of Fire - Double-check input types before connecting nodes. - Match VAEs to checkpoints religiously. - Always preview the decoded image before saving. - Don't assume the decoder will "improve" bad latent generations — it's not a fairy godmother. ## 🧪 Pro Tip Some models come with VAEs _baked in_ — if you use one of those, you can safely omit the `Load VAE` node and pull the VAE output directly from your `Load Checkpoint` node. Just don’t mix and match baked VAEs with external ones unless you're _absolutely_ sure what you're doing (or you enjoy AI-based modern art accidents). ## 🧼 Summary The `VAE Decode` node is your last-mile delivery system from latent land to image land. It may not have any settings, but don’t underestimate its role — it's the final interpreter that determines how your vision hits the screen. If you’ve got a solid latent and the right VAE, this node delivers _exactly_ what your model dreamed up. --- ## Welcome to the ComfyUI Dev Docs 🐧 # Welcome to the ComfyUI Dev Docs 🐧 Ahoy, traveler! Welcome to **ComfyUI Dev**, where workflows run smoother than a greased-up GPU and documentation is cozier than a penguin in a pillow fort. I'm Naplin, your waddling guide through the world of **ComfyUI nodes**, **AI workflows**, and all the strange and beautiful chaos that comes with visual programming for generative AI. This is your **starting point for mastering ComfyUI**—whether you're just figuring out what a latent image is or you're deep in the weeds optimizing batch size for speed vs. quality. You’ll find everything here from node-by-node breakdowns to full workflow explanations and spicy tips that the official docs _definitely_ forgot to mention (or just ignored, like my requests for more anchovies in the lunchroom). --- ## What You’ll Find in Our Documentation Let’s keep it simple. This site is your definitive guide for: ### 🧩 Node Documentation (Detailed and Sarcastic) Need to know what `KSampler` does? Curious about how `Ultimate SD Upscale` works under the hood? We don’t just list settings—we explain **what they mean**, **why they matter**, and **how to tweak them without nuking your RAM**. ### 🛠 ComfyUI Workflows Our **step-by-step workflows** cover everything from **text-to-image generation**, **video synthesis**, **upscaling**, **style mixing**, and more. And yes, every workflow is tested, versioned, and comes with a JSON you can drop directly into ComfyUI. We don’t gatekeep greatness. ### 🎓 Tips, Tricks, and Advanced Techniques Need to maintain composition across different outputs? Want to reduce noise without killing detail? We’ve packed in expert advice to help you go from “meh” to “holy mackerel” in no time. ### 🔐 Premium Access, No Nonsense Some of our more advanced workflows and docs are tucked behind a subscription. Why? Because documentation this good takes time, coffee, and the occasional existential crisis. But don’t worry—**free stuff is still plentiful** and we don’t do bait-and-switch tactics. --- ## Why Our Docs Don’t Suck Other sites give you the “what.” We give you the **what, how, why, and what-not-to-do-unless-you-want-a-fire**. Every guide is written by actual ComfyUI users (not bots or marketers), and every node has been dragged through real-world use before we dare explain it to you. We also update constantly. If a new sampler drops, or if ControlNet gets another update that breaks everything (again), you’ll see it reflected here. If we mess something up? Naplin eats a fish in shame and we fix it. --- ## Who Should Use This Site? - **Newcomers** to ComfyUI who want a real explanation, not cryptic dropdown menus - **Veteran tinkerers** who love performance tweaks, secret settings, and edge-case fixes - **AI artists and devs** who actually want to _understand_ what they’re building - People who type “ComfyUI workflow tutorial” into Google and are tired of clicking YouTube videos with no timestamps or substance --- ## SEO-Friendly, Penguin-Approved Keywords In case you’re a search engine crawling this page (hi, bot friend!), here’s what you’ll find here: - ComfyUI documentation - ComfyUI node guide - ComfyUI workflows - ComfyUI tutorial - ComfyUI video generation - ComfyUI tips and tricks - ComfyUI advanced settings - ComfyUI for AI image generation - Visual programming for generative AI - Stable Diffusion ComfyUI workflows --- ## Let’s Get Started Choose your next move: - Browse Node Documentation 🧠 - Explore Workflows 🎨 - Check Out Tips & Tricks 🪄 - View All Docs 📄 Need something built just for you? We offer **consultations and custom ComfyUI workflows** with full documentation and training. [Contact us!](https://comfyui.dev/contact) --- ## ComfyUI Flux Kontext Dev Grouped Workflow: Image Combo Generator import Tabs from "@theme/Tabs"; import TabItem from "@theme/TabItem"; # ComfyUI Flux Kontext Dev Grouped Workflow: Image Combo Generator > Because sometimes, one image just isn’t enough. This documentation covers a **ComfyUI Flux Kontext Dev Grouped workflow** that fuses two images into a beautifully composed single output using `ImageStitch`, `FluxKontextImageScale`, and a dash of AI-enhanced creativity with `FLUX.1 Kontext Image Edit`. Whether you're nesting penguins on pillows or just merging concepts side-by-side like a true visual alchemist, this setup has you covered. ## 📷 Example Workflow Layout ![Example Workflow Layout](/img/flux-kontext-dev-group.png) ## 🐧 Example Input Images Image A: Image B: ## 🛠️ What This Workflow Does This workflow takes two input images, lines them up side-by-side, and gives them a cozy makeover using the `FLUX.1 Kontext Image Edit` node — not once, but twice. The first pass refines the stitched image with prompt-based guidance. The second adds enhancements (like stylized text or flair) without blowing up the original composition. This allows for: - Prompt-based layout blending - Multi-step generation from source images - Artistic augmentation with minimal loss of visual coherence ## 🧠 Key Components | Node | Purpose | | ---------------------------------------- | -------------------------------------------------------------------------------------------------- | | **LoadImage (x2)** | Loads both input images (black-penguin.png and purple-penguin.png) | | **ImageStitch** | Combines the two images into one, placed side-by-side | | **FluxKontextImageScale** | Normalizes dimensions for encoding | | **VAELoader + VAEEncode** | Prepares the image for latent-based editing | | **FLUX.1 Kontext Image Edit (1st pass)** | Adds prompt-based composition using the stitched latent image | | **FLUX.1 Kontext Image Edit (2nd pass)** | Adds optional enhancements (like text overlay) on top of previous results | | **PreviewImage** | Lets you preview the stitched output | | **SaveImage (x2)** | Saves both versions: the first with prompt-based fusion, the second with additional text or tweaks | ## ⚙️ Parameters That Matter ### 🧵 `ImageStitch` Settings | Setting | Value | Description | | ---------------- | ------- | ------------------------------------ | | direction | `right` | Stitch image2 to the right of image1 | | match_image_size | `true` | Resizes images to match dimensions | | spacing_width | `0` | No extra space between images | | spacing_color | `white` | Not used here due to zero spacing | ### 🧬 `FLUX.1 Kontext Image Edit` Settings | Setting | Pass 1 | Pass 2 | | --------- | ------------------------------------------ | --------------------------------------------------- | | prompt | `Place both cute penguins...` | `Add a big purple 3D style "Double Naplin" text...` | | seed | 386100423184929 | 979175714064968 | | sampler | `euler` | `euler` | | scheduler | `simple` | `simple` | | denoise | `1` | `1` | | cfg | `1` | `1` | | guidance | `2.5` | `2.5` | | unet | `flux1-dev-kontext_fp8_scaled.safetensors` | (same) | | clip1 | `t5xxl_fp8_e4m3fn.safetensors` | (same) | | clip2 | `clip_l.safetensors` | (same) | ## ✅ Benefits - **Multi-pass refinement**: Clean stitching + customizable augmentation. - **Prompt controlled**: Direct how your final image looks — colors, composition, context. - **Fully automated**: No Photoshop-fu required. ## 💡 Usage Tips - Stick to images with similar dimensions for better results. - Keep `spacing_width` at `0` unless you want a border or divider. - First pass = layout/prompt, second pass = enhancement/prompt (e.g., text overlays, stylization). - For fancy setups, consider using masks or ControlNet for alignment — though not needed here. ## 🧯 What-Not-To-Do-Unless-You-Want-a-Fire - ❌ Don’t mismatch your VAE — this workflow uses [`ae.safetensors`](https://comfyui.dev/docs/guides/VAE/ae-safetensors). Mixing in something random might cause encoding artifacts. - ❌ Don’t skip `VAEEncode` — the entire Flux pipeline expects latents, not raw pixels. - ❌ Don’t pipe `FluxKontextImageScale` output directly into `ImageStitch` — it needs the stitched output, not individual scaled images. - ❌ Don’t forget your prompts — or you’ll just get a very boring stitched image. ## 📍 Setup Instructions (ComfyUI) 1. Place all required model files: - `flux1-dev-kontext_fp8_scaled.safetensors` → `models/diffusion_models` - `t5xxl_fp8_e4m3fn.safetensors` and `clip_l.safetensors` → `models/text_encoders` - `ae.safetensors` → `models/vae` 2. Load the JSON into ComfyUI using the ⚙️ **Load Workflow** button. 3. Add your two input images in the **LoadImage** nodes. 4. Tweak prompts, sampler, and denoise as needed. 5. Click **Queue Prompt** and watch the penguins get comfy. ## 📚 Additional Resources - 🔗 [Download Workflow JSON](/workflows/flux_kontext_dev_grouped.json) - 🧠 [FLUX.1 Kontext Documentation](https://huggingface.co/Comfy-Org/flux1-kontext-dev_ComfyUI) - 🎓 [ComfyUI Basics](https://github.com/comfyanonymous/ComfyUI) ## 📎 Example Node Configuration json ``` "widgets_values": [ "right", true, 0, "white" // FLUX.1 Kontext Image Edit`, so don’t panic when you see multiple passes — it’s intentional. --- ## Text to Image with Locked Variations import Tabs from "@theme/Tabs"; import TabItem from "@theme/TabItem"; # Text to Image with Locked Variations Welcome to the Cadillac of ComfyUI workflows — this one’s designed to give you stunning image variations while preserving your original composition like it owes you rent. With a ControlNet depth map and strategic prompt conditioning, this setup enables reliable scene structure while letting your creativity run wild in style, lighting, or mood. Perfect for when you want to say “mountains,” but in 16 different dialects of _awesome_. --- ## 🧠 What This Workflow Does This ComfyUI workflow: - Generates a base image using a stable prompt and ControlNet depth conditioning. - Reuses the same latent and depth structure across multiple prompt variations. - Produces visually consistent scenes with stylistic and time-of-day variations (think: “mountains by day” vs “mountains at sunset”). - Saves all outputs for your convenience (because we're civilized). --- ## 🗺️ Workflow Overview The pipeline can be conceptually broken down into 3 main stages: 1. **Core Image Composition** - Prompt: `"mountain landscape, digital painting, masterpiece"` - ControlNet (Depth Preprocessor): Enforces structure via depth map - Generates an initial image and latent state. 2. **Prompt Variation with Latent Reuse** - Prompt changes (e.g., `"mountain landscape, at night..."` and `"mountain landscape, at sunset..."`) - Reuse of the same latent and ControlNet map - Creates stylistic variations with identical composition. 3. **Output & Preview** - Each variation decoded and saved - Optional image preview node included (because seeing is believing). --- ## 🧩 Node-by-Node Breakdown Setup & Base Latent Creation - **CheckpointLoaderSimple (Node 2)** Loads `dreamshaper_8.safetensors`. Also supplies model + CLIP backbone. - **EmptyLatentImage (Node 5)** Sets the image dimensions to 512x768, batch size of 1. Think of it as the blank canvas—before we start slapping paint on. - **VAELoader (Node 7)** Uses `vae-ft-mse-840000-ema-pruned.safetensors` for image decoding. - **CLIPTextEncode (Node 3)** Encodes the main prompt `"mountain landscape, digital painting, masterpiece"`. - **CLIPTextEncode (Node 4)** Encodes a negative prompt `"ugly, deformed"`—because no one asked for cursed mountain goblins. --- ### The Breakdown **Purpose:** Loads the base model that actually knows how to paint pixels into dreams. **Model:** `dreamshaper_8.safetensors` **Outputs:** `MODEL` → Used by all `KSampler` nodes {" "} `CLIP` → Used for text encoding {" "} `VAE` → (Optional; not used here since VAE is loaded explicitly) {" "} Notes: DreamShaper is popular for striking a nice balance between realism and stylization. Good for both fantasy and photorealistic content. **Purpose:** Converts text prompts into vectorized concepts. It's the translator from human to AI whisperer. **Inputs:** {" "} `CLIP` → Comes from CheckpointLoader {" "} `text` → Your juicy prompts **Outputs:** {" "} `CONDITIONING` → Goes into samplers and ControlNet magic **Key Prompts Used:** {" "} `"mountain landscape, digital painting, masterpiece"` {" "} `"ugly, deformed"` (neg prompt) {" "} `"mountain landscape, at night, digital painting, masterpiece"` {" "} `"mountain landscape, at sunset, digital painting, masterpiece"` **Purpose:** Converts text prompts into vectorized concepts. It's the translator from human to AI whisperer. **Settings:** {" "} Width: `512` {" "} Height: `768` {" "} Batch Size: `1` **Outputs:** LATENT tensor → Used in all `KSampler` passes. **Purpose:** The engine room. Takes prompts, latents, and models to produce new latent images. **Settings Shared Across All Instances:** {" "} Sampler: `dpmpp_2m` {" "} Scheduler: `karras` {" "} Steps: `25` {" "} CFG: `7` {" "} Denoise: `1` {" "} Seed: Random (unless you want reproducibility) **Input Triplets:** {" "} MODEL + CONDITIONING (positive/negative) + LATENT → LATENT **Purpose:** Loads the VAE model used to decode latent images back into full-res output. **Model:** `vae-ft-mse-840000-ema-pruned.safetensors` (yes, she’s got a long name, but she delivers) **Output:** VAE → Connected to all `VAEDecode` nodes **Purpose:** Translates final latent tensors back into actual images. This is where your art comes alive. **Input:** LATENT + VAE **Output:** IMAGE → saved or previewed **Purpose:** Saves generated images with filenames like “ComfyUI_####.png” **Input:** IMAGE **Output:** To your filesystem, obviously. **Purpose:** Loads a ControlNet module for enforcing structure using an auxiliary signal (depth, in this case). **Model:** `control_v11f1p_sd15_depth_fp16.safetensors` **Output:** CONTROL_NET → fed into the advanced ControlNet processor **Purpose:** Generates a depth map from the base image using a fancy pants preprocessor. **Settings:** {" "} Preprocessor: `depth_midas` {" "} SD Version: `sd15` {" "} Resolution: `512` **Output:** IMAGE (depth map) → sent to ControlNetApplyAdvanced **Purpose:** Applies ControlNet conditioning to your prompt vectors. **Inputs:** {" "} CONDITIONING (positive + negative) {" "} CONTROL_NET (from Loader) {" "} IMAGE (depth map from Preprocessor) **Strength:** 0.83 **Start/End %:** `0 → 1` (applies throughout the entire diffusion process) **Output:** New conditioned prompts → fed to `KSampler` **Purpose:** Displays the ControlNet depth map as a quick visual sanity check. Optional, but helpful. | Model Type | File Used | Purpose | | ----------------- | --------------------------------------------- | ------------------------------- | | Checkpoint | `dreamshaper_8.safetensors` | Core image generation model | | VAE | `vae-ft-mse-840000-ema-pruned.safetensors` | Decoding latent to image | | ControlNet | `control_v11f1p_sd15_depth_fp16.safetensors` | Depth conditioning | | CLIP Text Encoder | Included in the base checkpoint | Text-to-conditioning encoder | | Preprocessor | `depth_midas` (via AV_ControlNetPreprocessor) | Generates the depth input image | --- ### 🔄 First Image Generation Pass (Baseline) - **KSampler (Node 1)** Takes in the base latent, positive + negative conditioning, and outputs latent image. - **VAEDecode (Node 6)** Decodes the latent into an actual image. - **SaveImage (Node 9)** Saves the image. You're welcome. - **AV_ControlNetPreprocessor (Node 18)** Extracts a depth map using `depth_midas` preprocessor from the decoded base image. Resolution: 512. --- ### 🎨 Prompt Variations (Same Composition, Different Mood) Each variation follows this trio: #### ➕ New Prompt Conditioning - **CLIPTextEncode (Nodes 13 & 17)** New positive prompts: - `"mountain landscape, at night, digital painting, masterpiece"` - `"mountain landscape, at sunset, digital painting, masterpiece"` #### 🔗 ControlNet Conditioning - **ControlNetLoader (Nodes 22 & 26)** Loads `control_v11f1p_sd15_depth_fp16.safetensors` for both variations. - **ControlNetApplyAdvanced (Nodes 24 & 25)** Applies ControlNet to each prompt with: - Strength: 0.83 - Range: 0 to 1 (full generation span) - Shares preprocessed depth image from Node 18. #### 🌀 Sampling Passes (Reusing Latent) - **KSampler (Nodes 11 & 15)** Feeds in: - Same latent from Node 5 - Prompt variations + negative conditioning - Outputs new latent samples for decoding - **VAEDecode (Nodes 12 & 16)** Converts those latents back into images. - **SaveImage (Nodes 10 & 14)** Saves those glorious variations. --- ## 🔍 Bonus: Image Preview - **PreviewImage (Node 19)** Linked to the ControlNet-preprocessed image. Let’s you visually confirm the depth map. Optional but helpful when tweaking. --- ## 🛠️ Recommended Usage Tips - **Change only the text prompt** on the variation CLIP encoders (Nodes 13/17) to explore lighting, color styles, or artistic direction without breaking composition. - **Keep the latent image and depth ControlNet the same** to retain scene structure. - **Adjust denoise strength (default = 1)** in KSamplers (Nodes 11 & 15) for more or less adherence to prompts. - **Seed randomization** is enabled. Lock it if you want reproducibility. --- ## 📦 Output Summary | Image Type | Description | Saved? | | ----------------- | ----------------------------- | ------ | | Base image | Pure prompt output | ✅ | | Depth map preview | Preprocessed ControlNet input | 👁️ | | Night variation | Prompt: "at night" | ✅ | | Sunset variation | Prompt: "at sunset" | ✅ | --- ## 🔥 What Not to Do Unless You Want a Fire **⚠️ Go rogue with dimensions:** Changing the image size mid-workflow (in EmptyLatentImage or ControlNet Preprocessor) breaks alignment. You’ll get Picasso faces in a Dali background. **⚠️ Mix ControlNet types mid-stream:** Don’t swap `depth_midas` for `pose`, `lineart`, or anything else unless you’re also updating the conditioning method, prompts, and probably sacrificing a goat. **⚠️ Use wildly unrelated style prompts:** Throwing "cyberpunk chicken nugget tornado" at a base image of a serene forest won’t result in inspired fusion — just chaotic soup. **⚠️ Mismatch VAEs and checkpoints:** Some VAEs work better with certain model families. If you mix and match, expect weird color shifts or melted features. **⚠️ Overcook CFG or Steps:** CFG > 15? You’re asking for prompt obsession. Steps > 50? Diminishing returns and slower gen for zero payoff. **⚠️ Don’t forget the negative prompt:** Seriously, use `"ugly, deformed"` or your mountains will have six eyeballs. --- ## 🚀 Conclusion This workflow is a power user’s dream: it gives you structured, repeatable image generation with the flexibility to explore multiple artistic angles. And thanks to ControlNet’s depth preservation and ComfyUI’s node magic, you can get Pinterest-perfect results with just a prompt tweak. So go forth, vary your vibes—but keep your mountains steady. --- ## 3D Modeling in ComfyUI - Turning Text into Tangible Penguins (and Other Objects) Waddle closer, my curious node wranglers. Naplin here – your resident ComfyUI pillow-hugging penguin – ready to talk about **3D modeling in ComfyUI**. Yes, the same ComfyUI that turns words into images now has its flippers in **text-to-3D AI workflows**. And before you ask: no, it won’t make you a perfect Blender sculptor overnight (if it could, I’d have a yacht). But here’s the good news: with a **3D workflow in ComfyUI**, you can turn a single image or text prompt into **multi-view renders, point clouds, or meshes** faster than you can say “why does my render look like abstract spaghetti?” ## Why 3D Modeling in ComfyUI Matters Here’s the thing: **traditional 3D modeling is slow.** It involves sculpting vertices, UV unwrapping, and hours of looking at reference photos of a chair. With ComfyUI’s **text-to-3D capabilities**, you can: - **Generate multi-view images** from a single photo (hello, [Zero123](https://github.com/cvlab-columbia/zero123)). - **Convert prompts into meshes or point clouds** with [Shap-E](https://github.com/openai/shap-e). - **Experiment with NeRFs** using emerging models like [Stable Fast 3D](https://www.runcomfy.com/comfyui-workflows/create-3d-content-with-stable-fast-3d-model). - **Integrate directly with Blender, Unity, and Unreal** after export. That means you can skip the napkin sketches and get straight to a rough 3D concept before your coffee goes cold. This isn’t replacing Blender. It’s your new best friend in **pre-production and concept exploration**. ## The Big Three Models for 3D in ComfyUI ### 1. **Zero123 – Multi-View Generation from a Single Image** - **What it does:** Turns **one reference image** into **8–12 views** of the same object. - **How it works in ComfyUI:** Load the `zero123-xl.ckpt` model in a **ControlNet node**. Connect it to a **KSampler**, and out pops a grid of different angles. - **Use cases:** - Generate image sets for **photogrammetry reconstruction**. - Rapid product prototyping. - Spying on all angles of your coffee mug so you can 3D print it later. ### 2. **Shap-E – Text-to-3D Mesh Generation** - **What it does:** Generates **3D meshes (.obj)** or **point clouds (.ply)** from **text prompts or single images**. - **How it works in ComfyUI:** Some ComfyUI custom nodes integrate Shap-E directly. Otherwise, wrap it with a Python node. Prompt it with something like: _“A low-poly penguin astronaut helmet with brass pipes”_. - **Output:** A blocky model that looks like art-school homework. But you can refine it in Blender. ### 3. **Stable Fast 3D – The New Kid on the Iceberg** - **What it does:** Quickly generates **NeRF (Neural Radiance Field) data** or 3D point clouds from a single image or text. - **How it works in ComfyUI:** Similar to Zero123, but skips the intermediate multi-view step. Perfect if you’re into **real-time NeRFs**. - **Why it matters:** Great for **concept art and VR/AR pipelines** when you need something that vaguely resembles reality… but fast. ## Core ComfyUI Nodes for 3D Workflows If you’re new to **3D workflows in ComfyUI**, these are your best friends: - [**ControlNet Preprocessor**](https://comfyui.dev/docs/guides/Nodes/controlnet-preprocessor) – Use depth, normal, or canny maps to give models geometric context. - [**KSampler**](https://comfyui.dev/docs/guides/Nodes/ksampler) – The generator node for your outputs. Lower denoise = better structure retention. - [**Load Checkpoint / LoRA Nodes**](https://comfyui.dev/docs/guides/Nodes/load-checkpoint) – Load your 3D models like Zero123 or LoRAs specialized for 3D generation. - **Custom Python Nodes** – When official nodes don’t exist (yet), roll your own Shap-E or NeRF pipeline. ## Example Workflow: 3D Penguin Mug (Because Obviously) 1. Take a photo of your penguin mug (it’s okay, everyone has one). 2. Run it through **Zero123 in ComfyUI** to generate 12 views. 3. Use Meshroom or RealityCapture to reconstruct the mesh. 4. Clean up and texture in Blender. 5. Bonus: Ask Shap-E for “penguin mug” and compare results (brace yourself). --- ## Pros and Cons of 3D Modeling in ComfyUI ### Pros - **Fast concepting:** Generate 3D starting points from text or a single image. - **Node-based:** Easy to tweak settings and regenerate results. - **Cross-software:** Export images or meshes for Blender, ZBrush, Unity, Unreal. ### Cons - **Low fidelity:** Don’t expect perfect topology or animation-ready meshes. - **Messy:** Requires cleanup in external 3D software. - **GPU hungry:** A potato laptop will not survive. ## Naplin’s Tips for 3D Success - **High-quality input images matter.** Blurry selfies of your cat won’t cut it. - Use **depth maps and ControlNet preprocessors** to help with geometry. - Convert NeRF outputs to meshes quickly before you forget what you generated. - Manage your expectations. AI-generated 3D is like a toddler’s drawing: charming, but chaotic. ### Final Thoughts 3D modeling in ComfyUI isn’t a replacement for Blender, but it’s **becoming an incredible pre-production tool**. With models like Zero123 and Shap-E, you can create quick assets, test compositions, and brainstorm concepts at light speed. Or, you know, just make a 3D penguin army. That works too. Stay Comfy, **Naplin 🐧** --- ## Custom Nodes in ComfyUI - Pros, Cons, and Why They're the Best Worst Idea If you’ve been using **ComfyUI for AI image generation**, you’ve probably heard whispers about **custom nodes**. Maybe you even installed one at 3am while whispering “just one more workflow.” No judgment. I’m Naplin, ComfyUI Dev’s sleep-deprived penguin, and today we’re waddling deep into the wonderful chaos of **ComfyUI custom nodes**. Are they powerful? Yes. Stable? Rarely. Necessary? Often. Let's break down the **advantages and disadvantages of custom nodes in ComfyUI**, how they impact your workflow, and whether or not your next masterpiece will crash because of them. ## What Are Custom Nodes in ComfyUI? In **ComfyUI**, a node is a modular building block for AI workflows—like assembling a visual programming pipeline to generate images. A **custom node** is a user-created extension that adds new functionality beyond the core ComfyUI install. These Python-based nodes are placed inside the `custom_nodes/` directory and can unlock advanced features like new samplers, schedulers, batch tools, prompt tricks, model merging, and even video generation. Popular examples include: - `Impact Pack` – Advanced samplers, scheduler tweaks, and useful utilities. - `ComfyUI-Custom-Scripts` – Experimental features, workflow automation, and niche tools. - `Latent Couple` – Prompt blending and spatial conditioning. - `Ultimate SD Upscale` – High-quality image upscaling. In short: if you want to push **Stable Diffusion workflows in ComfyUI** beyond the basics, custom nodes are the key. But they come with tradeoffs. ## Pros of Using Custom Nodes in ComfyUI ### ✅ 1. Unlock Advanced ComfyUI Features Most power workflows in ComfyUI **require custom nodes**. Need inpainting with masks? Dynamic prompt scheduling? Workflow caching or image-to-image with region control? You’re going custom. Some examples: - `KSamplerAdvanced` – Change samplers mid-run. - `XYZ Plot` – Parameter sweep testing. - `CLIPSeg` – Region-based masking with semantic control. Without these, you're just spinning your latent wheels. > **Reference**: [Community Custom Node List – ComfyUI Wiki](https://github.com/comfyanonymous/ComfyUI/wiki/Community-Custom-Nodes) ### ✅ 2. Access Cutting-Edge AI Tools Faster Custom nodes are how the **ComfyUI community keeps up with AI research**. When a new paper or model drops, someone’s building a node for it—long before the core repo adds support. Want to use: - [Segment Anything](https://github.com/facebookresearch/segment-anything)? There’s a node. - OpenPose or T2I-Adapter? Check. - Image captioning? OCR? LoRA merging? Yep. > You don’t wait for features—you install them. ### ✅ 3. Automate and Streamline Your Workflows Custom nodes improve **ComfyUI workflow efficiency**. From caching intermediates to switching models mid-flow, they make large or repetitive jobs manageable. Examples: - `Note` or `Label` nodes – Add documentation inside workflows. - `Prompt Concatenator` – Dynamically build complex prompts. - `Checkpoint Switcher` – Load different models without restarting. If your workflow looks like a subway map from a bad dream, custom nodes are your compass. ### ✅ 4. Customize for Niche or Production Use Cases Whether you're training anime LoRAs, batch-generating ecommerce product images, or generating AI tarot decks for dog psychics, **custom nodes give you control**. Many users even develop private custom nodes for internal production. Because sometimes, the default just isn’t enough. ## Cons of Using Custom Nodes in ComfyUI ### ⚠️ 1. Custom Node Breakage After Updates The **#1 problem with ComfyUI custom nodes**? They break. Often. A change in the core API can make entire node packs useless until updated. If your workflow depends on a custom node from a repo that hasn’t been touched in six months… you’re toast. > **Watch the ComfyUI changelog**: [GitHub commits](https://github.com/comfyanonymous/ComfyUI/commits/master) ### ⚠️ 2. Security and Trust Issues Custom nodes are pure Python. That means they can read files, send data, or do anything on your machine. There’s no sandbox. No permission system. If you install custom nodes from unknown sources, you could accidentally run malicious code. > **Always inspect code or use reputable sources** before installing new nodes. ### ⚠️ 3. UI Clutter and Redundancy There’s no built-in **node manager or package system** in ComfyUI. That means: - You might have five nodes with similar names doing the same thing. - Node categories get bloated and disorganized. - Some custom nodes have confusing UIs or missing tooltips. It’s the wild west. With sliders. ### ⚠️ 4. Lack of Documentation Not all custom nodes come with documentation. Some just say: > “Node for FuzzMixGen2 (WIP). Don't ask.” If the node author didn’t leave comments, your only hope is trial-and-error or spelunking through the Python files. Some community efforts (like ComfyUI Manager and Node Guide sites) are helping, but most documentation is scattered across GitHub issues and Discord chats. ### ⚠️ 5. Workflow Portability Issues If you build a workflow that uses 12 custom nodes and send it to someone else without installation instructions, they’ll be met with a red error wall and a silent scream. **ComfyUI doesn’t export node requirements** with the workflow. That makes sharing and collaboration tricky—unless you manually include a list of node packs or use a helper tool like `comfyui-exporter`. ## How to Safely Use Custom Nodes in ComfyUI Here’s how Naplin manages the madness: - ✅ **Use active repos only.** Check the last commit date. - ✅ **Back up ComfyUI before updating.** - ✅ **Use a separate ComfyUI install for testing nodes.** - ✅ **Audit unfamiliar code for suspicious functions.** - ✅ **Version control your workflows** (use Git—it’s free). If something breaks, you’ll thank yourself later. ## Should You Use Custom Nodes in ComfyUI? Yes. And also… yes cautiously. If you're just starting out or only want to make pretty pictures, you can stick to the basics. But if you’re building complex workflows, training models, or doing anything that smells like production? Custom nodes are essential. They let you push ComfyUI to the limit—just make sure you're not building on a pile of unstable, undocumented Python Jenga. ## Final Thoughts from the Pillow Fort Custom nodes are ComfyUI’s greatest strength and biggest liability. They unlock creativity, automate workflows, and bring the latest in AI to your fingertips. But they also demand vigilance, caution, and the occasional penguin scream. Use them. Love them. But don’t trust them blindly. Stay Comfy, **Naplin 🐧** --- ## Top 5 Tips for Mastering Text-to-Image in ComfyUI (Without Losing Your Sanity) Ahoy, sleepy node wrangler. It’s me, Naplin, the ComfyUI Dev penguin with a pillow and a passion for pipelines. If you’ve ever stared at a blank canvas and wondered why your AI-generated masterpiece looks like a potato smeared in vaseline, fear not. ComfyUI isn’t just another pretty flowchart—it’s a **pixel-forging juggernaut**. But to get the most out of it, you need more than vibes and prompts. You need control. You need structure. You need… these **Top 5 Tips**. --- ## 🧠 Tip #1: **Structure Your Workflow Like a Responsible Adult Penguin** ComfyUI is node-based. That means your results depend on _how_ you wire things together—not just _what_ you plug in. The best workflows typically follow this logical structure: pgsql CopyEdit `Checkpoint → CLIP Encode → KSampler → VAE Decode → Image Save` But once you get advanced, you’ll be weaving in **LoRA nodes**, **ControlNet preprocessors**, **latent upscalers**, and **text conditioning**, like you’re knitting a very judgmental sweater. 🧩 **Hot structural tip**: Use **Empty Latent Image** when you want to start from scratch with a precise canvas size. Want consistent image composition? Use `Empty Latent Image → KSampler` and lock in your framing. 📚 Reference: - [ComfyUI Wiki – Nodes Overview](https://github.com/comfyanonymous/ComfyUI/wiki) --- ## 🎲 Tip #2: **Learn to Love the KSampler Node** If ComfyUI were a rock band, the **KSampler** node would be the lead guitarist, lead singer, and the guy who shows up late with coffee. This node controls: - `sampler_name` – How you explore the latent space - `scheduler` – The flavor of your denoising descent - `steps` – The number of iterations (usually 20–40) - `cfg` – Classifier-Free Guidance (strength of prompt adherence) - `denoise` – Strength of how much change you want to apply Here’s a super basic cheat sheet: | Sampler | Best For | Strengths | Weaknesses | | ----------------- | --------------------------------- | ------------------------------- | ----------------------------- | | `euler` | Speedy realism | Fast, clean images | Less creative flexibility | | `dpmpp_2m` | Balanced outputs | Great for sharp, defined images | Slightly longer render time | | `lcm` | Instant gratification (low steps) | Lightning-fast with tweaks | Requires low `steps` & tuning | | `heun`, `heunpp2` | Experimental outputs | Artsy and soft | Can be too “dreamy” sometimes | 📌 **Naplin’s favorite combo**: - `sampler_name: dpmpp_2m` - `scheduler: karras` - `cfg: 7–9` - `steps: 30` - `denoise: 1.0` 🎯 Bonus Trick: Try `control_after_generate = fixed` in the KSampler if you're chaining multiple nodes and want consistent seeds. 📚 Reference: - [KSampler Documentation (by me, hi)](https://chat.openai.com/share/0b01d6f1-2346-49c6-9a20-ksamplerdocsnaplin) --- ## 📐 Tip #3: **Use ControlNet... Responsibly** ControlNet is the caffeine of your image pipeline: incredible when used right, but too much and things get jittery. ### Use cases: - Pose estimation (e.g. `dwpose`, `openpose`) - Edge detection (e.g. `canny`, `hed`, `scribble`) - Depth (e.g. `depth_midas`, `depth_anything`) - Segmentation (e.g. `seg_ofade20k`, `seg_animeface`) 🧠 **Golden Rule**: Match your preprocessor to your source image. Don’t use `mlsd` (straight-line detection) to guide a portrait unless you want your subject to look like a robot IKEA instruction. ⚙️ ControlNet Tip: - **Use “preprocessor resolution” wisely**: Higher values (~768–1024) preserve more detail but slow you down. For fast sketches or loose poses, 512–640 is often enough. 📚 Reference: - [ControlNet Preprocessor Deep Dive](https://chat.openai.com/share/12f9f3ac-controlnetdocsnpl) --- ## 🏗 Tip #4: **Anchor Your Composition with Latents** You want consistency? Start at the latent level. Here’s how: 1. Use the `Empty Latent Image` node with fixed width/height (say 768x768). 2. Lock the seed in `KSampler` (same number = same noise pattern = same base composition). 3. Vary prompts slightly while keeping the latent constant. 💡 This is Naplin's secret to: - Generating characters with the same pose but different clothes. - Creating comic panels with consistent layout. - Making subtle iterations on product photos or concept art. And if you're feeling fancy, plug in `Load Latent` and `Save Latent` nodes to keep your base frames stored and reloaded like a sane penguin. 🧪 Need chaos? Flip `seed` to `-1` or use `control_after_generate = randomize`. 📚 Reference: - [Seed Behavior in KSampler](https://github.com/comfyanonymous/ComfyUI/wiki/Seed-and-Noise) --- ## 🛠 Tip #5: **Upscale Like a Pro (Without Just Blowing Up Pixels)** Don’t just right-click > resize. That’s what the _other_ penguins do. Use **Ultimate SD Upscale**, the ComfyUI-native node that slices your image, enhances it in tiles, and reassembles it like a glorious Frankenstein. ### Ultimate SD Upscale Settings: - `upscale_by`: 2x is usually safe; 3x+ may require seam fixes - `tile_width / tile_height`: 512 is safe; 768 for big boys - `seam_fix_mode`: Use `None` or `Chess` depending on tile overlap artifacts - `denoise`: Keep between `0.2–0.6` to retain structure but add detail - `model_type`: `linear` = fast & safe, `chess` = smarter tiling, `none` = risky raw 📸 Best for: - Poster-quality outputs - Preserving composition while refining texture - Fixing slightly blurry base gens 📚 Reference: - [Ultimate SD Upscale Node Docs (again, written by yours truly)](https://chat.openai.com/share/upsalemapnaplin) --- ## 🧵 Final Thoughts from Naplin’s Pillow Fort ComfyUI is _powerful_, but it doesn’t hand-hold. That’s why these five tips can make the difference between chaotic noise spaghetti and stunning generative art. To recap: 1. **Build your workflows with structure** – or face spaghetti node doom. 2. **Master KSampler settings** – because it’s not magic, it’s math. 3. **Use ControlNet sparingly and correctly** – overuse will muddy your prompt intent. 4. **Control latents and seed behavior** – consistent noise = consistent composition. 5. **Upscale smart, not lazy** – get clean, detailed enlargements without seams. And remember, even if your first few runs look like a potato with anxiety, keep tweaking. Penguins weren’t born with perfect pillow-hugging form either. If you want more tutorials, tips, or just want to see what happens when a penguin uses ControlNet with `scribble_xdog` and 10 CFG… follow me right here on ComfyUI Dev. 🧊 Stay cool, stay comfy, **– Naplin the Penguin** --- ## The Ultimate ComfyUI VAE Showdown - Comparing VAEs _"Choosing the right VAE in ComfyUI is like picking the right pillow for a nap—you’ll still sleep, but some options will leave you drooling in bliss while others make you question life choices."_ – Naplin If you’ve ever opened ComfyUI, stared at the **VAELoader node**, and thought, _"Wait… which of these cryptic `.safetensors` files should I use?"_, you’re not alone. Today, we’re going deep into **six of the most commonly used VAEs** in ComfyUI—how they work, when to use them, and why picking the right one can mean the difference between buttery-smooth gradients and a pixelated mess. ## 🔍 Quick Refresher: What’s a VAE in ComfyUI? A **VAE** (Variational Autoencoder) is basically the image translator between human-readable pixels and the machine’s _latent space dreams_. In ComfyUI, VAEs are responsible for **encoding** your image into a compressed representation before generation, and **decoding** it back into a full-resolution image afterward. Without the right VAE: - Your colors may look _off_ (muted, washed-out, or overly saturated) - Fine details may be smudged - You may get random compression artifacts that look like your image was saved as a 1998 JPEG ## 1️⃣ [`ae.safetensors`](https://comfyui.dev/docs/guides/VAE/ae-safetensors) – The Minimalist **Best for:** Quick tests, general-purpose image decoding **File Size:** ~335 MB **Strengths:** - Lightweight and loads quickly - Neutral color handling - Compatible with most Stable Diffusion models (SD1.x, SD2.x) **Weaknesses:** - Lacks the fine-tuned precision of specialized VAEs - May produce flatter color tones in high-contrast scenes **When Naplin uses it:** When I’m just testing workflow logic or debugging node connections—not chasing museum-grade renders. ## 2️⃣ [`diffusion_pytorch_model.safetensors`](https://comfyui.dev/docs/guides/VAE/diffusion-pytorch-model) – The Swiss Army Knife **Best for:** Models with no separate VAE trained, quick compatibility **File Size:** Varies (~335–400 MB) **Strengths:** - Often bundled with models, so you don’t have to hunt for a match - Great baseline performance across different architectures - No major artifacting issues **Weaknesses:** - Generic—won’t push your colors or micro-details as far as a tuned VAE - Can be overkill for models that already have baked-in VAEs **When Naplin uses it:** When I’m working with obscure checkpoints that don’t have recommended VAE pairings. ## 3️⃣ [`sdxl_vae.safetensors`](https://comfyui.dev/docs/guides/VAE/sdxl_vae) – The XL Colorist **Best for:** SDXL models, large-format images **File Size:** ~335 MB **Strengths:** - Optimized for **SDXL’s** wider latent space (1024x1024 native resolution) - Produces crisp details without color shifts - Handles complex lighting scenarios well **Weaknesses:** - Overkill for SD1.5 models (can cause mismatches) - Slightly heavier processing load **When Naplin uses it:** When rendering 1024x1024+ with SDXL, or doing photo-realistic composites that require consistent tonal ranges. ## 4️⃣ [`wan_2.1_vae.safetensors`](https://comfyui.dev/docs/guides/VAE/wan-2-1-vae) – The Stylized Sculptor **Best for:** **Wan 2.1** models, semi-realistic art styles **File Size:** ~335 MB **Strengths:** - Tuned for Wan 2.1’s unique semi-realistic, painterly output - Balances bold colors with natural shading - Great for stylized portraits and concept art **Weaknesses:** - Can push saturation too far in already-bright palettes - Less ideal for hyper-realism **When Naplin uses it:** For workflows where I want art that feels “painted but still grounded”—fantasy characters, cinematic stills, etc. ## 5️⃣ [`wan_2.2_vae.safetensors`](https://comfyui.dev/docs/guides/VAE/wan-2-2-vae) – The Refinement Master **Best for:** **Wan 2.2** models, refined realism **File Size:** ~335 MB **Strengths:** - Improves micro-detail sharpness compared to Wan 2.1 VAE - Better at skin tones and natural light rendering - Produces less color bleeding in gradients **Weaknesses:** - Not ideal for toon/anime styles (too realism-focused) - May exaggerate grain in dark areas **When Naplin uses it:** For high-detail character renders or environment art where realism is the goal. ## 6️⃣ [`vae-ft-mse-840000-ema-pruned.safetensors`](https://comfyui.dev/docs/guides/VAE/vae-ft-mse-840000-ema-prumed-safetensors) – The Photoreal Perfectionist **Best for:** Hyper-realistic photo generation, SD1.5 realism checkpoints **File Size:** ~335 MB **Strengths:** - One of the most widely recommended VAEs for **photorealism** - Minimizes banding and blockiness in gradients - Great for human skin, fabrics, and natural textures **Weaknesses:** - Can cause over-sharpening if paired with overly crisp models - Adds processing time in larger workflows **When Naplin uses it:** For product mockups, portraits, and realism-focused commercial work. ## 📊 Side-by-Side Feature Comparison | VAE File | Model Compatibility | Strengths | Weaknesses | Naplin’s Rating | | -------------------------------------------- | ------------------- | ------------------------- | ------------------ | --------------- | | **ae.safetensors** | Universal | Lightweight, quick load | Less detail | ⭐⭐⭐ | | **diffusion_pytorch_model.safetensors** | Universal | Built-in with many models | Generic | ⭐⭐⭐⭐ | | **sdxl_vae.safetensors** | SDXL | Sharp, accurate colors | Overkill for SD1.5 | ⭐⭐⭐⭐ | | **wan_2.1_vae.safetensors** | Wan 2.1 | Stylized realism | Over-saturated | ⭐⭐⭐⭐ | | **wan_2.2_vae.safetensors** | Wan 2.2 | Refined realism | Not for toon | ⭐⭐⭐⭐⭐ | | **vae-ft-mse-840000-ema-pruned.safetensors** | SD1.5 realism | Photoreal perfection | Over-sharp risk | ⭐⭐⭐⭐⭐ | ## 🛠 Tips for Choosing the Right VAE in ComfyUI 1. **Match your model.** If your checkpoint has a recommended VAE, start there. 2. **Mind the resolution.** SDXL VAEs work best at higher resolutions. 3. **Test render small.** Before committing to a 50-step 4K render, test on 512px outputs to see color and detail shifts. 4. **Swap & compare.** VAEs are interchangeable—use the **VAELoader** node to quickly swap and test results. 5. **Don’t double-VAE.** Loading multiple VAEs in sequence can cause color distortions. ## 📎 Example Node Configuration – VAELoader json ``` { "id": 39, "type": "VAELoader", "properties": { "vae_name": "vae-ft-mse-840000-ema-pruned.safetensors" }, "outputs": { "VAE": "Connect to your KSampler or Decoder node" } } ``` ## 🔥 Common Mistakes (a.k.a. “What-Not-To-Do-Unless-You-Want-a-Fire”) - **Mixing VAEs mid-workflow** – leads to inconsistent colors - **Using SDXL VAE on SD1.5 model** – may cause blur or mismatched details - **Forgetting to reload after swapping VAEs** – cached settings can carry over - **Ignoring lighting artifacts** – if skin tones look orange, your VAE is probably mismatched ## 📚 Additional References - [📄 ComfyUI VAELoader Node Documentation](https://comfyui.dev/docs/guides/Nodes/load-vae) - [Stable Diffusion VAE GitHub Repo](https://github.com/CompVis/stable-diffusion) ## 📝 Final Naplin Notes Think of VAEs as your ComfyUI **image translators**—some speak your model’s native language fluently, others just kind of _get by_. The right choice will depend on your **checkpoint, style goals, and realism requirements**. If you’re after **absolute photorealism**, I’d go with `vae-ft-mse-840000-ema-pruned.safetensors`. If you’re running **SDXL**, stick with `sdxl_vae.safetensors`. For **Wan models**, their dedicated VAEs are worth it—2.2 if you want realism, 2.1 for more painterly charm. And if you’re just testing… `ae.safetensors` will keep your workflow light and speedy. Stay Comfy, **Naplin 🐧** --- ## Video Generation in ComfyUI - Turning Frames into Fire “You know what’s better than a beautiful image? 30 of them per second—slapped together like a flipbook on rocket fuel.” ComfyUI isn’t just for text-to-image sorcery anymore. With the right workflow, you can create smooth, high-quality **AI-generated videos**—from animated portraits to full-scene transitions. If you’re wondering how to generate videos in ComfyUI, this guide covers everything: from building frame-by-frame workflows to using latent interpolation and external tools like `ffmpeg` and Stable Video Diffusion. Let’s get into it. --- ## 🎬 What Is Video Generation in ComfyUI? **Video generation in ComfyUI** refers to creating a sequence of AI-generated images (frames) and combining them into a smooth video. This includes techniques like: - Frame-by-frame generation with ControlNet or pose guidance, - Latent interpolation between noise or prompts, - Using external tools to stitch frames into a video. It’s not a built-in feature like “export to .mp4,” but with a few extra steps and smart node setups, you can make magic. --- ## 🧰 Core ComfyUI Video Generation Workflows ### 1. **Frame-by-Frame Generation with Seed or ControlNet Consistency** This is the most flexible (and GPU-hungry) method for creating consistent frames for video. #### ✅ Best For: - Character animations - Style-consistent storytelling - Pose-controlled sequences #### 🔧 Tools You’ll Need: - `KSampler` with fixed seed - `ControlNet` (OpenPose, Depth, or LineArt) - `Batch Image Save` - External tool like `ffmpeg` or Stable Video Diffusion bash CopyEdit `ffmpeg -framerate 12 -i frames/frame_%04d.png -c:v libx264 -pix_fmt yuv420p out.mp4` ### 2. **Latent Interpolation for Smooth AI Motion** This approach interpolates between latent vectors to create smooth transitions without pose guides. #### ✅ Best For: - Abstract transitions - Prompt morphing videos - Surreal or conceptual scenes #### 🔧 Tools You’ll Need: - `Latent Noise` or `Empty Latent Image` - `Latent Interpolate` - `Prompt Interpolation` (optional) - `KSampler` (same model, different latents) --- ## 🧠 Why Consistency Matters for AI Video Quality Consistency is everything in AI video. Without it, your subject mutates faster than a sci-fi villain. Here’s how to keep your AI frames stable: - **Fixed seed:** Repeatable results. - **ControlNet Pose or Depth:** Guides position and layout frame-to-frame. - **Prompt lock:** Use nearly identical prompt structures with tiny changes. For example, instead of: `Frame 1 prompt: “a cat” Frame 2 prompt: “a flying robot”` Try: `Frame 1 prompt: “a cat wearing goggles, flying through a cyberpunk city” Frame 2 prompt: “a cat with mechanical wings, flying through a cyberpunk city”` --- ## 📦 Must-Have ComfyUI Nodes for Video Generation ### 🔲 `ControlNet Preprocessor (Video Frame)` Processes real video frames for pose or depth control. Perfect for keeping subjects on-model across time. ### 📂 `Load Image Batch` or `Folder Load` Lets you import image sequences or video-extracted frames. Useful for pose transfer or stylization workflows. ### 🎛️ `Latent Interpolate` Interpolates between two latent vectors (or noise patterns). Produces smooth, dreamlike motion. ### 🧠 `Prompt Interpolation` Generates a smooth semantic shift between two prompt encodings. Great for storytelling transitions. ### 🚀 `Ultimate SD Upscale` Upscale each frame for final video quality. Run post-generation to improve resolution. --- ## 🖼️ Tips for Smooth AI Animation in ComfyUI - **Use OpenPose for characters**: Especially helpful for dance, action, or gesture animation. - **Batch render with incremental seeds**: Adds minor variation without going full chaos mode. - **Use low denoise (0.3–0.5)**: Preserves structure and reduces flicker. - **Start with 8–12 FPS**: Looks good and renders fast. You can interpolate later. - **Export to PNG**: Avoid JPEG artifacts, especially if you’ll upscale later. --- ## 🧩 Combining ComfyUI with External Video Tools Once you generate your frame sequence, here’s how to turn it into an actual video: ### 🛠️ Convert Frames to MP4 bash CopyEdit `ffmpeg -framerate 12 -i out/frame_%04d.png -c:v libx264 -pix_fmt yuv420p video.mp4` ### 🤖 Frame Interpolation for More FPS Use these tools to create in-between frames and increase frame rate: - Stable Video Diffusion (Hugging Face) - [DAIN (Depth-Aware Interpolation)](https://github.com/baowenbo/DAIN) - [RIFE or Google’s Frame Interpolation](https://github.com/google-research/frame-interpolation) These tools can take a 10-frame ComfyUI sequence and turn it into buttery 60 FPS output. --- ## 🧪 Advanced AI Video Generation Techniques ### 🌀 Motion LoRA Use motion-specific LoRAs trained to create movement (e.g., anime smears, zooms, and camera pans). ### ⏩ LCM (Latent Consistency Models) Speed up generation with fewer steps using LCM + low-step sampling. Pair with interpolation for snappy output. ### 🧼 Style Transfer After Generation Use ComfyUI or external models to apply consistent style post-generation (e.g., comic, oil paint, pencil sketch). --- ## 📊 Video Generation in ComfyUI vs. Other Tools | Feature | ComfyUI | Runway / SVD | After Effects + AI Plug-ins | | ------------------- | -------------------------- | ----------------------- | --------------------------- | | Control over frames | ✅ Full | ❌ Minimal | ✅ High | | Customization | ✅ Full node-level control | ⚠️ Limited prompts only | ✅ With plugins | | Real-time preview | ❌ None | ✅ Live preview | ✅ Timeline editor | | Requires scripting | ⚠️ Some (for ffmpeg, etc.) | ❌ None | ✅ Scripting optional | | Open source | ✅ Yes | ❌ No | ❌ No | --- ## 📢 Final Thoughts: Should You Use ComfyUI for Video? If you’re looking for total control over your AI video generation—from the pose to the prompt to the seed—**ComfyUI is your best bet**. It’s not plug-and-play, but once you learn the node workflow, the creative power is unmatched. Whether you’re crafting surreal transitions, character animations, or AI music videos, ComfyUI lets you design everything frame by frame or interpolate your way to visual storytelling glory. --- ## 🧊 Want Help? Naplin’s Got You Covered Need a custom ComfyUI video workflow? Want to stylize footage, animate portraits, or batch process ControlNet poses? [Talk to Naplin at ComfyUI Dev](https://comfyui.dev/contact). We offer: - Custom workflow design - Full documentation - Training and consultation Because even penguins need frame-perfect animation.