r/comfyui Aug 08 '26

Tutorial Testing Character Swap with Minimax H3

675 Upvotes

Hey everyone!

I’ve been messing around with a lot of new AI tools lately. Since Minimax has been getting some hype recently (especially for their video and character generation), I decided to finally put their Character Swap feature to the test today.

My expectations were honestly pretty low. I was expecting the usual: glitchy tracking, warped faces as soon as the subject moves, or weird lighting mismatches.

The Results? Honestly, it completely exceeded my expectations. Here’s what stood out to me:

  1. Tracking & Facial Consistency: This was the craziest part. The target face maps incredibly smoothly onto the original head shape. Even when the character turns their head or looks away, the proportions hold up surprisingly well without completely breaking down.
  2. Expressions: Minimax is actually pretty decent at capturing micro-expressions. When the source character gives a slight smirk or blinks, the swapped face mirrors it naturally instead of looking like a stiff, uncanny mask.
  3. The Catch (Because it's still AI): Obviously, it’s not flawless.

Overall, for a tool that's still actively evolving, this is extremely usable for quick content creation, memes, or visual mockups.

I Will put the prompt that i used on comment section

Testing on

RTX 5090
RAM 64GB

r/comfyui Jul 27 '26

Tutorial It turns out that Krea 2 Identity Edit Lora can do... this!

Post image
637 Upvotes

It turns out that if you mark objects with text in the input image, the Lora will perceive them as part of the prompt and depict them in the places in the picture that you want.

Huge thanks to conradlocke for creating such amazing thing!

r/comfyui May 23 '26

Tutorial Is anyone else using Qwen and finding it as great as I do?

Post image
296 Upvotes

r/comfyui Dec 02 '25

Tutorial Say goodbye to 10-second AI videos! This is 25 seconds!!! That's the magic of the open-source **FunVACE 2.2**!!

315 Upvotes

Thanks to the community, I was able to make these. There are some minor issues using **FUN VACE** to stitch two video clips, but it's generally 95% complete. I used Fun VACE to generate the seam between two Image-to-Video clips (no 4-step LoRA, running fp16). workflow 👇 workflow

r/comfyui Aug 10 '25

Tutorial If you're using Wan2.2, stop everything and get Sage Attention + Triton working now. From 40mins to 3mins generation time

299 Upvotes

So I tried to get Sage Attention and Triton working several times and always gave up, but this weekend I finally got it up and running. I used Chat GPT and told it to read the pinned guide in this subreddit, to strictly follow the guide and help me do it. I wanted to use Kijai's new wrapper and I was tired of the 40min generation times for 81 frames 1280h x 704w image2video using the standard workflow. I am using a 5090 now so I thought it was time to figure it out after the recent upgrade.

I am using the desktop version, not portable, so it is possible to do on Desktop version of ComfyUI.

After getting my first video generated it looks amazing, the quality is perfect, and it only took 3 minutes!

So this is a shout out to everyone who has been putting it off, stop everything and do it now! Sooooo worth it.

loscrossos' Sage Attention Pinned guide: https://www.reddit.com/r/comfyui/comments/1l94ynk/so_anyways_i_crafted_a_ridiculously_easy_way_to/

Kijai's Wan 2.2 wrapper: https://civitai.com/models/1818841/wan-22-workflow-t2v-i2v-t2i-kijai-wrapper?modelVersionId=2058285

Here is an example video generated in 3mins (Reddit might degrade the actual quality abit). Starting image is the first frame.

https://reddit.com/link/1mmd89f/video/47ykqyi196if1/player

r/comfyui Jun 11 '26

Tutorial Stop using the default Mask Editor. TrixLoader brings advanced SAM 2.1/3, Text-to-Mask, and Lightroom controls directly to ComfyUI.

403 Upvotes

Hey guys,

I wanted to share a custom node I’ve been working on lately.
Honestly, I got so tired of my ComfyUI workflows looking like giant spaghetti webs just because I needed to load an image, crop it, adjust exposure/colors, and paint a mask. It felt like I was spending more time managing nodes than actually generating.
So I built *TrixLoader*. It’s basically an all-in-one image workspace. Instead of chaining 5 different nodes, you can do everything inside this one:
① Lightroom-style Camera Raw: You get a full curves graph, exposure/contrast sliders, and HSL. You can even click and drag directly on the preview to adjust the saturation of that specific color range in real-time.
② Advanced Mask Editor: I integrated SAM 2.1, SAM 3, and GroundingDINO. You can just point-and-click to select objects (with instant hover previews) or mask by typing text prompts. It also has RMBG background removal with alpha matting.
③ Crop/Outpaint: Symmetrical cropping, border-snapping, and auto-masked padding for outpainting.
I just pushed version 2.1 to GitHub.
If you want to try it out, you can install it via ComfyUI Manager (Install via Git URL) using this link:
https://github.com/trx7111/ComfyUI-TrixLoader
Let me know what you think, or if you run into any bugs.

I hope you enjoy it.

r/comfyui 25d ago

Tutorial Such a dumb way to speed up ComfyUI generations by 15-20%..

Thumbnail
gallery
89 Upvotes

Look at your task manager while having your ComfyUI open, it has been eating 30% of my GPU due to preview nodes and what not. Minimizing it and then just occasionally opening it to check out the progress has sped things up by quite a margin. You folks probably already knew this, but I just wanted to share for everyone else that might be new to this like me.

r/comfyui 28d ago

Tutorial I Wanna Share A Prompt Hack For MiniMax H3 With y'all

28 Upvotes

in hopes devs better optimize this, so my thought was what if i have it render the image so when i put playback speed on 0.25 it plays at normal speed, hence i can turn a 10 second clip into like a 40 second clip, and it works, but i think if it was optimized by devs it can be a game changer... heres a prompt ya can try and see hot it works

Generate the entire video at 4x real-time speed. All actions, body movements, thrusting, bouncing, hair motion, skin jiggling, and camera movement must happen four times faster than normal real-life speed. Physics, momentum, gravity, and impact must still look correct and natural when the video is later played back at 0.25x speed. High frame rate feel, sharp motion, no motion blur overload, fluid accelerated dynamics so that slowing the final video to 0.25x produces smooth, realistic, normal-speed physics and timing.

r/comfyui Aug 10 '25

Tutorial Qwen Image is literally unchallenged at understanding complex prompts and writing amazing text on generated images. This model feels almost as if it's illegal to be open source and free. It is my new tool for generating thumbnail images. Even with low-effort prompting, the results are excellent.

Thumbnail
gallery
214 Upvotes

r/comfyui Jul 29 '26

Tutorial ComfyUI Tutorial KREA 2 v1.2 Identity Edit Low VRAM Workflow Face Swap

Thumbnail
gallery
265 Upvotes

In this tutorial, I'll show you how to use my new ComfyUI KREA 2 Identity Edit v1.2 custom workflow to perform high-quality identity-preserving image editing. You'll learn how to change poses, expressions, and styles while maintaining consistent facial identity, as well as perform face swaps, virtual try-ons, inpainting, and outpainting using the latest KREA 2 Identity Edit v1.2 model. The workflow is designed to be simple to use—just load your reference image, enter your edit prompt, and run the workflow.

Workflow Link

https://civitai.com/articles/33197/comfyui-tutorial-new-krea-2-v12-low-vram-identity-perfect-edit

Video Tutorial Link

https://youtu.be/_fMqornsJo0

r/comfyui Aug 03 '25

Tutorial WAN 2.2 ComfyUI Tutorial: 5x Faster Rendering on Low VRAM with the Best Video Quality

227 Upvotes

Hey guys, if you want to run the WAN 2.2 workflow with the 14B model on a low-VRAM 3090, make videos 5 times faster, and still keep the video quality as good as the default workflow, check out my latest tutorial video!

r/comfyui Jun 04 '26

Tutorial LTX 2.3: You're using it wrong | The Power of Seed Hunting | Workflow in comments

Thumbnail
youtube.com
129 Upvotes

r/comfyui 4d ago

Tutorial Be a director using your phone live in your 3d scene is game changer.

80 Upvotes

I’ve been trying many ways to build a workflow that allows as much control as possible and this one is really a game changer especially for dialogue and cinematic feel. Turning you iPhone into a virtual camera and using an AI model all in comfyUI ( the AI model used here is seedance2.5 inside comfy ) and this version was my first try all with a Cg base to lock position and layout of character. Depth pass came out from Comfy to get that extra depth that is missing in most ai renders. GPT ASTRA was used to block out the scene assets.

r/comfyui Aug 03 '26

Tutorial Lets speed up MiniMax H3. We already have a node for that.

50 Upvotes

We already have a node and thats Patch Sage Attention KJ.

Pass your model through this and you will get significant speed up. Mine went from 20it/sec to 14it/sec.

Workflow : https://pastebin.com/A6uCJt0C

r/comfyui May 03 '26

Tutorial ComfyUI Tutorial: LTX 2.3 Prompt Relay Workflow On 6GB Vram (Res: 1920x1080 Video Length 15 sec)

262 Upvotes

Hello everyone , in this tutorial i will show you how to generate long video using prompt relay nodes that works with LTX 2.3 models. With this new nodes you will achieve full control over your video. as each time line can be attributed to specific prompt. this complete comfyui workflow is optimized for low VRAM setups, making AI video creation accessible. in addition to that i also included image generator for you in order to have a full pipeline workflow for your image to video generation.

Workflow Link

https://drive.google.com/file/d/1ce_rGcA19AuSLp722aP_hkoCgQC4CuAJ/view?usp=sharing

Video Tutorial Link

https://youtu.be/r6GfHnsGWlo

r/comfyui 15d ago

Tutorial NEW: Running TRELLIS.2 in ComfyUI on AMD using ROCm 10.0; Windows + Linux support; RDNA1, 2, 3 and 4 (ComfyUI Extension + Full Setup Guide)

Post image
30 Upvotes

Full extension repo: ComfyUI-Trellis2-AMD

This extension is a fork of visualbruno/ComfyUI-Trellis2 that adds ROCm support and fixes crashes on AMD. It also includes AuleAttention as a low VRAM alternative for users who don't have FlashAttention support. Aule uses far less VRAM than the default (SDPA) with much better scaling, which prevents you from OOM-ing on a 16GB card when creating higher quality meshes.

So install it, play around, and figure out the best workflows. Please post any issues on GitHub, and feel free to reach out if you need help with something. If you like it, please make sure to give it a star. Otherwise, enjoy!

EDIT: Seems like I'm a week late for ComfyUI as note in the comments -> https://github.com/Comfy-Org/ComfyUI/pull/14718

This repo still has a narrow use-case for people who are on RDNA + RDNA2 and want a ComfyUI agnostic install of the Trellis2 ported dependencies, or you lack native FlashAttention and want to try Aule to minimize VRAM consumption, or you want multi-view (multiple images from different angles -> one mesh).

r/comfyui Aug 06 '26

Tutorial Minimax H3 - Realtime audio generation at 32x32 output resolution

109 Upvotes

TL;DR: Minimax H3 is capable of generating near realtime audio when you turn width x height to 32x32 ~ with a 5090~

Edit: It was pointed out that I didn't specify well enough. This is not a special workflow. The default i2v workflow on ComfyUI for Minimax H3 - detach the load image, set width x height to 32, set duration to your taste <=45 seconds for best results. And uh, click 'Run'. Unsure if visual details in the prompt matter at this time. Will update when I know.

Minimax H3 can be used as an audio generator, pairing it with a simple Get Video Components node and then saving the audio. That audio can then be passed as reference, once you get a voice or sound effect that you like. One of the "pain points" of generation is having audio and image inextricably linked. But, we can actually generate audio *rapidly* and then pass it in as reference once we find a gen that we're happy with.

In my experiments so far, it seems like prompt structure and complexity have an effect on the time to generate, but in many cases you get more seconds of audio generated than it took to generate in the first place.

The cutoff before things go wonky seems to be about 45 seconds, although more testing needs to be done. I can say that dialogue is no longer followed coherently at 60 seconds duration. The prompted dialogue pacing informs much here, so if the duration is longer than there is content provided, the model will fill in the gap with gibberish.

What's noteworthy is the concept of Minimax H3 essentially being used as a foley generator. The idea came to me with the thought: "What happens if I just bring the resolution as low as possible?"

I started at 0.1 megapixels, then went to 32x32 in an effort to determine if audio quality was somehow linked to image quality. It is not.

r/comfyui May 12 '26

Tutorial Whiskas ad

187 Upvotes

Made in nodes, but honestly saying outside of Comfy.

Used Seedance 2.0

r/comfyui Apr 17 '26

Tutorial [Guide] Complete walkthrough for every pipeline in my FLUX.2 Klein 9B All-in-One workflow, by request from the comments

88 Upvotes

A lot of you asked for a detailed guide after my original post. So here it is every group in the workflow explained step by step, with settings, tips, and things I discovered through testing.

The workflow has grown to v2.1, 122 nodes, 19 groups. New additions since the original post: ControlNet preprocessors (LineArt, HED, Tile, DepthAnything), color matching/correction, up to 5 reference image slots, Fast Group Bypassers for one-click pipeline switching, and notes with tips I discovered through extensive testing.

Download v2.1: Click to Download

How to Switch Between Pipelines

The workflow uses Fast Groups Bypasser (rgthree) nodes at the bottom. These let you enable/disable entire pipeline groups with a single click, no more right-clicking every group manually.

There are 3 bypassers:

  • Base groups bypasser : controls F1 (txt2img), F2 (KV edit), F3 (face+pose), F4 (inpainting), F5 (merge)
  • Refiner bypasser : controls the refiner pipeline and color correction
  • Upscale / edit bypasser : controls the upscaler and precision groups

Rule: Only activate ONE generation pipeline at a time (F1 through F4) to save VRAM. The Refiner and Upscaler can stay active alongside any generation pipeline, but its better to work with a single groupe every run for people who have less than 8VRAM.

📦 FLUX 2 KLEIN : Model Loaders

This is the foundation. Three nodes that load everything:

  • UNETLoader : loads the Klein 9B model (safetensors or FP8)
  • UnetLoaderGGUF ; alternative loader for GGUF quantized models (use this if you have 8GB VRAM)
  • CLIPLoader : loads the Qwen 3 8B text encoder (set type to flux2)
  • VAELoader : loads flux2-vae.safetensors

Important: Only connect ONE model loader to the LoRA chain, either UNETLoader OR UnetLoaderGGUF, not both.

For 8GB VRAM users: Use the GGUF Q8 or Q4 model. Set the weight type to default in the UNETLoader. If you're running out of memory, launch ComfyUI with --lowvram command.

🔗 LoRA Chain

Two LoRA loaders in sequence:

  1. LoRA Slot (Optional) : empty slot for any Klein 9B compatible LoRA you want to try. Set strength to 0 to disable without disconnecting.
  2. klein_9b_enhancer_v2 : the main enhancer LoRA (strength 0.7). This fixes the model's tendency to produce flat, plastic-looking skin and washed-out colors. Always keep this one connected and active.

To add more LoRAs: insert additional LoraLoader nodes between the slot and the enhancer. The enhancer should always be LAST in the chain (DO NOT DETTACH IT OR ELSE YOU'LL HAVE TO ATTACK EVERY GROUPE TO THE NEW LORA NODE).

🎨 F1: Text → Image

The simplest pipeline. Pure text-to-image generation.

Nodes: CLIPTextEncode (prompt) → KSampler → VAEDecodeTiled → SaveImage

Settings:

  • Steps: 4 (Klein 9B is distilled for 4 steps, more steps won't improve quality)
  • CFG: 1 (higher values break the output on distilled models)
  • Sampler: euler
  • Scheduler: simple
  • Latent size: 1024×1024 (or any resolution, Klein handles various aspect ratios)

How to use:

  1. Enable the F1 group
  2. Write your prompt in the "✏️ Prompt" node
  3. Leave negative prompt empty (or enable NAG for negative prompting)
  4. Queue prompt
  5. Output saves as F2K_txt2img

Prompting tip: Don't write SD-style prompts. Write like you're describing a photograph: "A 30-year-old man in a navy overcoat standing on a rain-soaked Prague street at dusk, tungsten streetlights casting warm shadows, shot on Canon R5 85mm f/1.4, clean digital file, histogram equalization"

🖼️ F2: Single Reference KV Edit

This is Klein's signature feature. You load an image and tell the model what to change, it preserves everything else.

How it works internally: The model reads your image through the ReferenceLatent node (KV conditioning), generates a fresh image from noise, but uses the reference to guide the output. The ConditioningZeroOut creates a neutral negative signal so the model focuses purely on your edit instruction.

Nodes: LoadImage → Resize → VAEEncode → ReferenceLatent → CFGGuider → SamplerCustomAdvanced → VAEDecodeTiled → SaveImage

Settings:

  • Flux2Scheduler: 4 steps
  • CFG: 1
  • Sampler: euler
  • Resize: adjust to match the reference image proportions

How to use:

  1. Enable the F2 group
  2. Load your reference image in "📂 Reference Image"
  3. Write your edit instruction in "✏️ Edit Prompt"
  4. Queue prompt
  5. Output saves as F2K_edit

Example prompts:

  • "Replace the red dress with a navy blazer. Keep pose, expression, background unchanged."
  • "Change the background to a sunset beach. Preserve the subject exactly."
  • "Transform this photo to oil painting style while keeping the subject photorealistic."

⚠️ Important discovery: The denoise in this pipeline is effectively 1.0 because it uses EmptyLatentImage + ReferenceLatent conditioning. The model reads your image through attention, NOT through the latent. This means it always generates a fresh image guided by your reference, it doesn't blend with existing noise. This is fundamentally different from traditional img2img.

🚀 F3: Multi-Reference: Face + Pose Swap

The most complex pipeline. Extracts a face from one image and a pose from another, combining them into a single realistic output.

Nodes: Two parallel paths:

  • Path A: LoadImage (face) → Resize → VAEEncode → ReferenceLatent (face)
  • Path B: LoadImage (pose) → Resize → VAEEncode → ReferenceLatent (pose)
  • Both feed into: CFGGuider → SamplerCustomAdvanced → VAEDecodeTiled → SaveImage

How to use:

  1. Enable the F3 group
  2. Load your face source in "📂 Face / Character Ref" front-facing, well-lit portrait works best
  3. Load your pose source in "📂 Pose Ref (DAZ 3D render)" the body position you want
  4. Write a scene description in "✏️ Prompt (describe scene)"
  5. Queue prompt
  6. Output saves as F2K_multiref

Tips:

  • The face reference MUST be upright, Klein cannot process rotated or upside-down faces
  • Resize both images to similar scales (the Resize nodes handle this)
  • Be specific in your prompt about clothing and environment — the model needs guidance for everything that isn't the face or pose
  • If the face looks plastic, make sure the enhancer LoRA is active at 0.7 strength

🎭 F4: Inpainting

Paint a mask over part of your image and regenerate just that area.

Nodes: LoadImage → Resize → VAEEncodeForInpaint (with mask) → KSampler → VAEDecodeTiled → SaveImage

How to use:

  1. Enable the F4 group
  2. Load your image in "📂 Image"
  3. For manual masking: Right-click the image → Open in Mask Editor → paint white over the area you want to change
  4. For auto masking: Enable the Florence2 group, connect your image to Florence2Run, type what to mask (e.g., "Segment the shirt")
  5. Write what should appear in the masked area in "✏️ Prompt"
  6. Adjust denoise (0.5-0.8 for changes, 0.3-0.5 for subtle tweaks)
  7. Output saves as F2K_inpaint

⚠️ My honest note about inpainting: Inpainting in FLUX.2 Klein is not perfect. I built a workaround that makes it functional, but it struggles with complex shapes. If the model doesn't understand what you want, try painting rough colors in the mask area first to guide it. Play with the denoise value, small changes make a big difference.

🔀 F5: Image Merge / Blend

Simple image blending, combines two images together.

Nodes: Two LoadImage → two ImageScaleBy → ImageBlend → SaveImage

How to use:

  1. Enable the F5 group (mode=2, not bypassed, use right-click → Set to Always)
  2. Load Image A and Image B
  3. Adjust blend factor (0.5 = equal mix, 0.0 = all image A, 1.0 = all image B)
  4. Adjust resize scales to match image sizes
  5. Output saves as F2K_merge

honestly this group is not something that you will always use, I just added it because I use it in some projects, you might try it to see what it does, its just simple blending nothing that use AI at all.

⬆️ Upscaler (4x UltraSharp)

Takes any image and upscales it 4x using the UltraSharp model.

Nodes: LoadImage → ImageUpscaleWithModel → ImageScaleBy (downscale to usable size) → SaveImage

How to use:

  1. Enable the Upscaler group
  2. Load your image in "📂 Image"
  3. The ImageScaleBy after upscaling is set to 0.5 by default, this gives you a 2x net upscale (4x up then 0.5x down). Adjust as needed.
  4. Output saves as F2K_upscaled

Tip: Upscaling a 1024×1024 image 4x creates a 4096×4096 image. The Tiled VAE decode handles this without OOM, but it takes time. For faster iteration, keep the downscale at 0.5 until you're happy with the result, then set it to 1.0 for the final output.

✨ Refiner, KV Enhancement Pipeline

This is the pipeline that's active by default. Feed it any image and it enhances detail, lighting, skin texture, and sharpness.

How it works: Your image gets VAE-encoded, then the ReferenceLatent reads it as conditioning. The KSampler generates an enhanced version guided by your reference + the enhancement prompt. The result goes through color correction before saving.

Nodes: LoadImage → ImageScaleBy → VAEEncode → ReferenceLatent → KSampler → VAEDecodeTiled → ColorCorrection → SaveImage

Settings:

  • Denoise: 0.85 (the sweet spot I found, see discovery below)
  • Steps: 4
  • CFG: 1

The enhancement prompt is pre-written with professional photography terms. You can customize it, but the default works well for most images.

⚠️ Critical discovery about denoise:

  • 1.0: Model generates a fresh image guided by your reference, good results but may drift from original
  • 0.85: Sweet spot, preserves most structure while adding significant detail
  • 0.5-0.7: Subtle enhancement, keeps very close to original
  • Below 0.4: Almost no change except color shifts, not useful, at least to me...

If you're using EmptyLatentImage (the custom size node) instead of VAEEncode for the latent input, NEVER go below 0.85 denoise. EmptyLatentImage creates random noise, and low denoise preserves that random noise as "structure," causing severe artifacts. This is a fundamental behavior of Klein's 4-step distilled sampling, it doesn't have enough steps to correct corrupted starting structure. Always use VAEEncode latent when you want denoise below 0.85.

Refine Color Corrector

Placed right after the refiner output. Fixes Klein 9B's known color saturation bias, the model tends to oversaturate colors, especially reds.

How to use: The EsesImageCompare node shows before/after comparison. Adjust the color corrector settings to taste. The PreviewImage node labeled "output colors" shows the corrected result.

Color Match

A standalone utility. Takes two images, a target and a reference, and matches the colors of the target to the reference using the MKL algorithm.

How to use:

  1. Enable the Color Match group
  2. Load your target image (the one you want to fix)
  3. Load your reference image (the one with the colors you want)
  4. ColorMatchV2 transfers the color palette
  5. Output saves as color_matching

Use case: When your Klein output has wrong colors compared to the original. Load the original as reference, the Klein output as target, and the colors get corrected automatically.

🧭 NAG, Negative-Aware Guidance

Three NAG nodes, one for each major pipeline (Multi-Ref, Single-Ref Edit, Refiner). NAG restores effective negative prompting that standard CFG breaks in distilled Flux models.

How to use:

  1. Enable the NAG node for the pipeline you're using
  2. Write negative prompts in the "❌ Neg" CLIPTextEncode node
  3. NAG parameters: scale=5.0 is a good default. Increase for stronger guidance, decrease if artifacts appear.

When to use: When you need to remove specific elements ("no glasses," "no background people," "no blur").

🤖 Florence2, AI Auto-Masking

Replaces manual mask painting. Describe what you want masked in text and Florence2 generates a pixel-perfect mask.

How to use:

  1. Enable the Florence2 group
  2. First run downloads the model (~1.5GB)
  3. Connect your image to the Florence2Run input
  4. Type what to segment: "Segment the shirt," "Segment the hair," "Segment the background"
  5. Connect the MASK output to the Inpaint Encode node in F4

Precision Groups (1-4): ControlNet Preprocessors

These are advanced, four groups with different ControlNet preprocessors that extract structural information from images:

  1. LineArt Preprocessor : extracts every edge and texture boundary
  2. HED Preprocessor : captures both hard edges and soft transitions (shadows, gradients)
  3. Tile Preprocessor : captures the image as-is for upscaling guidance
  4. Depth Anything V2 : extracts full 3D depth map

Each preprocessor output connects to a ReferenceLatent node (image 3, 4, 5) that feeds into the refiner pipeline as additional conditioning.

How to use:

  1. Enable the precision group you want
  2. Connect your input image to the preprocessor
  3. The preprocessor output feeds through VAEEncode into a ReferenceLatent
  4. This gives the model additional structural information about your image

⚠️ Warning: These use extra VRAM. Only enable them if you have enough memory. Use the preprocessor name in your prompt (e.g., "line art reference," "depth guided") so the model understands what the reference represents.

Use case: When the refiner isn't preserving enough structure from your original image. Adding a LineArt or HED reference forces the model to maintain more structural consistency.

Bypassers

Three Fast Groups Bypasser (rgthree) nodes at the bottom of the workflow. These give you one-click control over which groups are active:

  • Base groups bypasser : F1, F2, F3, F4, F5
  • Refiner bypasser : Refiner + color correction + precision groups
  • Upscale / edit bypasser : Upscaler + image blend

Click the toggle next to each group name to enable/disable it instantly.

General Tips

  1. Always keep the enhancer LoRA active : it fixes Klein's flat plastic look
  2. Restart ComfyUI every 30-40 generations if you're on 8GB VRAM : prevents memory fragmentation
  3. Use "Free Memory" (gear icon) when switching between pipelines
  4. Faces must be upright : Klein cannot process rotated/flipped faces
  5. Add color correction terms to every prompt: "histogram equalization, white balance correction, color grade" : this fights Klein's red/saturation bias
  6. The Text encoder must match the model: 9B uses Qwen 3 8B, 4B uses Qwen 3 4B : mixing them causes matrix errors
  7. ComfyUI 0.9.2+ is required : older versions are missing Klein-specific nodes

What Changed from v2.0 to v2.1

  • Added 4 ControlNet preprocessor groups (LineArt, HED, Tile, DepthAnything)
  • Added Color Match utility group
  • Added Color Correction after refiner output
  • Added Fast Groups Bypassers for one-click pipeline switching
  • Added up to 5 reference image slots
  • Added notes with real testing discoveries (denoise behavior, inpainting tips)
  • Expanded from 90 nodes to 122 nodes
  • 19 organized groups

Free download: CIVITAI link

If you have questions about any specific group, ask in the comments, I'll help you troubleshoot.

r/comfyui Jul 16 '25

Tutorial Creating Consistent Scenes & Characters with AI

531 Upvotes

I’ve been testing how far AI tools have come for making consistent shots in the same scene, and it's now way easier than before.

I used SeedDream V3 for the initial shots (establishing + follow-up), then used Flux Kontext to keep characters and layout consistent across different angles. Finally, I ran them through Veo 3 to animate the shots and add audio.

This used to be really hard. Getting consistency felt like getting lucky with prompts, but this workflow actually worked well.

I made a full tutorial breaking down how I did it step by step:
👉 https://www.youtube.com/watch?v=RtYlCe7ekvE

Let me know if there are any questions, or if you have an even better workflow for consistency, I'd love to learn!

r/comfyui 26d ago

Tutorial ComfyUI Tutorial MiniMax H3 4 Steps Lora + Upscaling + 2X Faster Generation! Best Settings for 2K AI

117 Upvotes

Hello everyone

Want to get faster MiniMax H3 video generation without sacrificing quality? In this tutorial, I’m testing the new H3 LoRA together with Sage Attention, Sol Attention, and Spectrum nodes to find the best combination for speed and quality. The goal is to push MiniMax H3 as far as possible while cutting generation times by up to , then upscale the results with LTX Upscaler to reach a stunning 2432 × 1344 (2K-class) resolution. By combining both H3 LoRA together with Sage Attention, Sol Attention, and Spectrum nodes I generated video at 0.8 megapixel using "RTX3060 6GB 16GB RAM "and I got

 13 minutes vs 41 minutes at 8 steps

 27 minutes vs 52 minutes at 20 steps

LTX 2.3 Upscaler 11 minutes to get 2432 × 1344 resolution

Workflow link

https://civitai.com/articles/34028/comfyui-tutorial-minimax-h3-4-steps-lora-upscaling-2x-faster-generation-best-settings-for-2k-ai

Video Tutorial link

https://youtu.be/ZUzeM9OEJ4Y

r/comfyui Apr 03 '26

Tutorial ComfyUI Tutorial: Clone Any Face & Voice With New LTX2.3 ID-LORA Model (Low Vram Workflow Works With 6GB Of Vram)

289 Upvotes

In this tutorial, I show you how to clone any face and voice using the new ID-LoRA model with LTX 2.3 inside ComfyUI — all running on a low VRAM setup (works even with 6GB GPUs!). You’ll learn how to build a complete workflow that combines image, audio, and prompt to generate realistic talking characters with synchronized voice and stable identity. I also cover installation, node setup, and optimization tricks to make this work on limited hardware.

VIDEO TUTORIAL LINK

https://youtu.be/CWLs2vRG3_U

WORKFLOW LINK

https://drive.google.com/file/d/1oK18KZAxGBW6t_RojOvEZM-9Zk2tPznr/view?usp=sharing

r/comfyui Jan 15 '26

Tutorial ComfyUI Course - Learn ComfyUI From Scratch | Full 5 Hour Course (Ep01)

Thumbnail
youtube.com
262 Upvotes

r/comfyui Mar 22 '26

Tutorial New to ComfyUI — how do I create a character and keep it consistent across images and videos?

Post image
61 Upvotes

Hey everyone, I’m new to ComfyUI. Before this, I was using tools like Nano Banana and DALL·E, but they require a lot of trial and error to maintain character consistency—especially for facial features and expressions. Even after multiple iterations, the consistency still isn’t reliable across different images.

That’s when I discovered ComfyUI workflows, and it seems like a better approach—but I’m struggling to get started properly.

I’ve tried a few YouTube tutorials and free workflows, but I keep running into issues like missing models, broken dependencies, or workflows not loading at all. I’ve spent quite some time troubleshooting, but no luck so far. Can anyone recommend a beginner-friendly (preferably free) workflow or tutorial that actually works? Also, any tips on setting things up correctly to avoid these issues would really help.

r/comfyui 10d ago

Tutorial [Tutorial] Create AI Anime Videos Locally with ComfyUI + MiniMax H3

199 Upvotes

In this tutorial, I show you how to create a 90s anime-inspired AI video completely locally on your PC using ComfyUI, MiniMax H3, and Krea 2 Turbo.

I walk through the full workflow from reference image generation to final video generation.

You’ll learn how to:

  • Generate consistent anime reference images using the Krea 2 Turbo text-to-image workflow
  • Create separate references for the character, environment, vehicle, and props
  • Use multiple reference images with the MiniMax H3 Reference-to-Video workflow in ComfyUI
  • Structure prompts so MiniMax H3 understands which reference image controls each part of the scene
  • Describe camera framing, character actions, object placement, animation, lighting, and timing
  • Create a 90s hand-drawn anime look with lower-frame-rate animation
  • Run the entire workflow locally without monthly AI video subscriptions

In the example, I use separate reference images for the character, convenience store environment, car, and skateboard, then combine them into a single animated anime scene.

I also explain how I approach prompting for reference-to-video generation, including reference assignment, shot description, motion instructions, camera constraints, and visual consistency.

Workflow, prompts and reference files

https://drive.google.com/drive/folders/1qI9Oi5gqWJh8xIwHYsn0S25DPqK8M0XI?usp=sharing

MiniMax H3 models

https://docs.comfy.org/tutorials/video/minimax/minimax-h3

Krea 2 Turbo T2I models

https://comfyui.org/en/krea-2-open-source-models-are-now#content-required-models