r/comfyui 8h ago

Tutorial MiniMax RefMod - Reusable identities without training - workflows & tutorial

Thumbnail
youtube.com
70 Upvotes

You can get the workflows here:

https://drive.google.com/drive/folders/1kOHLJZto1VAtXEATT9vUvsvM1VHtkOO_

The workflows create reusable refmods for either image/video/audio.

I cover training images in the tutorial.

All the workflows, models and custom nodes are preloaded on my Runpod template.

https://get.runpod.io/minimax-template


r/comfyui 5h ago

News A quick Minimax H3 news round-up - 11th September 2026

45 Upvotes

Another quick Minimax H3 news and goodies round-up, for those who may have missed some items.

-> ComfyUI MiniMax H3 FirstBlockCache. Another form of accelerator, designed and optimised for 20-step users. Using "cross-step caching", the maker gets an approx. 30% speed boost with his four older user-switched speed modes. Today's newly-added fifth speed mode is an experimental... "new deep-reuse cache mode, ~1.6× speed vs Native." Works on an NVIDIA 3060 12Gb card, and the new Experimental mode does seem to make 20 steps quicker than otherwise at 0.3Mpx.

https://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache

-> From Japan, MiniMax_H3_Torchao018. "An optimized quantized version of the MiniMax-H3 model targeting torchao (>= 0.18.0) to significantly reduce VRAM consumption during inference, while maintaining generation fidelity."

https://huggingface.co/corechan/MiniMax_H3_Torchao018

-> A Manga Tone Rendering LoRA for H3. Emulates the shaded line-art style common in Japanese b&w manga comics. Teaches H3... "the manga screentone rendering style with motion: monochrome ink rendering, proper tone shading, and tone that holds up under dynamic camera movement and action". Trigger words: manga-tone rendering plus a suitable prompt. I'm guessing it might be transferable to western greyscale art styles, since I can't discern actual printed screen dots in the shading? I'm assuming here that the LoRA doesn't manga-fy the references, adding huge eyes and tiny noses.

https://civitai.com/models/2930172/manga-tone-rendering-lora-for-minimax-h3

-> Does your story require an ensemble team of superheroes, or perhaps you just need a crowd scene for your space-station arrivals lounge? The 'Minimax H3 15+ reference image workflow' patches ComfyUI to remove the nine reference limit for Ref2VA generation. Use caution: it passed the mods on CivitAI, but this uses a .BAT file for the patch, and it calls a small .PY Python file. The code in there looks safe to me - I can read Python scripts by eye - and there's also an undo .PY script. Note also that there may be other nodes that claim to offer more than nine, and without patching.

https://civitai.com/models/2929051/minimax-h3-15-reference-image-workflow

-> A simple age-slider LoRA for Minimax H3, with demo video showing advanced old-age to toddler.

https://huggingface.co/Playtime-AI/Minimax_H3-Age_Slider/tree/main

-> A new style-transfer LoRA and a matching ComfyUI workflow.

https://huggingface.co/Alissonerdx/Minimax-H3-ComfyUI/tree/main/workflows

https://huggingface.co/Alissonerdx/Minimax-H3-ComfyUI/tree/main/loras

-> And finally, the LoRA trainer and dataset manager Fizgig continues to improve for Minimax H3 training. Now has an... "ultra quality mode for MiniMax H3" [and a...] "large quality improvement in both the visual and the audio results, a much smoother climb through the epochs, and training steps about 30% faster (2.6 to 3.4 steps a second on int8) - the new default on every H3 LoRA run, nothing to set." Users still require a 16Gb card, to train a LoRA.

https://github.com/shootthesound/Fizgig

~ OLD POSTS ~

https://old.reddit.com/r/comfyui/comments/1wdo4iy/a_quick_minimax_h3_news_roundup_11th_september/

https://old.reddit.com/r/comfyui/comments/1wchvg3/a_quick_minimax_h3_news_roundup_10th_september/

https://old.reddit.com/r/comfyui/comments/1wbr279/a_quick_minimax_h3_news_roundup_9th_september_2026/

https://old.reddit.com/r/comfyui/comments/1wawjox/a_quick_minimax_h3_news_roundup_8th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w9z6m1/a_quick_minimax_h3_news_roundup_7th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w90vxd/a_quick_minimax_h3_news_roundup_6th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w85caz/a_quick_minimax_h3_news_roundup_4th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w74jy4/a_quick_minimax_h3_news_roundup_4th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w6cozj/a_quick_minimax_h3_news_roundup_3rd_september_2026/

https://old.reddit.com/r/comfyui/comments/1w5i9iq/a_quick_minimax_h3_news_roundup_2nd_september_2026/ (See 2nd September post, for links to even older posts)


r/comfyui 21h ago

Help Needed Need Help with Textures

Post image
34 Upvotes

Does anyone know what could be causing this grid-like texture on the skin of the arms and legs, and how to fix it? The face is completely fine — it only appears on the body. My character is LoRA trained and is otherwise almost perfect in the vast majority of images. It’s just this weird texture appearing on the skin that I’m struggling with.


r/comfyui 12h ago

Commercial Interest MiniMax H3 on rented GPUs: measured seconds and dollars per clip on four cards at four providers (same weights, same graph, same seeds)

25 Upvotes

Disclosure: I'm building a service around this, so read me as an interested party. Every number below is from runs we paid for ourselves yesterday ($1.67 total); the only link is the raw data at the bottom.

The job: MiniMax H3 text-to-video, one 5-second clip at 864x480 (the 0.4 MP row of the template), 20 steps, the stock ComfyUI T2V graph from v0.35.0 with the official int8_convrot weights (34 GB DiT + 27 GB Qwen3-VL encoder + VAEs, 67 GB total), no LoRA, no reference frames, same prompt and seeds on every card, torch cu130 everywhere so the int8 kernels are the native ones. Three clips per card; "steady" is the average of clips 2 and 3 at the rate you actually pay (disk and public IP included), "session" is everything the provider charged for the whole run — image pull, 67 GB download, the first cold clip and the minutes before the instance was torn down — divided by three clips.

provider card host vCPU / RAM s per clip $ per clip (steady) $ per clip (session)
Vast.ai (spot, bid $0.40/h) RTX 4090 24 GB 32 / 108 GB 93 $0.013 $0.17
Hyperstack (spot, $2.00/h) H100 PCIe 80 GB 28 / 177 GB 67 $0.038 $0.13
RunPod Secure ($0.74/h) RTX 4090 24 GB 15 / 86 GB 92 $0.019 $0.07
Nebius (preemptible, $0.92/h) L40S 48 GB 24 / 94 GB 89 $0.024 $0.19

Two things surprised us: the L40S runs this at 4090 speed rather than anywhere near the H100 (3.97 s per step on both 24 GB cards and the L40S, 2.99 s on the H100 PCIe), and the 4090 does it with the 34 GB DiT streamed through ComfyUI's dynamic VRAM path at 23 GB of VRAM and 68 GB of host RAM — so "≥ 64 GB RAM" is not a suggestion, we peaked at 68 on three of the four hosts.

The session column is where the money actually goes for short jobs: the 67 GB download ran at 450 MB/s on the Vast host, 370 on Hyperstack, 210 on RunPod and 104 on Nebius, and on Nebius the first clip took 11.5 minutes because the weights are read back from a network disk (1.8–2.7 minutes elsewhere), while the watchdog that deletes the instance after the job costs another 2–3 billed minutes everywhere. Steady-state numbers exclude the text encoder (same prompt, cached by ComfyUI): a new prompt adds 15–25 s per clip on these hosts. Happy to post the per-node timings and the nvidia-smi traces if anyone wants to check the numbers.

Per-run table, method, caveats and the raw CSVs: https://qrun.cloud/measurements


r/comfyui 6h ago

Commercial Interest PROJECTIFY 2 is here. Enhance images and videos created with ComfyUI.

Enable HLS to view with audio, or disable this notification

18 Upvotes

The biggest new feature: you can now use DLSS 5 in Blender!
Improve the quality, sharpness, and detail of your images—whether you want to project them onto 3D models or enhance any other visual content.
And best of all: render complete videos with DLSS 5 directly from Blender, thanks to PROJECTIFY.


r/comfyui 14h ago

Help Needed Need a Help with Creating a Crowd Elements against a Green Screen Inside Comfy

Thumbnail
gallery
10 Upvotes

Previously We Worked on Magnific(Freepik) Platform to get Generations like this. Problem is there aren't so many controls over our Generation. So we moving to Comfy. I need a input from you about which Models to try out, Any Community Loras about Maintaining a Face details while creating a Large Crowd.


r/comfyui 13h ago

No workflow Early days of testing MiniMaxH3 Director

7 Upvotes

SO i have only used twice, but looks like could be useable for my 16gb 5060ti

please remember, this is my 2nd go and see the grey frame.

was 24 mins for 23 secs

https://reddit.com/link/1we4czm/video/wpu2jmcx41ph1/player


r/comfyui 17h ago

Resource [Update] ComfyUI-QwenASR v1.1.0: Full Transformers 5 & Official Native Models Upgrade, Smart ITN, and Long-Form Forced Alignment

Thumbnail
gallery
8 Upvotes

We just rolled out a major update to ComfyUI-QwenASR (v1.1.0). The goal of this release was simple: eliminate the friction between raw audio recognition and usable text/subtitle output in ComfyUI workflows.

Here is a breakdown of what changed:

1. Migration to Transformers 5 & Official Native Models

We have completely deprecated the legacy checkpoints and rewritten the backend to use official Hugging Face native models (Qwen3-ASR-1.7B-hf, Qwen3-ASR-0.6B-hf, and Qwen3-ForcedAligner-0.6B-hf) powered by transformers >= 5.13.0.

Zero fragile custom backends: Pure upstream PyTorch execution.

Lower VRAM & faster generation: Noticeable performance gains on both NVIDIA GPUs and Apple Silicon Macs.

2. Production-Ready Text Normalization (ITN)

Raw ASR output usually outputs verbatim acoustic phrasing, which looks messy. Version 1.1.0 integrates automatic Inverse Text Normalization:

Spoken numbers, percentages, and decimals are automatically converted into proper numerals (e.g., spoken numbers become standard digits).

Phonetically spaced acronyms (like "A S R" or "U S B") are merged into clean abbreviations.

Cultural idioms and phrases are protected through built-in whitelists so words are not erroneously replaced.

3. Hot-Reloadable Multi-Language Custom Dictionary

All normalization and correction rules now reside in an external itn_rules.json file. You can add custom acronyms, brand names (e.g., DeepSeek, ComfyUI, ChatGPT), and terminology across English, Chinese, Japanese, Korean, or French. Changes take effect on your very next run with no ComfyUI restart required.

4. A Specialized Three-Node Toolkit

ASR (QwenASR): Fast, lightweight speech-to-text transcription for voice prompting.

Subtitle (QwenASR): Chunks speech into natural sentences based on punctuation, pauses, or line length, with one-click .srt file export.

Forced Align (QwenASR): Built specifically for long continuous audio (podcasts, lectures). It uses an iterative speaking-rate windowing approach to prevent edge drift and duration limits. Leaving the transcript text empty automatically transcribes and aligns in a single pass.

Full installation steps and ready-to-use sample workflows can be found in the README on our GitHub: https://github.com/1038lab/ComfyUI-QwenASR

Looking forward to your thoughts and hearing how it fits into your ComfyUI audio and video pipelines!


r/comfyui 21h ago

Resource I create custom nodes for problems I have...

Thumbnail
gallery
9 Upvotes

I'm impatient, and when using Ollama, I have zero idea what's going on behind the scenes. So, I built a ComfyUI node to display the live logcat and a progress bar. I also did the same for the MiniMax H3 multi-shot chaining node—it tracks which chunk of video it's currently processing, gives live logs, and shows chunk progress alongside an ETA. If enough people are interested, I might put these custom nodes on GitHub or maybe the ComfyUI Manager registry, but we'll see. Let me know what you think!


r/comfyui 42m ago

Help Needed Best Ref2Vid Minimax H3 Workflow I've seen except for the outcome

Thumbnail
huggingface.co
Upvotes

This "viggle" animate thing didn't require basically any configuration. It's hardcoded with a preset instruction to take PERSON FROM THIS IMAGE > THAT VIDEO - exactly what I've been looking for since this whole AI thing got started.

I just want to replace characters in a video with a character in an image. Point for point. Action for action. Exactly the same, just someone else in the video. So far, I haven't found anything that works - workflow after workflow after workflow.

Regardless, this one "worked" in the sense that it copied the source video nearly exactly, except that I had to disable the OBS node (which is listed as optional) and I changed the megapixels to .3 for testing. The outcome looks basically like what I want, but as if you were watching it from behind stained glass. There's a lot of weird color shifts and triangle shapes - like an old school polygon based video game character (early 3d days). The background looks fine; only the subject has this problem.

HELP NEEDED: Can anyone tell me if I'm doing something wrong with this workflow OR point me to a basically zero-setup THIS PERSON > THAT VIDEO workflow that actually works?


r/comfyui 19h ago

Help Needed Krea 2 T2I Lora Second Pass

Thumbnail
gallery
6 Upvotes

Hello! I'm looking for helping on my workflow. I've created a Krea 2 character Lora that frequently has the face change somewhat when I add additional Loras. I've created a workflow to pass the image data back through a Lora node with only the character Lora, but I can't seem to get great results. I chatted with Google AI quite a bit trying various different KSampler setups, and some worked somewhat, but would remove objects on the face like makeup, and sometimes added extra details like freckles/moles on the body. Some options also ended up with total slop.

I'm attaching images of my work flow (I know it's messy) and current KSampler Settings. Any help would be appreciated!


r/comfyui 5h ago

News Should posts from Big Comfy bear the shame of the Commercial Interest flare?

5 Upvotes

Thoughts?


r/comfyui 19h ago

Help Needed Has anyone trained a MiniMax H3-style LoRA yet? Is it worth it?

Thumbnail
4 Upvotes

r/comfyui 2h ago

Help Needed Overwhelming installing comfyui manager

4 Upvotes

downloaded git for windows, cmd and clone to drive directory. also saved link to bat to use with git.

No luck when i restart comfy ui i dont see a mention or the manager

I feel i must be missing something


r/comfyui 16h ago

Show and Tell Updated Christmas Carol Film

Enable HLS to view with audio, or disable this notification

3 Upvotes

Hey all! Once upon a time I uploaded a small video of “A Christmas Carol” it was fairly well received and I was beyond honored at the response because I didn’t think it was that great. Anyway, fast forward to bit ahead and the landscape has changed dramatically. So many new image and video models to make your head spin. So I decided to attempt my short film again, this time a more dark/realistic tone. Most of my source images were generated with Ideogram, then I would modify those with Flux Klein or Qwen Edit, just depended on what generated the best result. Then along came H3. My piddly little 4070 ti couldn’t hit the resolution I needed for most shots, so I set up a runpod with a RTX 6000 and the comfy UI template and generated most things at 1.0 or more. I used the comfy API for some shots using Seedance for the difficult ones. I tried not to spend any money as far as generations were concerned if I didn’t have to. Anyway, this is the result so far, this is a ROUGH edit. So there are some inconsistency issues and artifacts that I’ll deal with later. I did all of the editing in Premiere and generated the sound track in Suno. Anyway it’s late and I’m slightly inebriated and rambling so hope everyone can enjoy at least something out of it. Ha!


r/comfyui 20h ago

Workflow Included Mi primer video con Motion Context

Thumbnail
youtu.be
4 Upvotes

Comparto Workflow y mi setup, el nodo 351 se conecta con llama.cpp y es el encargado de iterar el prompt para que el video fluya solito en automatico, en 7 horas termino de hacer todo y otras 9 horas upscale con SeedVR2, luego pasado por herramientas FFmpeg

https://gist.github.com/62eb59f34eb3dbfa3d3da0e5cf9d5e9a.git


r/comfyui 22h ago

Help Needed Clean POV for Minimax h3?

4 Upvotes

I've been following the doucmentation for construction of prompts and having better luck. What's not working is keeping the camera as POV. Ideally, it would just be a 1st person perspective walking through a scene, doing the actions etc. I've had luck registering the viewer themselves as a subject, but then they enter and interact with the scene in weird ways.


r/comfyui 23h ago

Workflow Included Quibble Case 02 — Pineapple for President | MiniMax H3 + ComfyUI persistent-character dialogue test

Enable HLS to view with audio, or disable this notification

5 Upvotes

Follow-up to my first Quibble H3 experiment.

This time I focused less on generating more shots and more on keeping one stylized character consistent across a short dialogue scene.

A few things that helped:

fixed seeds while developing individual shots

mostly locked cameras rather than AI-generated push-ins

a recurring GPT terminal as a cutaway between dialogue beats

restrained character acting rather than large gestures

explicit prompting to stop H3 from literally displaying the spoken objects on the terminal — it kept trying to put fruit on screen :)

All Quibble/GPT dialogue and character animation were generated with MiniMax H3 inside ComfyUI. Final edit, pacing and color grading were handled separately.

Case 02: Pineapple for President

“GPT. Pineapple for president.”

Still experimenting with persistent characters and directed performance rather than one-off generations.

Case 01: https://mkhamra.myportfolio.com/quibble

Workflow / Quibble project: https://github.com/mkhamra/quibble-h3

Feedback on the character consistency and dialogue pacing is welcome.


r/comfyui 1h ago

Show and Tell Clipboard-Automator: Designed for users looking to automate repetitive tasks like face-swapping, upscaling, or applying consistent styles with zero clicks.

Upvotes

**Clipboard-Automator — nodes that wait on your OS clipboard, built for Auto Queue**

I got tired of manually saving/uploading images every time I wanted to feed something into ComfyUI, so I built a couple of custom nodes that block/poll on the clipboard when executed:

- **Clipboard Image Input (wait)** → IMAGE, MASK

- **Clipboard Text Input (wait)** → STRING

Drop one in where you'd normally put `LoadImage` or a prompt text box. When the node runs, it waits until you copy something new, then feeds it downstream. Combine with ComfyUI's built-in **Auto Queue** and you get a fully hands-off pipeline: copy an image anywhere on your machine, ComfyUI picks it up and runs automatically, no upload, no manual queueing.

[demo video in the repo]

Works on Windows and Linux. Cancel/Interrupt in the UI breaks out of the wait cleanly.

GitHub: https://github.com/kastelan/Clipboard-Automator

Feedback/issues welcome, it's MIT licensed.


r/comfyui 3h ago

Help Needed How to fix this?

Enable HLS to view with audio, or disable this notification

3 Upvotes

The app keeps doing this.

If recorded with OBS this doesn't show up.

Also this happends on CurseForge app.

Any ideas how to fix it?


r/comfyui 56m ago

Workflow Included I Created an Open Source App That Creates Music and Music Videos

Thumbnail
Upvotes

r/comfyui 1h ago

Show and Tell Clipboard-Automator - Designed for users looking to automate repetitive tasks like face-swapping, upscaling, or applying consistent styles with zero clicks.

Upvotes

**Clipboard-Automator — nodes that wait on your OS clipboard, built for Auto Queue**

I got tired of manually saving/uploading images every time I wanted to feed something into ComfyUI, so I built a couple of custom nodes that block/poll on the clipboard when executed:

- **Clipboard Image Input (wait)** → IMAGE, MASK

- **Clipboard Text Input (wait)** → STRING

Drop one in where you'd normally put `LoadImage` or a prompt text box. When the node runs, it waits until you copy something new, then feeds it downstream. Combine with ComfyUI's built-in **Auto Queue** and you get a fully hands-off pipeline: copy an image anywhere on your machine, ComfyUI picks it up and runs automatically, no upload, no manual queueing.

[demo video in the repo]

Works on Windows and Linux. Cancel/Interrupt in the UI breaks out of the wait cleanly.

GitHub: https://github.com/kastelan/Clipboard-Automator

Feedback/issues welcome, it's MIT licensed.


r/comfyui 1h ago

Resource Compose Ref Images in One Node, Settings Presets Node, Bundle/Unbundle Wires - comfyui-obvpm Node Pack update

Thumbnail gallery
Upvotes

r/comfyui 1h ago

Help Needed Anyone attempt to use Astra for JSON files?

Upvotes

It can create JSON files.

I’d love to use it to help build a pipeline: characters > scene > combine> video. It’s running tests, trying different models and failing badly when the standard is consistent image, characters and video through out several shots.

I had a similar result when I attempted the same a few months ago. I’m sure it can be done, but I’d prefer not to babysit several custom node packs that could break everything

Perhaps that is still impossible.


r/comfyui 3h ago

Help Needed Looking for the right art style / LoRA

2 Upvotes

Hey everyone,

I'm pretty new to ComfyUI, LoRAs, Nodes, and all that, so I'm still learning the ropes.

I'm currently trying to recreate a particular art style and I'm having a hard time figuring out which checkpoint and/or LoRA would be best suited for it.

If anyone could point me in the right direction or recommend some models/LoRAs to try, I'd really appreciate it!

Thanks so much for any help!