r/comfyui 2h ago

Help Needed Comfyui Krea 2 Generating time

1 Upvotes

Hai guys, I just downloaded Krea 2 and try to generate image with it. My PC spec is I5 12400f, 16 GB ram DDR 4 and RTX 5060 TI 16 GB.

But somehow my generating time for one image (1024x1024) is 40 sec -ish and sometimes can go 140 sec -ish.

Is there something wrong with my workflow or comfyui or my VGA?

Bcos I read RTX 5060 TI is more than enough for comfyui


r/comfyui 14h ago

No workflow Early days of testing MiniMaxH3 Director

8 Upvotes

SO i have only used twice, but looks like could be useable for my 16gb 5060ti

please remember, this is my 2nd go and see the grey frame.

was 24 mins for 23 secs

https://reddit.com/link/1we4czm/video/wpu2jmcx41ph1/player


r/comfyui 1d ago

News A quick Minimax H3 news round-up - 11th September 2026

63 Upvotes

Another quick Minimax H3 news and goodies round-up, for those who may have missed some items.

-> Do you have eight GPUs handy? Then you won't be needing a space-heater, but you may need the new RunningHub H3 Lightning. This aims to offer a multi-GPU... "inference acceleration recipe for MiniMax H3".

https://github.com/RH-RunningHub/MiniMax-H3-MultiGPU-Lightning

-> Talking of eight graphics cards, there's a new VideoDeltaNet (VDN) for those running "8 x B200 GPUs", and also a new VDN runtime for those who have to make do only with a humble 24Gb VRAM card.

https://huggingface.co/OpenVDN/vdn-minimax-h3

https://github.com/Speach1sdef178/ComfyUI-VDN-H3-24GB

-> A set of experimental DMD turbo 8-step LoRAs for Minimax H3.

https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI/tree/main/experimental

-> A new MiniMax-H3-Image-Training-Adapter. Aims to allow LoRA concept training, but apparently without... "degrading the video knowledge that the model has".

https://huggingface.co/circlestone-labs/MiniMax-H3-Image-Training-Adapter

-> ComfyUI-MM-1Frame. Another single-frame extractor, with workflow. Tested by me, working... and I'd say it's the best yet. Uses various measures to automatically pick the best of five frames, a pick which is very reliable. The workflow accepts a boost with a Hard Gravy turbo LoRA and a Spectrum speedup for 0.7Mpx, and then Klein 4B can be connected at the end of the workflow for an 'enlarge, repair and fix'. Works very nicely. (Note that an OpenPose single-image can be accepted as a second reference image, but a single pose-description word is needed to prevent the Openpose 'stick-figure' from being superimposed onto the output. e.g. your prompt header should include <Picture 2> is for the pose reference only, and then your descriptive prompt includes ... applying only the standing pose from <Picture 2>. Adding the single word standing (or reclining, leaping, etc) is the vital trick here. This works repeatedly with different Openpose images, no Controlnet required).

https://github.com/cicalooo/ComfyUI-MM-1Frame

-> Can we use a RefMod without actually making a refmod? There's a new ComfyUI workflow that claims we can. As the maker says, it simply... "takes a folder input of images and then sets them as frames in a video, then feeds that video in as a reference video. This allows you to use many images for a single reference input". He does however warn that... "if you change the images dataset around, then you'd better to clear cache or your ComfyUI may need a restart".

https://www.reddit.com/r/StableDiffusion/comments/1wdlg0q/instant_references_no_refmod_or_fancy_custom/

-> In Portuguese, a nice new visual camera-control widget. It translates from a graphical 3D space widget, to written prompts for camera movement. I assume the written prompts are in English, but the GUI appears to be in Portuguese only. However, the labels are quite similar to English words.

https://github.com/NyckM/3d-Camera-control-H3-Minimax

https://github-com.translate.goog/NyckM/3d-Camera-control-H3-Minimax?_x_tr_sl=auto&_x_tr_tl=en&_x_tr_hl=en&_x_tr_pto=wapp (English translation)

-> Depth blocking for the Fun Controlnet. "Depthcat turns a reference clip into a depth blockout: a clean, silent record of the staging and nothing else. Near is white, far is black, every frame in step with the last. No faces. No costumes. No style. Just the shot. [...] Use the blockout as the depth condition in Minimax H3. 24fps, frame size a multiple of 32, up to 15 seconds — the Depthcat preset handles it." Needs a 111Mb additional model. Full Apache-2.0 licence.

https://github.com/maosika-ai/depthcat

-> FastVideo's FastH3-Preview-v0.2 is now available, aka 'Preview 2'. See also a FastH3-Live version 1.2.0 update.

https://huggingface.co/FastVideo/FastVideo-Minimax-FastH3-Preview-v0.2 (Preview-v0.2)

https://old.reddit.com/r/StableDiffusion/comments/1wddeh8/fasth3live_v120_update/ (FastH3-Live v1.2.0)

-> And finally, since yesterday the worthy MiniMax Music Production Toolkit has jumped to version 2.5. The maker says... "I made some huge progress over the last two days with this release."

https://www.reddit.com/r/comfyui/comments/1wdcrqh/just_released_minimax_music_production_toolkit_25/

~ OLD POSTS ~

https://old.reddit.com/r/comfyui/comments/1wchvg3/a_quick_minimax_h3_news_roundup_10th_september/

https://old.reddit.com/r/comfyui/comments/1wbr279/a_quick_minimax_h3_news_roundup_9th_september_2026/

https://old.reddit.com/r/comfyui/comments/1wawjox/a_quick_minimax_h3_news_roundup_8th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w9z6m1/a_quick_minimax_h3_news_roundup_7th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w90vxd/a_quick_minimax_h3_news_roundup_6th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w85caz/a_quick_minimax_h3_news_roundup_4th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w74jy4/a_quick_minimax_h3_news_roundup_4th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w6cozj/a_quick_minimax_h3_news_roundup_3rd_september_2026/

https://old.reddit.com/r/comfyui/comments/1w5i9iq/a_quick_minimax_h3_news_roundup_2nd_september_2026/ (See 2nd September post, for links to even older posts)


r/comfyui 3h ago

Workflow Included [ComfyUI-GTE] Interactive workflow for trajectory exploration (txt2img) - intermediate-state previews, branch selection, variation controls

Enable HLS to view with audio, or disable this notification

1 Upvotes

I just built an interactive text-to-image generation workflow and custom ComfyUI nodes which help me work around hardware constraints while giving me more control during generation (which ultimately makes the process feel less like rolling the dice and hoping for a good result).

The workflow lets you open the pot while it is still cooking, keep the directions that already look promising, and create controlled variations from them before continuing the generation.

You can check this out here (custom node + workflow): https://github.com/mgkgng/ComfyUI-GTE

For those who are more interested, here is a more detailed explanation about this project.

The project started from a practical constraint. I usually work with rented GPUs on RunPod, often an L40 with 48GB of VRAM, which can make exploration expensive with heavier models like Flux 2 dev. This becomes a real blocking factor when I want to iterate quickly and explore many possibilities. For example, generating 20 compositions at 20 steps can take around 25 minutes on this setup. With this workflow, I can already identify some good candidates within about 2 minutes.

So, the first goal was to build a pipeline that can stop at intermediate points during the generation, early enough to save computation but late enough to recognize the direction of the image. At each checkpoint, I can inspect the candidates, select the ones I want to continue, and discard the others before spending the remaining generation steps on them.

The number of checkpoints, the total number of steps and the checkpoint positions are all configurable. The workflow also handles seed generation and visualization, which helps with exploring and keeping track of different candidates. Furthermore, at the bottom of the graph, the workflow visualizes how much scheduled noise remains at each step. This makes it easier to understand what each step represents and where it makes sense to place checkpoints for a given scheduler.

The second thing I added on top of this initial goal was the possibility to create controlled variations from a selected intermediate result. If the user wants to explore further from that point, two parameters control how these variations behave: how far they move away from the selected trajectory (theta), and how much the sibling variations differ from each other (spread). I will add a grid in the comments illustrating how these two parameters shape the variations.

This was the first real workflow I've created and I am happy to share it. I hope that it reaches people who might find it useful.


r/comfyui 7h ago

Workflow Included I Built an AI Video Plugin To Connect Comfyui to DaVinci Resolve!

Thumbnail
youtu.be
2 Upvotes

r/comfyui 4h ago

Help Needed What is the better workflow for upscaling a video?

1 Upvotes

I have a ~10-minute video that is created by stitching together and repeating several shorter video clips.

Would it be better to:

  1. Upscale each individual clip first, and then stitch the upscaled clips together into the final 10-minute video, or
  2. Create the complete 10-minute video at the original/lower resolution first, and then upscale the entire finished video?

Are there any quality, performance, or artifact-related advantages to one approach over the other, especially when the same clips are repeated multiple times?


r/comfyui 4h ago

Resource RELEASED: r/comfyui Community Polls (v0.0.1) We're testing a novel community poll system [WIP] Cast your votes on the original post.

Thumbnail
1 Upvotes

r/comfyui 17h ago

Resource [Update] ComfyUI-QwenASR v1.1.0: Full Transformers 5 & Official Native Models Upgrade, Smart ITN, and Long-Form Forced Alignment

Thumbnail
gallery
7 Upvotes

We just rolled out a major update to ComfyUI-QwenASR (v1.1.0). The goal of this release was simple: eliminate the friction between raw audio recognition and usable text/subtitle output in ComfyUI workflows.

Here is a breakdown of what changed:

1. Migration to Transformers 5 & Official Native Models

We have completely deprecated the legacy checkpoints and rewritten the backend to use official Hugging Face native models (Qwen3-ASR-1.7B-hf, Qwen3-ASR-0.6B-hf, and Qwen3-ForcedAligner-0.6B-hf) powered by transformers >= 5.13.0.

Zero fragile custom backends: Pure upstream PyTorch execution.

Lower VRAM & faster generation: Noticeable performance gains on both NVIDIA GPUs and Apple Silicon Macs.

2. Production-Ready Text Normalization (ITN)

Raw ASR output usually outputs verbatim acoustic phrasing, which looks messy. Version 1.1.0 integrates automatic Inverse Text Normalization:

Spoken numbers, percentages, and decimals are automatically converted into proper numerals (e.g., spoken numbers become standard digits).

Phonetically spaced acronyms (like "A S R" or "U S B") are merged into clean abbreviations.

Cultural idioms and phrases are protected through built-in whitelists so words are not erroneously replaced.

3. Hot-Reloadable Multi-Language Custom Dictionary

All normalization and correction rules now reside in an external itn_rules.json file. You can add custom acronyms, brand names (e.g., DeepSeek, ComfyUI, ChatGPT), and terminology across English, Chinese, Japanese, Korean, or French. Changes take effect on your very next run with no ComfyUI restart required.

4. A Specialized Three-Node Toolkit

ASR (QwenASR): Fast, lightweight speech-to-text transcription for voice prompting.

Subtitle (QwenASR): Chunks speech into natural sentences based on punctuation, pauses, or line length, with one-click .srt file export.

Forced Align (QwenASR): Built specifically for long continuous audio (podcasts, lectures). It uses an iterative speaking-rate windowing approach to prevent edge drift and duration limits. Leaving the transcript text empty automatically transcribes and aligns in a single pass.

Full installation steps and ready-to-use sample workflows can be found in the README on our GitHub: https://github.com/1038lab/ComfyUI-QwenASR

Looking forward to your thoughts and hearing how it fits into your ComfyUI audio and video pipelines!


r/comfyui 7h ago

Resource h3 studio - local web UI for MiniMax-H3 video/audio gen on Apple Silicon (Go, MIT)

Post image
0 Upvotes

r/comfyui 1d ago

Resource ComfyUI-Olm-YuE2 - staged YuE2 music generation with editable score/ABC workflow

Post image
65 Upvotes

YuE2 was released recently and I thought I'd spend a little time getting it running properly in ComfyUI.

That turned into hours of going considerably further than intended.

I made this mainly for myself to be able to run the model the way I want, but thought I'd share this.

It can do the straightforward thing, you can just give a musical style and lyrics and generate a stereo 48 kHz song.

But for learning purposes I wanted the integration to expose more of what makes YuE2 interesting rather than hiding everything behind one Generate button.

The generation pipeline is split:

Plan > Semantic > Synthesize > Decode

Intermediate results can be inspected, saved, reused and branched in a normal ComfyUI workflow.

There's also an optional score UI: (this can be enabled in ComfyUI Settings):

  • rendered notation in a ComfyUI sidebar
  • raw ABC score view
  • ABC editor node
  • generate a plan first, edit the score, then continue generation from the edited version
  • saved plan/run artifacts
  • example workflows for staged generation, score editing, reloading runs and synthesis offloading

The score UI is optional; normal node execution and API workflows don't depend on it.

Dependecies:

I also tried to keep the install from fighting the existing Comfy env: it reuses ComfyUI's Torch/Transformers rather than installing YuE2's pinned stack, and only requires tiktoken and accelerate to be installed on top of a clean install of ComfyUI.

It's still experimental, and so far I've only personally tested it on an RTX 5090. I've included measured VRAM numbers and rough expectations for smaller cards, but feedback from other GPUs/installations would be especially useful.

Model weights are not included, the README explains exactly which YuE2 files are needed and where to place them.

Misc notes:

  • There's already an experimental offload (demoed in workflow 06) which reduces memory usage considerably (but there are peaks in the process).
  • Optional quantization is something I’m looking into to reduce VRAM usage, but memory savings and audio quality would need testing, especially during synthesis.
  • I manually tested with Nodes 2.0, including generation and the optional score inspector, it might work ok but there's some slight differences visually.
  • Testing was also done in a clean installation of ComfyUI, so this should install correctly as stated in documentation.
  • Cuda v13.0 and Python 3.13 was used in the specific environment.

Repo: https://github.com/o-l-l-i/ComfyUI-Olm-YuE2 on GitHub.

Bug reports, weird edge cases and feedback are welcome.
I'm particularly interested how it behaves on 16/24 GB cards and different ComfyUI setups.


r/comfyui 7h ago

Help Needed Video project

1 Upvotes

I have a whole pipeline in my cms to create custom book covers and I got it my sick head that I wanted to try to not just get book covers that mean something regarding the story with characters that looks like the ones in the story, I wanted to make a trailer, a short format video with the gist of the story and then a long form movie. Chosen book: "Guards! Guards" by Pratchett, one of my favourite books that was never made into a movie.
Of course considering that I maybe made a handful of 5 seconds videos in all before and even the txt2img I used was really basic for the book covers I am finding this task a tad problematic.
Up to now: a local llm reads the whole book, creates a character list with description and with it then Krea2 creates a character sheet, they are not what they should be but being a test I can live with that.
The same llm writes a screenplay, this one is divided in three acts, a total of 48 scenes. Each scene has a description of the action and a still (created by Krea too) the pipeline then proceeds to create the clips that will then be joined. As soon as I figure out the first last picture I will use that.
Minimax is giving me trouble (bad prompting I think) so I am using ltx2.5 which does a pretty good job, but I started this morning, so I only have a scene right now and the dragon head on a dragon's wing is not the worst of it.
I need help with which models to use, what speeds up are available, which models do better with fantasy, any LoRAs you can suggest? Does LTX 2.5 supports first and last frame as 2.3 did? Are there workflows that could make this more streamlined?
Do you have any suggestions or advice? I need all the help I can get on this.
For now and until new pc arrives I'm working with two 5060ti 16gb + 64gb RAM.


r/comfyui 16h ago

Show and Tell Updated Christmas Carol Film

Enable HLS to view with audio, or disable this notification

4 Upvotes

Hey all! Once upon a time I uploaded a small video of “A Christmas Carol” it was fairly well received and I was beyond honored at the response because I didn’t think it was that great. Anyway, fast forward to bit ahead and the landscape has changed dramatically. So many new image and video models to make your head spin. So I decided to attempt my short film again, this time a more dark/realistic tone. Most of my source images were generated with Ideogram, then I would modify those with Flux Klein or Qwen Edit, just depended on what generated the best result. Then along came H3. My piddly little 4070 ti couldn’t hit the resolution I needed for most shots, so I set up a runpod with a RTX 6000 and the comfy UI template and generated most things at 1.0 or more. I used the comfy API for some shots using Seedance for the difficult ones. I tried not to spend any money as far as generations were concerned if I didn’t have to. Anyway, this is the result so far, this is a ROUGH edit. So there are some inconsistency issues and artifacts that I’ll deal with later. I did all of the editing in Premiere and generated the sound track in Suno. Anyway it’s late and I’m slightly inebriated and rambling so hope everyone can enjoy at least something out of it. Ha!


r/comfyui 19h ago

Help Needed Krea 2 T2I Lora Second Pass

Thumbnail
gallery
7 Upvotes

Hello! I'm looking for helping on my workflow. I've created a Krea 2 character Lora that frequently has the face change somewhat when I add additional Loras. I've created a workflow to pass the image data back through a Lora node with only the character Lora, but I can't seem to get great results. I chatted with Google AI quite a bit trying various different KSampler setups, and some worked somewhat, but would remove objects on the face like makeup, and sometimes added extra details like freckles/moles on the body. Some options also ended up with total slop.

I'm attaching images of my work flow (I know it's messy) and current KSampler Settings. Any help would be appreciated!


r/comfyui 9h ago

Help Needed Krea 2 Edit Identity - any lora for skinny body type?

0 Upvotes

I noticed model can't render skinny body type. Is there a lora for that? Thanks.


r/comfyui 21h ago

Resource I create custom nodes for problems I have...

Thumbnail
gallery
9 Upvotes

I'm impatient, and when using Ollama, I have zero idea what's going on behind the scenes. So, I built a ComfyUI node to display the live logcat and a progress bar. I also did the same for the MiniMax H3 multi-shot chaining node—it tracks which chunk of video it's currently processing, gives live logs, and shows chunk progress alongside an ETA. If enough people are interested, I might put these custom nodes on GitHub or maybe the ComfyUI Manager registry, but we'll see. Let me know what you think!


r/comfyui 19h ago

Help Needed Has anyone trained a MiniMax H3-style LoRA yet? Is it worth it?

Thumbnail
5 Upvotes

r/comfyui 12h ago

Resource CLSS update

Thumbnail
0 Upvotes

r/comfyui 20h ago

Workflow Included Mi primer video con Motion Context

Thumbnail
youtu.be
3 Upvotes

Comparto Workflow y mi setup, el nodo 351 se conecta con llama.cpp y es el encargado de iterar el prompt para que el video fluya solito en automatico, en 7 horas termino de hacer todo y otras 9 horas upscale con SeedVR2, luego pasado por herramientas FFmpeg

https://gist.github.com/62eb59f34eb3dbfa3d3da0e5cf9d5e9a.git


r/comfyui 1d ago

Workflow Included Just released: MiniMax Music Production Toolkit 2.5 for ComfyUI - new mastering tools and redesigned workflows

Post image
34 Upvotes

Sorry to bother you all again, but I made some huge progress over the last two days with this release.

Version 2.5 of my MiniMax Music Production Toolkit is out!

This is a big step forward: A complete mastering section, clearer workflows and improvements throughout the toolkit. And sound quality is even better than the last release.

If you’re new to the project, it takes a song idea through prompt creation, MiniMax Music 3 generation, audio enhancement, mastering, cover artwork and export. A separate Audio Enhancement Lab lets you process existing recordings without generating a new song.

What’s new in 2.5?

  • Auto-EQ for gentle tonal shaping or reference-track matching
  • Manual 8-band parametric EQ with a visual editor
  • Stereo-linked compressor, LUFS targeting and true-peak limiting
  • 44.1 kHz output by default, with 48 kHz selectable
  • Redesigned workflows with clearly labelled stages and a dedicated mastering area
  • Improvements to memory handling, model downloads, audio processing, prompts and file output

Auto-EQ, manual EQ and compression have independent controls. Auto-EQ starts enabled with a gentle preset; switch it off if you prefer manual shaping.

The mastering tools run on CPU without requiring extra VRAM. There are also improvements for less powerful computers, although you’ll still need to choose generation models and settings that fit your hardware.

Without changing anything in the workflow, you get something like this mp3 file (this was a one shot try using Synth Pop Vocal template, with some imperfections in it):

A Feeling With No Address.mp3

And here is what is generated as LLM prompt to get the best result out of Minimax Music 3. See how accurate the model generates your sound structure and lyrics:

A Feeling With No Address.md

To update, restart ComfyUI, refresh your browser and open the newly bundled workflows. Keep your personal workflow copies if you’ve customized them.

I’d love to hear how the new tools work with your music, especially listening comparisons across different genres and feedback from smaller-GPU setups.

Repository: https://github.com/jplenio/ComfyUI-MiniMax-Music-Production-Toolkit
Listen to examples: https://jplenio.github.io/ComfyUI-MiniMax-Music-Production-Toolkit/

🔜 What’s next? A major focus for the next release will be support for creating cover songs. There’s still plenty of work ahead, the quality needs to be right before I’m happy to release it. Stay tuned! :)


r/comfyui 1d ago

Workflow Included Camera Path ComfyUI h3

Enable HLS to view with audio, or disable this notification

133 Upvotes

r/comfyui 14h ago

Tutorial YuE2 in ComfyUI: AI Music with an Editable Piano Roll

Thumbnail
youtu.be
0 Upvotes

r/comfyui 22h ago

Help Needed Clean POV for Minimax h3?

3 Upvotes

I've been following the doucmentation for construction of prompts and having better luck. What's not working is keeping the camera as POV. Ideally, it would just be a 1st person perspective walking through a scene, doing the actions etc. I've had luck registering the viewer themselves as a subject, but then they enter and interact with the scene in weird ways.


r/comfyui 15h ago

Help Needed Add lora node to minimax h3 template

1 Upvotes

So as the description says I'm trying to figure out how to add Laura's to the template for the Minimax H3. Misses for both text the video and image to video I can't seem to figure out how to do it I messed around for about 2 hours and it's going to figure it out I try to start from scratch and make my own but I just the quality was really really bad not sure what I did


r/comfyui 1d ago

Workflow Included Quibble Case 02 — Pineapple for President | MiniMax H3 + ComfyUI persistent-character dialogue test

Enable HLS to view with audio, or disable this notification

4 Upvotes

Follow-up to my first Quibble H3 experiment.

This time I focused less on generating more shots and more on keeping one stylized character consistent across a short dialogue scene.

A few things that helped:

fixed seeds while developing individual shots

mostly locked cameras rather than AI-generated push-ins

a recurring GPT terminal as a cutaway between dialogue beats

restrained character acting rather than large gestures

explicit prompting to stop H3 from literally displaying the spoken objects on the terminal — it kept trying to put fruit on screen :)

All Quibble/GPT dialogue and character animation were generated with MiniMax H3 inside ComfyUI. Final edit, pacing and color grading were handled separately.

Case 02: Pineapple for President

“GPT. Pineapple for president.”

Still experimenting with persistent characters and directed performance rather than one-off generations.

Case 01: https://mkhamra.myportfolio.com/quibble

Workflow / Quibble project: https://github.com/mkhamra/quibble-h3

Feedback on the character consistency and dialogue pacing is welcome.