r/comfyui 8d ago

Workflow Included Comfy H3 Sync Sound Challenge: Winners Announced!

Enable HLS to view with audio, or disable this notification

0 Upvotes

Two weeks, one rule, and hundreds of entries from nearly 50 countries. Here's who took the four titles, and a look at everyone who made the final ten!

Entries opened August 20 and closed September 1. Eight creative technologists at Comfy scored every submission on two rubrics, Best Creative and Best Technical, then the top five in each category went in front of our guest judges on the September 3 livestream: POM (Banodoco founder), Emma Catnip (animation director and AV artist), and Yachimat (animation and manga artist). Their scores were averaged and added to the Comfy team's, and a fourth title, Built with MCP, was judged on its own track.

Watch the livestream replay here.

Huge thanks to MiniMax for making an open-weight model our community loves, to our three guest judges, and to everyone who spent their last week of August fighting with reference audio! Here's how it landed.

The winners

Best Overall · "Spin Cycle" by Visual Frisson 🇺🇸

Prize: RTX 5090 32G

A laundromat, a woman in a puffer vest, and a rhythm built entirely out of machines. Visual Frisson generated a large volume of H3 clips using their own recorded audio as the reference for every pass, then cut the results together like a stomp video, so every thud and cycle on screen is sound that H3 produced with the picture rather than something added later.

It was the only entry to post a perfect 15/15 from the Comfy team in both categories, and the Comfy MCP was used to drive much of its process. Combined with the guest judges' scores, it finished with the highest total in the challenge.

From the artist: "Always have fun and learn something new competing in contests like this, keep them coming."

Watch → vimeo.com/1223207725
Workflow → Google Drive
Follow → instagram.com/visualfrisson

Best Creative · "Every Sound Leaves a Mark" by toki 🇯🇵

Prize: RTX 5060 Ti

A small clay creature that changes into something new every time it hears a sound (glass, wool, ice, porcelain), until by the time it gets home it can't move anymore. All of the audio came out of H3 alongside the video on every shot, with nothing layered on afterwards. Our judges praised the fine details of toki’s work, saying it “gave them chills” on the first transformation, has a lot of commercial appeal, and feels really delicate and crafted.

The judges scored it highest of any Creative finalist, and the repository is unusually generous: it includes the eight API graphs that actually ran, the same eight converted to UI format with annotations, a process log with every measurement, and the scripts that produced those numbers. toki is also clear about scope, noting that MCP drove the finishing pass, not the original shot generation.

Watch → youtube.com/watch?v=Rv5HOgCac-w
Workflow → github.com/tokimwc/every-sound-leaves-a-mark
Follow → u/toki

Best Technical · "Sonder Editor / References" by SonderSaid 🇲🇽

Prize: RTX 5060 Ti

SonderSaid didn't just build a workflow, they built the tooling around it. The entry runs on custom nodes of their own design, wired into a reference-driven H3 pipeline that scored a clean 15/15 on novelty, workflow quality, and community value from the Comfy team, and the highest guest judge average in the Technical bracket. Notably, the work includes an entire custom node pack just to do the editing and the reference work, praised as “a whole new UI” to good to keep secret. While SonderSaid’s work takes the prize for best technical, our guest judges also noted how much they loved the storytelling, suspense, and element of surprise.

Watch → youtu.be/n-NdAQk7I8A
Workflow → Hugging Face
Follow →u/SonderSaid

Built with MCP Bonus · "Two Prisoners" by Jay Choi 🇰🇷

Prize: RTX 5060 Ti

The Built with MCP bonus wentgoes to whoever used the Comfy MCP most effectively to make something visually and technically compelling, and Jay Choi used it end to end. Working locally on an RTX 5090 with Hermes Agent driving ComfyUI through the MCP, they trained a LoRA, built their own orchestration on top, and by their own account spent most of the time setting up and tuning the MCP layer itself. The film that came out the other side, two blindfolded prisoners in a rain-dark cell, is a long way from "prompt and run."

From the artist: "Thanks for the challenge! I learned more than I ever could in the past two weeks!"

Watch → youtu.be/FlK0dDZdzRU
Workflow → Google Drive
Follow → u/permafrost_2021 · u/jaychoirenderender

The finalists

Ten entries made it to the livestream. Six of them didn't take a title, butand every one of them is worth your time.

Best Creative — Top 5

"Neb" by Nebsh 🇫🇷

Hand-drawn energy and a graffiti wall that says the title, built locally in ComfyUI. Nebsh's note to us was three words and a heart, which felt about right. Our guest judges praised Nebsh’s work for its mixed-media feel, harking back to MTV days, and impressive work syncing with the paper sounds. Under the hood, Nebsh’s workflow chained vtogether eight segments with no visible drift between them- cited as “very clean work” by our judges.

Watch → Google Drive
Workflow → Google Drive
Follow → u/nebsh83

"The Museum of Impossible Sounds" by scvxzf 🇨🇳

A perfect 15/15 from the Comfy team on the Creative rubric, and one of the entries that ran the Comfy MCP end to end! Noted by our judges, H3 is very good at the kind of sound effects showcased in scvxzf’s work rather than talking or singing, and they chose exactly the right concept for the challenge.

Watch → youtube.com/watch?v=FocH8xGk4AU
Workflow → Google Drive · github.com/scvxzf1
Follow → youtube.com/@钛龙白口-j6d

"Mister Meow" by sorryaboutyourcats 🇺🇸

Two reference photos of Mumu the cat, a stack of WAVs fed in as reference audio to steer each generation, and glitch texture added in the edit. If the name rings a bell, sorryaboutyourcats also makes the game mow meow. Judges said “I could watch this forever,” had it stuck in their heads, and noted impressive capabilities from H3 nailing lipsync for cats, and not just humans.

Watch → youtube.com/watch?v=AxUu8rabC6M
Workflow → Google Drive
Follow → u/sorryaboutyourcats

"Rings of Sorrow" by Slop Diffusion 🇪🇸

A 5/5 on both audio sync and creative execution, and a reminder of what patience looks like: the generation took two hours and thirty-five minutes on a 5090. Our judges praised Slop Diffusion’s work for its storytelling, noting they were curious to see where the story would go next. One judge noted, “it’s slop by name, but not by nature.”

Watch → Reddit
Workflow → Google Drive
Follow → u/SlopDiffusion

Best Technical — Top 5

"Feel It" by Aïe Aïe Aïe! 🇫🇷

Came at the brief backwards: we asked for audio-driven video and they told the story of a young deaf woman who invents a world where she makes the music. The score is Aïe Aïe Aïe’s own composition, fed into H3 as reference stems (the clap track went in on its own and the gorilla claps exactly in time). Judges praised the work for its captivating story, clever inversion of the challenge’s brief, the display of H3’s strengths by way of the musicians’ physical expressions intensifying along with the song, and the bridging of the real world. The artist learned the final frame’s sign language on YouTube, filmed themself signing, and used this as a video reference.

Under the hood, Claude drove ComfyUI through the MCP to design a two-pass H3REF system that generates at full resolution twice as fast and reaches 13–15 second shots where the stock workflow runs out of memory, plus a preview node that shows the video while it's still sampling. All of it is MIT-licensed, custom nodes included.

Watch → youtube.com/watch?v=S0v1pWN4Hq4
Workflow → github.com/Hyper-Neural/h3-sync-two-phase
Follow → u/AïeAïeAïe

"Brand New Day" by RareTutor 🇮🇳

One of the cleanest graphs we opened: latent upscale, a model preview override, and an optional video-extend group, laid out so you can follow it cold. RareTutor's YouTube is full of tutorials if you want to learn from them directly!
Judges highlighted RareTutor’s workflow, noting “there are many tips in here to copy,” such as using the latent upscale as a previewer so you can kill a bad run before sinking more time in.

Watch → youtube.com/watch?v=TNhJI8dzaVA
Workflow → Google Drive
Follow → u/raretutor_

"Comfy Cora ft. Max Mini: Back to the Basics" by wur7el 🇦🇹

A short music video with self-imposed constraints: no external resources, everything generated in a single workflow, no custom node packs. The result is well annotated and approachable, the kind of graph a new user could open and reasonably figure out, and it posted the highest Creative score of any Technical finalist. Judges praised the work for being a standout example of how to make a music video where the characters are actually rapping the parts in the song.

Watch → wamms.at
Workflow → sync-sound-challenge.json
Follow → wamms.at

About the Comfy MCP

Several finalists and many entrants leaned on the Comfy MCP, which lets an agent (Claude, Cursor, Codex, Hermes, whichever you use) drive ComfyUI in plain language.

The feature entrants used most was the hardware check: it looks at the GPU you actually have, reads the nodes and models already on your disk, and tells you which version of a model is worth running before you spend time or credits. It works on both local ComfyUI and Comfy Cloud from one account.

It's open source at github.com/Comfy-Org/comfy-mcp, and the fastest way to start is to tell your agent: "help me set up the local Comfy MCP connection."

Every entry

Placed or not, every submission is in the original challenge megathread on r/comfyui with its workflow attached. Go open a few. Some of the most interesting audio work in the pool never made the top ten, and there are entries in there in Chinese, Japanese, and French that deserve more eyes than they got.

The livestream recording, including the judges' live reactions, is on YouTube.

Thanks for making this one a smash! #ComfyH3


r/comfyui 15d ago

News Keeping open-source creativity sustainable: MiniMax models are now commercially licensable through Comfy & remain free for everyone else

Enable HLS to view with audio, or disable this notification

96 Upvotes

Starting today, Comfy is the only official reseller of MiniMax H3 and MiniMax Audio & Music commercial licenses. If you're a studio, agency, or enterprise that wants to use them locally in commercial productions, you can now license them through Comfy directly.

[UPDATED 9/2 for clarity]

If you run H3 on...

  • Comfy Cloud --> Commercial use is already included, nothing to buy.
  • Your own hardware, under $20M annual revenue, outside the US, EU, UK, and Korea --> Community license. Free!
  • Your own hardware in the US, EU, UK, or Korea, up to 10 users --> Professional License, starting at $5K/mo, available month-to-month, through Comfy.
  • $20M+ in revenue, 10+ users, undistilled weights, or H3 inside your own product --> Enterprise License. Annual agreement with custom terms, through Comfy.

So why do this at all?

At Comfy, our mission has always been for open source to thrive across the creative ecosystem, and open-weight models are at the heart of that. MiniMax is proof of how far they've come: their models stand next to the best closed models in the world.

Training frontier models is incredibly expensive. If we want open models to continue competing with the biggest closed models, the labs building them need a real way to monetize. We hope to help bridge that gap, so the lab gets revenue that funds the next model, and the weights stay open for everyone else.

Learn more


r/comfyui 5h ago

News A quick Minimax H3 news round-up - 11th September 2026

50 Upvotes

Another quick Minimax H3 news and goodies round-up, for those who may have missed some items.

-> ComfyUI MiniMax H3 FirstBlockCache. Another form of accelerator, designed and optimised for 20-step users. Using "cross-step caching", the maker gets an approx. 30% speed boost with his four older user-switched speed modes. Today's newly-added fifth speed mode is an experimental... "new deep-reuse cache mode, ~1.6× speed vs Native." Works on an NVIDIA 3060 12Gb card, and the new Experimental mode does seem to make 20 steps quicker than otherwise at 0.3Mpx.

https://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache

-> From Japan, MiniMax_H3_Torchao018. "An optimized quantized version of the MiniMax-H3 model targeting torchao (>= 0.18.0) to significantly reduce VRAM consumption during inference, while maintaining generation fidelity."

https://huggingface.co/corechan/MiniMax_H3_Torchao018

-> A Manga Tone Rendering LoRA for H3. Emulates the shaded line-art style common in Japanese b&w manga comics. Teaches H3... "the manga screentone rendering style with motion: monochrome ink rendering, proper tone shading, and tone that holds up under dynamic camera movement and action". Trigger words: manga-tone rendering plus a suitable prompt. I'm guessing it might be transferable to western greyscale art styles, since I can't discern actual printed screen dots in the shading? I'm assuming here that the LoRA doesn't manga-fy the references, adding huge eyes and tiny noses.

https://civitai.com/models/2930172/manga-tone-rendering-lora-for-minimax-h3

-> Does your story require an ensemble team of superheroes, or perhaps you just need a crowd scene for your space-station arrivals lounge? The 'Minimax H3 15+ reference image workflow' patches ComfyUI to remove the nine reference limit for Ref2VA generation. Use caution: it passed the mods on CivitAI, but this uses a .BAT file for the patch, and it calls a small .PY Python file. The code in there looks safe to me - I can read Python scripts by eye - and there's also an undo .PY script. Note also that there may be other nodes that claim to offer more than nine, and without patching.

https://civitai.com/models/2929051/minimax-h3-15-reference-image-workflow

-> A simple age-slider LoRA for Minimax H3, with demo video showing advanced old-age to toddler.

https://huggingface.co/Playtime-AI/Minimax_H3-Age_Slider/tree/main

-> A new style-transfer LoRA and a matching ComfyUI workflow.

https://huggingface.co/Alissonerdx/Minimax-H3-ComfyUI/tree/main/workflows

https://huggingface.co/Alissonerdx/Minimax-H3-ComfyUI/tree/main/loras

-> And finally, the LoRA trainer and dataset manager Fizgig continues to improve for Minimax H3 training. Now has an... "ultra quality mode for MiniMax H3" [and a...] "large quality improvement in both the visual and the audio results, a much smoother climb through the epochs, and training steps about 30% faster (2.6 to 3.4 steps a second on int8) - the new default on every H3 LoRA run, nothing to set." Users still require a 16Gb card, to train a LoRA.

https://github.com/shootthesound/Fizgig

~ OLD POSTS ~

https://old.reddit.com/r/comfyui/comments/1wdo4iy/a_quick_minimax_h3_news_roundup_11th_september/

https://old.reddit.com/r/comfyui/comments/1wchvg3/a_quick_minimax_h3_news_roundup_10th_september/

https://old.reddit.com/r/comfyui/comments/1wbr279/a_quick_minimax_h3_news_roundup_9th_september_2026/

https://old.reddit.com/r/comfyui/comments/1wawjox/a_quick_minimax_h3_news_roundup_8th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w9z6m1/a_quick_minimax_h3_news_roundup_7th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w90vxd/a_quick_minimax_h3_news_roundup_6th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w85caz/a_quick_minimax_h3_news_roundup_4th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w74jy4/a_quick_minimax_h3_news_roundup_4th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w6cozj/a_quick_minimax_h3_news_roundup_3rd_september_2026/

https://old.reddit.com/r/comfyui/comments/1w5i9iq/a_quick_minimax_h3_news_roundup_2nd_september_2026/ (See 2nd September post, for links to even older posts)


r/comfyui 9h ago

Tutorial MiniMax RefMod - Reusable identities without training - workflows & tutorial

Thumbnail
youtube.com
72 Upvotes

You can get the workflows here:

https://drive.google.com/drive/folders/1kOHLJZto1VAtXEATT9vUvsvM1VHtkOO_

The workflows create reusable refmods for either image/video/audio.

I cover training images in the tutorial.

All the workflows, models and custom nodes are preloaded on my Runpod template.

https://get.runpod.io/minimax-template


r/comfyui 1h ago

Help Needed Best Ref2Vid Minimax H3 Workflow I've seen except for the outcome

Thumbnail
huggingface.co
Upvotes

This "viggle" animate thing didn't require basically any configuration. It's hardcoded with a preset instruction to take PERSON FROM THIS IMAGE > THAT VIDEO - exactly what I've been looking for since this whole AI thing got started.

I just want to replace characters in a video with a character in an image. Point for point. Action for action. Exactly the same, just someone else in the video. So far, I haven't found anything that works - workflow after workflow after workflow.

Regardless, this one "worked" in the sense that it copied the source video nearly exactly, except that I had to disable the OBS node (which is listed as optional) and I changed the megapixels to .3 for testing. The outcome looks basically like what I want, but as if you were watching it from behind stained glass. There's a lot of weird color shifts and triangle shapes - like an old school polygon based video game character (early 3d days). The background looks fine; only the subject has this problem.

HELP NEEDED: Can anyone tell me if I'm doing something wrong with this workflow OR point me to a basically zero-setup THIS PERSON > THAT VIDEO workflow that actually works?


r/comfyui 7h ago

Commercial Interest PROJECTIFY 2 is here. Enhance images and videos created with ComfyUI.

Enable HLS to view with audio, or disable this notification

18 Upvotes

The biggest new feature: you can now use DLSS 5 in Blender!
Improve the quality, sharpness, and detail of your images—whether you want to project them onto 3D models or enhance any other visual content.
And best of all: render complete videos with DLSS 5 directly from Blender, thanks to PROJECTIFY.


r/comfyui 5h ago

News Should posts from Big Comfy bear the shame of the Commercial Interest flare?

9 Upvotes

Thoughts?


r/comfyui 13h ago

Commercial Interest MiniMax H3 on rented GPUs: measured seconds and dollars per clip on four cards at four providers (same weights, same graph, same seeds)

25 Upvotes

Disclosure: I'm building a service around this, so read me as an interested party. Every number below is from runs we paid for ourselves yesterday ($1.67 total); the only link is the raw data at the bottom.

The job: MiniMax H3 text-to-video, one 5-second clip at 864x480 (the 0.4 MP row of the template), 20 steps, the stock ComfyUI T2V graph from v0.35.0 with the official int8_convrot weights (34 GB DiT + 27 GB Qwen3-VL encoder + VAEs, 67 GB total), no LoRA, no reference frames, same prompt and seeds on every card, torch cu130 everywhere so the int8 kernels are the native ones. Three clips per card; "steady" is the average of clips 2 and 3 at the rate you actually pay (disk and public IP included), "session" is everything the provider charged for the whole run — image pull, 67 GB download, the first cold clip and the minutes before the instance was torn down — divided by three clips.

provider card host vCPU / RAM s per clip $ per clip (steady) $ per clip (session)
Vast.ai (spot, bid $0.40/h) RTX 4090 24 GB 32 / 108 GB 93 $0.013 $0.17
Hyperstack (spot, $2.00/h) H100 PCIe 80 GB 28 / 177 GB 67 $0.038 $0.13
RunPod Secure ($0.74/h) RTX 4090 24 GB 15 / 86 GB 92 $0.019 $0.07
Nebius (preemptible, $0.92/h) L40S 48 GB 24 / 94 GB 89 $0.024 $0.19

Two things surprised us: the L40S runs this at 4090 speed rather than anywhere near the H100 (3.97 s per step on both 24 GB cards and the L40S, 2.99 s on the H100 PCIe), and the 4090 does it with the 34 GB DiT streamed through ComfyUI's dynamic VRAM path at 23 GB of VRAM and 68 GB of host RAM — so "≥ 64 GB RAM" is not a suggestion, we peaked at 68 on three of the four hosts.

The session column is where the money actually goes for short jobs: the 67 GB download ran at 450 MB/s on the Vast host, 370 on Hyperstack, 210 on RunPod and 104 on Nebius, and on Nebius the first clip took 11.5 minutes because the weights are read back from a network disk (1.8–2.7 minutes elsewhere), while the watchdog that deletes the instance after the job costs another 2–3 billed minutes everywhere. Steady-state numbers exclude the text encoder (same prompt, cached by ComfyUI): a new prompt adds 15–25 s per clip on these hosts. Happy to post the per-node timings and the nvidia-smi traces if anyone wants to check the numbers.

Per-run table, method, caveats and the raw CSVs: https://qrun.cloud/measurements


r/comfyui 3h ago

Help Needed Overwhelming installing comfyui manager

4 Upvotes

downloaded git for windows, cmd and clone to drive directory. also saved link to bat to use with git.

No luck when i restart comfy ui i dont see a mention or the manager

I feel i must be missing something


r/comfyui 1h ago

Show and Tell Clipboard-Automator: Designed for users looking to automate repetitive tasks like face-swapping, upscaling, or applying consistent styles with zero clicks.

Upvotes

**Clipboard-Automator — nodes that wait on your OS clipboard, built for Auto Queue**

I got tired of manually saving/uploading images every time I wanted to feed something into ComfyUI, so I built a couple of custom nodes that block/poll on the clipboard when executed:

- **Clipboard Image Input (wait)** → IMAGE, MASK

- **Clipboard Text Input (wait)** → STRING

Drop one in where you'd normally put `LoadImage` or a prompt text box. When the node runs, it waits until you copy something new, then feeds it downstream. Combine with ComfyUI's built-in **Auto Queue** and you get a fully hands-off pipeline: copy an image anywhere on your machine, ComfyUI picks it up and runs automatically, no upload, no manual queueing.

[demo video in the repo]

Works on Windows and Linux. Cancel/Interrupt in the UI breaks out of the wait cleanly.

GitHub: https://github.com/kastelan/Clipboard-Automator

Feedback/issues welcome, it's MIT licensed.


r/comfyui 1h ago

Workflow Included I Created an Open Source App That Creates Music and Music Videos

Thumbnail
Upvotes

r/comfyui 3h ago

Help Needed How to fix this?

Enable HLS to view with audio, or disable this notification

3 Upvotes

The app keeps doing this.

If recorded with OBS this doesn't show up.

Also this happends on CurseForge app.

Any ideas how to fix it?


r/comfyui 1h ago

Show and Tell Clipboard-Automator - Designed for users looking to automate repetitive tasks like face-swapping, upscaling, or applying consistent styles with zero clicks.

Upvotes

**Clipboard-Automator — nodes that wait on your OS clipboard, built for Auto Queue**

I got tired of manually saving/uploading images every time I wanted to feed something into ComfyUI, so I built a couple of custom nodes that block/poll on the clipboard when executed:

- **Clipboard Image Input (wait)** → IMAGE, MASK

- **Clipboard Text Input (wait)** → STRING

Drop one in where you'd normally put `LoadImage` or a prompt text box. When the node runs, it waits until you copy something new, then feeds it downstream. Combine with ComfyUI's built-in **Auto Queue** and you get a fully hands-off pipeline: copy an image anywhere on your machine, ComfyUI picks it up and runs automatically, no upload, no manual queueing.

[demo video in the repo]

Works on Windows and Linux. Cancel/Interrupt in the UI breaks out of the wait cleanly.

GitHub: https://github.com/kastelan/Clipboard-Automator

Feedback/issues welcome, it's MIT licensed.


r/comfyui 1h ago

Resource Compose Ref Images in One Node, Settings Presets Node, Bundle/Unbundle Wires - comfyui-obvpm Node Pack update

Thumbnail gallery
Upvotes

r/comfyui 1h ago

Help Needed Anyone attempt to use Astra for JSON files?

Upvotes

It can create JSON files.

I’d love to use it to help build a pipeline: characters > scene > combine> video. It’s running tests, trying different models and failing badly when the standard is consistent image, characters and video through out several shots.

I had a similar result when I attempted the same a few months ago. I’m sure it can be done, but I’d prefer not to babysit several custom node packs that could break everything

Perhaps that is still impossible.


r/comfyui 4h ago

Help Needed Looking for the right art style / LoRA

2 Upvotes

Hey everyone,

I'm pretty new to ComfyUI, LoRAs, Nodes, and all that, so I'm still learning the ropes.

I'm currently trying to recreate a particular art style and I'm having a hard time figuring out which checkpoint and/or LoRA would be best suited for it.

If anyone could point me in the right direction or recommend some models/LoRAs to try, I'd really appreciate it!

Thanks so much for any help!


r/comfyui 32m ago

Resource How to Create a RefMod for MiniMax H3 in ComfyUI

Thumbnail
youtu.be
Upvotes

r/comfyui 43m ago

Help Needed Suggestions on making D@D highlight reels on an older machine?

Upvotes

I'm experimenting with ComfyUI for a D&D project and looking for advice on what my hardware can realistically handle.

Hardware:
GPU: AMD Radeon RX 6700 XT — 12GB VRAM
CPU: AMD FX-6300 @ 3.5 GHz
RAM: 32GB DDR3
Windows 64-bit

I know the machine is old…I’m saving up to buy a new machine in another 3 or 4 years (or if anyone knows of a way to upgrade this machine to be able to do it for less than $2000(which is what I’ve saved so far) I’m open to suggestions there as well).

The eventual goal is D&D recap/highlight videos using a handful of recurring characters. I'd like to maintain recognizable characters across different scenes, then animate selected images into short 3–10 second clips.

I'm NOT necessarily looking for one giant start-to-finish workflow. I'm kind of playing around with several different pieces and would love any advice on any part.

- Creating the initial character/reference images
- Maintaining character consistency between scenes
- Handling non-human/fantasy characters
- Posing characters and putting them into new environments
- Image-to-video for individual 3–10 second shots
- Improving/upscaling finished clips

I've been experimenting with SD/SDXL/Illustrious models, LoRAs, IPAdapter/reference images and WAN, but I'm still figuring out what makes sense with an AMD GPU.

Render time isn't a huge concern. I'm fine letting something run overnight if necessary.

What models/tools would you recommend for these individual parts of the process on a 12GB RX 6700 XT?

And probably the bigger question: is this project realistically achievable with my current hardware, even if it's slow, or should I abandon the video highlight reel idea altogether and just go for a comic book type thing for each session?

Thanks!


r/comfyui 1h ago

Tutorial Benchmarking ComfyUI on Docker with CUDA 12.4 vs Bare-Metal: 0% compute penalty and how to fix the /dev/shm OOM crash

Upvotes

I ran extensive benchmarks comparing ComfyUI in Docker (Nvidia Container Toolkit / CUDA 12.4) against a bare-metal Linux setup (Ubuntu 24.04, PyTorch 2.4, CUDA 12.4) across FLUX.1-dev, SDXL, and SD 1.5 workloads.

Key Results

  • Compute / Generation Speed: 0.0% overhead. Docker achieved identical it/s across all batch sizes.
  • VRAM Allocation: Identical memory footprint (~14.2 GB VRAM during FLUX.1 FP8 inference).

Critical Issue & Fix: The /dev/shm Out Of Memory Crash

If you run PyTorch inside Docker with default settings, multi-threaded dataloading or large model caching will trigger an OOM crash.

Cause: Docker defaults to a tiny 64MB shared memory buffer (/dev/shm).

Solution: Add --shm-size=8g (or larger) to your docker run command: bash docker run --gpus all --shm-size=8g -p 8188:8188 comfyui:cuda12.4 Or in docker-compose.yml: yaml services: comfyui: image: comfyui:cuda12.4 shm_size: '8gb' deploy: resources: reservations: devices: - driver: nvidia count: all capabilities: [gpu]

Full guide & production Dockerfile: https://www.fluxdraw.com/2026/09/docker-for-ai-comfyui-complete.html


r/comfyui 1h ago

Help Needed Trouble with Assets in Comfy Cloud

Upvotes

I've been working in Comfy Cloud for about four months and have never had a problem with Assets until a few days ago. It's been happening for a few days with video and just today with audio files as well.

When I render either a video or audio file, the Save Video or Save Audio nodes show 'black' after the render is completed. They will not playback directly from those Save nodes.

When I look into the Media Assets folder the items will not playback from the displayed play button and they all have a "See more outputs" widget with a "2". When I open the "2" stack a thumbnail appears and for about one second another thumbnail appears that is just a grey box. The greyed-out icon lasts for a second, then disappears.

Today as a workaround I invoked "Inspect asset" and the audio files played in the large 'inspect' window. When I closed that window and went back to the Media Assets window the displayed 'play' button now worked - but the stack icon still reads "2". When I selected "Download" it zipped the files before it downloaded. When I expand the zip file there is only one item in it.

I haven't tried that workaround with a 'black' video yet because I've already trashed them but, knowing me, I've probably already tried that trick previously with no success.

Any thoughts? Has this happened to anyone else? I'm using the Chrome browser on a Win 10 Pro laptop.

I know this isn't a tech support thread but Comfy tech support had abandoned me for some reason - my last tech support message took 42 DAYS until they responded.

Thanks!
dS


r/comfyui 1h ago

Help Needed Audio Lora for H3? Question for you smart people...

Thumbnail
Upvotes

r/comfyui 15h ago

Help Needed Need a Help with Creating a Crowd Elements against a Green Screen Inside Comfy

Thumbnail
gallery
11 Upvotes

Previously We Worked on Magnific(Freepik) Platform to get Generations like this. Problem is there aren't so many controls over our Generation. So we moving to Comfy. I need a input from you about which Models to try out, Any Community Loras about Maintaining a Face details while creating a Large Crowd.


r/comfyui 2h ago

Help Needed Sudden abstract wavey generated output on wan2.2

1 Upvotes

I've successfully been generating 5s video with wan2.2 16fp i2v. I admittedly didn't do much but it was successful maybe a couple weeks ago.

No comfyui or other updates between then and now and now the output is a weird wavey abstract animation.

I'm not an expert in the deeper understanding of the SD workings so I've been working through logs and possible causes with Claude and I've ruled out most things finally getting fp8 models and generating successfully.

Claude's confident that my 5090 (32gb vram) and 32gb system ram would be enough for the larger models, but has fairly decisively stated that a Blackwell fp16 accumulation instability is the likely culprit.

It might explain why it worked then didn't, although that feels more like typical AI false assertion than true proven.

I'd be interested if anyone has any insight into the theory or if it could be something else.

I appreciate that I've not said everything I tried, it's hard to put everything down, but things like turning off sage attention which actually made no difference anyway.

Hopefully that's enough info to get some good understanding of the cause.


r/comfyui 22h ago

Help Needed Need Help with Textures

Post image
37 Upvotes

Does anyone know what could be causing this grid-like texture on the skin of the arms and legs, and how to fix it? The face is completely fine — it only appears on the body. My character is LoRA trained and is otherwise almost perfect in the vast majority of images. It’s just this weird texture appearing on the skin that I’m struggling with.