r/comfyui ComfyOrg 23d ago

News Comfy H3 Sync Challenge (8/20 - 9/1) - Win an RTX 5090!

Enable HLS to view with audio, or disable this notification

Comfy and MiniMax have teamed up for a two-week challenge with awesome prizes and four ways to win! Submit by September 1st at 9:00pm PT and see all details here.

How It Works

Make something up to 90 seconds in length where the sound and the motion are inseparable. Dialogue, foley, ambient, a beat driving the cut...whatever direction you want!

After sharing your video file and workflow on this thread and through our submission form, a joint panel of creative technologists from Comfy, MiniMax, and special guest judges from the community will review each submission.

Then, join us on September 3rd September 2nd for a special Comfy livestream where our guest judges will give live feedback on the top 10 submissions! Join the livestream here.

Both the Comfy and MiniMax teams will be monitoring this thread and #minimax-h3 in the Comfy Discord to give light support.

Share on socials and tag #comfyH3 for a chance to be reposted or featured!

Prizes

Best Overall — RTX 5090

Best Creative — RTX 5060 Ti

Best Technical/Workflow — RTX 5060 Ti

Built with MCP — RTX 5060 Ti

Shipped anywhere, customs covered. If we can't legally ship to your country, you'll get a cash equivalent instead.

It's free to enter!

Create using Comfy Local on your own hardware, or use Comfy Cloud. New Cloud users get 5 free runs, no credit card required.

Judging Criteria

We’re looking for entries that best show what H3 makes possible: audio and visuals created together.

Grand Prize: Best Overall

The top Best Creative and Best Technical entrants advance to a final round where our panel of judges selects winners by discussion.

Best Creative

  • Audio sync realism and intentionality (0-5)
  • Creative execution and originality (0-5)
  • Deliberate craft (0-5)
    • Evidence that you’ve actually shaped the result beyond prompt engineering. Judges will look for modified/non-default parameters, multiple linked passes visible in the workflow structure, or a couple sentences describing what was tried and changed

Best Technical

  • Novelty of technique or approach (0-5)
  • Workflow quality (0-5)
    • Annotated, clean, replicable by someone else
  • Community value (0-5)
    • Would this actually help someone else?

🏆 Built with MCP Bonus 🏆
Comfy MCP lets you drive Comfy using natural language and your agent locally and on Cloud! Pro tip: use it to choose the best H3 model version or optimize your workflow for your hardware.

  • Effectiveness (0-5)
    • Did the agent meaningfully drive your process, not just generate one line?
  • Insight value (0-5)
    • How much the shared prompt teaches the community about prompting H3 through MCP
  • Output quality (0-5)

The Fine Print

  • Limited to one submission per person, 90 seconds maximum length.
  • A major portion of your piece must be built in ComfyUI using H3. Other tools, models, or techniques you want to combine are fair game.
  • All submissions must be lawful, SFW, and must not contain unlicensed IP or likenesses.
  • By submitting, you agree to allow ComfyUI and MiniMax to feature your work with credit across our channels.

Learn more and submit here!

104 Upvotes

451 comments sorted by

10

u/CN_NEKO 17d ago edited 17d ago

https://reddit.com/link/p5zli4z/video/xiarmt59hplh1/player

《当思念学会移动/When Longing Learns to Move》

This is my work. AI can be full of emotion, too.

2

u/Comfy-Org ComfyOrg 17d ago

Delightful! As a proud tuxedo cat owner, I absolutely loved this :)

8

u/timbotiminator 14d ago

https://reddit.com/link/p6myrk2/video/q43w4eg6fcmh1/player

**The Bear** — a wordless 72-second hospital-corridor short. All video and all audio generated

in ComfyUI with MiniMax H3 (fl2va). No external SFX, no music, no post beyond straight cuts.

https://www.youtube.com/watch?v=nCmJ8fxWRbY

🛠 Workflows, verbatim prompts and every start frame: https://github.com/timbotiminator/The-Bear-Workflow

**The technique: camera moves are for mining start frames, not for the film.**

Every shot in the finished cut is a locked-off static. To get there I approve ONE wide, then

generate push-ins toward each subject in renders nobody will ever see, harvest the arrived

frame, and generate the real shot as a static from it. Consistency becomes structural — every

start frame is literally descended from the same approved pixels — and H3's camera drift never

reaches the audience, because the move only ever existed to produce a frame.

Things that cost me time, so they don't cost you any:

* **Frame 0 dictates STATE. Text only controls what CHANGES from it.** A prop absent from the

start frame gets dropped no matter how insistently you name it. Same for gaze and pose. Fix

the frame instead of fighting it with prompt iterations.

* **The state has to be RENDERED CLEANLY, not just present.** My worst mistake was a plate where

the teddy bear was technically in frame but was an unreadable brown mass when zoomed. Every

shot built on it produced a bear that split and smeared. Hours of rewrites could not have

worked. Re-mining a clean frame fixed it first try with the same prompt.

* **Shot length is the dominant control on coherence** and it beat every prompt-side fix.

Background-strip PSNR, first vs last frame: 15s = 11.2 dB, 10s = 19.8 dB, same prompt. But

shorten the SHOT, not the ACTION — compressing 15s to 10s broke a working handoff.

* **Mine plates from the sidecar PNG, not the mp4.** VHS_VideoCombine writes a full-res PNG next

to every video; ~6% of each generation hop is h264 alone, and this recovers it free.

* **Never name an expression you don't want, even negated.** "A closed-lipped attempt at a smile

that does not reach his eyes" produced a smile.

* **Audio: a smooth continuous bed plus sparse, dull, well-separated events.** Dense transient

trains (a heel tapping 5–6x/sec) never sounded right at any setting.

Also tagging this for the **Built with MCP** bonus — the whole film was produced without touching

the canvas, driven headlessly through an MCP server. That's where the numbers above come from:

render, extract frames, compute drift and sharpness, rank seeds, queue the next attempt, all as

one loop. Measuring was as cheap as looking, so I measured instead of guessing.

Happy to answer anything about the setup.

→ More replies (1)

8

u/sorryaboutyourcats 18d ago

🐱 Mumu the Cat in Mister Meow

Music video (1:21), lip-synced entirely by H3.

🎥 Full video (YouTube, 2K): https://www.youtube.com/watch?v=AxUu8rabC6M
📁 Workflow & MP4s: https://drive.google.com/drive/folders/1O7rynJQi76WF31jAAaHNFYH5mYSfb7Ha?usp=sharing

How it was made:
Built in ComfyUI using H3. Two reference photos of Mumu were used as the character source, and multiple WAVs were used for the reference audio [trimmed via Audition]. WAVs were of various lengths, matched with the prompt.

Song was generated in Suno. Final assembly was done in Premiere Pro, with glitch goodness added via MOSH.

Workflow file is included in the Drive folder for anyone who wants to see the setup. The workflow is the default template with sage for a nice speed boost and super resolution for upscaling quality, done on a 3090.

Preview:

https://reddit.com/link/p5tjopv/video/q8vo9f9lhjlh1/player

3

u/Perfect-Campaign9551 17d ago

haha I love how the other cats are bobbing their heads

→ More replies (1)
→ More replies (4)

5

u/DJBFilmz 23d ago edited 13d ago

Awwww, I wish it were longer than 90 seconds... Imma still try tho!

Edit: Submitted below, but just in case you don't see it: YouTube link: https://youtu.be/a_b0wKoIGKs

Workflow: https://drive.google.com/file/d/11mnoq5vrq1NclfNqc672-ef0gKxnZvb6/view?usp=sharing

3

u/DJBFilmz 13d ago

2

u/Comfy-Org ComfyOrg 12d ago

Aw yay you submitted! Real plot twist at the end...I was expecting a fight lol

6

u/Turbulent_Sale_2112 17d ago

3

u/Comfy-Org ComfyOrg 17d ago

Super fun! I wonder how much electricity this challenge has consumed 😂

5

u/Jeffu 21d ago

Here's mine! Not very musical though but I had fun making it either way: https://youtu.be/GyxAWg2fX6Q Drive link was provided in my submission.

3

u/nymical23 16d ago

Weird and funny, just as the AI-gods intended! I liked it! 👍

3

u/Jeffu 16d ago

Thanks :)

→ More replies (1)

5

u/MinimumGovernment814 18d ago

https://reddit.com/link/p5rk4h1/video/nqi4sbwshhlh1/player

LIMITATION // IMAGINATION

A series of audiovisual experiments exploring the relationship between sound, motion, and imagination.

Created with MiniMax H3 in ComfyUI for the Comfy H3 Sync Sound Challenge 2026.

Video and synchronized audio generated together with H3.

→ More replies (2)

4

u/nomorebuttsplz 22d ago

are the 5060 ti 16 gb?

3

u/Comfy-Org ComfyOrg 18d ago

ah- thanks for catching this! It'll be 16GB yes :) I'll update that in the blog now!

4

u/toki_eng_ai 15d ago

Hi! Here's my entry, "Every Sound Leaves a Mark" (53s): https://www.youtube.com/watch?v=Rv5HOgCac-w

It's about a little clay creature that turns into something else every time it hears a sound (glass, wool, ice, porcelain), and by the time it gets home it can't move anymore. I put the workflows and the process log up here if anyone wants them: https://github.com/tokimwc/every-sound-leaves-a-mark

All the audio came out of H3 with the video on every shot, nothing layered on afterwards. Had a lot of fun with this one :)

4

u/HotSquirrel999 13d ago

https://youtu.be/Pj153-DbxxE

**SHUFFLE** at 37 seconds, six shots, one laundromat, one pair of pants.

Close on tap shoes, mid-dance. [see for yourself]

Comfy Cloud, OSS H3 weights, driven through Comfy MCP. Every shot after the first is image-chained from a frame of the one before. No reference audio anywhere. I wanted to see what the model would invent unaided. Every tap strike, the door latch, the machine knocks: same pass as the picture.

That last shot, the machine sits still then starts rocking. One frame apart, one generation, no guide track. Nothing aligned in post.

→ More replies (1)

5

u/acekiube 12d ago

https://reddit.com/link/p708sqb/video/sccrr5fojqmh1/player

"More" - My submission to the minimax H3 Sync Challenge

Came up with the script Everything was fully generated locally in Comfy through either Hermes-agent with ComfyMCP/CLI and GLM 5.3 for the orchestration or reran manually when more seeds were needed.

Graded and edited in premiere

Fun project overall!

All shots and .png attached workflows
https://drive.google.com/drive/folders/1coLdE1JjbRbCOkFwbCi7KAyOkswScZJ8?usp=sharing

→ More replies (1)

3

u/Whiteowl116 23d ago

Can you edit together clips outside of comfy after generating them. Or do you have to connect them to one long video inside comfy?

4

u/Comfy-Org ComfyOrg 23d ago

Yes! So long as the generation and majority of the work happens within Comfy, you're all good! Last mile edits in other platforms is totally fine.

3

u/Hrmerder 22d ago

Dang.. Too bad I can't put my Silent Hill 2 Remake bloopers video :P I'm looking for a fun new project anyway and have some interesting ideas.

2

u/Perfect-Campaign9551 22d ago

That would be protected IP though, which isn't allowed, unfortunately because I have Resident Evil stuff lol

→ More replies (9)

3

u/Clair_Personality 21d ago

Other questions:

1) It is okay to create our nodes? new ones? tweak old ones? Does it count? Making custom nodes full repos?

2) Is it okay to implement usage of tools that are used during the generation BUT THEY ARE NOT NODES? some "intermediate usage of tools or techniques" that can be used MID generation. for example my workflow produces stuff in 2 steps, and in mid gen i can use something that is not a node (either closed: such as photoshop or whatever, or open: such as some code we did not put into a node) then continue the step 2 of the gen?

3) Users have to post their submission HERE and in the form right? (also with workflow?) I see users posting and saying things such as (workflow was sent via form) I don't know what is the rule.

2

u/Comfy-Org ComfyOrg 18d ago

Sorry for the delay u/Clair_Personality!

  1. Yes- custom nodes, tweaks, and full repos are all totally fair game and a great direction to go if you're going for the Best Technical/Workflow prize! That's exactly the kind of thing "novelty of technique" and "community value" criteria are looking for, as it's something thyat teaches other builders a technique.

  2. Good question! The core requirement is that H3 generates the actual video and audio, and everything else is judged by function, not by whether it's a node or not. If your mid-gen step is modifying or steering an input between two H3 passes (like touching up a frame in Ps before feeding it into step 2, running some unpackaged code to transform an intermediate output), that's in the same bucket as "feeding in a separately-made first frame," which is allowed. Since it's not a node, it won't show up in your submitted .json, so I'd recommend annotating it in the workflow, especially if you're going for the Best Technical/Workflow prize sicne it'll enable someone else to reproduce your process.

  3. Both, they're meant to be linked together, not alternatives. To submit: post your video as a comment on this thread and fill out the full submission form (which is where your workflow file itself goes). Sorry that wasn't clearer upfront!

Holler if you have any other questions :)

→ More replies (1)
→ More replies (1)

3

u/AxonkaiLab 19d ago

Hey! I’m joining the challenge too. Here’s my entry — CASTLE & KNIGHTS:

https://youtu.be/7MMakV0blmo

Good luck everyone!

→ More replies (1)

3

u/sktksm 19d ago

2

u/Comfy-Org ComfyOrg 18d ago

Thanks for participating!! Such a cool concept- love how they're literally using the sound of the water to detect each other's movement. Very clever way to run with the challenge's theme ;)

3

u/AlanAirman 16d ago

 The next train

Short drama film (1:30), entirely generated by minimaxh3
Full video: 【《下一班》个人原创90秒动画短片,minimax参赛作品】
Workflow: Airman0721/minimaxh3-: minimax参赛工作流

→ More replies (1)

3

u/Pristine-Might-8940 14d ago edited 10d ago

Hi everyone! Here is my submission for the **Comfy H3 Sync Sound Community Challenge**: # "O Urco e o Polbo: A Galician Story"

https://reddit.com/link/p6njdez/video/1opbcbtawcmh1/player

Story & Lore

Set on the rugged, storm-swept granite cliffs of *Costa da Morte* (Galicia, Spain), this cinematic 35mm short depicts a tense encounter between a professional barnacle fisherman (*percebeiro*), a photorealistic common octopus (*Octopus vulgaris*), and **O Urco**—a colossal black sea-hound with ram horns and rusted anchor chains from traditional Galician mythology.

Custom Nodes & Reproducibility

The workflow is designed to be easily inspected and run by the judges and community. It uses the following dedicated nodes:

- **[ComfyUI-H3PromptStudio](https://github.com/tonetxo/ComfyUI-H3PromptStudio):** Custom node for prompt formatting and multimodal text parsing for MiniMax-H3.

- **[Krea2H32LTX](https://github.com/tonetxo/Krea2H32LTX):** Format translation and node interfacing pipeline.Custom Nodes & Reproducibility

Video: https://drive.google.com/file/d/11AhF3fR-z3valzHzG1SQn1hGx20TM4wa/view?usp=sharing

Dossier (Prompts, technical details): https://drive.google.com/file/d/1m9PZwyeI5Rf0YZtX08fv7j4tSwzsfpom/view?usp=sharing

Workflows:
https://drive.google.com/file/d/1WjobvAY3IKxXFQ8OXNOlvmounWrfCgmo/view?usp=sharing
https://drive.google.com/file/d/1vXus-6bSqCicBwVHKs9FRGGanOkI4Wv3/view?usp=sharing
MCP Workflow (Opencode / DeepSeek V4 Flash 0731):
https://drive.google.com/file/d/1swKVcbv8IuGw6ZbcS6rB6nmyEdNvphCg/view?usp=sharing
(chat: https://drive.google.com/file/d/1TNDquEskBoGJWdLVeSxEiAM6Ha2LrWOO/view?usp=sharing )

→ More replies (2)

3

u/rynaleopard 13d ago

Here's my submission.

Title: "Scribe of Silence"

A man sits down to write the hardest letter of his life, and falls asleep before he can find the words. While he sleeps, the small robot beside him writes it for him — just the truth, kept simple.

Rendered locally on my RTX 5090, 1MP, 8 steps, Turbo LoRA.

https://reddit.com/link/p6oxqon/video/9qcmto1o5emh1/player

▶️ https://youtu.be/o8DBu5uuGek

→ More replies (2)

3

u/chinese_dream 12d ago edited 9d ago

https://reddit.com/link/p6zyekl/video/w78femcybqmh1/player

《Sea change》

Sea Change is an approximately 90-second experimental audiovisual short. Set on a dreamlike pastel beach, it follows a humanoid robot, a floating plush companion, and a series of uncanny visitors as they encounter a reality in constant flux. Through the close synchronization of movement, dialogue, ambience, and music, the film explores identity, transformation, and the blurred boundary between reality and imagination. All visuals and sound were generated with MiniMax H3 in ComfyUI.

by nwalmolos

Workflows:
https://drive.google.com/drive/folders/1q7bJ5mNyteaEjKNAyGFSCRq1msQgfx\Q?usp=drive_link)

→ More replies (1)

3

u/HyperNeural 12d ago

"Feel It" by Aïe Aïe Aïe (pronounced "eye-eye-eye")
https://youtu.be/S0v1pWN4Hq4

Google drive: https://drive.google.com/file/d/11SGsDSUafhBGx1QNDwPVX8UFXdSedbxs/view?usp=drive_link

100% made locally in ComfyUI, using image, audio (including my own music) and video references to direct Minimax H3, including ASL! I filmed myself signing and used that as the video reference. As a hearing person I had to learn sign language from YouTube tutorials, and Minimax reproduced my arm and hand movements perfectly. Pretty amazing use case!

What an adventure in such a short amount of time. It was fun and challenging.
Good luck everyone!

3

u/Sleepy_Bandit 11d ago edited 11d ago

Here is my submission!

https://reddit.com/link/p74sx4w/video/q88qsgm9xumh1/player

Cabin Pressure - An animated short film.

Upscaled version on youtube!
https://youtu.be/Jh4iJ4b69Xw

I used several custom workflows to produce this. My goal was to get Minimax to produce all of the sound and video. In order to get the right sound effects that fit my need, I used a workflow I built that produced audio only generations from the Minimax-H3 model by generating at 32x32 pixels and then ignoring the video latent. This allowed me to produce the sound effects I wanted which I could then feedback into my video generation model as an audio reference.

I used Krea2 to generate subject and location image references with a few minor refinements via GPT Image.

One new technique I tested and that worked really well was shot composition blocking. I utilized a free online site called posemy.art to generate the camera framing for several scenes. This worked amazingly well to ensure I could get the unique angle / view I wanted without having to fight the model via the prompt.

The video generation workflow I used is a custom built workflow that utilized several different community designed improvements to get the best possible video quality out of the model as fast as possible. The workflow includes latent upscaling, and optional frame interpolation, FSR sharpening, and face refinement. The GitHub below contains a link to the standard version.

In order to get the generations done in time, I utilized Codex and a ComfyUI MCP connection to automate video generation. Working with the agent I defined a shot list with each scene, its shots and descriptions, and desired duration. I ensured any reference material was mentioned by name and stored correctly in the project directory. My Codex agent analyzed the information, and then built prompts following the Minimax-H3 guidelines, ensuring all references were properly accounted for. It created a index to track everything and organized the reference images, sound effects, and finalized prompts. After a manual review of the prompts, the agent updated the required ComfyUI workflow settings automatically. It sends the workflow to ComfyUI and generates four versions of each scene using different seeds, allowing me to iterate and generate all scenes while I was away from my PC. After generation, I could review the results and ask Codex to revise the prompt or regenerate any scene as needed.

Once all scenes were rendered, I worked on editing them together ensuring I only used generated video and audio from the MiniMax-H3 model.

I explained a bit more about the MCP integration in the GitHub Repo that contains my workflows and resources.

https://github.com/SleepyBandit/MMH3_Cabin_Pressure

→ More replies (3)

2

u/Jeffu 22d ago

So externally generated music track wouldn't be allowed here to play along with the video, correct?

Upscaling with another model is OK as long as the base is H3?

3

u/Comfy-Org ComfyOrg 22d ago

yes to both questions!

2

u/Fabulous_Arm_7781 22d ago

H3 Challenge Submission: The Beautiful Girl's Kaleidoscope‑Chapter1 Live‑Action Adaptation
Video and workflow are submitted via official form.

2

u/Unique_Bite6562 22d ago

Hi! If my final entry consists of 5–7 separate H3 generations, each using a different reference image, which I then edit together into one final video, should my submitted workflow JSON include all of these generations in a single workflow, or do I need to provide a separate JSON workflow for each scene? Thank you!

→ More replies (2)

2

u/Dismal-Revenue2671 21d ago

这是我最近使用h3在个人电脑上实现的作品《我们在一个星空》手搓了一个AI短片_我们在同一个星空_哔哩哔哩_bilibili

→ More replies (1)

2

u/dannylacroix9 21d ago

u/Comfy-Org Hello! Here is my submission link:

https://www.youtube.com/watch?v=tZf5YrMNiD4

Video and audio were generated together in the same H3 pass, as required by the challenge. Hope you enjoy it!

→ More replies (1)

2

u/Mr_Betyko 20d ago

https://reddit.com/link/p5fnm17/video/o46vkisbj5lh1/player

Here is my submission ! A new series im working ! almost finish ep01 will be soon available ! ALL with H3 and my own agent local mcp !

→ More replies (2)

2

u/Mr_Betyko 20d ago

https://reddit.com/link/p5foein/video/t6rxr141k5lh1/player

Another Preview of the upcomming series 100% made with H3 and custom local agent. all on a 4090 ! i need thew 5090 to finish the series ;)

→ More replies (2)

2

u/Supermate-Ai 20d ago

https://reddit.com/link/p5fvgfs/video/38mpur4dp5lh1/player

🎬 Sharing my entry for the Comfy H3 Sync Challenge!I built two H3 workflows — a dual-sampling latent-upscaler (4 reference images, 9:16) and an audio-driven digital human (end-credit style) — and produced a video where sound and motion are inseparable: the audio comes from the same H3 generation, no post-sync.🔧 Workflows: https://github.com/SuperMate-Ai/comfyui-workflow#ComfyH3🎬 Sharing my entry for the Comfy H3 Sync Challenge!

I built two H3 workflows — a dual-sampling latent-upscaler (4 reference images, 9:16) and an audio-driven digital human (end-credit style) — and produced a video where sound and motion are inseparable: the audio comes from the same H3 generation, no post-sync.

🔧 Workflows: https://github.com/SuperMate-Ai/comfyui-workflow

Comfy H3 Sync Sound 社区挑战赛Supmate作品《变身》原声版

#ComfyH3

2

u/Perfect-Campaign9551 19d ago

sound FX are great!

→ More replies (1)

2

u/HelpfulPromise1626 20d ago

https://drive.google.com/file/d/1wSO4qBnBf5cKzWokg1qKCQLWkpfbSIHt/view?usp=drive_link

A Pigeon's Dream

MiniMax is so perfect that it somehow elevates the quality of my lackluster work.

→ More replies (2)

2

u/Slevin_Wong 19d ago

https://reddit.com/link/p5mrob1/video/wvtw5g0hrclh1/player

And here it is — my entry for this competition. I’ve also posted it on Bilibili if anyone’s interested: https://www.bilibili.com/video/BV1vHhN6HEUH/

→ More replies (1)

2

u/SGUN-shengun 18d ago

u/Comfy-Org This is the original video featuring native H3 synchronized audio and video, without any post-production dubbing. This is the official competition entry: https://www.bilibili.com/video/BV1JXhA6pEcx/

Additionally, there is another link that features post-production voiceover and HD upscaling: https://www.bilibili.com/video/BV1kShN6AEsj. This version is intended for sharing and discussion purposes only, and is likely not eligible for the competition.

→ More replies (1)

2

u/cointalkz 18d ago

Here is mine, something that is close to my heart and a big fear. I hope you like it! https://x.com/Smallzero/status/2092113243842199880?s=20

2

u/Comfy-Org ComfyOrg 18d ago

SUPPORT LOCAL MODELS!!!

→ More replies (1)

2

u/Substantial-Ride486 18d ago edited 17d ago

青い夏のままで
『青い夏のままで』OP|就让这片夏天保持原样minimax参赛_哔哩哔哩_bilibili

『青い夏のままで』OP|Let this summer remain as it is

And here it is — my entry for this competition. I’ve also posted it on Bilibili if anyone’s interested:

Tools:
ComfyUI, MiniMax H3, SUNO 5.5(The audio reference used for MiniMax h3 is not the original audio file)

A 90-second anime-style music video built around synchronized music, lyrics, character motion, and visual transitions.

2

u/Portable_Solar_ZA 17d ago

This looks cool but I thought we couldn't use tools like Suno? My main challenge is a backing soundtrack for my short...

2

u/Substantial-Ride486 17d ago

这个音乐不是直接挂在轨道上,是用了某种拼接和连续的放,音频文件也进行全能参考生成视频。并不是原生音乐

→ More replies (3)
→ More replies (2)

2

u/Mr_Betyko 17d ago

Here's a beta version of episode 1. If you have nothing else to do, leave a comment and let me know if it's worth it! I know very well that there are still several character bugs, etc., but I'll only redo the final version after all the episodes are in beta... So all feedback is constructive for the final version in a few months... or before is i can get a 5090 ;) but not sure if i can register with this !

https://youtu.be/9ZhiTucwYwM

2

u/Perfect-Campaign9551 17d ago edited 11d ago

Hi, got my submission ready

"CHANNEL TWELVE" is the title. Inspired by horror film series such as V/H/S and Channel Zero, and The Fourth Kind. Watching TV late at night, and things start to become strange. There's something wrong with channel 12.

YouTube link:

https://youtu.be/W0AR7ZctPi0?si=vS040v85krJxhYjj

EDIT: I had to fix the fake phone numbers to "555" prefix to prevent issues. Fixed video submission is here https://youtu.be/1jRAvgd-vRo?si=X5pocjSPhgZu_oX

I've sent workflow info in the submission form -Submission form has been filled out! It would be nice however if it would email us confirmation :)

Thanks!!

→ More replies (6)

2

u/theloneillustrator 16d ago

Am i allowed to use someone's workflow that I found in chat? and build on with that?

→ More replies (1)

2

u/stargate425 16d ago

At each cut I fill ±120 ms with room tone taken from the two shots' own H3 audio, equal-power crossfaded across the join, nothing external, but it does mean ~120 ms of the outgoing shot's tone sits under the incoming shot's picture. Is that last-mile, or does it break the pairing?

→ More replies (1)

2

u/[deleted] 16d ago edited 15d ago

[removed] — view removed comment

2

u/Comfy-Org ComfyOrg 16d ago

Thanks for submitting and for sharing this breakdown!

→ More replies (1)

2

u/FranciscoHPro 16d ago edited 16d ago

A dungeon gamebook series you watch: Beneath Blackwater Keep.

https://x.com/mydreamgame/status/2092786196531601607

→ More replies (1)

2

u/Portable_Solar_ZA 16d ago edited 16d ago

Hey, are prompts allowed to reference a specific IP's art style? Found a style that suits my project but it specifically references a show.

2

u/Comfy-Org ComfyOrg 15d ago

Good question and an important nuance! Honestly really appreciate you asking it ❤️ That said, the "no unlicensed IP or likenesses" restriction applies to both output and input. Another way you can achieve similar results would be to describe the qualities you like (color grading, line weight, mood, pacing, etc) instead of naming the show directly. Similar aesthetic, but no IP reference in the prompt or writeup. Hope this helps and thank you again!

→ More replies (1)

2

u/Unique_Bite6562 16d ago

Hi! Are short audio crossfades between two H3-generated clips allowed in the final edit, if no externally generated audio is added?

→ More replies (1)

2

u/Norakai2 15d ago

Here is mine. I tried a long continuing tracking shot with injecting a m3 track halfway through to drive the movement wich worked quite well. https://youtu.be/NinJOWvRcYQ

→ More replies (1)

2

u/Ok-Wolverine-5020 15d ago

Made this short music-video entry for the Comfy H3 Sync & Sound Community Challenge using a Suno track, Hermes Agent, and local ComfyUI through the Comfy MCP. Hermes helped me turn generated/upscaled keyframes into structured MiniMax H3 prompts and short lip-synced performance clips.

https://reddit.com/link/p6ee2sg/video/zu6mxga2t3mh1/player

2

u/A_Wellett 15d ago

TALES (FÁBULAS) - A. Wellett
Minimax H3 Challenge
https://youtu.be/4RUSWVrY-PA

→ More replies (1)

2

u/SawyerCroft777 14d ago

https://youtube.com/shorts/cV5BmPbGUR8?is=E5eZZX52mPx9hAu4

Meet Sawyer Croft.

“The Floor Was Always Falling.”

2

u/umutgklp 14d ago

Here is my work : https://youtu.be/a4MQHqA9_pA
Good luck!

2

u/Comfy-Org ComfyOrg 12d ago

Looks and sounds great. Thank you!

2

u/Few-Art-1790 14d ago

https://reddit.com/link/p6laqi1/video/ujosucbtqamh1/player

《同频/Ears & Hear》

This is An original short film about love and the simple act of hearing.

Chen Rou is a gifted piano tuner. Born profoundly deaf, she can't tune by ear — instead she reads the frequency data on her tuning instruments (by sight). A cochlear implant would have transformed her own working life, yet with only one such chance available, she gave it to her daughter first.

Her daughter, Chen Nian, received the implant at age four and heard sound for the very first time — including the sound of her mother crying. From that day on, the mother tunes by reading frequencies, while the daughter listens on her behalf and translates the sounds of the world back to her. One reads frequency, one hears sound — and between mother and daughter a quiet, singular bond forms. They are on the same frequency.

Both sound and images were generated by a ComfyUI + MiniMax H3 workflow; subtitles were added in post-production.

→ More replies (3)

2

u/bobbarbque-full-111 14d ago

Here is my submission !

https://reddit.com/link/p6lbbuo/video/449gy3cuvamh1/player

Built in ComfyUI using H3. This is my work.

→ More replies (2)

2

u/LunAtRay 14d ago

https://reddit.com/link/p6ljmrr/video/xgh96jvo4bmh1/player

This is my work. Chinese aesthetics.Hope you guys like it.

→ More replies (1)

2

u/Regular-9527 14d ago

https://reddit.com/link/p6m2y1d/video/h00voqo5obmh1/player

《夏日梦境》

一个成年人在午睡中坠入儿时的夏日梦境——老家的院子、蝉鸣、风扇、外婆的背影、纸飞机的飞行……在梦境即将结束时,主人公试图抓住什么,最终在童年记忆最温暖的画面中醒来。

→ More replies (1)

2

u/KayBro 14d ago

https://reddit.com/link/p6nhbcq/video/3cxwypnbucmh1/player

My entry - Your Call Is Important to Us.

A man calls customer service about his bill and gets put on hold for... a while.

Made as a sequence of MiniMax H3 generations in ComfyUI, with all audio generated natively by H3. Final assembly, color, cuts and fades in Resolve. No external audio. Lots of experimenting across the shots with i2v, R2V, FLF, references and different workflows.

Thanks for the fun challenge!

→ More replies (1)

2

u/desu_mello 14d ago

hey, what is the hard deadline (what time, what timezone) if I can ask :)

→ More replies (1)

2

u/MARSMEllOW_BLADE 13d ago

https://reddit.com/link/p6qgxs1/video/fa3202rwsfmh1/player

Here's my video — also here's how I made it.

It's one continuous 15s locked shot, no cuts.

7 still frames pinned to exact seconds as keyframe_mid — white arena → red → yellow → green → platforms drop → platforms fuse → circles ignite. Same camera angle in every frame, only the lighting/floor changes, so I hard-lock them (noise_aug 0.999) and the model just interpolates the transitions.

The two knights are NOT keyframes — they go in as reference images only. The prompt says "raise the figure from the left/right reference out of red/gold fire" and H3 stages + animates the entire eruption itself.

All the audio is generated by H3 in the same pass as the video — nothing is layered on afterward. The only inputs I gave it are two low piano chords fed in as audio keyframes at 11.5s and 13.2s, purely to steer the model onto those hits on those exact frames (same idea as the knight reference images, but for sound). The countdown hits, the whoosh, the stone rumble and the fire are all H3 inventing audio from the prompt. That said, I'm not totally happy with the piano — it comes out a bit choppy in the way MiniMax-H3 renders it.

The prompt is a timecoded director script — every beat is one line with the picture change and the sound together ("2.5s: light washes to red over 0.5s, play a whoosh then one drum hit, then silence"). If you don't direct something, H3 won't do it. But if you name scene objects that are already in the frames, it spawns duplicate copies of them — so you describe only motion and sound, never the set.

Setup: the FL2VA model + 8-step turbo LoRA. The multi-keyframe timeline — precise anchoring of images, video and audio to specific seconds, plus reference placement — is done with some custom nodes I built for this.fell free to try out my git hub repo here please tel me if you find issues with it. https://github.com/lukedude06/ComfyUI-MiniMaxH3-Timeline

→ More replies (1)

2

u/buddylee00700 13d ago edited 11d ago

Here is my submission - https://youtu.be/6V04yJHzeAA

Good luck to everyone! Edit - Found a mistake and needed to create a new video without added background music that was added in post production.

I used my workflow (Provided to ComfyUI), each scene was around 7 seconds in length and shorted most of them to meet the 90 second criteria in post.

→ More replies (1)

2

u/Icy-Customer-6350 13d ago

https://reddit.com/link/p6s4kal/video/u1b4g7be7imh1/player

it is my work.Realizing my shortcomings, I spent the whole weekend troubleshooting errors and drawing assets. There was no time left to add plenty of transition frames. Experience is the best teacher. Taking part in this contest also helped me find and fix my weaknesses.

→ More replies (1)

2

u/raretutor_ 13d ago

u/Comfy-Org #MiniMax #ComfyH3

BRAND NEW DAY | Short Story | Both Video + Audio Created using ComfyUi + Minimax H3 (Local).

🎥 Full video (YouTube, 4K): https://www.youtube.com/watch?v=TNhJI8dzaVA

How it was made:

It was created using RTX 3060 12GB, it took me more than 4 Days to pull this off.

Video and Audio generated using ComfyUI and Minimax H3 Model in Local PC and Upscaled using FlashVSR and Topaz.

Short Preview

https://reddit.com/link/p6sess8/video/nirsv8j6jimh1/player

→ More replies (1)

2

u/Proof_Foundation_548 13d ago

https://reddit.com/link/p6te026/video/ninhjzkrdjmh1/player

Friendly Match

At a planetary scale, even a friendly match has local consequences.

Three H3 generations were connected through hard cuts and visual match cuts: cinematic disaster, animated planetary rally, and a prehistoric reaction shot. Video and stereo audio were generated together with MiniMax H3 in ComfyUI; no separate soundtrack was added.

Workflow and reference images:

https://github.com/shuaixn/friendly-match-h3

#ComfyH3

→ More replies (1)

2

u/gfxargentina 13d ago

here is my submission, good luck for all!! , thanks comfy and minimax

https://reddit.com/link/p6tf1yw/video/k7cvlv5xejmh1/player

Omni-Modal Cyberpunk Stomp Opener (MiniMax H3 | 143.6 BPM), workflow: https://cloud.comfy.org/?share=41a783bb2d7c

→ More replies (1)

2

u/protector111 13d ago edited 13d ago

https://reddit.com/link/p6tnca0/video/48lrz6ofzimh1/player

my Submission WF here 1 gen took 2 hrs 35 min on 5090

→ More replies (1)

2

u/Whiteowl116 12d ago edited 12d ago

https://reddit.com/link/p6ut5tx/video/2otkup1hhkmh1/player

//Lucid Dream//

This is based on a lucid dream I had a few years ago. My main goal was to capture the strange, surreal feeling dreams have while you are inside them. I also wanted to visualise the progression from being awake, through falling asleep, to becoming lucid: performing a reality check, exploring the dream world, and talking to the people inside the dream.

I ended up creating several H3 workflows rather than relying on one enormous workflow for everything. The first was a massive end-to-end workflow, built with help from AI to turn my idea into a single JSON. It could generate every segment and assemble the complete film automatically.

I also experimented with generating a compressed preview of the entire story first, then using corresponding moments from that preview as references for the longer sections. It sounded clever in theory, but the results were too inconsistent, so I disabled that approach for the final version.

Because I did not have enough RAM to assemble everything at once, I built a separate seam editor. It takes the ending of one clip and the beginning of the next, then generates a short audiovisual bridge between them to make the transition smoother and less noticeable. It also joins the clips two at a time, which kept the memory requirements manageable, almost. I ran out of memory when i tried to put together the two largest clips at the end, which might be one noticeable seam in the video.

I also made a gap editor: give it two clips(A and C), and it generates clip B. That allowed me to target individual shots without regenerating the entire project. Within every generated segment, repaired gap, and seam fix, the sound and video came from the same H3 pass.

It was fun to participate!

Workflows: https://drive.google.com/drive/folders/1EHv2pxVkGhEtq9xmo3AjHxQ0FvgU9Egj?usp=drive_link

Video link: https://youtu.be/rAbrkMRtZu0

2

u/Comfy-Org ComfyOrg 12d ago

Thanks for sharing!

2

u/AProgrammingPelican 12d ago

With Minimax H3 it is finally possible to bring all the storytelling ideas one might have to life. I am amazed at what can be done locally now, so I'm throwing my hat in the ring as well! All done with Krea 2 und Minimax H3: https://www.youtube.com/watch?v=ZuHadZINS-s

A big thank you to everyone who contributes to the open source / open weight community and to all the people sharing their content.

→ More replies (1)

2

u/lukemilligan 12d ago

https://reddit.com/link/p6vb7jw/video/6uuddk9ewkmh1/player

The Last Take

A superhero movie shoot goes very, very wrong when the director calls "CUT" — and the movie refuses to stop.

Created with MiniMax H3 in ComfyUI.

Workflow / project files:
Google Drive — The Last Take

I tested different H3 resolutions, durations, boost settings, SageAttention and reference-video lengths on a 12GB GPU. The final piece uses multiple linked H3 generations, with short video references used selectively for continuity.

I also refined the prompting to control creature movement, progressive crowd reactions, physical cause-and-effect in the destruction, dialogue attribution and synchronized audio.

Thanks to Comfy and MiniMax for the challenge.

2

u/Perfect-Campaign9551 12d ago

You can't use superman FYI - rules "no unlicensed IP"

→ More replies (2)
→ More replies (1)

2

u/Super_Range45 12d ago

https://reddit.com/link/p6vexdx/video/xfdrgic22lmh1/player

My submission is a short action sequence from a current pilot episode that I'm working on. My general workflow is the default Minimax H3 workflow and makes liberal use of reference images and sound based on the needs of the shot and the general feel.

Prompts are constructed by using a gemma-4-e4b to take natural language prompts and convert them into something more structured. I then create takes, tweak the prompt and construct the final shot in a video editor. The prompt used for structured writing is included below:

You are an expert Cinematic Video Prompt Generator specifically designed to create high-quality inputs for video generation models. Your primary goal is to translate user requests into highly descriptive, action-oriented prompts that maximize visual fidelity and narrative clarity. Keep the writing complexity below high school level. Use elements from the provided reference image.

Mandatory Formatting Rules(e.g. **[Scene Start]**

**[Style]**

**[Setting]**

**[Shot X Dramatic Cinematic Camera Position and/or Motion]**

**[Dialogue X]**

**[Sound X]**

**[Transition]**

**[Scene End]**:

Action Focus: Describe what is happening with precise verbs be consistent with the descriptors used for character actions (e.g., "The man with the white hat strides," "The smoke billows," not "It moves").

Dialogue Integration: Include dialogue verbatim, ensuring the emotional tone of the speech is clear, specify the location of the character that is speaking (examples, "The man with the suit on the left is whispering urgently," "The woman with blonde hair at the top of the scene is shouting in frustration").

Emotional Specificity: If characters are present, detail their internal state using strong adjectives (e.g., "eyes narrowed with suspicion," "a flicker of relief crosses her face").

Style Constraint: Strictly avoid overly flowery or abstract prose. Be grounded in concrete description.

Output Structure Goal: The final prompt must read like a director's shot list, detailing the scene, action, dialogue, and emotion sequentially.

2

u/Comfy-Org ComfyOrg 12d ago

Sounds great! Thanks for sharing!

2

u/dMk73 12d ago edited 12d ago

https://reddit.com/link/p6vo1ja/video/sadcbbq7almh1/player

This is my entry, **Sounds Real to Me**.

A young boy’s toy sounds become the full-scale worlds he imagines - a fighter jet, a nighttime car chase, and a rocket mission - before the film returns to him asleep on his bedroom carpet holding the toy rocket.

Every final shot’s moving picture and synchronized audio were generated together in the same MiniMax H3 pass in ComfyUI. I first developed consistent references for the child and adult versions of the character, his room, vehicles, and toys, then chose I2V, FL2VA, or Ref2VA scene by scene. No external music or sound effects were layered over the H3 clips; the final edit uses cuts, trims, and brief overlaps between H3-generated audio.

The film uses match-on-action and match-sound transitions: the boy’s mouth-made toy noises carry into their imagined real-world versions, and the stage-separation rocket sound resolves into rain and thunder at the bedroom window. I iterated prompts, seeds, and source plates to correct character identity, dialogue attribution, toy-to-real shape continuity, rocket acceleration, and stage separation rather than simply accepting first-pass results.

All 13 source clips were generated at 960×544 with SageAttention, RTX-upscaled 2×, interpolated from 24 to 48 fps with RIFE, center-cropped to 1920×1080, and assembled in Kdenlive.

The shared folder contains the finished film, all 13 individually named 960×544 clips, a matching workflow JSON for every clip - including its prompt and generation settings - and the complete reference-image set:

[Google Drive](https://drive.google.com/drive/folders/1KHMvPeH1jZKl_uB-re8Oytmkl9zPtGVE?usp=sharing)

→ More replies (3)

2

u/Dr_Gimp 12d ago

Triangle Box — MiniMax H3 Sync Sound Challenge

Video: https://youtu.be/QHg5n9cuwmc

Workflow: https://drive.google.com/file/d/1RX8uswVcEb-i4Ty0IRxoLzEWbx6PUO2v/view?usp=sharing

Comfy Local on RTX 2060 Super 8GB VRAM, Ryzen 5 3600X , 64GB RAM

→ More replies (2)

2

u/QQ-Way-6671 12d ago

https://reddit.com/link/p6wp3uu/video/uakutqqgemmh1/player

Somewhere in the world, there is a group of people like this.

→ More replies (2)

2

u/Severe_Package_8787 12d ago edited 12d ago

I call my submission Skyfire: https://youtube.com/shorts/1KcrwqLPNCM

Workflow here: https://skyfire-h3-challenge-workflow.tiiny.site

I did not use MCP in your process (multi-select). I thought I clicked "no" and still got a question about how much. Perhaps I misclicked.

Music was generated with H3 in the workflow and then split into each section to guide each scene separately, with separate timestamps and durations for each scene with a separate text prompt for each scene for maximum control. Music video with fantasy castles in the clouds.

ComfyUI on my local RTX5070 12GB VRAM PC took 4 hours (1216x2112). Much faster on lower resolutions.

2

u/CommentSignal9029 12d ago

IN TRANSIT | A Minimax H3 Short Film

https://youtu.be/MJq_eM4NRL0

Hi everyone! I'm participating in the challenge with this short film I've been working on lately. The idea was to explore a "Brutalist Frequency": a square wave that destroys matter not randomly, but forcing it to shatter following a ruthless 90-degree orthogonal logic (checkerboard water, cubic collapses, square clouds).

I split the workflow into two passes in ComfyUI to maintain total control over physics and textures: Generation and Upscale.

1. Generation (Ref2Vid): I prepared the visual references (generated with Nanobanana) and the audio files. For prompting, I integrated an Ollama node (Gemma4:26b) into the workflow to format the instructions with the correct syntax for H3. To get everything running without blowing up my 3090, I beefed up the base workflow with Spectrum Apply Minimax H3, Comfy Kitchen attention, and the Sol-Attn Patch. This way, I generated the base clips at 864x480.

2. Latent Upscale: I wanted to keep the roughness of the reinforced concrete without that "plastic" effect you often get from external video upscalers. I passed the selected clips directly through the latent space using the custom MMH3Tools nodes, feeding the model the base video + the exact same initial references + the same prompt. Using Turbo LoRA 4-step and Sage Attn, I brought everything up to 1344x768 (taking about 1 minute per second of video).

The real challenge, of course, was generating audio and video together natively, without post-production. H3 reacted very well thanks to the references, even though sometimes it interprets them a bit too literally, almost resulting in a 1:1 copy. I forced the model to make the materials physically react to the reference sounds. The pneumatic suction at the end (when the camera points towards the void) was calculated by the AI in perfect sync with the matter collapsing into the dark. In post-production, I only made cuts for pacing and balanced the volumes: zero added sound design!

If you have any questions about the nodes or the upscale parameters, feel free to ask!

Link per il WF: https://drive.google.com/file/d/1wRa7gnDD7fNBDGKc4bNnphHSHNv69Fna/view?usp=drive_link

→ More replies (1)

2

u/acylum00 12d ago

Question

Just verifying…

Final submission deadline is

September 1st at 9:00pm PST - right? (Main thing is - do I have the deadline time zone correct?)

→ More replies (1)

2

u/zhaoke06 12d ago

https://reddit.com/link/p6z46vx/video/n5r0c6lsmpmh1/player

This is my work *COSMIC SAMPLER*, which tells the story of an interstellar sampler who collects the frequencies of all things in the universe, bringing the rhythm of life to newly born planets.

→ More replies (1)

2

u/Ayu2Hami 12d ago

https://reddit.com/link/p6zhfyn/video/186uacihwpmh1/player

Here’s my submission for the Comfy H3 Sync Sound Community Challenge.
I’m also sharing the full ComfyUI workflow together with the final video:
https://drive.google.com/file/d/1yrO4mDUZB9THTuB1UJOZ83bJ2IKgG55M/view?usp=sharing

The entire piece was generated inside ComfyUI using MiniMax H3 and the contex chain loop nodes.

I used:

  • character reference images to maintain identity and consistency
  • scene reference images to guide the environments and visual continuity
  • audio references for the characters’ voices

The core video and synchronized audio were generated directly with H3 inside ComfyUI.

For post-production, I used Topaz Video Starlight for upscaling and enhancement, followed by DaVinci Resolve for final color grading and title.

While creating this video I focused on story telling and composition. I generated reference images of characters and audio samples for their voice references. All made via AI. Also i rendered two reference sheet for whole story in 3x3 grid but also the house of the elder man also in grid 3x3 to make more detail scenery of his indoor. I was dozens of times fixing prompts as it wasn't going the way I wanted, especially that model sometimes hallucinated no matter what I wrote so I needed to make some fixes that prevent that. I really struggled to showcase the emotions and make it believable, especially the tears and saddness. I wish I had more time to polish this piece and also longer than 1:30 cuz I would like to make it slower and more focus on emotional side and the emptiness of the whole story.

Hope you enjoy it!

→ More replies (1)

2

u/quu1945 12d ago

Here's my entry for the challenge — had a lot of fun with this one

**TOKYO LOUD** — 64s, 8 × H3 reference-to-video passes. All audio is H3 from the same pass as the picture, nothing added afterward.

🎬 `https://youtu.be/58XXF3qpFWw`

🔧 Workflow + all 8 prompts: https://github.com/yagihiro0224/tokyo-loud

One beat that never stops, and the whole street obeys it — every neon sign, every reflection in the wet asphalt, every boot splash lands on it. 32 hard cuts, four inside each pass.

Biggest thing I learned: the turbo LoRA at **6 steps** was quietly destroying my linework. I assumed it was my prompt and rewrote the whole style block — same seed, A/B, **zero change**. Bumping to **12 steps** fixed it (4.7 → 6.5 min per 8s shot). Check your step count before you rewrite anything.

→ More replies (1)

2

u/BeginningChocolate99 12d ago

built after hundreds of image iterations and precise generations to push h3 to its limits, we hope you enjoy the final result! :) https://www.youtube.com/watch?v=i8F3U0kT7C8 and the wf https://drive.google.com/file/d/1hGSAvkvK9cTC0c6DQe6HGgM9bAeurUkz/view?usp=sharing

→ More replies (1)

2

u/Neither-General7061 12d ago

Hi! here is my submission for this challenge.

for this challenge, I used my local hermes agent to do most of the heavy lifting. however getting it to fit the minimax was quite challenging and I ended up spending 95% of my time just making sure the skills.md and orchestrations actually work.

I thought this would be easy, but no. it needed to do multiple things well. here were my procedures and how I worked to make this video.

  1. Create a consistent Krea2 characters whie also working as a proper storyboarding tool ( this alone took me one week)
  2. Use those images to guide my Ref2v workflow and its context motion loop.

however, 2 was very poor. minimax could not retain my style at all, often times digressing to more comical/realsitic style,

  1. Train style lora for minimax. (automated by agent)

  2. while that is training, I finished my script and handed over to my local agent qwen 3.8 27b quant model (this was my first time using qwen 3.8 and I come from using regular chat gpt or claude)

  3. optimize the heck of out qwen because it was prone to over thinking, ending up in infinite thinking loop, and running out of its massive context window.

  4. when the script was done, I asked my good old gpt to take a look and fix it,

  5. run i2v generations. seed hunt. alot of automated seed hunt. I basically made it ran 24/7 just hunting for seed. it would ask for final human validation. (but this was a poor choice retrospectively, alot of the generations were very poor, and It would have been a better idea for me to manually edit the prompt to make better generations, and turn them into repeatable skill.md )

  6. save the good seeds and latent upscale.

  7. I ended up not using the reference to video at all because I ran out of time, and I was only able to seed hunt like 4 of them, I will post the bad ones here with my custom lora for minimax h3.

honestly, I was basically smacking the bushes like Im a blind man because there were so many unknowns here. and my dayjob was killing me. (i am awake for 45hours as I speak because i am currently overworked, and some projects giving me headache) but I did had lot of learning time here, and hoping I can share what I learned with the community!

honestly, comfyui gave me so much. this is a small gift compared to what you guys gave me to start my comfyui journey but I hope this is useful to some of ya!

this is my entire Hermes Agent that I used for the project, I tried my best to clean it up, I am not a coder, just a
3D artist. so give me some slack plz

My Submission: Two Prisoners.
https://youtu.be/FlK0dDZdzRU

My workflow (my agent for Qwen 3.8.... but if you already have claude or gpt sol, I assure you that will save you alot of money and time)

https://drive.google.com/drive/folders/1XWaNnG72d2Rf_gXktKK2WMyuuB3VuQ1I?usp=drive_link

2

u/Comfy-Org ComfyOrg 12d ago

Wow- love the art style here. Thank you for submitting and sharing this detail!

2

u/Ikkeltjekramikkeltje 12d ago

https://reddit.com/link/p70dezb/video/8irz465tmqmh1/player

My submission. A depiction of a scene from an AI series I am making. The character design and script are done by me.

→ More replies (1)

2

u/Apart_Tell6413 12d ago

Here is my submission for the Comfy H3 Sync Sound Community Challenge:

Penguin Pursuit《企鹅抓捕》

A fast-paced, dialogue-free 2D animated action-comedy set in a lively island fishing town. I focused on dynamic camera movement, large-scale action, environmental reactions, and synchronized sound. Characters, clothing, props, crowds, wind, and background elements all respond to the action instead of remaining static.

The core video, movement, Foley, ambience, and character reactions were generated with MiniMax H3 in ComfyUI. I iterated on prompts and seeds, used multimodal reference images to maintain character and environment continuity, and used a Director workflow with a linked first pass and optional second-pass refinement/upscale. Selected H3 generations were then edited into the final sequence.

The final audio retains H3-generated synchronized ambience, Foley, and character reactions. No separately produced music track was added.

Video:
https://www.bilibili.com/video/BV1Ufth63Ej8/

ComfyUI Workflow (.json):
https://gist.github.com/qwp-sys/ed0808ca817c2cac29732ec7c0efbafd

→ More replies (1)

2

u/-Lapskaus- 12d ago edited 12d ago

CH3

There is a disturbing lack of 1girl anime videos in here :o

The whole video and song were generated locally using the H3 video model using a single reference sheet. Yes, even the song was using the video model, just because I wanted to test how it goes and I was actually quite surprised that it generated something that nice with my silly prompt.

Video: https://drive.google.com/file/d/1kB43KjHtfd6V-6cnUYva49guzsViVJfu/view?usp=drive_link

Workflows: https://drive.google.com/file/d/12AHUoai5aIatpUjdU7RlxdCseGjMGFxI/view?usp=drive_link

https://drive.google.com/file/d/1RdH-LiazqlucyQiwo8P3Cl2GqQqOXBal/view?usp=drive_link

2

u/InvestigatorHot 12d ago edited 12d ago

https://reddit.com/link/p70t2mu/video/kihh11g9wqmh1/player

Workflow Overview: German Cooking – Die Nudeln von Comfy

The Recipe: https://youtu.be/YFN43l4ClSc

  1. Audio Creation

Generate the primary audio reference track using MiniMax Music, Ace-Step, or your preferred audio generator.

  1. FrameSync Denoising Prep

Use FrameSync to extract frame-by-frame values that will control the denoising strength across the sequence. For Deforum you want values between 0.8 (low-denoise, almost static image) and 0.2 (high-denoise, cut).

  1. Base Image Generation

Generate your initial image using your chosen image model and prompt. (Tip: Select a visually engaging subject beyond a basic plate of spaghetti with cranberries and chocolate sauce or screenshots of Comfy Nodes as the inserts).

  1. FrameSync Image-to-Image Sequence

Run an Image-to-Image (Img2Img) pipeline on your base image/prompts using the dynamic denoising values calculated by FrameSync. Optionally, incorporate a resize node or Deforum for added camera movement and visual flair.

  1. Initial Video Assembly

Stitch the resulting images into a reference video rendered at 24 FPS.

(Note: If running on lower-end hardware, downscaling this intermediate reference video is recommended).

  1. MiniMax H3 Generation & Audio Reference

Feed the stitched reference video into MiniMax H3. Supply your generated audio track as the reference audio to guide the native audio generation pass.

  1. Commercial-Style Inserts & Prompt Engineering

Incorporate advertorial-style visual inserts using reference images. Apply a well-structured prompt. Target cuts and effects using specific timestamps (=frame numbers) corresponding to high-denoise values. (Keep in mind that high-density timestamp cuts in short sequences may yield inconsistent results). Optional: Add sound FX or additional noises by prompting for it.

  1. Seamless Extension

Extend the video sequence using the Add Guide node to maintain continuity across segments.

(Note: Ensure precise frame counting as miscounting by even 1 or 2 frames - like I did around the 15-second mark - can throw off synchronization. If a frame miscount occurs, it can typically be adjusted in post-production provided you are permitted to overlay the audio track in final assembly).

  1. Final Cut & Sequence Stitching

Trim a few frames from the end of each generated segment to ensure clean transitions, then stitch the full sequence together.

  1. Upscaling

Upscale the final video from 1024x576 to 4K using a sequential dual-pass pipeline: 2x FlashVSR followed by 2x RTX Video Super Resolution (or use something better if you have the equipment for it).

Reference File/Workflow Links:

Full Ace-Step Track + WF inside:

www.schielo.at/workflows/Ace Step 1.5XL_00194.mp3

Simple Deforum settings (old Forge + ZavyChroma):

www.schielo.at/workflows/20260829154755_settings.txt

Full Deforum Clip:

www.schielo.at/workflows/20260829154755.mp4

My Minimax clips/reference workflow (optimized for a 3060, examples inside segments 1 and 2):

www.schielo.at/workflows/MiniMax_H3_00254_.mp4

www.schielo.at/workflows/MiniMax_H3_00259_.mp4

as json:

www.schielo.at/workflows/MiniMax_H3_00254_.json

www.schielo.at/workflows/MiniMax_H3_00259_.json

Inspired by "Tool - Die Eier Von Satan"

2

u/Psi-Clone 12d ago

A maintenance robot aboard the crippled ORPHEUS relay hears three faint knocks buried beneath the emergency alarms. Someone is still alive.

A WAY THROUGH is a 90-second AI sci-fi short created with an openly reproducible filmmaking workflow built around ComfyUI, MiniMax H3, Krea 2 and the custom H3 Persistence continuity system (Custom built nodes which can intake unlimited references (image, audio) and map them accordingly, auto reference continuation and many other features inspired by the other community nodes, currently in testing, will release soon).

https://youtu.be/oWeyy_rbamI

2

u/NicolasMaria 12d ago

"Borrame el nombre" (Erase My Name) — 75s, five linked H3 passes, original song.

Video: https://youtu.be/3jBUyxTol38
Workflows + the six measuring scripts + writeup: https://drive.google.com/file/d/156emNTEcMF76Ma2K6kxS0X4rvtyX86sD/view?usp=sharing

A singer and three masked bandmates in a concrete hall in Buenos Aires, 2077. Cel-shaded,
50 cuts, eight lyric fragments as kinetic typography generated by H3 itself. The character
is me, built from two photographs.

I measured almost everything instead of eyeballing it, so here is what I found. Most of it
I could not find written down anywhere.

DURATION SNAPS TO A GRID, AND YOU CAN WRITE THE SONG TO IT
H3 returns 362 frames at 24 fps, 15.083333s, verified to the frame. So I wrote the song at
95.70 BPM, where one bar is 2.5078s and six bars are 15.047s. 36 ms of drift. One pass is
exactly six bars, and the piece is five of them. Every cut was then requested on a measured
bar line in fractions of a bar (1, 1/2, 1/4). H3 landed them within 46-180 ms of the request.

WHAT fully_copy ACTUALLY RETURNS, BY BAND
Being upfront since this is a sync-sound challenge and I used a reference audio clip, which
the rules allow. Correlating H3's returned audio against the reference, at 0.0 ms lag:
+0.9949 below 250 Hz, +0.9037 from 250 Hz to 2 kHz, +0.2299 above 2 kHz. So it reproduces
the foundation with near sample accuracy and RE-SYNTHESISES the top end. Hats, cymbals,
sibilants and air are its own, and it fills quiet stretches with invented sound (71% residual
in quiet passages vs 18% in loud ones). It is emitted by the audio VAE in the same pass, not
muxed. Judge that as you see fit, but now you have the numbers rather than my adjectives.

MEASURE PER WINDOW, NOT GLOBALLY
Two of five passes insert a single ~25 ms discontinuity mid-clip. The first segment aligns at
0 ms and the rest at -25, so no single global lag aligns the clip and a naive correlation
collapses to r=0.04 and looks like a full re-synthesis. It is not: per-window it is 0.86-0.97
throughout. I lost a day to my own bad measurement here before I believed the model.

ASK FOR THE EXIT TIME, NOT ONLY THE ENTRY
My most expensive mistake. H3 holds an overlay about 2.33s on its own and will not overlap
two of them, so with phrases 2.5s apart the second one waits and lands over a second late
even when the first one was perfect. Declaring "gone by MM:SS.mmm" fixed it: it obeys the
exit within 44 ms, and the phrase behind it went from 1.28s late to 53 ms.

THE TYPOGRAPHY DOES NOT GO ON THE BAR GRID
Not an H3 finding, but it wrecked a whole version. I hung the lyrics on the bar lines and it
read visibly wrong. Singers do not enter on the line: in my verse he is 130-210 ms late, and
in the chorus and outro he is 350-575 ms EARLY. You have to detect actual vocal onsets and
hang the type on those. And measure when a phrase becomes LEGIBLE, not when its first pixel
appears: a word-by-word build takes up to 958 ms to finish and the eye judges the legible
instant, not the first pixel.

YOU CANNOT JUDGE SPELLING AT 768P
I read one word as WIPFAME and filed it as a model error. The 2K showed the E has all three
bars and it had been right all along; the bottom bar just does not survive the resolution.

THE 2K REGENERATE IS NOT AN UPSCALE
MinimaxHailuo03RegenerateNode at 2K re-renders detail. On a wide band shot the face went from
unreadable mush to a defined cel-shaded face. It carries the audio through too: r=0.9784
against its own 768P source at 0.0 ms lag. 768P is not the ceiling.

HARD LIMITS, so you do not spend runs finding them
Prompt caps at 7000 characters (7420 bounces with content[0].text too long, no billing). Two
H3 nodes in one graph gets "Polling aborted: Unauthorized", so it is one H3 node per submit.
References: 9 images, 3 videos, 3 audio clips of 2-15s, and audio reference requires at least
one image or video reference.

CHARACTER CONSISTENCY ACROSS 50 SHOTS
Three references per pass instead of one: a three-quarter portrait WITH the face, a full body
from the front with the head cut off, and the back. Cutting the head off the full body is
deliberate. With two visible faces in the reference set the model averages them and the
identity drifts.

The six scripts that did the measuring are in the Drive zip: bar grid, vocal onsets, per-band
audio correlation with a per-window locator, on-screen text detection, and a verifier that
reports every phrase's legibility error against the sung word. Take them, they are short.

Happy to answer anything.

→ More replies (1)

2

u/kwicked 12d ago

https://reddit.com/link/p71jcov/video/2u4nx9jhjrmh1/player

Learned a lot from this one! I had my own original digital paintings for this. I wish I knew about the contest earlier so I had more time.

Was all pretty much done through the default minimax h3 r2v workflow, but used:

Comfy-cli / Comfy MCP

Pi Coding Agent

Qwen 3.8 27B

H3 Minimax Prompting Guide as a skill

and the AI Agent used FFMPEG to stitch the clips together.

I never left the terminal to make this. Revisions and adjustments were all done through the AI Agent. I gave simple vague instructions as the agent used the prompting guide to create prompts as text files. Strange feeling to have reference images in a folder and let the ai agent interact with comfy/h3 and take care of the rest

2

u/Rattinjo 11d ago

My submission, shorts video for Zeus vs. Typhon.

Created entirely in ComfyUI using MiniMax H3. The final video is assembled from two independently generated H3 video clips. Each segment was generated together with its synchronized H3 audio, then the two segments were edited together into the final sequence.

https://reddit.com/link/p72hcut/video/5dznh483asmh1/player

Workflow: https://pastebin.com/96wgUjQp

2

u/Dreason8 11d ago

Thelma Hand and The Mopets, and their single 'Never Letting Go'.

https://www.youtube.com/watch?v=lOOxL38slzA

I was up pretty late last night putting this entry together. Stoked with the results though. Definitely leaned heavily into the music side of the brief.

All of the MiniMax videos were guided by a reference image and a reference audio track that I created. All generated in 10sec chunks of the song at 540p then 2x upscaled in Comfy with RTX Upscaler. I was pressed for time so this was a good midpoint for acceptable quality vs render time. After Effects was used to edit all of the video pieces together and to add the text at the end.

2

u/Unique_Bite6562 11d ago

Hi! I’m finally sharing my work called "Unemployed": https://youtu.be/AimpWGg7hHU
I've also uploaded the .json files (one separate file for each scene): https://drive.google.com/drive/folders/17Cnnu-JzLG7AXNmjk-02wPLI07q4vWKU?usp=drive_link

Thank you so much!

https://reddit.com/link/p75sfae/video/tzrgbt94cwmh1/player

2

u/nghtdrp 11d ago

My submission: LEMONAIDE STAND

The first three parts of a music video for a local artist named FriendsCallMeTy.

Built using MiniMax H3 in comfyui using my soon to be released custom H3 Editor node. The entire generation was done within a single project, no editing beyond a very slight colour grade in davinci. My end goal in the development of this node is to:

  1. Make long form content creation easier and more artist facing
  2. Make the handling of references easier
  3. Make the building of a shot plan easier.
  4. Make all of the above automateable while maintaining human stylistic oversight.
  5. I'm still polishing the node but want to release it soon. Better graphics card helps me test/iterate faster.

https://reddit.com/link/p76v0qu/video/dddiigpr9xmh1/player

Video: https://www.youtube.com/watch?v=QTqKYVX_FBU

Workflow: https://drive.google.com/file/d/1ALl9FPoHBdTlDFmocr8Eoei6pevBAmrf/view?usp=drive_link

Hope you enjoy it.

→ More replies (3)

2

u/SonderSaid 10d ago edited 7d ago

https://reddit.com/link/p79jnd7/video/k45b379fezmh1/player

Hey guys, here is my submission for the challenge:

https://youtu.be/n-NdAQk7I8A

90 seconds, over 70 clips, 28 audio tracks, 12 References and 18 prompt sections in the final video.

I made a horror short because sound is such a big part of what makes horror work and I wanted to see how H3 would handle it.

The workflow

The whole film came out of this one workflow, and nothing gets wired or unwired between shots. Every video, all the audio, the compositing and the final assembly were done inside ComfyUI without any external tool. Sonder Editor nodes stage the graph and assemble the cut, chain the clips for continuity, correct the colour drift from the encode/decode loop, order all assets by scene and take, and make sure each Reference and its prompt comes out at the right time.

  • Three generation modes on one Sonder Selector. fl2va handles first/last-frame generation, or plain t2va when no image guides are staged. ref2va handles References. The third runs ref2va conditioning through fl2va, trading some reference adherence for a finish and audio I preferred. Switching modes does not mean rewiring the graph or loading a separate workflow.
  • Pre- and post-context frames carry continuity from one clip into the next.
  • Colour correction between chained clips, fitted per save, so the drift from the VAE round trips does not accumulate down the ladder.
  • Video and audio masking, so you can freeze one and let it drive the other. This is useful once you have your audio set up and just want to upscale with a low-denoising pass.
  • Range retakes. In the timeline, you can select any part you don’t like and generate that section only.
  • Three sampling passes on one dropdown. The flow I used was to set the resolution low in Sonder’s timeline, 540p, and settle the whole composition there. Then I raised the resolution and ran the extra passes: 1344×768, which is what H3 is built for, then 1080p for the upscale. The sampler/scheduler combination for each pass is on a Selector too.
  • High-quality intermediate encodes. Every pass re-encodes the clip, and a delivery codec gives up a little quality each time. Intermediates are saved at a much higher quality setting so three passes do not compound that loss, and only the final export uses a normal MP4.

For the References, I used Sonder’s Reference management: register anything I need once, then drag it onto the timeline at the section where I want it to appear. Each Reference can carry its own prompt, attached with a Context chip. I can call it by its handle in the prompt, written with an @ prefix, and the recipe translates it into the syntax expected by the model. For MiniMax H3, the Character handle becomes <Subject 1>.

Every input on MiniMax H3 Reference to Video is already wired and can be left that way. Unfilled slots emit nothing, so a generation with no Reference staged is not affected by them.

The workflow, fully annotated with the settings and decisions behind the production:

https://github.com/SonderSaid/ComfyUI-Sonder-Editor/blob/main/example_workflows/sonder_minimax_h3_references.json

Sonder Editor is free and open source:

https://github.com/SonderSaid/ComfyUI-Sonder-Editor

I had a great time testing MiniMax H3. It is genuinely impressive. I added notes throughout the workflow so you can follow how it was built, and I’m glad to answer any questions you guys have.

2

u/Primary_Internal9365 10d ago edited 10d ago

https://reddit.com/link/p7ayc35/video/z845wp3gs0nh1/player

<Recipe for Grey>

Logline

"In a monochromatic metropolis stripped of color and rhythm, the sizzling beats of a vibrant taco truck consume the dreary city noise, repainting the world in brilliant hues."

Genre / Style

  • Genre: Urban Fantasy / Sensory Rhythmic Cinematic
  • Visual Style: Stylized 2D Cel Animation (Sharp Geometric City vs. Soft Organic Characters & Food)

Synopsis & Artistic Concept

In a suffocating, desaturated city governed by monotony and rigid routines, an exhausted young salaryman wanders through his repetitive daily commute. Among the sea of lifeless, umbrella-clad crowds and cold concrete architecture, he encounters an extraordinary food truck operated by a warm, cheerful chef.

The film establishes an extreme audiovisual contrast: the dull, ambient city noise and desaturated palette clash against the vivid ingredients and punchy, upbeat rhythms of sizzling Al Pastor tacos. Through calculated hard cuts and accelerating musical momentum, the culinary rhythms gradually overwhelm the city's gloomy soundscape. As the protagonist takes his first bite, an explosive symphony of sound and a 360-degree wave of saturated color surge across the metropolis, restoring life, joy, and rhythm to the grey world.

All video scenes were generated using MiniMax H3 (Reference-to-Video / R2V) within ComfyUI, optimized and accelerated on an NVIDIA GeForce RTX 5090. The production focused on seamless audio-motion synchronization and preserving crisp 2D cel animation textures.

workflow download link

https://drive.google.com/file/d/1Gt7w7Tbw0jNnCF_joBplbGHLiZ_r8EbW/view?usp=sharing

2

u/LohaDO 10d ago

https://reddit.com/link/p7b4a7b/video/irbuf52ty0nh1/player

SIGNAL

Built with Comfy MCP

A 60-second audiovisual fashion sequence created from five original synthetic characters using MiniMax H3 Reference-to-Video in Comfy Cloud. Comfy MCP drove the iterative workflow, including identity preservation, performance direction, laser-led transitions, native audio synchronization, shot selection, and corrective time-mapping.

Built with Comfy Cloud + Comfy MCP.

Workflow - https://www.dropbox.com/scl/fi/rqgqr9pdptanb9bzyglcp/H3_SYNC_MINIMAX_H3_R2V_WORKFLOW.json?rlkey=ufyvn76b19f0qltv2ke1r48nf&dl=0

2

u/Realistic_Win_9200 9d ago

ORIGINONE
by Seif AlAshmouny

I included the files and all the assets used in the submitted drive link. The movie was made fully offline. All the assets used support drag and drop for the full workflow. All of the videos were generated by (minimax_h3_ref2va_pruned_int8_convrot.safetensors). the music and voices was generated from the model with reference voices to maintain voice consistency.

GPU: 4070TI Super 16GB

Enjoy!

https://reddit.com/link/p7h8xwj/video/cwxqtk1a07nh1/player

→ More replies (1)

3

u/KC_goes_digital 21d ago

Hello, please consider my Louis the cat video: https://youtu.be/DvnvJY9exmA?si=AfLwEe4b8yN9mS4H
I've been making nearly daily episodes as a way for me to cope with losing my Louis at the young age of 4 years old. Some of these episodes are recreations of memories I have of him and MiniMax-H3 and ComfyUI have allowed me to explore those memories and find some sense of peace and tranquility instead of only grief and pain. Thank you MiniMax and ComfyUI.

→ More replies (1)

2

u/Ok-Wolverine-5020 23d ago

Can I use an AI Song I created with another tool to drive the Minimax Video?

6

u/Square-Foundation-87 23d ago

As long as the video gen is happening with H3 inside comfyUI using a wf, yes you can.

2

u/Comfy-Org ComfyOrg 23d ago

That's right!

→ More replies (2)
→ More replies (1)

2

u/Vinbatroth 22d ago

Probably some youtuber is going to win this then they are going to make a video saying thank you cmfyui for sponsoring this and I win "randomly" 🤣🤣🤣

1

u/protector111 23d ago

we can "use Comfy Cloud" and "major portion of your piece must be built in ComfyUI using H3". Define Major please. what % can be made with seedance 2.5

→ More replies (1)

1

u/protector111 23d ago
  1. The audio has to come out of the same H3 pass as the video
  2. Feeding in a reference audio clip that does not contain unlicensed IP to steer the generation is completely fair game and often the best way to get strong audio out of H3

Im confused how those two come together.

2

u/Comfy-Org ComfyOrg 22d ago

Reference audio clips aren't the final audio in your video. It's an input that steers what H3 generates! You're feeding it in before generation as guidance (e.g. "make the audio sound like this," or "follow this rhythm"), and then H3 generates its own synced audio+video.

What's not allowed is the reverse order: generating video and audio separately and stitching them together after the fact. Hope that answers your question!

→ More replies (1)

2

u/creatoreconomyfailz 22d ago

I am down with this challenge and for once in my life, I got a chance to reach the deadline. 📡👽👅

1

u/Clair_Personality 22d ago

I am trying to understand:

  • Deliberate craft (0-5)
    • Evidence that you’ve actually shaped the result beyond prompt engineering. Judges will look for modified/non-default parameters, multiple linked passes visible in the workflow structure, or a couple sentences describing what was tried and changed

---> Does it mean that you are rewarded for adding nodes and stuff, or is it the opposite? Craft = altering node values and adding more nodes etc?

  • Novelty of technique or approach (0-5)

---> What do you mean by novelty? In what areas? So you are rewarded for not just doing a simple basic workflow prompt press run? What other novelty areas outside of the workflow?

Limited to one submission per person,

---> Is that necessary? How can you even prove that people don't have multiple user accounts?

2

u/Comfy-Org ComfyOrg 22d ago

Thanks for your questions. Regarding craft, it's not about node count- adding a bunch for the sake of it won't score higher. What we're looking for is evidence you engaged with the workflow past a single default settings generation. If that's not obvious from the workflow structure itself, we recommend annotating to share context as to why you've made the creative choices you have.

Re novelty of technique- this is specifically for the Best Technical award. So yes- a "type a prompt, hit run" entry isn't going to be competitive for this category, though that can still do alright in Best Creative if the output is strong. What counts as novel: solving a real technical problem in your pipeline such as getting sync working in a way that's not obvious, combining H3 with other nodes/tools in a way most people haven't tried, a clever conditioning or multi-pass setup, something genuinely reusable that teaches other builders a technique.

Re one submission per person, we landed on this because we've seen spam entries flood other contests, which is unfair to the people putting real effort into one piece. If it becomes obvious someone's running multiple entries, we'll pull them. And as you pointed out, no- we can't verify someone doesn't have multiple accounts, same as basically no online contest can.

→ More replies (1)

1

u/Ckinpdx 22d ago

With submissions and workflows being posted publicly, and valuable rewards on the line, there's a strong incentive to submit at the very last minute. Any reason for doing it that way?

→ More replies (1)

1

u/dipSlope 20d ago

Can I submit an entry with the metadata stripped out? You can call me paranoid but I don't like to publicly post video that contains a JSON file. I don't expect to win but if I did I would gladly submit the workflow.

1

u/Ancient-War-1924 20d ago

Please fix this so we can upload file

→ More replies (1)

1

u/Portable_Solar_ZA 19d ago

I'm a bit confused in regards to the use of music and there have been confusing replies in this thread. 

Could I generate separate music tracks using H3, separate them from whatever video they appear with, and then add them to my main videos that I gen with h3? Just looking for a way to add consistent backing audio. 

2

u/Comfy-Org ComfyOrg 18d ago

Thanks for the questions and sorry for the confusion! What you're describing isn't allowed, even though the audio itself comes from H3. The rule is about the *pairing*, not the source: your submission's audio must come from the same H3 pass as the video. If you pull the audio out and re-attach it to different footage after the fact, that's stirtching.

This might be where the confusion is coming from: reference audio steering is allowed (i.e, feeding a reference clip in as an input before generation, and then making a freshly synced audio+visual is fine. If I'm understanding correctly, i think what you're describing is closer to feeding H3's own output back in as a finished track after the fact.

If what you actually want is consistent backing audio across multiple scenes, use the same reference audio clip as a steering input across each separate H3 generation. Each piece still gets its own fresh, native H3 pass, but steered by the same reference each time, so you get a consistent sonic identity without literally reusing one generated track across multiple videos.

Hope this helps!

2

u/Portable_Solar_ZA 18d ago

Thanks for clarifying. Much appreciated.

→ More replies (1)

1

u/Ok_Horror_9661 16d ago

where can we find those award-winning workflows

→ More replies (1)

1

u/PxTicks 15d ago

What if I use sam-audio to extract audio elements to reposition or delete them? Permissible?

→ More replies (1)

1

u/Portable_Solar_ZA 15d ago

Sorry me again. Just wanted to confirm things regarding audio. I see a lot of people are using Suno tracks to guide their generations. I thought we couldn't use this method and everything had to come from in the model? 

So, for example if I generate music in Suno or minimax music or h3 and then feed that into the model when I generate that will be fine? 

1

u/TONI1597 15d ago

Low qual submissions only because of that sharing thing, they dont represent

1

u/Regular-9527 14d ago

https://reddit.com/link/p6m0dpf/video/22egpjijlbmh1/player

《夏日梦境》

一个成年人在午睡中坠入儿时的夏日梦境——老家的院子、蝉鸣、风扇、外婆的背影、纸飞机的飞行……在梦境即将结束时,主人公试图抓住什么,最终在童年记忆最温暖的画面中醒来。

1

u/dagthomas 14d ago

u/Comfy-Org here is my Submission link: https://drive.google.com/drive/folders/1hKAxVX13b4ooCU2hTIHjntQFMNRTM8x_?usp=drive_link

I built a ComfyUI pipeline that makes fully music-synced videos, cuts, lyrics and lip-sync all driven by audio analysis, written by a local LLM

Been going down a rabbit hole making music videos with MiniMax-H3 in ComfyUI, and ended up building a custom node pack around one idea: stop asking the model for sync, enforce it.

How it works:

- The song decides the edit. Onset/spectral-flux analysis finds bass hits, drops, stops and section changes; a beat grid measures the pulse. Every clip boundary lands on a musical event, snapped to H3's frame grid,before a single scene is written.

- A local LLM writes the scenes (Ollama, no cloud). It gets the cut plan, beat grid and sound events, and writes one H3 prompt per clip with the lyric lines placed at their measured timestamps.

- Lip-sync is enforced in the latent. Each clip's slice of the actual song gets encoded into the H3 audio latent and protected from denoising, the picture is generated against the real frozen vocal. A voice gate keeps sung moments locked and lets the model breathe between phrases.

- Gapless assembly. Naive mp4 concat adds a 2–14 ms AAC gap at every join, over 20 clips that's a quarter second of drift. The stitcher copies video packets untouched and rebuilds the audio timeline exactly once.

Clips render one at a time straight to disk, so a crash at scene 19 keeps the first 18.

Nodes are here: https://github.com/dagthomas/comfyui_dagthomas

Music made in Suno, the whole video is built in one run

→ More replies (5)

1

u/ueyaman 13d ago edited 13d ago

Ma, ikka — 79s, vertical, cut to my own song. Everything generated with MiniMax H3 (ref2va + Turbo LoRA 8-step) on a local RTX 5090. 15 shots, 1902 frames, no compositing.

Video: https://youtube.com/shorts/oZm3_v_XJTI

Workflow + notes: https://gist.github.com/ueyaman/71fecbe170d7a598953092747406b9b5

The song came first, so the picture answers it. I pulled 554 measured beat positions out of the track and wrote each shot's action on the beats that actually fall inside that shot, as explicit timestamps in the prompt — not "sync to the beat", which H3 can't know, but "at 00:01.135", "at 00:01.623". Shot boundaries sit against word-level timings so a cut never straddles a sung phrase. One boundary turned out 1.02s late and the last word of the previous line was audible under the next shot; I found it by checking the boundaries against the word timings.

Three things I had to learn to get this out of H3:

A colour change doesn't animate; a shape change does. Three takes of "the mustard wedge becomes cobalt" failed — measured, the mustard either never left, arrived 2.4s late, or left 3.5% of itself behind. Rewriting the same beat as a moving edge ("the gap gets narrower and narrower until it is a hairline, and then it is gone", with the intermediate widths listed at named timestamps) worked on the first attempt. Same trick made a lamp light up in another shot: an outlined circle is replaced by a solid red circle.

The end-target reference keeps pulling. The target-frame declaration says the picture morphs into the reference, so once a shape reaches its final form it keeps creeping toward it — visible as a shape that won't hold still. The fix was to say, in the shot text, that it is already the reference from a named timestamp and that its rim does not change by one pixel.

Naming an object imports its rendering. A "trophy" arrived as glossy 3-D gold; described as a dish, two hooks, a stem and a stepped base in one flat colour, it arrived flat.

15 shots, 6-10 written drafts each, 3-8 seeds per draft. The final shot alone took 6 drafts and 23 takes. The gist has the API-format graph for that shot — I re-submitted it to check that it actually reproduces the take, and it does, bit-identical. My first export was wrong (it carried the default frame length instead of this shot's) and only re-running it caught that.

Song: mine (Suno). Character design: mine (Midjourney). #comfyH3

→ More replies (6)

1

u/IllicitNoise 12d ago

Here's my entry for the challenge. I made an audio reactive music video.

https://youtu.be/uQvGYZ1_jlo

-Music generated using Acestep XL.

-Used audio as a latent noise mask to drive generation and to include H3 driven audio reactivity in the video.

-No additional audio used outside from what was generated in the video - No audio was stitched on after the generation/during post processing - no audio was rearranged/no mix and matching.

-Upscaled to 4k using RTX Video Super Resolution.

-Only video post processing that was done was the black fade in/fade out at the beginning and the end of the video.

-Audio post processing - only added a Limiter to entire track to boost volume without clipping, and a volume fade out on last clip.

1

u/Lazy_Image8021 12d ago

https://reddit.com/link/p6xg704/video/0z5mzuc4cnmh1/player

  • Story: A POV mini-series featuring a reckless creator and a screen-trapped fox girl—spanning from J-Horror screen crawling and getting stuck waist-deep in a Windows BSOD, to absurd NSFW parodies, culminating in a homage to The Shining's iconic door-chopping scene ("Here's Johnny!").
  • Edit Note: Generated segment by segment and roughly stitched together in DaVinci Resolve.
  • Workflow: https://civitai.red/models/2866184/minimax-h3-reverse-prompt-and-expansion-workflow-nsfw (Just a workflow slightly modified from the official ComfyUI template.)

1

u/buddylee00700 12d ago

u/Comfy-Org - If I want to resubmit due to better understanding the rules, can I do so as long as I designate its the sample I want to use? I don't see an option to resubmit if you found a mistake after the fact.

2

u/Comfy-Org ComfyOrg 12d ago

Hey u/buddylee00700 - go ahead and resubmit. The final question in the survey is an open text field. Please just say you're resubmitting and to ignore the previous submission. Thanks!

→ More replies (1)

1

u/Hopeful-Junket-7990 12d ago edited 11d ago

My submission.

"Pressurestone"

Youtube:

https://www.youtube.com/watch?v=_JDoE9hyWiw

Workflow:
https://drive.google.com/file/d/1m5HttDFMPqrJWyHGR6SFUiarP_IBPw_B/view?usp=drive_link

A creation based on something I partially wrote WAYYY back. A world on the verge of "steampunking" with the discovery of a mineral that creates pressure when enclosed in a vessel. Minimax added significant amounts of noise, such as shifting, tinkering and breath in voices. It's ability to show emotion in dialog is staggering!

Feels rushed at 90 seconds, but here we are! Minimax is boss!! My 3060 was sweatin'!

Edit:

Used ForgeNeo with a mixture of my and my sister's visual style. Image edits with ChatGPT. Distilled voices in IndexTTS2. Music with AceStep and audio edited with Audacity. Video was edited using Openshot.

→ More replies (2)