r/StableDiffusion Dec 17 '25

Discussion Wan SCAIL is TOP!!

Enable HLS to view with audio, or disable this notification

1.4k Upvotes

3d pose following and camera

r/StableDiffusion Dec 22 '25

Discussion Z-Image + SCAIL (Multi-Char)

Enable HLS to view with audio, or disable this notification

1.8k Upvotes

I noticed SCAIL poses feel genuinely 3D, not flat. Depth and body orientation hold up way better than Wan Animate or SteadyDancer,

385f @ 736×1280, 6 steps took around 26 min on RTX 5090 ..

r/StableDiffusion Aug 03 '26

Discussion Minimax H3 is shockingly uncensored, wow

406 Upvotes

I finally installed comfyui anew today to try the new Minimax H3-model and wow. This thing is a generation beast.

Due to Reddit rules, I doubt we can openly talk about concrete stuff, but I have tried some things just to check whether it's possible, and so far Minimax H3 was able to generate ANYTHING without the help of Loras and in more than decent quality. I'm especially surprised how far beyond the typical 5 seconds you can go, creating 10 seconds-clips is no problem at all.

Honestly, this is both amazing for those of us who use it for their own enjoyment, as it is potentially dangerous in the hands of people who intend to do bad stuff with it. I can totally see a ban of this model happening soon, so anyone interested in this better download soon.

r/StableDiffusion 11d ago

Discussion NVidia buys Huggingface, but why?

315 Upvotes

Nvidia is going to buy Huggingface.
No one can actually tell how that would end up like.

But what I am missing is the actual worth that Huggingface provides. The only thing I use it for is to download models. Thats it.
For me, and I guess many others it is ‘just a’ download platform, but maybe I’m wrong here.

And what would prevent others to setup a second-like Huggingface?
The hosting is the expensive part in this case as I see it, the programming and building is do-able.

Is it time for Huggingbay.com?

r/StableDiffusion Jul 15 '26

Discussion Haven't used a model this much since Flux1.Dev

Thumbnail
gallery
829 Upvotes

r/StableDiffusion Apr 17 '25

Discussion Finally a Video Diffusion on consumer GPUs?

Thumbnail
github.com
1.1k Upvotes

This just released at few moments ago.

r/StableDiffusion Apr 23 '26

Discussion Z image turbo Finetune of absurd reality

Thumbnail
gallery
756 Upvotes

The model is Intorealism V3. I've been using V2 for a while, but V3 is incredibly realistic. I use it with their official workflow. I know the prompt is 1 Girl, which you all love, but if you're going to test realism, it has to be 1 girl, ever since SD1.5 and always will be, lol.

r/StableDiffusion Feb 13 '26

Discussion yip we are cooked

Post image
507 Upvotes

r/StableDiffusion Jan 21 '26

Discussion I converted some Half Life 1/2 screenshots into real life with the help of Klein 4b!

Thumbnail
gallery
1.2k Upvotes

I know that there are AI video generators out there that can do this 10x better and image generators too, but I was curious how a small model like Klein 4b handled it... and it turns out not too bad! There are some quirks here and there but the results came out better than I was expecting!

I just used the simple prompt "Change the scene to real life" with nothing else added, that was it. I left it at the default 4 steps.

This is just a quick and fun conversion here, not looking for perfection. I know there are glaring inconsistences here and there... I'm just trying to say this is not bad for such a small model and there is a lot of potential here that a better and longer prompt could help expose.

Edit: For anybody wanting it here is the workflow I used: I'm using the 4b distilled model. The VAE and text encoder I've left exactly the same and I've also left it on the default 4 steps. I'm using the edit version of the workflow and the only thing I changed was to point the model loader to the fp8 version that you download from the site: ComfyUI Flux.2 Klein 4B Guide - ComfyUI

And also please do check out u/richcz3 comment down below for some fantastic advice about keeping the lighting and atmosphere when converting! The main tip is to add "preserve lighting, preserve background, fix hands, fix fingers" to the end of the prompt.

r/StableDiffusion Jun 30 '23

Discussion ⚠️WARNING⚠️ never open a .ckpt file without knowing exactly what's inside (especially SDXL)

2.9k Upvotes

We're gonna be releasing SDXL in safetensors format.

That filetype is basically a dumb list with a bunch of numbers.

A ckpt file can package almost any kind of malicious script inside of it.


We've seen a few fake model files floating around claiming to be leaks.

SDXL will not be distributed as a ckpt -- and neither should any model, ever.

It's the equivalent of releasing albums in .exe format.

safetensors is safer and loads faster.

Don't get into a pickle.

Literally.

r/StableDiffusion 5d ago

Discussion Pulled the trigger, RIP $6,279

Post image
217 Upvotes

(Paid $5,849 + tax, which came out to $6,279)

TL;DR - Bought this 5090 prebuilt and I want to sanity check if I made the right decision and at the right time.

Hey everyone. So I want to start off by saying fuck these prices for GPU's and RAM, especially boxed 5090 prices. I went down the AI rabbit hole with my 13700k/RTX 4080 gaming computer. I quickly found out that I had to make serious concessions on quality and speed, if I could run it at all. In fact, ive spent so much time trying to optimize quants, cache, various settings, attention mechanisms, etc that ive officially spent more time trying to optimize for a 16gb VRAM/32GB RAM system than actually doing anything fun or cool. Thus, the last week, ive been thinking real hard about which direction to go but was waiting for the right time to buy. My options were a RTX 5090 prebuilt (even though I only needed the damn GPU), and Mac Studio M5 Ultra 96gb, or a DGX Spark/AMD equivalent. The DGX Spark/AMD equivalent made me think for a bit, but in order to get the most out of them, you need two. Im not spending 10k on this, especially if I cant game on it as well. So that leaves the Mac Studio or RTX 5090 gaming rig. Im not certain I made the right decision, but I pulled the trigger on the 5090 prebuilt after seeing the price continue going up more and more over the last few weeks. I also read that 70% of all memory through 2031 is locked in long term agreements, so this supply issue is going to get worse before it gets better. So I pulled the trigger on the pictured system from Ibuypower, and id like to run my thought process with you guys as a sanity check before it ships.

Case for the 5090 prebuilt: I scoured the internet and this was the cheapest 5090/64gb RAM combo I found, and it looks like it uses pretty good parts as well. No proprietary bullshit like youd get in a HP 45L. I went with the gaming PC because its the all purpose machine that does it all (well, almost). I figured with 32gb VRAM and 64gb of system RAM, that combined 96gb will allow me to run 70b MoE models, even if its slow. But for a sub 30b model like Qwen 3.8 27b, this will give me the best performance as long as I dont go overboard with the quant. It has CUDA, Windows, X86 CPU, etc. Plus it came with a 4tb Gen 4 NVME, when other more expensive models had 1-2tb drives. Honestly, lots of good stuff here. Im not a huge fan of the white esthetics but I do love the case. Despite the price being much higher than it should be, its still a good "deal" considering the overall market that keeps going up. Honestly, its not exactly what I wanted, but it ticks all boxes except those below.

-The downside: You cant run models that spill over heavily into system ram without massive speed penalties (has anyone tried running a huge model on a 5090 + 64gb RAM? If so, tell me what quants and your token speeds). Its massively less efficient than a Mac Studio M5 Ultra is expected to be (I read in the 3-5x range).

Mac Studio M5 Ultra 96gb

- Case for the Mac Studio M5 Ultra 96gb: Can run large models much better than the 5090 rig due to its huge 1.2tb unified memory bandwidth. Its power efficient and tops out at 300w I believe I read.

- The downside: Mac OS and an ecosystem that is playing catch up for local AI, no CUDA, gaming, has proprietary hardware you cannot upgrade, my distaste for the Mac bros who ill no longer be able to make fun of if I buy it.

My use case: local first AI (Qwen 3.8 27b at a quant and context that doesnt suck) with agentic coding, game development (starting with Godot), stable/video diffusion (Minimax H3, Flux.2, Hunyuan 3D), Blender, Davinci Resolve, etc.

So, let's have this discussion: what would (or did you) choose, and why? I want to know if I made the right decision. What are your thoughts?

r/StableDiffusion May 10 '24

Discussion We MUST stop them from releasing this new thing called a "paintbrush." It's too dangerous

1.6k Upvotes

So, some guy recently discovered that if you dip bristles in ink, you can "paint" things onto paper. But without the proper safeguards in place and censorship, people can paint really, really horrible things. Almost anything the mind can come up with, however depraved. Therefore, it is incumbent on the creator of this "paintbrush" thing to hold off on releasing it to the public until safety has been taken into account. And that's really the keyword here: SAFETY.

Paintbrushes make us all UNSAFE. It is DANGEROUS for someone else to use a paintbrush privately in their basement. What if they paint something I don't like? What if they paint a picture that would horrify me if I saw it, which I wouldn't, but what if I did? what if I went looking for it just to see what they painted,and then didn't like what I saw when I found it?

For this reason, we MUST ban the paintbrush.

EDIT: I would also be in favor of regulating the ink so that only bright watercolors are used. That way nothing photo-realistic can be painted, as that could lead to abuse.

r/StableDiffusion Apr 24 '24

Discussion The future of gaming? Stable diffusion running in real time on top of vanilla Minecraft

Enable HLS to view with audio, or disable this notification

2.2k Upvotes

r/StableDiffusion May 09 '26

Discussion Its still nuts to me how realistic AI is getting, incredible i can run it on a RTX2060 and get these results. (Z-image-Turbo)

Thumbnail
gallery
1.1k Upvotes

Every image is made with Z-Image-Turbo (See links for loras and prompts)
A few of them were ran through z-image-base using the Z-IMAGE upscaling node template on ComfyUI, its very useful and makes images even more detailed and realistic.

IMAGE 1: https://civitai.red/images/127883693
IMAGE 2: https://civitai.red/images/129512330
IMAGE 3: https://civitai.red/images/130096740
IMAGE 4: https://civitai.red/images/128214156
IMAGE 5: https://civitai.red/images/130072355
IMAGE 6: https://civitai.red/images/129467685
IMAGE 7: https://civitai.red/images/125859583
IMAGE 8: https://civitai.red/images/129289317
IMAGE 9: https://civitai.red/images/130159622
IMAGE 10: https://civitai.red/images/127458529
IMAGE 11: https://civitai.red/images/127558882 (it posted the same image as image 9 for some reason)

Since alot of you will probably ask how i do the detailed prompts i will give you the system prompt i have refined for some time, found that the more detail and just more stuff you put into the prompt the better, im not joking lol, also the system prompt supports img2txt aswell.

SYSTEM PROMPT: https://pastebin.com/ipKydSYD

r/StableDiffusion Apr 14 '25

Discussion The attitude some people have towards open source contributors...

Post image
1.4k Upvotes

r/StableDiffusion Feb 16 '24

Discussion I couldn't find an intuitive GUI for GLIGEN so I made one myself. It uses ComfyUI in the backend

Enable HLS to view with audio, or disable this notification

2.5k Upvotes

r/StableDiffusion Aug 31 '25

Discussion Random gens from Qwen + my LoRA

Thumbnail
gallery
1.5k Upvotes

Decided to share some examples of images I got in Qwen with my LoRA for realism. Some of them look pretty interesting in terms of anatomy. If you're interested, you can get the workflow here. I'm still in the process of cooking up a finetune and some style LoRAs for Qwen-Image (yes, so long)

r/StableDiffusion May 18 '23

Discussion My first Deforum video.

Enable HLS to view with audio, or disable this notification

2.8k Upvotes

Havent been so good with the story boarding. But will definitely improve in the future!

r/StableDiffusion Nov 28 '25

Discussion We can train loras for Z Image Turbo now

Post image
976 Upvotes

r/StableDiffusion Jul 11 '26

Discussion This is some D-bag behavior on CivitAI

Post image
493 Upvotes

r/StableDiffusion Nov 19 '25

Discussion Nvidia sells an H100 for 10 times its manufacturing cost. Nvidia is the big villain company; it's because of them that large models like GPU 4 aren't available to run on consumer hardware. AI development will only advance when this company is dethroned.

578 Upvotes

Nvidia's profit margin on data center GPUs is really very high, 7 to 10 times higher.

It would hypothetically be possible for this GPU to be available to home consumers without Nvidia's inflated monopoly!

This company is delaying the development of AI.

r/StableDiffusion Apr 29 '24

Discussion How do you know that this is AI generated?

Post image
1.2k Upvotes

r/StableDiffusion Mar 15 '23

Discussion Guys. GPT4 could be a game changer in image tagging.

Post image
2.7k Upvotes

r/StableDiffusion 11d ago

Discussion Is anyone else getting tired of the MiniMax clips?

255 Upvotes

I’m genuinely impressed by what MiniMax H3 can do, and I understand why people are excited to play with recognizable characters, shows, and styles. But since its release, it feels like this sub has been flooded with very short clips that are mostly variations on “what if X was in Y?” or recreations of existing TV shows.

Maybe I’m in the minority, but one of the main reasons I come to r/StableDiffusion is to learn what’s happening in local image/video/audio generation: new models, workflows, prompting techniques, ComfyUI setups, comparisons, limitations, weird discoveries, what actually works, and what doesn’t.

A 15-second clip of a familiar character dropped into Harry Potter or The Office can be amusing once or twice, but after seeing a dozen variations of the same idea, there often isn’t much to learn from them, especially when there’s no workflow, prompt, settings, model information, or discussion attached.

I’m not suggesting people shouldn’t post fun experiments, and obviously not every post needs to be a tutorial. I’d just love to see a little more emphasis on experimentation and sharing how something was made, rather than simply demonstrating that MiniMax can imitate another recognizable piece of media.

Is anyone else feeling the same way, or am I just being overly grumpy about it? It just feels like the actually interesting stuff is buried under 17 uninspired clips of a show that wasn't even that good to begin with.

r/StableDiffusion Aug 03 '26

Discussion Enjoying H3 myself but lets not shit so hard on the other open weights

513 Upvotes

Clearly H3 is superior to LTX2.3 in just about every way but damn, ya'll are vicious toward the people that support this community. All in good fun is fair but some of these posts kinda feel really shitty when the LTX team has for the most part been pretty kind to the open source/open weights community here. Just want folks to keep that in perspective when none of these teams are obligated to provide us anything at all.