old but relevant, it's like a production system. We can continue to tune and squeeze 1mp pictures out of it in high quality on relevant localized hardware without breaking a sweat. 1-3 seconds on a m3 system, and then cascade of tooling below stays the same.
SV_BubbleTime 23 hours ago [-]
A1111 also had its time. Comfy takes a couple hours to learn, and then it’s vastly superior. Like, if you are already on Mac there is no reason to also kick yourself in the face trying to go back to 1111.
qclibre22 23 hours ago [-]
There's also sdcpp ( https://github.com/leejet/stable-diffusion.cpp ) . It is for people that know command line shells like bash or Powershell. SDCPP is like llama.cpp but it is for using text to image AI models. Not everyone likes the node/graph programming interface in Comfy or a Webpage interface. Some people need these GUI/web interfaces and get lost and frustrated in shell scripting.
dmikeyanderson 22 hours ago [-]
Totally and sdcpp and drawthings were inspiration for this work. It's not just about the UI, there's a whole stack of plugins, APIs etc. It would be a bigger shift, than just use other UI. Composition etc.
Plus this was fun to learn where the attention and time spent is eating up. I plan to git into the UNet, and then do some Multi-Model Memory Attention tuning as well, for Refining and Inpaint Swaps for the 8Gb Macs.
vunderba 22 hours ago [-]
A1111 (or rather Forge Neo at this point) can still be pretty useful (for a Windows user) if somebody just want a batteries-included approach because it makes it trivial to move back and forth between masking, manual inpainting, tacking on a high‑res fix and/or upscale, and swapping things out on the fly quickly.
Caveat: No idea how well A1111 and its variants function on Mac though.
These days I mostly use ComfyUI because after building out a node‑based workflow, it’s easy to export as a JSON file with websocket outputs that I can integrate into my programs.
I really miss the prompting that A1111 had in concert with the grid. It was so easy to do a grids changing variables per axis. e.g. prompt on vertical and model on horizontal axis. is there a comfy node that does this? i've looked and havent found one.
echelon 18 hours ago [-]
Comfy's UX sucks for most people.
Even today you have to plug nodes if you want to change the reference image count. That's unacceptable.
I've seen more people vibe coding their own UX than put up with Comfy. And these are not engineering minded people.
IronWolve 23 hours ago [-]
Automatic1111 wasnt being updated, people moved to Forge Neo, ComfyUI and Maestro.
vunderba 22 hours ago [-]
Invoke AI is also really good if you prefer working in a canvas-style environment, in a more traditional graphical way. It easily has the nicest workflows for outpainting/regional/inpainting by far.
What a crazy repo. Some cyber archeologists need to do a case analysis of the entire diffusion model ecosystem circa the original Stable Diffusion up to now.
jurgenburgen 24 hours ago [-]
Trigger warning: AI content.
trencedamp 23 hours ago [-]
If anyone is triggered by AI content they surely avoid HN
selectodude 22 hours ago [-]
I love AI content. I hate content written by AI.
It's not serious. It shows that the author doesn't care about showing what he's up to.
He cares about making content to post on a blog.
And that's disrespectful to the reader.
dmikeyanderson 22 hours ago [-]
I'm not an editorialist, I spend my time on the software side of things, but want a detailed technical journal of what happened to achieve the gains. Also why the github link is near the top of the intro which I wrote. I'm not trying to disrespect the reader, but the time of the user of the software. Appreciate your view. I've made adjustments.
jurgenburgen 21 hours ago [-]
I spend all day reading AI generated text in one form or other in my day job. When I come to unwind in HN I’m not interested of more of the same. Seeing “load-bearing”, “landing things”, “it’s not x, it’s y” makes me think of work. I honestly would rather read human text with typos and warts.
BoxOfRain 19 hours ago [-]
Load bearing is a funny one because before LLMs the only time I heard it was in the phrase 'load-bearing bug'.
trencedamp 6 hours ago [-]
It got a lot of use in comedy screenwriting. How many shows have you seen where someone says "don't touch that, that's a load-bearing [object that should never be load bearing]"
The fort episode of community springs to mind but there are plenty others
trencedamp 6 hours ago [-]
And my point is you'll find both on HN
smallerize 23 hours ago [-]
The speedup includes running LoRAs, right? Which ones are you using?
dmikeyanderson 22 hours ago [-]
Yes it includes LoRA composition as well, a bunch from the internet, and they work just fine!
yogorenapan 22 hours ago [-]
A1111 is such a throwback. It's been ages since that & AI image generation was cool. Back before the massive amount of slop and when things were just a fun experiment
Plus this was fun to learn where the attention and time spent is eating up. I plan to git into the UNet, and then do some Multi-Model Memory Attention tuning as well, for Refining and Inpaint Swaps for the 8Gb Macs.
Caveat: No idea how well A1111 and its variants function on Mac though.
These days I mostly use ComfyUI because after building out a node‑based workflow, it’s easy to export as a JSON file with websocket outputs that I can integrate into my programs.
https://github.com/Haoming02/sd-webui-forge-classic/tree/neo
Even today you have to plug nodes if you want to change the reference image count. That's unacceptable.
I've seen more people vibe coding their own UX than put up with Comfy. And these are not engineering minded people.
https://github.com/invoke-ai/invokeai
It's not serious. It shows that the author doesn't care about showing what he's up to.
He cares about making content to post on a blog.
And that's disrespectful to the reader.
The fort episode of community springs to mind but there are plenty others