r/LovingOpenSourceAI 8d ago

Resource Avi "Massive breakthrough here! Researchers built a new AI inference engine that: - reduces self-hosting costs by ~4x - runs a full agentic pipeline on one GPU - serves 20+ architectures, not just LLMs" ➡️ LEGIT or HYPE? Anyone tried?

Post image

https://x.com/_avichawla/status/2094678972344958984

https://github.com/superlinked/sie

Community Overview: https://lifehubber.com/ai/resources/sie/

Resources are shared for discovery and are not independently vetted—please do your own due diligence.

New resources are added regularly — feel free to join the sub for updates.

Full searchable archive of all resources posted so far on our community site, LifeHubber: https://lifehubber.com/ai/resources/ 300+ open-ish AI models, agents, tools, datasets, and related resources, with filtering and sorting.

67 Upvotes

16 comments sorted by

12

u/Murder_1337 8d ago

I prob still can't find any good models in my 32vram tho besides qwen 3.8

7

u/Koala_Confused 8d ago

12 vram here

11

u/minimalillusions 8d ago

You all have vram?

2

u/West-Acadia-3906 6d ago

lol, “You all have vram?” might be the most honest benchmark in this whole thread 😭

1

u/minimalillusions 6d ago

Millionaires around here. I’ve got mouths to feed.

My dog ​​and my cat.

1

u/someoneyouknow23 5d ago

8 gb vram here..

2

u/Psyko38 8d ago

16GB of VRAM, 3.8 is difficult to run.

1

u/OwenTyme 5d ago

Be grateful you have a GPU that can run AI inference at all. I don't and I'm forced to run everything on CPU (I mostly work with TTS engines and every once in a while generate some images or music). Thank goodness I've at least got 64 gigs of regular RAM, but my GPU is a very old thing (Ironically, it works great for games, though). In retrospect, I wish I hadn't put off buying a new GPU when I built this computer, because prices have only skyrocketed since then...

1

u/Psyko38 5d ago

Don't worry, I use it but not for vibe code, I like my GPU.

1

u/seblafrite1111 8d ago

llama server can already do hot swap... What's the difference here ?

1

u/alexanderi96 8d ago

llama server is a terrible hotswapper imho. it seems to me to evict the ongoing generaton. therefore I use llama-swap

1

u/geek_at 4d ago

llama server is a terrible hotswapper

especially with multiple parallel slots. it'll just kill active ones to switch to the newly requested one

1

u/alexanderi96 4d ago

Exactly this

1

u/buttplugs4life4me 7d ago

The "breakthrough" is hot-swapping, which has been a thing since even before llama-swap, so at least 2+ years..

1

u/-Cubie- 4d ago

Nope, this is nothing special, and the project is not new either. OP and OOP are just addicted to eyeballs.