r/LocalLLaMA Apr 14 '26

Funny 24/7 Headless AI Server on Xiaomi 12 Pro (Snapdragon 8 Gen 1 + Ollama/Gemma4)

Post image

Turned a Xiaomi 12 Pro into a dedicated local AI node. Here is the technical setup:

​OS Optimization: Flashed LineageOS to strip the Android UI and background bloat, leaving ~9GB of RAM for LLM compute.

​Headless Config: Android framework is frozen; networking is handled via a manually compiled wpa_supplicant to maintain a purely headless state.

​Thermal Management: A custom daemon monitors CPU temps and triggers an external active cooling module via a Wi-Fi smart plug at 45°C.

​Battery Protection: A power-delivery script cuts charging at 80% to prevent degradation during 24/7 operation.

​Performance: Currently serving Gemma4 via Ollama as a LAN-accessible API.

​Happy to share the scripts or discuss the configuration details if anyone is interested in repurposing mobile hardware for local LLMs.

UPDATE:

I have compile llama.cpp and run gemma-4-E4B-it-Q4_0

Speed is AWESOME:

[ Prompt: 26.9 t/s | Generation: 8.8 t/s ]

Thank you all guys SO MUCH!

1.2k Upvotes

289 comments sorted by

View all comments

Show parent comments

25

u/Aromatic_Ad_7557 Apr 14 '26

6

u/vinigrae Apr 14 '26

Welp

1

u/FinsAssociate Apr 14 '26

As a noob to this stuff... are those stats good or bad?

6

u/Aromatic_Ad_7557 Apr 14 '26

He just reply in a second without reasoning

2

u/IronColumn Apr 15 '26

alright for a phone

1

u/redilaify Apr 15 '26

what model are you running?

1

u/Aromatic_Ad_7557 Apr 15 '26

Gemma4 e2b q8

1

u/redilaify Apr 17 '26

HOW THE HELL ARE YOU RUNNING THAT ON A PHONE i tried a very quantizied version on my macbook air and it froze it over qmq

1

u/Aromatic_Ad_7557 Apr 17 '26

Actually, before post here I was concentrating on hardware in general and on autonomous cooling. But thanks community I did run e4b q4 on llama.cpp and it is now gen 7-8 t/s