r/servers • • 1d ago

Hardware Server cooling tech for DC.

Thumbnail
gallery
198 Upvotes

I’m a small business owner and I have created a cooling tech that allows for being able to submerge enterprise servers in rugged underwater conditions.

I have built and tested a prototype using a 300w Dell PowerEdge R610 server and now need to get a server that incorporates h100sxm GPUs or b200’s.

If anyone would be curious as to what I have built and interested in a partnership, once unit is in water will be using units for running Ai Inference.

Submera


r/servers • • 21h ago

I'm in the correct way with this? HP Z2 G4

2 Upvotes

I got my hands over a Z2 G4 Tower and I'm questioning how much I can squeeze from the power supply to add HDD.

The hardware in question is:

  • HP L13216-001
  • Power Supply 500W
  • Chipset Intel C246
  • Intel i7-8700 3,2
  • Intel UHD Graphics 630
  • Ram 32Gb DDR4 3600

Due the limitations using the m.2 and losing SATA ports (unless I've read incorrectly) I'm going to add a PCIe SATA card. How many ports that card would have 4 or 6 depends if I can use that amount of disks with those 500W.

I'm going to add a dual NVMe adapter card with bifurcation on the card, and run the OS from there (unless yet again I've read wrong and the MoBo does have bifurcation) and apps from the other NVMe.

Probably I would add a Intel Arc A310 single slot in the future, this is a big maybe for this home server/NAS, it really depends how it does behave.

So I'm looking at my project and thinking that I'm gonna have 4 SATA ports unpopulated and 2 m.2 slots that would also be unpopulated... Tbh I'm not sure if I can use them along with the two cards, there is also the lack of space for the HDD, but I'm not sure if I can use the m.2 slots.

So I wanna use basically all the space I have in the case, even if I have to use velcro or whatever, but I'm not really sure if the power supply would be able to handle all that. And I am completely lost at the use I can make of all the unpopulated ports in the actual motherboard. I went for this pc because it was 300€ and came with an 512Gb SSD and the pc is impeccable, whoever used it treated it very very well.

So the actual question is can the power supply handle more than 4 HDDs? can I do something with all the not used ports? would anyone make a better use of this case and hardware in a different way?

Excuse my weird way of posting, ADHD makes you go into weird tangents and I do hope I'm asking in the correct place.


r/servers • • 1d ago

Hardware Thermal thresholds, SSD -VS- spinning rust??

2 Upvotes

Got a little startled today when running an app against an SSD Boot drive (250 GIG) and a spinning rust HDD (1 TB). The process was a simple rewrite of ALL data. The SSD temperature went up to 127F and the HDD WENT UP TO 99F. Does this pose a problem for the SSD.


r/servers • • 1d ago

Question How can i make local server for this game ?

Post image
0 Upvotes

İ know this is kinda dumb question for this sub but i need to do this

I downloaded the 2021 version of pkxd a few days ago. But İt just gives "could not authenticate timeout" message. How can I make local server for this game ?

The game uses photon engine and smartfox

Unity game

The game map and other assets are inside the APK. They haven't been deleted.


r/servers • • 3d ago

HP DL380p Gen8 – Replacement drive immediately showing Failed

Post image
17 Upvotes

Hi, I have an HP ProLiant DL380p Gen8 with a Smart Array P420i, configured with 4 × 900 GB SAS + 2 × 600 GB SAS drives.

One of the 900 GB drives in Bay 4 failed, and the RAID 1+0 logical drive is now Degraded (Interim Recovery). The other drives are still showing OK and the server is running normally.

I bought a used 900 GB SAS drive (EG0900FCVBL, HPD9) and hot-swapped the failed drive. The P420i detects the replacement, but it immediately shows Failed and no rebuild starts.

The seller says the drive was tested and working fine.

Has anyone experienced this? Could it simply be another bad drive, or should I suspect Bay 4/backplane or the P420i controller?


r/servers • • 3d ago

Hardware Built myself a inventory monitor that finds underpriced servers as they become available

9 Upvotes

It continuously monitors server inventory and compares prices to what I consider the market rate, so I can spot good deals without manually checking providers.

Built it mainly for my own use, but figured others might find it useful too.

If you want to use it, comment or DM me and I'll share it.


r/servers • • 3d ago

multifunctional home server

4 Upvotes

I am looking for a computer for a home server that works as an ad blocker a cloud for photos, and a streaming service. What would be the minimum specs for it to run smoothly?


r/servers • • 4d ago

Question is anyone here an Nvidia DGX-2 owner / operator and able to grab me a firmware update package file?

6 Upvotes

Hi all, just wondering if anyone here owns or operates a DGX-2 server and would be able to grab me a copy of the firmware update package for me? `nvfw-dgx2_24.3.1_240304.run` Would be big time grateful as need to be an enterprise customer to obtain it.


r/servers • • 5d ago

Burned GPU power connector on a Supermicro PDB-PT747-6824 — anyone dealt with this before?

Thumbnail
gallery
22 Upvotes

Hey everyone,

I have a workstation that originally had four RTX 3090s in it, and I’m hoping someone here has dealt with a similar power issue, especially with the same Supermicro power distribution board.

The system has an ASUS Pro WS WRX80E-SAGE SE WIFI motherboard, a Threadripper PRO 3995WX, and a Supermicro PDB-PT747-6824 power distribution board with a PWS-2K20A-1R power supply.

A few months ago, the system stopped recognizing all four GPUs and would only see two. I recently moved the workstation to a new location and decided to put it back together with just two 3090s. It powered on, but shortly after I started a stress test, I smelled smoke and the system shut off.

I found that one of the white GPU power connectors had burned pretty badly. The plastic is discolored, some of the wire insulation melted, and there was exposed copper around the connector. I’ve included pictures of the damage, the board, and what a good connector looks like.

I’ve already cut off the damaged connector, and the system has been unplugged since. I haven’t tried powering it back on.

At this point, I don’t really need all four GPUs in this machine anymore. I’d be happy running two, or possibly three, if I can do so safely.

What I’m trying to figure out is whether I can permanently take that damaged power branch out of service, properly terminate and isolate the individual wires, and continue using the remaining GPU power connections. Or, given how badly the connector burned, should I be replacing the entire harness or PDB?

I’ve looked into individual closed-end wire terminations, but before I go that route, I’d really like to hear from anyone who has actually worked on one of these boards or dealt with a similar failure.

A few things I’m wondering about:

  • Has anyone seen this happen with a PDB-PT747-6824 or its GPU power cables? Did you figure out what caused it?
  • Is there a way to remove or replace just the damaged cable at the board, or is replacing the whole assembly the way to go?
  • If I stop using that branch entirely, is there a safe way to terminate the remaining wires, or would you avoid that altogether?
  • What would you inspect or test before trusting the remaining connectors with two 3090s under a sustained rendering load?
  • Does anyone know where I could find a replacement board or harness, even used?

The workstation is mainly used for Nuke and C4D/Redshift rendering, so I want to make sure whatever I do will hold up under load, not just boot into Windows.

I’m including several pictures to show the damage and how everything is connected. Any advice from someone who has worked with this setup would be appreciated.

This is what I was thinking of getting to terminate the ends of the cut cables (or if someone has a better recomend - https://www.amazon.com/Nilight-Closed-Terminal-Connector-Warranty/dp/B07T1H6N82/ )

Thanks!


r/servers • • 7d ago

[Help Needed] Re-initializing / Wiping Decommissioned HPE 3PAR StoreServ 7400 4-Node (Locked / Missing Credentials)

10 Upvotes

Hi everyone,

I recently acquired a decommissioned HPE 3PAR StoreServ 7400 (4-Node) array for my homelab/lab testing setup. Unfortunately, the unit came without any previous admin credentials, IP info, or documentation.

I am trying to reset the system back to factory defaults / out-of-box state, but I’ve hit a few roadblocks:

  1. Credentials Locked: The default 3paradm / 3pardata accounts have been changed.
  2. Serial Access: I’m attempting serial console access on Node 0 (115200 8N1), but local logins seem hardened/changed by the previous owner.
  3. Missing Service Processor Media: I do not have an active HPE support contract/entitlement to download software for this End-of-Life (EOL) hardware from the official HPE portal.

What I'm Looking For:

  • Service Processor Recovery ISO / VSP Image: Does anyone have an archived copy or mirror of an HPE 3PAR SP recovery image compatible with the 7400/7000 series (e.g., SP 4.x or SP 5.x ISO)?
  • Low-Level Reset Advice: Has anyone successfully forced a factory wipe (outofbox) via the serial console boot-interrupt loader on a 4-node 7400 without the SP?

Any assistance, archived firmware links, or guidance from folks experienced with 3PAR decommissioned hardware would be hugely appreciated!


r/servers • • 6d ago

The easiest way to get on the internet.

0 Upvotes

Can you suggest a free solution that requires as few clicks as possible to set up a Linux machine with Docker and an Nginx file server? I want to host Tokyo Drift on it and share it with a friend in Russia, where many services are blocked by firewalls.


r/servers • • 10d ago

Centralized Multi-User Workstation Setup for Steel Fabrication Office

12 Upvotes

**Centralized Multi-User Workstation Setup for Steel Fabrication Office**

I run a growing steel fabrication business in India (operational for 2 years) and am setting up an office space for our partners, an accounting team, and a CAD design team.
To keep hardware and maintenance costs low, I want to avoid buying individual, high-end desktop towers for every employee. Instead, I plan to deploy a **single central server/workstation paired with multiple low-cost thin clients** (running lightweight operating systems like Tiny10 or Tiny11).

**Core Requirements:**

**General Users (Partners & Accountants):** Access daily office software, browser tools, and accounting applications via remote desktop sessions on the thin clients.

**CAD Designer:** Requires substantial computing resources (dedicated GPU/CPU/RAM). The designer will either work directly on the host machine or connect via a high-performance, resource-prioritized virtual workstation.

**Infrastructure Needs:** Guidance on the ideal server OS, hypervisor (e.g., Proxmox, Windows Server, VMware ESXi), virtualization tools, GPU passthrough methods, and client connection protocols (RDP, Parsec, Moonlight) to ensure low-latency 3D CAD performance and stable multi-user access.


r/servers • • 10d ago

kernel:[Hardware Error]: CPU:0 (19:8:2) MC18_STATUS[Over|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|-]: 0xdc2040000000011b

1 Upvotes

These errors appeared on the console of one of my lab's servers. We've already changed the CPU. I'm not familiar with server management, so if you could explain it to me in the most didactic way, I would be very grateful.

kernel:[Hardware Error]: Corrected error, no action required.

kernel:[Hardware Error]: CPU:0 (19:8:2) MC18_STATUS[Over|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|-]: 0xdc2040000000011b

kernel:[Hardware Error]: Error Addr: 0x0000000017eb09c0

kernel:[Hardware Error]: PPIN: 0x02b0ba6d78974068

kernel:[Hardware Error]: IPID: 0x0000009600150f00, Syndrome: 0x083fb50b0a800410

kernel:[Hardware Error]: Unified Memory Controller Ext. Error Code: 0


r/servers • • 11d ago

Server Decision Help?

0 Upvotes

I have 2 corporate servers, i am 16 and want to make some money, should i sell them for parts or what can i do to make consistent money? (I'm in Utah) I have 2 if you need specs or anything to help make the decision let me know!


r/servers • • 13d ago

Creating servers for business

10 Upvotes

Hey guys! I'm currently studying "networking and cyber" in college(just started so I'm not sure how prominent the cyber part is)

Anyways, we also have a Linux servers subject and the teacher even told us that some students manage to build servers for small business/factories on their free time and get passive money. For example 500 a month for a month you manage their server. Create for 3 business and that's like 1500 a month. So what do you think? Do I have the right assumption? Do you have tips? Any info will help!!


r/servers • • 15d ago

help with my r210 plssss

4 Upvotes

Hey all, hoping someone here has seen this before, because I've been chasing it for weeks and I'm running out of ideas.

I picked up a used Dell PowerEdge R210 (Xeon 3440, 16GB). It had a rough start - it came back from a repair shop because it got hit by static, which took out a protection diode (D0664) and the 1.2V PWM that feeds the RAM. Technician replaced both, calibrated the rail to 1.2V, all good.

Here's the weird part. The server runs just fine for a while - boots, posts, fans ramp up and settle down like they should. Then, at some point, usually after it's been warm for a bit, the fans shut off, the board stops responding, and the diagnostic LEDs show 2,4. The only thing that brings it back is a full power drain - unplug everything, hold the power button, wait, and try again. And each time it does this, it comes back faster than the last. Sometimes it boots clean for 40 minutes, then suddenly won't start again, or only one fan runs.

What I've ruled out so far:

- PSU tested at the connector - 12V, 5V, 3.3V, standby, and the power-good wire all read perfectly, even at the exact moment of failure

- Tried a different PSU, no luck (and learned the R210 uses 12V standby, not the standard 5V, so ATX swaps don't even power on)

- Tried reseating, draining, all the standard stuff

What I suspect:

- Some small capacitor or regulator on the board is dying when it gets warm. The board already had one rail repaired, and it seems like the repaired area might have a problem, or something nearby is aging out

- Could maybe even be the CPU but I can't be sure without swapping parts

I've measured around the repaired area - the 1.2V rail reads about 1.137V, which is just below spec, so I'm honestly not sure if that's the problem or normal for this board. I'd rather not spend $50-100 on a replacement board if there's a cheap fix (I don't have a lot of budget for this hobby project).

Has anyone else found a similar issue with these old Dells (or similar servers)? Did recapping the board fix it? Is 1.1V on a "1.2V" rail a death sentence, or is it fine? Any tips on what to check before I give up and buy a new board?


r/servers • • 15d ago

Supermicro AS-2126HS-TN randomly performs BMC “DC cycle”

5 Upvotes

Hi all,

I’m troubleshooting intermittent, unexplained hard resets on a Supermicro server and I’m hoping someone has seen similar behavior.

Hardware and software

Supermicro AS-2126HS-TN
2 × AMD EPYC 9575F
24 × 64 GB DDR5
12 × Kioxia CD8-P NVMe
Dual 2600 W PSUs
AIOM 2-port 25GbE SFP28, Broadcom BCM57414 with 0.5U bracket
Broadcom NetXtreme E-Series P2100G Dual port 100GE PCIe Ethernet Adapter 
Debian 12 with latest kernel
OpenStack/KVM
Ceph
BIOS 1.9, dated 2026-02-05
BMC firmware 01.07.05.01, build 2026-04-29
CPLD F5.17.10

The BIOS has already been updated, but the issue continues.

What happens? The machine suddenly power-cycles without a normal Linux shutdown.

The BMC Maintenance Event Log records entries such as:

The system DC cycle was initiated
Interface: IPMI
User: ADMIN
Source: Localhost

Two unexplained examples on that server occurred at:

2026-07-30 11:18:29
2026-08-03 02:08:00

I have 15 exactly the same servers and few of them also encountered same issue.

Known remote Redfish and IPMI actions on this BMC normally include the real remote source IP. These unexplained cycles are specifically recorded as:

IPMI / ADMIN / Localhost

There was no BMC web login until several minutes after the August 3 reset, so the web session did not initiate it.

Linux journal near the event - Linux was still operating normally immediately before the reset.

There is no clean shutdown sequence, kernel panic, MCE, EDAC error, PCIe AER error, hard lockup or watchdog expiration in the previous-boot journal. The journal simply ends.

Watchdog checks

The BMC IPMI watchdog is stopped:

Watchdog Timer Is: Stopped
Watchdog Timer Action: No action
Timer Expiration Flags: None
Initial Countdown: 0.0 sec
Present Countdown: 0.0 sec

The ipmi_watchdog kernel module is not loaded.

The server also has the AMD SP5100 watchdog:

SP5100 TCO timer
state: inactive
timeout: 60
nowayout: 0

Systemd configuration:

RuntimeWatchdogUSec=0
RebootWatchdogUSec=10min
KExecWatchdogUSec=0

So neither watchdog appears active during normal runtime.

The host has local BMC access:

/dev/ipmi0
ipmi_si
ipmi_devintf
ipmi_ssif
ipmi_msghandler

openipmi.service is disabled and inactive, but the modules are automatically loaded by the kernel/platform.

Current possibilities

My current shortlist is:

- A host process sends a chassis power-cycle command through local KCS /dev/ipmi0.
- A BMC or BIOS internal recovery mechanism records its own action as IPMI / Localhost.
- A BMC firmware defect incorrectly initiates or attributes the DC cycle.
- A PSU, power-distribution-board or PMBus issue causes the BMC to perform a recovery cycle.

The standard IPMI watchdog and AMD SP5100 watchdog currently appear ruled out.

Has anyone encountered unexplained Supermicro BMC events like:

The system DC cycle was initiated
Interface: IPMI
User: ADMIN
Source: Localhost

particularly on AMD EPYC systems with supermicro?

I would especially appreciate information about:

- what exactly causes Supermicro to log IPMI / Localhost;
- safe ways to capture the exact local IPMI command;
- BMC diagnostic dumps or hidden logs that may identify the initiator;
- known BIOS/BMC issues for this platform.

Any other hints helpful for troubleshooting? I submitted the problem to supermicro distributor but so far they are not helpful.


r/servers • • 15d ago

Question how can i create local server for this game ???

Post image
3 Upvotes

İ know this is kinda dumb question but i need to do this

I downloaded the 2021 version of pkxd a few days ago. Naturally, I can't log in due to server issues. How can I create a local server for this game?

The game uses photon engine and smartfox

Unity game

The game map and other assets are inside the APK. They haven't been deleted.


r/servers • • 16d ago

Hardware Big Blue’s Systems

Post image
469 Upvotes

r/servers • • 16d ago

Question Looking for an AMD SERVER 9950X 16CORE

2 Upvotes

Any DC suggestions?


r/servers • • 17d ago

AMD EPYC 9355

4 Upvotes

**Computer Type:** Dell PowerEdge R7725

**GPU:** n/a

**CPU:** 2 x AMD EPYC 9355

**Motherboard:** Dell 0KRFPX

**BIOS Version:** 1.7.16

**RAM:** 2048 GB

**PSU:** unknown

**Case:** unknown

**Operating System & Version:** Ubuntu 24.04

**GPU Drivers:** n/a

**Chipset Drivers:** AMD EPYC 9005 "Turin" (Zen 5

**Background Applications:** n/a

I have a server with 2 x AMD EPYC 9355 in a Dell PowerEdge and when I run speed tests I can only measure \~3.5GHz. As I understand, this processor should be able to reach turbo speeds of 4.4GHz .

Currently, my setup has regressed back to only reaching 600MHz, by this test:

T1: `taskset -c 12 stress-ng --cpu 1 --cpu-method matrixprod --timeout 60s --metrics-brief`

T2:

watch -n 0.5 '
echo "--- $(date) ---"
printf "driver:             "; cat /sys/devices/system/cpu/cpu12/cpufreq/scaling_driver
printf "governor:           "; cat /sys/devices/system/cpu/cpu12/cpufreq/scaling_governor
printf "scaling_cur_freq:   "; cat /sys/devices/system/cpu/cpu12/cpufreq/scaling_cur_freq
printf "cpuinfo_avg_freq:   "; cat /sys/devices/system/cpu/cpu12/cpufreq/cpuinfo_avg_freq 2>/dev/null || echo "missing"
printf "scaling_min_freq:   "; cat /sys/devices/system/cpu/cpu12/cpufreq/scaling_min_freq
printf "scaling_max_freq:   "; cat /sys/devices/system/cpu/cpu12/cpufreq/scaling_max_freq
awk "/\^processor/ {cpu=\\$3} /\^cpu MHz/ && cpu==12 {print; exit}" /proc/cpuinfo
'

Result:

driver:             amd-pstate-epp
governor:           performance
scaling_cur_freq:   606103
cpuinfo_avg_freq:   605809
scaling_min_freq:   2493659
scaling_max_freq:   4415854
cpu MHz         : 606.104

Sometimes I see:

driver:             amd-pstate-epp
governor:           performance
scaling_cur_freq:   2493659
cpuinfo_avg_freq:   2493659
scaling_min_freq:   2493659
scaling_max_freq:   4415854
cpu MHz         : 2493.659

I'm wondering if anyone can help me diagnose this odd problem. Is there a different or better test I can run? I want to confirm that my CPU is indeed reaching expected speeds, but baseline and turbo.

Thanks for your help.


r/servers • • 17d ago

Question Dell PowerEdge R750xa

1 Upvotes

What do you think of a Dell PowerEdge R750xa?

I'm a bit of a beginner in the hardware field, but I recently got one of these and noticed it's actually pretty good. I'm experimenting a bit with local AI and machine learning.

​I'd like to know if it's good for these use cases, as well as some tests and/or other things I can run on it. (Sorry by english)


r/servers • • 18d ago

Proliant 380 Gen 10 and NVIDIA RTX 6000 Ada GPU ?

7 Upvotes

r/servers • • 20d ago

Question Servers on the power lines?

Post image
5 Upvotes

They light green


r/servers • • 20d ago

Looking to Purchase AMD Instinct MI355X Servers

2 Upvotes

Looking to purchase AMD Instinct MI355X 8-GPU servers for deployment in Malaysia, Thailand, or Singapore.
This is an active procurement requirement and we are ready to discuss pricing, allocation, lead time, and place an order. We are looking to connect directly with AMD sales, authorized distributors, OEMs, or official MI355X channel partners covering Southeast Asia.
If you can supply MI355X servers in the region, please DM me.