I bought an Optiplex 5040, with an i5-6500TE, and 8 GB DDR3L RAM.

When I bought it, I installed Fedora Server on it. It got stuck every few days but I could never see the error. The services just stopped working, I couldn’t ssh into it, and connecting it to a monitor showed a black screen.

So, I thought let’s install Ubuntu Server, maybe Fedora isn’t compatible with all of its hardware. The same thing is happening, now, but I can see this error. Even when there’s nothing installed on it, no containers, nothing other than base packages, this happens.

I have updated the bios. I have tried setting nouveau.modeset=0 in the grub config file. I have tried disabling and enabling c-states. No luck till now.

Would really appreciate if anyone helps me with this.

UPDATE:

  • I cleaned everything and reapplied the thermal paste. I did not see any change in the thermals. It never goes over 55°C even under full load.
  • I reset the motherboard by removing that jumper thing.
  • I ran memtest86, which took over 2½ hours. It did not show any errors.
  • I ran a CPU stress test for over 15 hours, and nothing crashed.
  • I also ran the Dell’s diagnostic tool, available in the boot menu of the motherboard. The whole test took over 2 hours but did not show any errors. It tested the memory, CPU, fans, storage drives, etc.
@nialv7@lemmy.world
link
fedilink
English
117M

Time to add a cron job to auto reboot it once a day

Possibly linux
link
fedilink
English
37M

Does it have a Nvidia GPU? If it doesn’t then the nouveau modest does nothing.

My guess it that you are using a badly support realtek device. However I would need to see the full dmesg

@rookbrood@lemmy.world
link
fedilink
English
67M

It’s a long shot, but I had something similar on one of mine servers once. It was fixed by installing irqbalance and starting that daemon at startup.

@nutbutter@discuss.tchncs.de
creator
link
fedilink
English
25M

This actually worked. The CPU has to get stuck, it will in a day of being turned on, or it will keep working for weeks.

Thanks a lot for this!

@rookbrood@lemmy.world
link
fedilink
English
25M

Haha completely forgot about this comment. But i’m very gald it helped! I remember struggling with that problem for a long time.

@solrize@lemmy.world
link
fedilink
English
137M

Try running memtest86 for a few days to test memory. That is fairly easy to do though it involves booting from a flash drive. Web search should find info.

@nutbutter@discuss.tchncs.de
creator
link
fedilink
English
17M

I just ran it. It took over 2 hours to finish. Showed no errors. Is there a benefit of running it for a few days?

@solrize@lemmy.world
link
fedilink
English
27M

If the problem is intermittent then longer run has better chance of catching it, but 2 hours with no errors is a good sign, with regard to the memory.

@bulwark@lemmy.world
link
fedilink
English
77M

I’ve never seen this particular error, but CPU stall warnings seem like a fairly common thing. I wouldn’t jump straight to hardware fault, but it’s a possibility.

https://docs.kernel.org/RCU/stallwarn.html

@catloaf@lemm.ee
link
fedilink
English
97M

I’d lean toward bad hardware.

Try stress testing the CPU and RAM. See if you can get it to happen more frequently. Also see if you can disable that CPU core, either in the BIOS or in the OS, to see if the problem goes away.

@loganb@lemmy.world
link
fedilink
English
87M

I’m with catloaf. Consistent CPU soft locks point to a possible bad memory module or CPU.

Clear CMOS.

Try removing one memory module at a time.

See if there is an option to disable hyperthreading in bios.

Another thing to try is to remove the CPU, careful not to damage the LGA pins on the motherboard, and clean the CPU contacts with alcohol. Take care to ground yourself out and the case before handling the CPU out of socket.

Possibly linux
link
fedilink
English
37M

Don’t try to clean CPU pins. That is a very bad idea

@loganb@lemmy.world
link
fedilink
English
47M

I mean speaking from experience, its resurrected a couple problematic CPUs for me. CPU pins no, pads on an LGA style CPU, sure.

Possibly linux
link
fedilink
English
17M

It is much harder to mess up pads

@admin@sh.itjust.works
link
fedilink
English
27M

The CPU in this has no pins, is just contacts on the chip. The pins are in the motherboard, like the new 7000 series Ryzen.

Possibly linux
link
fedilink
English
17M

Clean pads not pins

@ikidd@lemmy.world
link
fedilink
English
9
edit-2
7M

Can you sudo dmesg | grep microcode and see if you have any errors?

If so, I’d be inclined to sudo dnf reinstall microcode_ctl then do a sudo dracut -f to regenerate the initramfs, and reboot. Be sure to have a working fallback kernel like LTS installed so you can recover if need be.

Edit: I just read you changed to Ubuntu. I can’t be arsed to figure out how Ubuntu does this stuff so that’s on you to figure out. Alternatively, install Arch and use intel-ucode package.

@TheBigBrother@lemmy.world
link
fedilink
English
-3
edit-2
7M

Set watchdog to reboot every day and juice it until the last drop before it definetly crashes.

Edit: that’s just a workaround if you really want try to fix it my answer would be restoring BIOS defaults, then clean install of the OS and then check it everything works fine, if the error persist try installing another OS if it still fail then go to the first step.

Shimitar
link
fedilink
English
247M

Had the same issues, it was heat.

Cool down your server, add a fan or a cooler…

I added a usb-powered fan sucking cooler air from outside the server area directly blowing it on the chassis.

That fixed for me.

@nutbutter@discuss.tchncs.de
creator
link
fedilink
English
17M

I cleaned everything and reapplied the thermal paste. That did not solve the problem. Also, the CPU is only of 35 watts and never goes over 55°C.

tenchiken
link
fedilink
English
157M

That server sounds a bit older in the teeth… Has new thermal paste been applied to the cpu? Even if the reported temps are under 90c, you might be getting hot spots causing glitches inside the package.

Worth trying a couple of different generations of kernel as well, both newer and older. You might be hitting a regression somewhere.

@nutbutter@discuss.tchncs.de
creator
link
fedilink
English
27M

I cleaned everything and reapplied the thermal paste. That did not solve the problem. Also, the CPU is only of 35 watts and never goes over 55°C.

@lolonaut@discuss.tchncs.de
link
fedilink
English
6
edit-2
7M

I had problems with soft locks because somehow the PSU was in corrupt state, maybe through a black out or something. The problem persisted through reboots and power offs, only cutting power helped.

@Decronym@lemmy.decronym.xyz
bot account
link
fedilink
English
5
edit-2
7M

Acronyms, initialisms, abbreviations, contractions, and other phrases which expand to something larger, that I’ve seen in this thread:

Fewer Letters More Letters
LTS Long Term Support software version
PSU Power Supply Unit
RAID Redundant Array of Independent Disks for mass storage
SATA Serial AT Attachment interface for mass storage

4 acronyms in this thread; the most compressed thread commented on today has 8 acronyms.

[Thread #907 for this sub, first seen 6th Aug 2024, 10:15] [FAQ] [Full list] [Contact] [Source code]

@slacktoid@lemmy.ml
link
fedilink
English
17M

What kind of drives do you have in your RAID? Is it SMR?

@nutbutter@discuss.tchncs.de
creator
link
fedilink
English
17M

The boot drive is an SSD, which is not in any RAID. I have another HDD connected via SATA. Another HDD connected via USB.

@slacktoid@lemmy.ml
link
fedilink
English
1
edit-2
7M

So 2 HDDs one SATA and one via USB in RAID? Can you remove the RAID drives and test it out?

Also what size are the drives and what’s their capacity?

@nutbutter@discuss.tchncs.de
creator
link
fedilink
English
17M

No, they are not in RAID either.

@slacktoid@lemmy.ml
link
fedilink
English
17M

My bad. What are you using the drives for?

@nutbutter@discuss.tchncs.de
creator
link
fedilink
English
17M

Films and shows, via Jellyfin.

@slacktoid@lemmy.ml
link
fedilink
English
17M

Is it just you, any transcoding? Have you tried it without the USB drive?

removed by mod

@seaQueue@lemmy.world
link
fedilink
English
207M

Make sure the microcode package is up to date as well

@4am@lemm.ee
link
fedilink
English
87M

Yeah, always check all of this stuff. Server hardware gets a lot more updates than like gamer board BIOS, companies invest high millions, even low billions in this stuff and they expect problems to be address promptly for that kind of cash.

Check for any peripherals or cards, too. RAID, backplanes, networking cards; drivers, firmware, anything.

@nutbutter@discuss.tchncs.de
creator
link
fedilink
English
17M

I did reset it. It did not help. I ran memtest86 for over 2 hours and did a CPU stress test for over 15 hours. Nothing crashed during the testing.

Possibly linux
link
fedilink
English
57M

Poorly supported hardware will also do this

@seaQueue@lemmy.world
link
fedilink
English
87M

An Optiplex 5040 should be well and thoroughly supported for 6+y now

Create a post

A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don’t control.

Rules:

  1. Be civil: we’re here to support and learn from one another. Insults won’t be tolerated. Flame wars are frowned upon.

  2. No spam posting.

  3. Posts have to be centered around self-hosting. There are other communities for discussing hardware or home computing. If it’s not obvious why your post topic revolves around selfhosting, please include details to make it clear.

  4. Don’t duplicate the full text of your blog or github here. Just post the link for folks to click.

  5. Submission headline should match the article title (don’t cherry-pick information from the title to fit your agenda).

  6. No trolling.

Resources:

Any issues on the community? Report it using the report flag.

Questions? DM the mods!

  • 1 user online
  • 156 users / day
  • 570 users / week
  • 1.55K users / month
  • 4.16K users / 6 months
  • 1 subscriber
  • 4.28K Posts
  • 89K Comments
  • Modlog