S
  • Status
  • Events
  • Monitors
SGHVAIC
All Systems OperationalDegraded PerformanceDowntime PerformanceMaintenance
Apr 10, 2025
1 year ago
CUDA Errors and Degraded GPU Performance
Monitoring · April 10 at 10:40 AM (UTC)

The issue was caused by the Nvidia vGPU licensing server and has been resolved. We are closely monitoring the system.

Apr 8, 2025
1 year ago
Platform Outage: Storage Out of Space
Monitoring · April 8 at 10:29 AM (UTC)

The storage expansion is now complete. VMs are now online.

Identified · April 8 at 8:58 AM (UTC) (2 hours earlier)

Various VMs are currently offline due to the IO error caused by storage being out of space. We are attempting to expand the storage using spare disks.

Mar 8, 2025
1 year ago
Users Unable to Log in to Services
Monitoring · April 15 at 9:03 AM (UTC)

The cause appears to be a hardware/firmware issue in the physical host. The issue has been mitigated through forced regular synchronization with an NTP server. We will continue to monitor the system in case any time-related issue appears again. We apologize for any inconvenience caused.

Investigating · April 15 at 8:00 AM (UTC) (1 hour earlier)

The time synchronization issue appeared again. We are currently investigating the root cause.

Monitoring · April 11 at 9:19 AM (UTC) (4 days earlier)

The issue appeared again on VM-LINUX-CP (code-hvaclab.sghvaic.org) after a brief loss of connectivity to the domain controller, due to a bug in winbind.

The issue has been resolved after a restart of the winbind daemon. A permanent fix is currently underway.

Monitoring · March 8 at 8:21 AM (UTC) (1 month earlier)

The issue appears to be resolved at the moment.

Identified · March 8 at 4:00 AM (UTC) (4 hours earlier)

The cause is identified to be a time synchronization issue between the servers. We are currently resolving the issue.

View events history

powered by openstatus.dev