Current System Load
The plot below shows the status of the CPU nodes on the current Cirrus service for the past day.
A description of each of the status types is provided below the plot.
CPU

- alloc: Nodes running user jobs
- idle: Nodes available for user jobs
- resv: Nodes in reservation and not available for standard user jobs
- down, drain, maint, drng, comp: Nodes unavailable for user jobs
- mix: Nodes in multiple states
Service Alerts
| Status | Start | End | Scope | Impact | Reason |
|---|---|---|---|---|---|
| Planned | 2026-07-23 08:30 | 2026-07-23 12:30 | Compute | Maximum 256 nodes available to run work, extending queue times | Maintenance on power delivery to Cirrus expansion |
| Planned | 2026-07-02 08:00 | 2026-07-31 18:00 | Compute nodes | Possibility of intermittently longer queue times and reduced node availability | We are taking action to improve the resilience of our cooling infrastructure in hot weather. This will mean fewer episodes of reduced compute capacity in future, but while the work is ongoing it is possible that we may need to take nodes out of service at short notice, which will increase queue times. We apologise for any inconvenience. |
Recently Resolved Service Alerts
This table lists the last five resolved service alerts A full list of historical resolved service alerts is available.
| Status | Start | End | Scope | Impact | Reason |
|---|---|---|---|---|---|
| Resolved | 2026-07-20 08:30 | 2026-07-21 18:00 | Connectivity to service | Update 2026-07-20 13:30: Login access restored, further possibility of interruptions to access for further 1.5 days. Unable to remotely access the service for ~0.5 days with risk of issues or interruptions for up to two days. Scheduled work will continue to run throughout, though there may be issues with workflows that require access external to the service. | Upgrade of ACF firewalls and core network |
| Resolved | 2026-07-14 07:00 | 2026-07-14 08:00 | Connectivity to service | Unable to remotely access the service for ~1hr. Scheduled work will continue to run throughout, though there may be issues with workflows that require access external to the service. | Essential upgrade work to main router at the ACF |
| Resolved | 2026-07-08 08:00 | 2026-07-16 11:40 | Compute and Login | User login service has been temporarily suspended. User work is unlikely to start running before access is restored. | Completing activity to expand Cirrus to 640 nodes, upgrade CPE and apply security updates that will allow us to re-enable debugging at the end of this maintenance outage. |
| Resolved | 2026-07-03 10:45 | 2026-07-03 13:00 | login service and /home file system | Possible intermittent performance issues | Completing recovery actions after Ceph storage issues on 2 July |
| Resolved | 2026-07-02 14:50 | 2026-07-02 18:30 | /home file system | /home file system is currently running very slowly which may affect file access and prevent logging on. | Unexpected outages related to a problem with the CEPHFS /home disk system |
Service Maintenance Sessions
We keep maintenance downtime to a minimum on the service but do occasionally need to perform essential work on the system. Maintenance sessions are used to ensure that:
- software versions are kept up to date;
- firmware levels on HPE and third-party peripheral equipment are kept up to date; essential security patches are applied;
- failed/suspect hardware can be replaced;
- new software can be installed; periodic essential maintenance on HPE electrical and mechanical support equipment (refrigeration systems, air blowers and power distribution units) can be undertaken safely.
Additional maintenance sessions can be scheduled for major hardware or software updates; major upgrades to facility plant and infrastructure; acceptance testing following major service upgrades and statutory electrical testing.
No upcoming or ongoing maintenance sessionsA list of all previous maintenance sessions.