Job Monitor / Oil & Gas

Senior IT Systems Administrator (Linux & HPC)

KBR · United Kingdom · Posted Sep 18, 2026

Company KBR Location Leatherhead, United Kingdom Employment type Full-time Posted Sep 18, 2026 Listed via KBR
KBR is a global provider of differentiated, professional services and technologies delivered across a wide government, defense and industrial base. Drawing from its rich 100-year history and culture of innovation and mission focus, KBR creates sustainable value by combining engineering, technical and scientific expertise with its full life cycle capabilities to help our clients meet their most pressing challenges today and into the future. We deliver science, technology and engineering solutions to governments and companies around the world. KBR employs approximately 37,000 people worldwide with customers in more than 80 countries and operations in over 29 countries. KBR is proud to work with its customers across the globe to provide technology, value-added services, and long-term operations and maintenance services to ensure consistent delivery with predictable results. At KBR, We Deliver. KBR is looking for a Senior IT Systems Administrator (Linux & HPC). The Opportunity: The Senior IT Systems Administrator (Linux & HPC) is responsible for the administration, maintenance, security and operational reliability of Linux-based enterprise and high-performance computing platforms. The role will provide hands-on technical ownership across Linux operating systems, HPC compute, workload scheduling and associated storage and network services. A major focus will be the operation of cloud and co-located HPC platforms. The position requires strong Linux engineering experience, practical knowledge of HPC environments, including NVIDIA Base Command Manager, Azure CycleCloud and SLURM. The ability to diagnose complex cross-platform issues, and the confidence to work independently while acting as a senior technical resource for colleagues is also essential. This role does not include formal supervisory or people-management accountability Key Responsibilities Linux Systems Administration Administer, configure, patch, harden and upgrade enterprise Linux server platforms, with particular emphasis on Red Hat Enterprise Linux or comparable distributions. Manage core Linux services including identity and access integration, SSH, DNS client configuration, time synchronisation, software repositories, filesystems, logging and scheduled services. Automate repeatable administration using shell scripting and configuration-management or orchestration tooling. Monitor Linux performance, availability, capacity and security, and resolve complex operating-system and application integration issues. Maintain build standards, technical documentation, operational procedures and recovery runbooks. HPC Platform & Workload Scheduling Operate and support HPC clusters spanning management, login, compute and storage components. Administer SLURM, including queues and partitions, scheduling policies, job submission, accounting, fair-share, reservations and troubleshooting failed or poorly performing workloads. Support NVIDIA Base Command Manager and Azure CycleCloud for cluster provisioning, node lifecycle management, monitoring and integration with SLURM. Work with engineering and scientific users to diagnose job, compiler, library, MPI, resource-allocation and performance issues. Plan and execute maintenance activities while protecting service availability and active workloads. Compute, Storage & Infrastructure Integration Administer physical and virtual server infrastructure and support hardware lifecycle, firmware and operating-system maintenance. Support Cisco compute infrastructure and its integration with Linux and HPC management services. Operate and support NetApp file and data services used by Linux and HPC platforms, including provisioning, permissions, capacity, performance and availability. Apply a working understanding of high-speed networking, IP addressing, routing, DNS, NFS, network dependencies and storage connectivity to end-to-end troubleshooting. Collaborate with network, security, application and storage specialists, vendors and service providers to resolve cross-domain incidents. Security, Resilience & Operational Support Apply secure configuration, least privilege, vulnerability remediation and audit controls to Linux and HPC systems. Ensure backup, restore and disaster-recovery requirements are defined, implemented and regularly tested for supported platforms. Monitor service health and capacity, respond to incidents, identify root cause and contribute to problem management. Plan and deliver changes through established change-management processes, including risk assessment, testing, implementation and back-out planning. Participate in an appropriate operational support or on-call arrangement where required. Provide technical guidance, peer review and knowledge transfer to infrastructure colleagues and service-desk teams. Qualifications, Skills and Experience Essential Technical Skills & Experience Substantial hands-on experience administering Linux in a complex enterprise o
Explore more
1386 KBR jobs → 259 jobs in United Kingdom →
Oil & Gas Alert tracks 21+ employer career pages and delivers daily digests of new vacancies. Set your own keywords →