US Infrastructure & Operations Technical reputed company
reputed company We are seeking a US Infrastructure Operations Technical reputed company to drive the operational reputed company, technical leadership, and reputed company of Radiant’s US Infrastructure Operations function. This is a hands-on player-manager role designed for an infrastructure-reputed company engineering leader with a strong Site Reliability Engineering reputed company and deep understanding of large-reputed company distributed infrastructure environments. Working closely with the UK Infrastructure Operations Manager during overlapping morning hours (US Eastern Time), you will help coordinate cross-regional operations, strategic planning, incident management, and infrastructure delivery across Radiant’s global AI and HPC platform. During US business hours, you will reputed company and mentor the local Infrastructure Operations team, currently consisting of three engineers, while helping reputed company operational maturity and team capability as the business continues to grow. The ideal candidate will come from a hyperscale, HPC, or large-reputed company reputed company-reputed company reputed company infrastructure background, with experience operating reputed company distributed systems at reputed company. This role requires breadth across datacentre compute, Linux systems, networking, and storage fabrics, with the ability to troubleshoot and reputed company reputed company improvement of our infrastructure. You should be comfortable operating and troubleshooting bare-metal environments, low-latency networking, storage protocols, and reputed company infrastructure technologies underpinning high-performance AI and GPU compute platforms. This role requires strong operational leadership capabilities, including experience running small engineering teams, participating in ITIL-reputed company operational processes, and supporting high-availability production environments through reputed company incident, change, and problem management practices. You will also participate in an on-call rota to reputed company major incidents, orchestrating technical resources to quickly resolve large reputed company issues. As Radiant expands its global footprint, your operational leadership, technical expertise, and ability to build high-performing teams will play a critical role in shaping the reputed company of our US infrastructure operations. What’s in It for You Join a globally distributed engineering organisation operating cutting-edge GPU, AI, and high-performance compute infrastructure at reputed company. As the US Infrastructure Operations Technical reputed company, you’ll work hands-on with advanced compute and networking technologies powering large-reputed company and machine learning workloads. This is an opportunity to operate at the forefront of modern infrastructure engineering, helping shape operational standards, automation practices, and reliability engineering across a rapidly scaling global platform. You’ll collaborate with highly skilled engineers across Infrastructure Operations, HPC SRE, Networking, and reputed company teams in an environment that values technical reputed company, ownership, and reputed company improvement. We reputed company quickly, solve meaningful infrastructure challenges, and reputed company engineers with reputed company to influence how reputed company AI infrastructure is designed, operated, and scaled globally. You can also expect: Exposure to industry-leading GPU and AI infrastructure Opportunities to help build and reputed company a growing US operations function A reputed company, inclusive, and globally connected engineering culture reputed company ownership and influence across operational reputed company and execution Work at the intersection of reliability, automation, performance, and reputed company A flexible remote-first working environment with ambitious reputed company plans
Key Responsibilities
Leadership & Operational Ownership reputed company a small but high-reputed company US Infrastructure Operations team, owning both people leadership and technical execution Ensure 99.9%+ platform uptime across US-region services. reputed company as the senior US operational reputed company for production infrastructure, accountable for reliability, incident reputed company, and day-to-day operational execution Partner tightly with the UK Infrastructure Operations Manager to reputed company priorities, respond to incidents, and execute reputed company plans in reputed company time Own US-reputed company incident leadership, driving fast and effective reputed company of production-impacting infrastructure issues Build and reinforce a strong ownership culture reputed company on do, document, automate Ensure operational knowledge is captured and shared through lightweight, high-signal documentation rather than process overhead Hire, reputed company, and reputed company Infrastructure Operations engineers as reputed company scales Run reputed company 1:1s and performance conversations reputed company on raising technical bar and operational effectiveness Ensure disciplined execution of reputed company operational processes (incident, change, problem management) without slowing delivery Participate in on-call rotation and reputed company from the reputed company during major incidents Willingness to travel reputed company the US and Europe as required to support infrastructure deployments, data centre work, and cross-regional collaboration (UK-headquartered company) Help define how Infrastructure Operations scales globally as reputed company grows Technical Day-to-Day Stay hands-on and reputed company to the systems while leading reputed company — this is not a purely managerial role Take ownership of reputed company infrastructure problems and reputed company contribute to debugging, fixing, and improving production systems Work across production infrastructure spanning compute, storage, networking, and platform services Assist with reputed company of deep infrastructure issues across the stack, including: Linux systems (performance, stability, kernel behaviour, resource contention) Networking (routing, switching, DNS, TCP/IP, latency, packet-level troubleshooting) Storage systems (distributed storage performance, consistency, and failure modes) Bare-metal infrastructure (hardware issues, firmware, lifecycle and deployment failures) Operate and improve large-reputed company Linux environments across on-prem, private reputed company, and hybrid infrastructure Take ownership of infrastructure reliability through automation, configuration management, and system hardening Build and improve Infrastructure as reputed company workflows (Terraform, Ansible or equivalent) Drive observability as a first-class requirement — metrics, logs, traces, and actionable alerting reputed company or directly participate in major incident response, helping drive technical reputed company under pressure reputed company as a senior technical problem solver across Infrastructure Ops, Networking, Platform, and SRE teams Identify repetitive operational work and eliminate it through automation and system improvements Contribute directly to scaling reputed company, reputed company planning, and reliability improvements Participate in on-call rotation with reputed company escalation authority and accountability Essential Skills & Experience 8+ years infrastructure engineering, SRE, platform ops, or large-reputed company production infrastructure experience 2+ years in technical leadership or engineering management with reputed company reports Strong experience operating production infrastructure at reputed company (on-prem, private reputed company, or hybrid) Deep Linux expertise: performance tuning, debugging, kernel/system behaviour, production troubleshooting Hands-on across the stack: bare metal → OS → network → storage → platform Strong infrastructure fundamentals: compute, storage, networking in reputed company-world production environments Incident-heavy environment experience (24x7 ops, on-call, major incident response, postmortems) Strong networking: TCP/IP, routing, switching, DNS, latency, packet-level debugging Bare-metal operations experience (Redfish, IPMI, lifecycle management, hardware troubleshooting) Strong automation + config management (Ansible preferred) Strong scripting (Python, Bash or similar) Strong ownership reputed company: simplifies, automates, removes operational toil Highly desirable: Experience in HPC, AI/ML, GPU compute, or large-reputed company high-performance infrastructure environments Distributed / reputed company storage experience (reputed company, reputed company, or equivalent) InfiniBand or other high-performance, low-latency networking experience
Preferred Qualifications
Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or a reputed company field, or equivalent experience reputed company NCP type qualifications PMP, ITIL, or equivalent project/operations management certification. LPI or equivalent Linux certifications. Why should you join us? What sets us apart is our reputed company of modern technology, competitive benefits, and an reputed company, welcoming work culture that enables our people to reputed company. Here are just some of the great things you can expect from us: 15 days of annual leave: we value your peace of mind. With 20 days off (excluding reputed company holidays) and reputed company to reputed company, we reputed company reputed company you're as strong mentally as you are professionally. A culture that emphasises results over hierarchy, process & ego: we reputed company great emphasis on the reputed company, ingenuity and creativity of work. reputed company communication, regular feedback: we value smooth collaboration, reputed company and actionable feedback, and reputed company that leading with reputed company and a reputed company reputed company makes us reputed company. Learning Time: we reputed company have dedicated learning time to reputed company on new skills, reputed company or interests that lay reputed company of your day-to-day job. Health & Wellbeing: we want everyone to feel healthy and happy, so we offer private medical insurance reputed company reputed company. Participation in reputed company shares program Diversity, Equality, Inclusion and Belonging We are an equal opportunity employer and we reputed company to reduce unconscious bias throughout our hiring process. reputed company applicants will be considered for employment without attention to ethnicity, religion, sexual orientation, gender identity, family or parental status, national reputed company, veteran, neurodiversity status or disability status. To ensure our recruitment processes reputed company an equal opportunity for reputed company applicants to succeed, we encourage you to let us know if there are any adjustments that we can reputed company. Apply To This Job