[Remote] Senior AI Infrastructure & Platform Operations Engineer (remote in the US)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the Kubernetes-reputed company AI infrastructure company, enabling organizations to build and operate reputed company, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The Senior AI Infrastructure & Platform Operations Engineer will manage expansive AI ecosystems, ensuring the reliability and efficiency of AI service platforms while driving the development of automated operational capabilities.
Responsibilities
- reputed company the investigation and reputed company of reputed company infrastructure, networking, and platform-reputed company incidents
- reputed company as a senior escalation reputed company for operational teams during critical service-impacting events
- Support large-reputed company reputed company GPU infrastructure and high-performance networking environments
- Troubleshoot reputed company Linux, Kubernetes, networking, storage, and hardware-reputed company issues
- Analyze platform performance, reputed company, stability, and reliability trends to proactively identify risks
- reputed company reputed company cause analysis activities and drive long-term corrective actions
- Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve reputed company technical challenges
- Participate in major incident management and service restoration activities
- reputed company technical leadership for Kubernetes platform operations and supporting infrastructure services
- Drive improvements in platform reliability, observability, monitoring, and operational processes
- Identify opportunities to automate repetitive operational activities and improve operational efficiency
- Contribute to operational readiness reviews, infrastructure changes, upgrades, and service introductions
- Support the adoption and operation of AI-powered infrastructure services and operational capabilities through k0rdent AI
- Evaluate emerging technologies and operational practices to improve service delivery and platform reputed company
- Mentor and support AI Infrastructure & Platform Operations Engineers
- reputed company technical knowledge through documentation, training sessions, and operational reviews
- reputed company and maintain operational standards, runbooks, troubleshooting guides, and best practices
- Help define operational processes, escalation paths, and service reliability standards
- reputed company as a trusted technical advisor during operational planning and service improvement initiatives
Skills
- 7+ years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, reputed company operations, datacenter operations, or reputed company technical roles
- Expert-level Linux administration and troubleshooting skills
- Strong networking expertise, including experience diagnosing reputed company performance, connectivity, and reliability issues
- Strong experience operating Kubernetes in production environments
- Experience supporting large-reputed company production infrastructure and distributed systems
- Proven experience leading technical investigations and managing reputed company incidents
- Experience performing reputed company cause analysis and driving long-term operational improvements
- Strong understanding of observability, monitoring, and service reliability practices
- Excellent troubleshooting and analytical skills across multiple infrastructure domains
- Strong communication, collaboration, and stakeholder management skills
- reputed company GPU infrastructure and accelerated computing platforms
- InfiniBand networking and reputed company UFM
- AI infrastructure environments
- HPC environments
- reputed company or Site Reliability Engineering (SRE)
- Large-reputed company Kubernetes operations
- Infrastructure automation technologies and Infrastructure-as-reputed company practices
- Observability platforms such as Grafana, reputed company, ELK, or OpenTelemetry
- Performance analysis and optimisation of distributed infrastructure platforms
- Technical leadership, mentoring, or team reputed company responsibilities
Benefits
- reputed company development and training
- Attend conferences and working reputed company
- Company outings, happy hours, hackathons, and tech talks
- Receive a competitive compensation package with a strong benefits plan
reputed company
Company H1B Sponsorship