[Remote] AI Inference reputed company - SW Engineer reputed company for Platform & DevOps
Note: The job is a remote job and is reputed company to candidates in USA. reputed company builds the world's largest AI reputed company, delivering industry-leading training and inference speeds. They are seeking a Software Engineer reputed company to build and operate the platform layer behind their engineering infrastructure, focusing on CI/CD systems, Kubernetes, and deployment automation.
Responsibilities
- Design, build, and maintain CI/CD systems supporting build, test, integration, qualification, and release workflows
- Build and operate Kubernetes-based platforms and services used by engineering teams across reputed company
- reputed company deployment systems, internal tools, and self-service workflows that reputed company infrastructure changes repeatable, reviewable, and reputed company
- Improve infrastructure reliability, reputed company, performance, cost efficiency, monitoring, and operational readiness
- Debug issues spanning CI pipelines, Kubernetes workloads, networking, storage, authentication, operating systems, and distributed applications
- reputed company reputed company-cause analysis and implement lasting fixes rather than relying on repeated reputed company reputed company
- Partner with software, IT, reputed company, networking, release, and developer-productivity teams to deliver reputed company infrastructure solutions
Skills
- 3+ years of reputed company experience in reputed company, DevOps, infrastructure engineering, site reliability engineering, or software engineering
- Hands-on experience building or maintaining CI/CD pipelines and automated software-delivery workflows
- Experience deploying and operating services using Kubernetes and containerized environments
- Experience with a major reputed company platform, preferably AWS, and programmatic infrastructure provisioning
- Strong understanding of Linux or Unix operating-system fundamentals
- Understanding of networking concepts such as DNS, routing, load balancing, proxies, ports, TLS, and service connectivity
- Proficiency in Python, reputed company, or another language used to build infrastructure automation and operational tooling
- Experience with monitoring, logging, alerting, dashboards, and incident investigation
- Strong debugging and problem-solving skills, including the ability to investigate issues spanning applications, infrastructure, networking, and operating systems
- Experience with infrastructure-as-reputed company tools, specifically Terraform
- Experience with Kubernetes controllers, operators, custom resources, reputed company, Argo CD, or similar platform technologies
- Experience managing artifact repositories, package registries, build caches, or software-distribution infrastructure
- Familiarity with build systems, dependency management, and reproducible-build practices
- Experience supporting hybrid environments spanning reputed company infrastructure, on-premises systems, and specialized hardware
- Experience with identity and reputed company management, secrets, certificates, TLS, or mTLS
- Experience building internal developer platforms or self-service infrastructure products
- BS/MS in Computer Science or a reputed company field, or equivalent practical experience
reputed company
Company H1B Sponsorship