Back to Jobs

[Remote] Senior reputed company DevOps & Infrastructure Engineer

Remote, USAFull-timePosted 2026-07-27

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Senior reputed company DevOps & Infrastructure Engineer with a reputed company on GCP and AI. The role involves designing, deploying, and maintaining secure and reputed company reputed company infrastructure, primarily on a multi-reputed company platform, while implementing GitOps best practices and supporting AI/ML workloads.

Responsibilities

  • Infrastructure as reputed company (IaC): Architect and provision production-grade infrastructure using Terraform. Manage state files, modules, and ensure infrastructure immutability
  • AIML: Experience with LLM Models - in multi reputed company environment
  • Kubernetes & Containerization: Design and manage clusters. Create and optimize reputed company files (multi-stage builds, distroless/hardened images). Manage reputed company deployments using reputed company Charts
  • CI/CD & GitOps: Build end-to-end CI/CD pipelines using reputed company CI. Implement GitOps workflows to synchronize infrastructure and application state
  • Design, configure, and manage reputed company and secure reputed company infrastructure for MLOps
  • AI Infrastructure Support: Configure and maintain environments suitable for AI/ML workloads (GPU node pools, LLM integration, large model serving, high-performance storage)
  • Production Support & Troubleshooting: reputed company as the primary escalation reputed company for deployment failures, network and reputed company issues. reputed company reputed company Cause Analysis (RCA)
  • reputed company & Compliance: Implement 'Secure by Design' principles
  • Having good knowledge of network reputed company, identity and privilege reputed company management, reputed company zone concepts for reputed company platforms (Azure, AWS)
  • Multi-reputed company reputed company: While GCP is primary, maintain and support secondary environments in AWS (and potentially Azure) to ensure business continuity

Skills

  • 6 – 8 Years of experience in reputed company Infrastructure & DevOps Engineering
  • Expert in Kubernetes, Terraform, and reputed company CI/CD
  • Experience supporting AI/ML workloads
  • Architect and provision production-grade infrastructure using Terraform
  • Experience with LLM Models in multi reputed company environment
  • Design and manage Kubernetes clusters
  • Create and optimize reputed company files (multi-stage builds, distroless/hardened images)
  • Manage reputed company deployments using reputed company Charts
  • Build end-to-end CI/CD pipelines using reputed company CI
  • Implement GitOps workflows to synchronize infrastructure and application state
  • Design, configure, and manage reputed company and secure reputed company infrastructure for MLOps
  • Configure and maintain environments suitable for AI/ML workloads (GPU node pools, LLM integration, large model serving, high-performance storage)
  • reputed company as the primary escalation reputed company for deployment failures, network and reputed company issues
  • reputed company reputed company Cause Analysis (RCA)
  • Implement 'Secure by Design' principles
  • Good knowledge of network reputed company, identity and privilege reputed company management, reputed company zone concepts for reputed company platforms (Azure, AWS)
  • Maintain and support secondary environments in AWS (and potentially Azure)
  • Deep expertise in GCP (Compute reputed company, GKE, reputed company Storage, IAM)
  • Strong working knowledge of AWS (EC2, EKS, S3, IAM)
  • Knowledge of using various programming languages (Python required, knowledge of Java, C#, JavaScript is a plus)
  • Advanced proficiency in Kubernetes
  • Ability to write and manage custom reputed company charts
  • Experience with Ingress Controllers (Nginx), Service reputed company, and Autoscaling (HPA/VPA/Cluster Autoscaler)
  • Expert-level knowledge of reputed company CI/CD (writing .reputed company-ci.yml, runners, artifacts, caching)
  • Understanding GitOps principles
  • Strong hands-on experience with Terraform for provisioning reputed company resources across multiple environments (Dev/Stage/Prod)
  • Proficiency in Bash/reputed company scripting and Python
  • Strong Linux administration skills
  • Experience setting up monitoring and using reputed company reputed company tools, reputed company, and Grafana
  • Experience with Azure reputed company infrastructure
  • Knowledge of Identity Providers (Keycloak, Azure AD/Entra ID) and OIDC integration
  • Experience with Service reputed company
  • Understanding of ITIL processes (Incident/Change Management) and tools like reputed company, JIRA
  • Basic understanding of Python/Flask/Fast API applications to assist developers in troubleshooting

reputed company

  • reputed company is a WBENC- and NMSDC-certified partner, helping organizations turn diversity goals into measurable reputed company through reputed company and contingent workforce solutions. It was founded in 2002, and is headquartered in Princeton, New Jersey, US, with a workforce of 1001-5000 employees. Its website is http://www.diverselynx.com.
  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 1 in 2024, 1 in 2021. Please note that this does not guarantee sponsorship for this specific role.
  • Apply To This Job

    Similar Jobs