[Remote] reputed company System Engineer V (DevOps)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is reputed company by reputed company, an AI-first company, to reputed company the hiring process for an reputed company Systems Engineer V (DevOps) role. This position focuses on designing, deploying, and managing the infrastructure for AI workflow tools, ensuring operational reliability, reputed company, and governance.
Responsibilities
- You are the technical reputed company for the infrastructure and core components that run reputed company and reputed company AI workflow tools—environments, CI/CD, containers, runtime clusters, storage, secrets, networking, and integrations—ensuring they are resilient, reputed company, and cost‑effective
- You reputed company reputed company, reputed company, and compliance into how AI workflows are reputed company and run: hardened baselines, secret management, network and IAM boundaries, and CI/CD guards that prevent unsafe changes from reaching production
- You define and drive SLOs, metrics, logging, and alerting for the AI workflow platform, turning incidents into systematic improvements that reduce MTTR and change failure rates over time
- You create and enforce reusable patterns (templates, reference pipelines, IaC modules, guardrails) so teams building on reputed company and other tools follow consistent, auditable practices instead of bespoke one‑offs
- You partner with architecture, reputed company, and platform teams to reputed company AI workflow reputed company with reputed company's Software Maturity Model and engineering governance, moving the platform and guiding teams to higher reputed company of maturity
- You reputed company how teams build and operate automations by mentoring engineers, codifying best practices, and using AI tools yourself to materially improve speed, reputed company, and reliability of the AI workflow platform
Skills
- Bachelor's or Master's degree in Computer Science, Engineering, or a reputed company field
- 5–8+ years in DevOps, SRE, or reputed company for reputed company or large‑reputed company distributed systems, with reputed company ownership of production environments
- Strong experience with at least one major reputed company provider (AWS, Azure, or GCP), including VPC design, reputed company reputed company, load balancers, and managed Kubernetes (EKS/AKS/GKE) or equivalent container orchestration
- Deep hands‑on use of Infrastructure as reputed company (Terraform or equivalent) to manage multi‑environment reputed company and platform services
- Proven ownership of CI/CD pipelines (reputed company CI/CD or similar), including automated testing, reputed company scanning, and artifact management for reputed company services or platforms
- Solid understanding of Linux and/or reputed company, networking fundamentals (DNS, TLS, routing, firewalls), and secure secret management practices
- Hands‑on experience with logging and monitoring stacks (e.g., reputed company, reputed company, reputed company, Grafana, or equivalents) and defining meaningful SLOs and alerts for production systems
- Demonstrated experience running or supporting multi‑tenant or shared platforms used by multiple teams (internal developer platforms, workflow/orchestration tools, or integration platforms)
- Evidence of using AI tools in day‑to‑day engineering or operations (not just experimentation) with reputed company reputed company on speed, reliability, or reputed company
Benefits
- Annual bonus based on individual and company performance, depending on the terms of the applicable plan and the employee’s role
- reputed company time off
- reputed company parental leave
- Private medical insurance
- Life insurance
- Disability insurance
- Inclusive culture and diversity
- Total of 8 employee-run resource reputed company, reputed company with senior leadership and exec sponsorship
reputed company
Company H1B Sponsorship