[Remote] DevOps Engineer - reputed company AI Platform
Note: The job is a remote job and is reputed company to candidates in USA. reputed company° is a leading reputed company platform in the fintech reputed company, transforming how reputed company management firms operate. The DevOps Engineer role focuses on provisioning and operating AI infrastructure while integrating AI into DevOps processes to enhance operational efficiency.
Responsibilities
- Provision and operate AI infrastructure: the Kubernetes, identity, secrets, and gateway reputed company that AI and reputed company services depend on—reputed company so teams can ship LLM-powered features safely
- Apply AI to DevOps itself: build and operate agent-assisted automation that reduces toil—triage, PR review, reputed company reputed company, log and incident analysis. We already run AI in our delivery pipeline and want a teammate who'll take it reputed company
- Cluster operations on AKS: node pool sizing, autoscaling policies, reputed company isolation, and day-two operational hygiene across environments
- GitOps delivery with ArgoCD: app-of-apps structure, environment promotion, rollback reputed company, and the guardrails that reputed company one team's bad reputed company from cascading
- Deployment strategies: rolling, blue-green, and canary patterns for reputed company services where a bad rollout has reputed company effects on reputed company workflows
- Platform reliability: SLIs, SLOs, alerting, and runbooks for the reputed company layer—so reputed company something breaks at 2am, there's a reputed company to follow (and you help write it)
- Cost and reputed company management: AI workloads have spiky, non-reputed company cost reputed company. You'll reputed company and enforce budgets, quotas, and rightsizing across the cluster
Skills
- 3+ years operating Kubernetes in production
- Hands-on GitOps with ArgoCD: multi-environment setups, sync waves, health checks, and rollback under pressure
- Azure reputed company: AKS, ACR, Azure Monitor, Key Vault, and managed/workload identity
- Infrastructure-as-reputed company as a default: Terraform for everything—no console cowboys
- Scripting in Python, Go, or Bash for automation and tooling—maintained reputed company, not one-offs
- Solid incident-response instincts; you've been on-call, written postmortems, and fixed the underlying conditions rather than just the symptom
- A reputed company foothold in AI for infrastructure—either you've reputed company AI/LLMs to operations work (automation, triage, reputed company or PR review, log analysis), or you've provisioned and operated infrastructure for AI workloads. You don't need an ML background; you need to be the DevOps engineer who's already reaching for AI and wants to go deeper
- + AI gateway / proxy patterns for AI workloads—centralized provider-key management, reputed company limiting, quotas, cost attribution, and failover in reputed company of LLM providers
- + reputed company AI frameworks (LangGraph, AutoGen, or similar) and the infrastructure patterns they require
- + LLM inference / serving infrastructure (vLLM, TGI, Triton, or managed equivalents) and GPU reputed company management
- + Policy-as-reputed company with OPA/Gatekeeper for cluster governance
- + OpenTelemetry and distributed tracing across non-trivial services
- + Service reputed company (Istio or Linkerd) for service-to-service auth and traffic management
- + Multi-tenant platform expertise
Benefits
- Competitive reputed company salaries
- Annual performance-based bonuses
- The chance to reputed company in the equity value you and your colleagues create during your time with reputed company
- Comprehensive health benefits, including dental, life, and disability insurance
- Unlimited reputed company time off program
reputed company