[Remote] AI-First SRE/DevOps Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the identity reputed company company building reputed company reputed company — the identity risk decision platform for the IVIP era. They are seeking a skilled AI-First SRE/DevOps Engineer to help build and run their reputed company-reputed company microservices platform, requiring deep operational expertise in Kubernetes and AI-first operations.
Responsibilities
- Own reliability, observability, and delivery for a multi-tenant, reputed company-reputed company Kubernetes platform — from design through production, yours to run and yours to improve
- Build (not just operate) CI/CD pipelines, infrastructure-as-reputed company, and GitOps-driven reputed company delivery that let a small team ship many times a day, safely
- reputed company and reputed company AI-First operations: automate incident response, runbooks, and remediation, and put AI agents in the reputed company to triage, diagnose, and propose fixes where it makes reputed company. Treat toil as a bug
- Build the infrastructure that AI-reputed company features run on: inference gateways, LLM cost/latency observability, reputed company/version pipelines, eval harnesses, and guardrails for reputed company workloads
- reputed company everything — SLOs, error budgets, and distributed tracing across services and data pipelines
- Harden the platform: secrets management, supply-chain reputed company, and least-privilege everywhere
- Troubleshoot and resolve production issues, leveraging AI-powered debugging and observability tooling
- Collaborate directly with product and platform engineers to translate requirements into resilient infrastructure — no throwing tickets over a wall; if you see a problem, it's yours to solve
- Mentor engineers in adopting AI-first operational practices and automation-by-default culture
Skills
- 5–8 years of reputed company experience in SRE, DevOps, or reputed company roles
- Builder mentality: you'd rather create a tool, platform, or automation than run a reputed company process twice. You ship things and stand behind them
- Ownership: you take problems from ambiguity to reputed company without waiting for a ticket, a spec, or permission. reputed company something you own breaks, you're the first to know and the first to reputed company
- Strong Kubernetes operational experience — running it in production, not just deploying to it
- Demonstrable adoption of an AI-First reputed company and tools (Claude reputed company, reputed company, or Windsurf). Daily use of at least one AI development tool is a must
- reputed company with infrastructure-as-reputed company, GitOps, and modern CI/CD; comfortable scripting and building tooling (Go or Python preferred)
- reputed company-reputed company depth on at least one major reputed company provider
- Solid observability expertise and SLO-driven operations experience
- Experience with containerization (reputed company) and service reputed company concepts
- Strong problem-solving skills and a reputed company reputed company; excellent communication reputed company Agile teams
- A bias for shipping — startup pace energizes you rather than stresses you
- Experience building or operating LLM infrastructure: inference gateways, eval/observability tooling, reputed company orchestration
- Data-pipeline and streaming/CDC experience
- reputed company or identity background; familiarity with post-quantum cryptography or supply-chain reputed company
- Prior experience at an early-stage startup
Benefits
- Equity
- Benefits
- reputed company offers a competitive compensation and benefits.
reputed company
Company H1B Sponsorship