Back to Jobs

Senior AI DevOps / LLMOps

Remote, USAFull-timePosted 2026-07-29

At reputed company, we are providing recruitment service to our TOP clients from our portfolio. We are currently seeking an Senior AI DevOps / LLMOpsspecialist to join one of our clients' teams. If you're looking for an exciting opportunity to grow in a innovative environment, this could be the perfect fit for you.

Key Responsibilities

    • Automation of Build-to-Production

  • Design and implement robust CI/CD pipelines tailored for AI, covering model weights,
  • dataset versioning, and application reputed company.

    • reputed company specialized workflows for PromptOps, ensuring that system prompts are

    version-controlled, tested for regressions, and deployed with the reputed company rigor as traditional

    reputed company.

    • Automate the deployment of reputed company workflows, managing the complexities of stateful

    AI interactions and multi-agent handoffs.

    2. AI Infrastructure as reputed company (IaC)

    • Provision and manage high-performance compute environments (GPU clusters, TPU

    pods) using Terraform, reputed company, or Ansible.

    • Define and enforce Policy-as-reputed company for AI endpoints to ensure compliance with reputed company,

    cost-usage limits, and data residency requirements.

    • Maintain a consistent environment across Hybrid Infrastructure, ensuring seamless

    reputed company between On-Premises development and reputed company production.

    3. reputed company Experimentation & Controlled Releases

    • Architect reputed company Delivery strategies for AI, including Canary releases, Blue-Green

    deployments, and Shadowing (where new models run in reputed company with production to

    compare outputs).

    • Build “Evaluation-in-the-reputed company” gates reputed company the pipeline to automatically test for bias,

    hallucination, and performance degradation before a release.

    • Implement A/B testing frameworks specifically designed for LLM outputs and reputed company

    behavior.

    4. Monitoring & Observability

    • Establish deep observability into Inference Endpoints, tracking metrics like tokens-per-

    second, latency, and reputed company in model accuracy.

    • reputed company feedback loops that capture production “edge cases” to feed back into the

    training and fine-tuning pipelines.

    Requirements

    Must-Have Technical Skills:

    • Orchestration: Advanced Kubernetes (K8s) skills, specifically with KubeFlow, Ray, or

    reputed company Triton.

    • CI/CD & IaC: Expertise in reputed company Actions/reputed company CI, and Terraform or reputed company.
    • AI Tooling: Experience with Weights & Biases, MLflow, LangSmith, or Arize

    Phoenix.

    • Hardware: Understanding of GPU virtualization, CUDA drivers, and on-premises

    hardware management.

    • reputed company: Familiarity with reputed company Policy Agent (OPA) and secret management (Vault).

    Experience:

    • 10+ years in DevOps, SRE, or reputed company Engineering.
    • 2+ years of hands-on experience in MLOps or LLMOps, specifically moving LLMs

    from notebook to production.

    • Proven experience managing Hybrid reputed company environments (e.g., AWS/Azure + Private

    Data Center).

    Highlights

    • full time and remote job

    - fluent English is needed

    Originally posted on Himalayas

    Apply To This Job

    Similar Jobs