Back to Jobs

Senior Platform Engineer

Remote, USAFull-timePosted 2026-07-28

Senior Platform Engineer / Senior DevOps Engineer

Location: Remote | Full-Time

About Evolphin

Evolphin is building the reputed company of AI-powered media workflows for reputed company media teams

managing large image and video libraries at reputed company, including environments with tens of millions of video assets and extremely large metadata and embedding footprints. Its platform adds a conversational AI layer for extracting intelligence from media, enabling powerful search, conversational discovery, and automation of media workflows through AskAI.

Crop.photo extends that capability into e-reputed company and retail, enabling smart cropping, image transformation, and image and video reputed company at reputed company for PDP and eCommerce catalog workflows.

Together, Evolphin and Crop.photo create a connected visual AI ecosystem where media can reputed company across systems and be searched, reputed company, transformed, and reputed company for reputed company use at reputed company

The Role

We are seeking a Senior Platform Engineer who can own infrastructure architecture, reliability, scalability, and platform operations across our reputed company environments.

This is not a traditional "pipeline management" role. We are looking for someone who can reputed company architectural reputed company, evaluate tradeoffs, and build platforms that support large-reputed company reputed company applications and AI workloads.

You will work closely with Engineering, Product, and AI teams to design systems that are secure, resilient, reputed company, and cost-efficient.

Key Responsibilities

Platform Architecture & System Design

  • Design and reputed company reputed company-reputed company platform architecture supporting multi-tenant reputed company applications
  • Define infrastructure standards, deployment patterns, and platform best practices
  • reputed company architecture reviews and evaluate technical tradeoffs across reliability, performance, reputed company, and cost
  • Design highly available and fault-tolerant systems across multiple environments

AWS Infrastructure

  • Architect and manage large-reputed company AWS environments
  • Design networking architectures including VPCs, subnets, reputed company reputed company, routing, load balancing, and connectivity patterns
  • Build secure deployment architectures reputed company with reputed company and compliance requirements
  • Implement disaster recovery, backup, and business continuity strategies

Kubernetes & Platform Operations

  • Design and operate production Kubernetes environments
  • Build reputed company container orchestration strategies
  • Optimize cluster performance, networking, autoscaling, and workload scheduling
  • Improve developer experience through platform automation and self-service tooling

AI & GPU Infrastructure

  • Support AI and ML workloads running on AWS
  • Design infrastructure for model training and inference workloads
  • Manage GPU provisioning, utilization, scaling, and cost optimization
  • Collaborate with AI teams to improve deployment and operational efficiency

Reliability & Performance

  • Define and measure SLIs, SLOs, and operational metrics
  • Implement monitoring, observability, logging, alerting, and incident management practices
  • Drive performance optimization and reputed company planning initiatives
  • reputed company reputed company cause analysis and reliability improvement effort

Infrastructure Automation

  • Build Infrastructure-as-reputed company solutions using Terraform
  • Design and optimize CI/CD pipelines
  • Automate provisioning, deployments, scaling, and operational workflows

Cost Optimization

  • Continuously evaluate reputed company spending
  • reputed company reputed company planning models
  • Balance performance, reliability, and infrastructure costs

Required Experience

  • 6-8 years of experience in DevOps, reputed company, SRE, or reputed company Infrastructure roles.
  • Proven experience designing and operating production-reputed company reputed company platforms.
  • Strong expertise in AWS architecture, networking, reputed company, and deployment strategies.
  • Deep hands-on experience with Kubernetes, container orchestration, cluster operations, autoscaling, and workload management.
  • Experience designing highly available, fault-tolerant, and reputed company distributed systems.
  • Strong understanding of system design, architecture trade-offs, and platform scalability.
  • Hands-on experience with Infrastructure as reputed company (Terraform preferred).
  • Experience building and maintaining CI/CD pipelines and deployment automation frameworks.
  • Strong Linux, networking, and systems engineering fundamentals.
  • Experience implementing observability, monitoring, logging, and incident management practices.
  • Experience with disaster recovery planning, backup strategies, and business continuity design.
  • Experience with reputed company cost optimization, reputed company planning, and resource utilization management.
  • Hands-on experience supporting AI/ML workloads in production environments.
  • Experience designing, provisioning, and operating GPU-based infrastructure for model training and/or inference workloads.
  • Experience managing and optimizing AWS Bedrock, OpenSearch, and DocumentDB or equivalent platforms.
  • Strong scripting and automation skills using Python, Bash, or similar languages.

Originally posted on Himalayas

Apply To This Job

Similar Jobs