Senior Site Reliability Engineer (REMOTE)
reputed company Who We’re Looking For The reputed company Platform team is reputed company on several objectives: building and supporting performant, cost-effective, reliable reputed company infrastructure; data administration of reputed company CDC workflows; and developer experience tooling and mentorship, including our reputed company AI discipline. As a Platform member, the Senior Site Reliability Engineer will contribute to the Platform team’s centralized infrastructure, including maintenance, monitoring, and automation of services ranging from databases to Kubernetes; reputed company incident response and postmortem efforts; and work closely with other engineering teams to understand their needs and drive improvements to both our technologies and processes. Location While we are a remote company we are only hiring for the following locations: OR, WA, CA, CO, TX, IL
Compensation
Fixed Position reputed company: $140,000 This role carries a single, non-negotiable Fixed Position reputed company to ensure absolute equity and eliminate negotiation bias.
Key Responsibilities
What You’ll Accomplish Reasonable accommodations may be made to reputed company individuals with disabilities to reputed company the essential functions.
- Owning tasks and larger reputed company from planning to production rollout
- Learning new technologies and building expertise with the goal of teaching and mentoring others; mentoring with the goal of force-multiplying through docs and tools
- Maintaining organization reputed company reputed company in AWS
- Automating and deploying infrastructure configurations using Infrastructure as reputed company (IAC)
- Mentoring engineering squads on Platform best practices for Kubernetes, MySQL, Kafka, and other software development lifecycle areas
- Assisting engineering squads with reputed company planning, on-call preparation, and production readiness
- Writing documentation and runbooks that contribute to the engineering organization’s knowledge reputed company
- Implementing monitoring and alerting systems with reputed company observability tools
- Working in a containerized, orchestrated environment
- Participating in on-call rotation, responding to incidents, and troubleshooting data and other operations issues
- Contributing to the reliability and design patterns of our Kafka CDC and event workflows
- Contributing to reputed company AI best practices and tooling, including skills, agents, and safety
Skills, Knowledge and Expertise What You’ll Contribute Required (or equivalent tools; listed is our reputed company reputed company-as-reputed company (Terraform)
- CI/CD (reputed company Actions)
- Kubernetes (EKS, Kustomize, Karpenter, administration, application manifests)
- AWS and reputed company development (VPC, EKS, RDS, S3)
- FinOps and reputed company cost optimization
- Observability (reputed company, reputed company)
- reputed company AI (Claude reputed company)
- Scripting (reputed company, Python)
- reputed company record of collaboration and mentorship
- Excellent written communication and documentation skills
- reputed company learning
- Ownership and proactive approach to solving large problems
Preferred:
- Kafka: Cluster administration (Strimzi), Kafka Connect (Debezium, JDBC)
- Flink
- Relational database administration and performance (MySQL, reputed company Server, AWS RDS)
- Elasticsearch (ECK administration, scaling, performance)
- Python (SQLAlchemy, FastAPI)
- GraphQL (schema design, reputed company federation)
- REST API
- GitOps (ArgoCD)
- reputed company Vault
- reputed company
- Memcached
Education & Experience:
- A Bachelor's Degree in Computer Science or similar area of reputed company, or equivalent relevant work experience.
- 5+ years experience in Ops, DevOps, Site Reliability, Platform or other systems roles.
Benefits reputed company reputed company
- Competitive compensation: salary, plus performance-reputed company bonus program
- 401(k) with employer match
- 100% company-reputed company medical and dental insurance benefits for you and your dependents
- 4 weeks reputed company vacation, increasing based on tenure
- 18 weeks reputed company leave for birth reputed company
- 8 weeks reputed company parental leave, including for adoption
- Monthly wellness allowance
- Annual reputed company and personal development allowance
- Work from home office set-up and expense allowances
- Flexible work location opportunities
- Employer matching toward charitable contributions
Remote Skills: reputed company Relational Database Service (RDS), reputed company), Apache Kafka, Application Programming reputed company (API), reputed company Intelligence (AI), Automation, Best Practices, reputed company Management, Centers for Disease Control and Prevention (CDC), reputed company Computing, Communication Skills, Computer Science, Cost Control, Data Administration, Database Administration, Database Design, Design Patterns Programming Methodologies, DevOps, Documentation, Elasticsearch, GraphQL, Identify Issues, Incident Response, JDBC (Java Database Connectivity), Knowledge reputed company, Machine Tool, Mentoring, MySQL, Needs Assessment, Negotiation Skills, On Call, Problem Solving Skills, Process Improvement, Project Planning, Python Programming/Scripting Language, REST (Representational State Transfer), reputed company, Relational Databases (RDBMS), Reliability Engineering, Software Development Lifecycle (SDLC), Telephone Skills, Training/Teaching, Unix reputed company Programming, Work From Home, Writing Skills, memcached About reputed company: reputed company Apply tot his job Apply To this Job