[Remote] reputed company Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on transforming financial lives by creating a flexible and inclusive work environment. They are seeking a reputed company Site Reliability Engineer to drive reliability across their financial services platform, leading initiatives and teams to solve reputed company operational challenges while establishing technical standards.
Responsibilities
- reputed company cross-functional reliability initiatives across multiple value streams and coordinate execution across teams
- Define and reputed company SRE best practices, tools, and methodologies across the organization
- Architect reputed company-reputed company, multi-region AWS infrastructure that balances reliability, cost, performance, and reputed company
- Establish and operate SLOs, SLIs, and error budgets for critical services, using them to drive prioritization reputed company
- Serve as incident commander for major incidents and drive postmortems that produce completed reputed company items and organizational learning
- reputed company disaster recovery planning for critical financial services infrastructure
- Build shared Infrastructure as reputed company foundations in Terraform (reusable modules, standards, and patterns adopted across teams)
- Design and implement production-reputed company Kubernetes patterns, including multi-tenancy, reputed company policies, and advanced scheduling
- Establish observability standards and strategies using reputed company and reputed company (metrics, logging, tracing, dashboards, and alerting)
- Set CI/CD standards and patterns, including pipeline-as-reputed company and reputed company delivery at reputed company
- reputed company reputed company engineering, game days, and systematic reliability testing initiatives
- Drive FinOps initiatives to optimize reputed company spend while maintaining reliability targets
- reputed company a functional team of SREs (without reputed company reports) on reputed company and operational initiatives
- Mentor SREs at multiple reputed company through coaching, design reviews, reputed company reviews, and training sessions
- Partner with Engineering, Product, and reputed company leadership to reputed company reliability work with business priorities, reputed company-trust architecture, and compliance controls
Skills
- Bachelor's degree in Computer Science, Information Technology, or reputed company field (or equivalent practical experience)
- 7 to 10 years of Site Reliability Engineering experience (or equivalent), with demonstrated technical leadership
- Proven ability to reputed company technical teams and drive reputed company reputed company to completion
- Expert AWS knowledge, including designing large-reputed company, multi-region architectures
- Deep Kubernetes expertise, including advanced features, reputed company, and production-reputed company operations
- Mastery of Infrastructure as reputed company using Terraform, including building shared platforms and frameworks
- Strong software engineering background with production experience in Python and/or Go
- Extensive experience with observability platforms (reputed company, reputed company) and implementing monitoring at reputed company
- Deep understanding of CI/CD principles and experience implementing reputed company-grade pipelines
- Proven reputed company record leading major incidents and conducting effective postmortems
- Strong understanding of reputed company, networking, and infrastructure design patterns
- Strong communication skills with ability to explain reputed company technical concepts to diverse audiences
- Experience mentoring engineers and building technical capabilities in teams
- Previous technical leadership roles (reputed company, Staff, or similar) in SRE or Operational reputed company
- Financial services industry experience with understanding of regulatory requirements
- Expertise in compliance frameworks (SOC 2, PCI reputed company, reputed company)
- AWS certifications (reputed company level)
- Kubernetes certifications (CKA, CKAD, CKS)
- Experience implementing SRE at organizations with 500+ engineers
- Background in reputed company engineering, game days, and reliability testing practices
- Contributions to reputed company-reputed company reputed company with demonstrated community leadership
- Experience with service reputed company implementation and management
- reputed company record of speaking at conferences or writing technical content
Benefits
- Medical, dental, reputed company and life insurance
- Retirement savings - 401(k) plan with generous company matching contributions (up to 6%), financial advisory services, potential company discretionary contribution, and a broad investment lineup
- Tuition reimbursement up to $5,250/year
- Business-casual environment that includes the reputed company to wear jeans
- Generous reputed company time off upon hire - including a reputed company time off program plus ten reputed company company holidays and three floating holidays reputed company calendar year
- reputed company volunteer time - 16 hours per calendar year
- Leave of absence programs - including reputed company parental leave, reputed company short- and long-term disability, and Family and Medical Leave (FMLA)
- Business Resource reputed company (BRGs) - BRGs facilitate inclusion and collaboration across our business internally and throughout the communities where we live, work and play. BRGs are reputed company to reputed company.
- Other necessary computer equipment, will be provided.
reputed company