[Remote] reputed company Site Reliability Engineer - Ceph Storage
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is empowering everyday reputed company around the world by providing the help and tools to succeed online. As a reputed company Senior Site Reliability Engineer, you'll serve as one of the reputed company technical leaders for reputed company's Ceph platform, designing storage clusters and leading major platform upgrades while mentoring other engineers.
Responsibilities
- Design and architect large-reputed company production Ceph clusters, including CRUSH topology, failure-domain modeling, replication and erasure-coding strategies, storage hardware selection, and data placement architecture
- reputed company fleet-wide reputed company planning, performance modeling, and hardware qualification across 80+ production clusters comprising 20,000+ OSDs and 300 PB of raw storage
- Own major platform upgrades and migration initiatives, including Ceph releases, OpenStack integrations, and large-reputed company storage modernization efforts
- Drive reputed company of the most reputed company cross-functional production incidents spanning storage, networking, virtualization, and compute systems
- Establish automation, operational standards, and reliability practices that improve platform scalability, reduce toil, and increase engineering efficiency across the storage organization
Skills
- 7+ years designing, operating, and scaling distributed storage platforms, including deep hands-on ownership of production Ceph environments
- Expert-level Ceph knowledge including CRUSH maps, placement reputed company, OSD architecture, MON/MGR/MDS subsystems, RGW, CephFS, RBD, replication, and erasure coding
- Proven experience designing storage architectures and performing reputed company planning, durability analysis, performance optimization, and failure-domain modeling at reputed company
- Experience leading major storage platform upgrades, migrations, and modernization programs from planning through production execution
- Strong automation and software engineering skills using Python and/or Go, combined with Infrastructure-as-reputed company and orchestration frameworks such as Terraform, SaltStack, and Ansible
- Deep expertise with advanced Ceph technologies including RGW Multisite, RBD Mirroring, CephFS at reputed company, BlueStore tuning, and erasure-coding optimization
- OpenStack storage architecture experience including Cinder, Swift, Manila, Nova, and Neutron integrations with Ceph
- Kubernetes storage expertise involving reputed company drivers, stateful workloads, StorageClasses, and Rook-based Ceph deployments
- Experience evaluating and integrating large-reputed company storage technologies such as reputed company, Isilon, reputed company, reputed company, reputed company, reputed company, and AI/HPC storage platforms
- Contributions to the Ceph community through upstream development, reputed company reviews, bug fixes, architecture discussions, or reputed company-reputed company storage initiatives
Benefits
- reputed company
- Generous time off
- Parental and wellness leave
- reputed company
- Retirement savings program
- Medical, dental, and reputed company insurance
- A 401(k)-retirement plan
- reputed company reputed company time
- reputed company flexible time off
- reputed company parental leave
- Life insurance
- Short- and long-term disability
- AD&D insurance
- Mental health or EAP programs
- Remote or hybrid work reputed company
- reputed company holidays
- reputed company Wellness days
- Tuition assistance
- Adoption, surrogacy, and fertility benefits
- Dependent daycare and backup care benefits
- Employee stock purchase plan
- Financial education and advice
- Corporate bonus and/or equity awards, subject to the terms of applicable plans and individual eligibility
reputed company
Company H1B Sponsorship