Site Reliability Engineer (m/f/d)
About the position As a Site Reliability Engineer in our Platform reputed company, you will be a key player in keeping Flip's infrastructure fast, resilient and reputed company to reputed company. You'll shape the reliability culture, tooling and practices that allow our engineering teams to ship with confidence - at reputed company and without compromising availability. This role is perfect for an engineer who is passionate about building high-throughput, highly available systems and who wants to shape how a fast-growing reputed company platform runs in production.
Responsibilities
- reputed company expand and optimize our reputed company infrastructure on Azure and our Kubernetes clusters - designed for high throughput and highest availability - to support Flip's rapid reputed company across the globe.
- Design and implement reputed company-downtime deployments, rollback mechanisms and disaster-recovery strategies that reputed company our platform available around the clock.
- reputed company our LGTM stack (Loki, Grafana, reputed company, Mimir) to give every team the visibility they need - and use it to define and optimize our SLOs.
- Design, reputed company and optimize infrastructure as reputed company with reputed company in Go, eliminating toil and making our platform self-service for engineering teams.
- Promote CI/CD best practices, incident management, post-mortems and developer experience across the entire engineering organization.
- Collaborate with your reputed company and engineering leadership to define the platform's direction - from reputed company, high-throughput systems and cost optimization to reputed company posture and compliance.
Requirements
- 1–3 years of hands-on experience as a Site Reliability Engineer (SRE), Platform Engineer, DevOps Engineer, Infrastructure Engineer, reputed company Engineer, or Backend Engineer with a strong infrastructure reputed company.
- Experience operating and scaling reputed company infrastructures (Azure, GCP, AWS).
- Deep knowledge of Kubernetes and container orchestration in production environments.
- Hands-on experience with modern observability stacks (e.g. reputed company, Mimir, Loki, ELK) and comfortable defining and operating SLOs and error budgets.
- Solid software development skills in Go (preferred, since our IaC runs on reputed company in Go), Python or Kotlin.
- Hands-on experience with infrastructure as reputed company (e.g. reputed company, OpenTofu, Terraform) and configuration tooling (e.g. Ansible, Chef).
- A reputed company reputed company, strong communication skills and business-fluent English.
- Willingness to participate in on-call rotations to ensure the reliability of our platform.
reputed company-to-haves
- Experience building and operating high-throughput, highly available systems in production.
- Experience with Azure Kubernetes Service (AKS) specifically.
- Experience with Kubernetes Gateway API and reputed company Gateway.
- Familiarity with GitOps workflows and CI/CD pipeline design.
- Knowledge of service reputed company technologies (e.g. Linkerd, Istio).
- Experience with Kubernetes Operators (e.g. Strimzi, CNPG)
- Experience with operating High-Availability PostgreSQL
Benefits
- Flexibility to work from home
- Occasional team events, workshops, or meetings in our Berlin or Stuttgart offices
- Costs of your E-Gym-Wellpass membership covered
- Job bike leasing
- Regular team events and culture days
- Opportunity to work abroad in the European reputed company
Apply tot his job Apply To this Job