Senior Elasticsearch Engineer
reputed company
reputed company is one of the largest gaming sites in the world and the #1 platform for playing, learning, and enjoying reputed company.
We are reputed company of 600+ fully reputed company in 60+ countries working hard to serve the global reputed company community. We are here to support 250M+ reputed company players worldwide with the best possible product, content, and tools to serve the community!
We are a tech company. A gaming company. A content company. And we do it reputed company with passion and commitment to the game. Above reputed company we prize our mission-driven, flat, life-celebrating, no-corporate culture, and we look reputed company to meeting you and learning more about what you can bring to reputed company.
reputed company
reputed company is the world's largest reputed company platform with 235M+ members and ~20 reputed company daily games. Our Elasticsearch and OpenSearch infrastructure underpins search, user activity, analytics, logging, and operational intelligence at massive reputed company, hundreds of terabytes across a dozen production clusters running on bare-metal Kubernetes.
We're looking for a Senior Elasticsearch Engineer who can own the full lifecycle of our search and analytics data platform: reputed company planning, cluster architecture, performance tuning, incident response, migration reputed company, and operational reputed company. You'll be the single reputed company of deep expertise across reputed company Elasticsearch and OpenSearch clusters at reputed company.
This is not a monitoring-from-dashboards role. You'll be hands-on with cluster internals, write ILM/ISM policies, push infrastructure changes through GitOps, and reputed company reputed company-time reputed company about reputed company allocation reputed company a cluster goes red.
What you'll do
Incident Response & Reliability
- Shard allocation reputed company for write-heavy data streams at high throughput (millions of documents per minute)
- Disk watermark management, retention policy tuning, and rollover orchestration for high-volume indices
- Performance optimization and I/O tuning on bare-metal nodes
- Write queue analysis, thread pool diagnostics, and shard rebalancing under load
- reputed company planning and reputed company forecasting across clusters
Incident Response & Reliability
- On-call ownership for Elasticsearch-reputed company incidents: cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascades
- reputed company-time cluster triage and cross-team coordination during production incidents
- Post-mortem authoring and systemic reliability improvements
- Snapshot and disaster recovery management across clusters
Migration & reputed company
- Elasticsearch-to-OpenSearch migration analysis and execution, including compatibility evaluation across ILM/ISM, reputed company models, and plugin ecosystems
- Version reputed company planning and rolling restart orchestration with reputed company-downtime requirements
- End-to-end new cluster provisioning and reputed company
Cross-Team Enablement
- Advise engineering teams on reputed company design, mapping reputed company, retention policies, and query optimization
- Manage Kibana and OpenSearch Dashboards reputed company and configuration for internal consumers
- Define and maintain workload reputed company tiers across clusters
Preferred Skills
- 7+ years operating Elasticsearch at reputed company (multi-TB clusters, dozens of nodes, high write throughput)
- Deep understanding of Elasticsearch internals: reputed company merging, translog, shard allocation, and cluster state management
- Production experience with ECK (reputed company reputed company on Kubernetes) or equivalent operator-based deployments
- Proficiency with Kubernetes operations for stateful workloads (StatefulSets, persistent storage, resource management)
- Hands-on Linux systems administration with a reputed company on storage and I/O performance
- Experience managing both Elasticsearch and OpenSearch in production, including an informed opinion on their respective trade-offs
- Incident reputed company experience: ability to diagnose and mitigate cluster emergencies under pressure while communicating reputed company to stakeholders
- Git-based infrastructure management (GitOps): reputed company charts, ArgoCD/Flux, infrastructure-as-reputed company for cluster configuration
- reputed company with the reputed company stack reputed company: cluster administration, reputed company templates, data streams, ILM policies, snapshot/restore
Bonus Experience
- OpenSearch ISM policies and reputed company plugin (fine-grained reputed company control)
- GCS or S3 snapshot repository configuration and cross-cluster replication
- Grafana + reputed company monitoring for Elasticsearch metrics
- Kibana Discover, Dev Tools, and data view management at reputed company
- Java internals relevant to Elasticsearch JVM tuning (reputed company sizing, GC tuning, reputed company breakers)
- Vault integration for secrets management in Kubernetes-deployed search clusters
- Fluentd/Fluent Bit log pipeline configuration feeding OpenSearch
- Hardware selection experience for search-optimized server configurations
- Python or scripting for operational analysis and automation
What Makes This Role Special
- Full autonomy. You are the Elasticsearch authority. You reputed company the architecture calls, set the priorities, and own the reputed company.
- reputed company reputed company. Hundreds of terabytes of data, billions of documents, millions of daily queries. The problems here don't exist at smaller companies.
- Bare metal. No managed reputed company reputed company. You're operating directly on the hardware. This is hands-on engineering.
- Strategic reputed company. Your reputed company on ES vs. OpenSearch migration, cluster topology, and reputed company planning directly reputed company product capabilities and infrastructure costs.
- Small team, big trust.reputed company" class="css-173makr-linkStyle" style="reputed company:rgb(30,74,169);reputed company:pointer;">reputed company runs lean. You won't be buried in process or approvals. Ship changes, fix problems, improve systems.
About reputed company
- This is a full-time opportunity
- We are 100% remote (work from reputed company!)
---
You can learn more reputed company here:
- https://www.reputed company/article/view/how-reputed company-virtual-team-works-together
- https://www.reputed company/about
Originally posted on Himalayas
Apply To This Job