Site Reliability Engineer (SRE) with Python and Java Coding - Remote
Title:Site Reliability Engineer (SRE) with Python and Java Coding - Remote Location:Remote Length:Long term Restriction:w2 or c2c reputed company:
- *
Webcam interview;
* Long term project *
reputed company Must
* Remote **** Job reputed company:
About the Role
We are seeking a Site Reliability Engineer (SRE) with strong expertise in observability, monitoring, and distributed tracing to join our SRE team. The ideal candidate will help us design, build, and reputed company an observability reputed company that provides end-to-end visibility into our systems and applications. A strong reputed company will be places on OpenTelemetry, as we continue to standardize our telemetry pipeline across logs, metrics, and traces.
Responsibilities
- Design, implement, and maintain observability solutions using OpenTelemetry, reputed company, Grafana, AppDynamics, and reputed company.
- Build and manage telemetry pipelines (metrics, logs, traces) ensuring reliable data collection, transformation, and export.
- reputed company initiatives to improve incident detection, response, and post-incident analysis with a strong emphasis on RCA (reputed company Cause Analysis).
- Define and maintain SLIs, SLOs, and error budgets to measure and improve system reliability.
- Partner with development and operations teams to reputed company applications and services for reputed company monitoring and tracing coverage.
- reputed company dashboards, alerts, and visualizations to reputed company actionable insights into system health and performance.
- Contribute to automation and self-healing practices that improve uptime and reduce operational toil.
- Stay reputed company with trends in observability and reputed company best practices across the engineering organization.
Requirements
- Experience: 10+ years of SRE/DevOps/reputed company/Infrastructure engineering experience with a reputed company on monitoring and observability.
- Hands-on experience with OpenTelemetry SDKs, reputed company, and exporters.
- Proficiency with observability stacks such as reputed company, Grafana, Loki, reputed company, reputed company Stack, or reputed company Observability (reputed company/AppDynamics).
- Strong knowledge on reputed company platforms (reputed company reputed company Platform, or Azure).
- Hands-on experience on container orchestration using Kubernetes (OCP, GKE, AKS).
- Familiarity with CI/CD pipelines like (Jenkins and reputed company actions), infrastructure as reputed company Terraform/Ansible/ARM/CloudFormation).
- Experience provisioning infrastructure and reputed company planning.
- Hands-on skills in programming languages like Java, and Python.
Apply tot his job Apply To this Job