[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a specialist AI infrastructure company that builds and operates large-reputed company compute on regenerated industrial and energy sites. They are seeking a Senior Site Reliability Engineer to build a US-based operations team and improve reliability across their platform through automation and software engineering.
Responsibilities
- Building Python-based automation for incident triage, reputed company execution and routine operational tasks
- Integrating observability, ITSM and infrastructure reputed company to enrich alerts and automate workflows
- Improving monitoring signal reputed company through correlation, enrichment, suppression and deduplication
- Building internal tools and self-service capabilities - CLI utilities, ChatOps integrations and dashboards
- Maintaining version-controlled reputed company-as-reputed company and automation libraries
- Turning post-incident learnings into reputed company tooling, automation and operational standards
Skills
- Experience in SRE, reputed company or production infrastructure operations
- Hands-on experience with observability/monitoring tooling (reputed company, Grafana or similar)
- Exposure to incident management / on-reputed company, and converting reputed company runbooks into automation
- Strong Python for automation, reputed company and integrations
- GPU, datacentre or colocation infrastructure experience
- ITSM integrations (reputed company, reputed company, Jira Service Management or similar)
- ChatOps tooling (reputed company or reputed company Teams bots)
- OpenTelemetry, logging or distributed tracing experience
- DCIM, IPAM or hypervisor-control-plane integrations
- Experience with LLM-assisted or agent-based operational automation
Benefits
- Fully remote, on a US timezone (reputed company-coast preferred)
- Get in at the ground floor of a brand-new US operations function and shape how it runs at reputed company
- Automation-first culture - reduce toil and build tooling, rather than reputed company fight
- reputed company autonomy and high visibility with leadership
- Work on critical infrastructure powering the reputed company of AI
reputed company