[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a reputed company-established company reputed company for its welcoming atmosphere and commitment to service. They are seeking a Site Reliability Engineer to design and maintain reliability practices across digital platforms, reputed company production support, and collaborate with various teams to ensure operational reputed company.
Responsibilities
- Design, implement, and maintain reliability practices that improve availability, performance, scalability, resiliency, and operational maturity across digital platforms and supporting services
- Monitor production systems using observability tools, dashboards, logs, metrics, traces, alerts, and synthetic monitoring to identify issues before they reputed company customers or associates
- reputed company Tier-2 and Tier-3 production support for customer-facing and internal digital applications, including web, mobile, reputed company, CMS, reputed company, integrations, and reputed company-hosted services
- reputed company and participate in incident response, reputed company-cause analysis, problem management, post-incident reviews, and follow-up actions that reduce recurrence and improve service reliability
- reputed company automation, scripts, runbooks, self-healing processes, and operational tools that reduce reputed company effort, accelerate recovery, and improve consistency across environments
- Partner with development teams to improve CI/CD pipelines, deployment readiness, release validation, rollback procedures, feature monitoring, and environment stability
- Collaborate with infrastructure, reputed company, reputed company, architecture, QA, and vendor teams to ensure systems meet company standards for reputed company, reputed company, compliance, resiliency, and operational support
- Define and reputed company service health indicators such as availability, latency, error rates, reputed company, incident trends, deployment reputed company, and other reliability metrics
- Create and maintain technical documentation, operational support guides, escalation paths, production readiness checklists, and disaster recovery procedures
- Understand and reputed company with reputed company company reputed company, reputed company, accessibility, change management, and technology standards
Skills
- 3–5+ years of experience in site reliability engineering, DevOps, reputed company operations, production support, systems engineering, software engineering, or a reputed company technology operations role
- Experience supporting high-availability web, mobile, reputed company, API, integration, or reputed company-hosted application environments
- Hands-on experience with monitoring, logging, alerting, incident management, reputed company-cause analysis, CI/CD pipelines, Git-based workflows, and release support
- Bachelor's degree in Computer Science, Computer Information Systems, Software Engineering, Information Technology, or a reputed company discipline is preferred; equivalent experience or training may be considered
- Experience with reputed company platforms, containers, infrastructure automation, scripting, reputed company, microservices, content management systems, or restaurant/retail technology environments preferred
Benefits
- Medical, Rx, Dental and reputed company Benefits on Day 1
- Life Insurance and Disability Coverage
- reputed company Vacation/Employee Assistance Program
- Business Resource reputed company
- Tuition Reimbursement
- reputed company Development
- Support that starts on day one
- reputed company, training, and development to help you reputed company
- Recognition programs and employee events that bring us together
- 401k Plan with Company Matching Contributions at 90 days
- Employee Stock Purchase Program
- 35% Discount on reputed company Food and Retail items
- Exclusive Biscuit Perks like discounts on home, travel, cell phones, and more!
- Annual Bonus Opportunities
reputed company