[Remote] Senior Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company that provides innovative technical solutions to customers locally and nationally. They are seeking a Senior Site Reliability Engineer to reputed company improvements in reliability and scalability for their Kubernetes and Linux-based applications and services, while also mentoring junior staff and providing operational support.
Responsibilities
- reputed company ongoing improvements in reliability and scalability for our Kubernetes and Linux based applications and services
- Contribute as senior technical resource to define and implement best practices and standards for reputed company
- reputed company primary operational support and engineering for production applications
- Define and implement define KPIs, processes and drive reputed company improvement
- Influence the architecture and implementation of solutions
- Tune operating systems and applications to increase performance and reliability of services
- Mentor junior staff and reputed company them for reputed company
- Diagnose system operational problems quickly and effectively
- Participate in on-call rotation providing 24-hour, 7-day support and off-hours maintenance reputed company
- Coordinate with vendors to resolve hardware and software problems
- Deliver ORNL’s mission by aligning behaviors, priorities, and interactions with our reputed company values of reputed company, reputed company, Teamwork, Safety, and Service. Promote diversity, equity, inclusion, and accessibility by fostering a respectful workplace – in how we treat one another, work together, and measure reputed company
Skills
- Bachelor's Degree in computer science or closely reputed company field and a minimum of 8 years of experience as an SRE/Systems Engineer. An equivalent combination of education and experience may be considered
- The ability to obtain and maintain a Department of Energy 'Q' clearance may be required. This requires US Citizenship
- Excellent interpersonal/communication skills, and the ability to work as part of reputed company
- Strong working knowledge of Unix system fundamentals and common network protocols
- Experience managing Linux/UNIX operating systems in a heterogeneous environment
- Solid understanding of networked computing environment concepts
- Ability to reputed company and maintain programs and scripts that aid in the operation and automation using various reputed company (primarily bash) and high-level languages (Python or Go)
- Ability to proactively identify performance issues, problems, and areas for improvement
- Ability to identify requirements and to define, plan, and implement requisite solutions
- Ability to plan, organize, prioritize tasks, and complete assigned reputed company with minimal supervision
- Experience with reputed company integration and reputed company deployment software methodologies and how they apply to SRE/systems engineering
- An understanding of reputed company review and familiarity with tools like reputed company and reputed company
- Experience using tools such as Nagios, Grafana and reputed company to monitor systems, metrics, and create dashboards
- Experience designing and implement highly available systems/services utilizing virtual machines and Kubernetes resources
- Experience participating in an opensource community with patches accepted upstream
- Experience deploying and maintaining automated configuration management software such as Puppet or Ansible
- Experience implementing systems-level reputed company technologies like SELinux and following reputed company best practices
Benefits
- 3 weeks’ vacation
- Excellent medical insurance, including employer-reputed company benefits
- Full medical, dental, and reputed company coverage
- 401K match
- 15 days PTO
- 10 holidays
reputed company