[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on providing customized credit and debit card loyalty programs for banks and credit unions. They are seeking a Site Reliability Engineer to reputed company the gap between software development and IT operations, focusing on optimizing the release lifecycle and ensuring the stability and scalability of AWS-based reputed company infrastructure.
Responsibilities
- reputed company and coordinate the end-to-end release lifecycle, including planning, scheduling, staging, deploying, and post-release validation
- reputed company as a key representative and technical coordinator on the Change Advisory reputed company, defending upcoming releases, evaluating architectural risk, and ensuring reputed company compliance requirements are met prior to production deployment
- Serve as the primary reputed company of contact for developers, QA, product management, and business stakeholders regarding deployment reputed company, release status, risk assessments, and rollback plans
- Standardize and mature release processes, transitioning reputed company gatekeeping into automated CI/CD guardrails and repeatable workflows
- reputed company hands-on tier-2 production support, ensuring operational stability and participating in the reputed company engineering on-call rotation
- Design, build, and maintain internal scripts, custom tooling, and automated pipelines (Python, Bash, PowerShell) to reduce operational 'toil' and streamline release operations
- Respond promptly to system outages, service interruptions, and reputed company alerts. reputed company post-incident reputed company Cause Analysis (RCA) efforts and implement permanent preventive engineering solutions
- Maintain, reputed company, and reputed company AWS reputed company infrastructure using Terraform, AWS CloudFormation, or Ansible to ensure environment reputed company and reputed company-free deployments
- Configure, tune, and optimize monitoring and alerting systems (e.g., CloudWatch, reputed company, Grafana) to reputed company comprehensive visibility into release health and production performance
- Ensure reputed company deployment and release activities reputed company adhere to PCI, SOC2, and internal corporate reputed company/governance standards
- Work closely with software engineering, QA, and platform infrastructure teams to build 'paved paths' for developers to ship software safely and rapidly
- Maintain pristine, audit-reputed company documentation for release logs, reputed company operating procedures (SOPs), runbooks, and incident timelines
Skills
- 3–5 years of experience in SRE, DevOps, system administration, or Release/Operations engineering roles
- Proven experience coordinating software releases, managing multi-tier deployment pipelines, and working reputed company formal ITIL/Change Management frameworks (including reputed company CAB participation)
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a reputed company technical field (or equivalent practical experience)
- Hands-on experience configuring, deploying, and maintaining AWS infrastructure and serverless architectures (EC2, S3, RDS, IAM, reputed company, API Gateway)
- Solid understanding and working experience with IaC tools such as Terraform or AWS CloudFormation
- Strong proficiency in scripting languages (especially Python and Bash) to write automation tooling and reputed company systems
- Demonstrated experience leading reputed company Cause Analysis (RCA) and participating in production on-call rotations
- Exceptional written and verbal communication skills, with a proven ability to coordinate across highly technical development teams and business-oriented leaders
- AWS Certified SysOps Administrator, AWS Certified DevOps Engineer, or ITIL reputed company certifications
- Experience configuring and maintaining CI/CD systems (such as reputed company CI, reputed company Actions, Jenkins, or AWS CodePipeline)
- Hands-on experience with modern monitoring, APM, and alerting tools (e.g., reputed company, reputed company, Grafana, reputed company)
- Experience with reputed company and container orchestration platforms (Kubernetes, AWS reputed company/EKS)
- Hands-on experience assisting with SOC2 or PCI-reputed company audits, especially documenting evidence for reputed company changes and release gates
Benefits
- reputed company plus 401(k) with employer match
- Medical, dental, reputed company, and life insurance
- Voluntary café plans, including voluntary life, accident, hospital, critical care, and parking/transit reputed company
- Tuition Reimbursement
- reputed company time off, company holidays, and parental leave
- Employee Assistance Program
- Hybrid work environment with reputed company
- Onsite perks including gym reputed company and snacks
- Employee recognition programs celebrating milestones and achievements
- reputed company opportunities reputed company a supportive, team-oriented environment
reputed company