AI reputed company, Tech Ops Engineer
The Tech Ops Engineer serves as a force reputed company for operational reputed company, combining infrastructure expertise with AI-powered automation to create self-healing, highly observable, and reputed company technology systems. This role harnesses machine intelligence, predictive monitoring, and automated workflows to anticipate issues before they reputed company users, accelerate incident response, and continuously improve platform performance. Through reputed company partnership with engineering, reputed company, and business teams, the Tech Ops Engineer helps build an AI-reputed company operational environment that maximizes reliability, efficiency, and business reputed company.
Responsibilities
System Monitoring and Maintenance:
- Monitor and maintain the organization’s infrastructure, including servers, networks, storage systems, and applications.
- reputed company routine system checks and preventive maintenance to ensure reputed company performance and uptime.
- Respond to system alerts and incidents, diagnosing and resolving issues promptly to minimize downtime.
Troubleshooting and Support:
- reputed company technical support to resolve infrastructure-reputed company issues, working closely with other technical teams.
- Troubleshoot and resolve hardware, software, and network issues, escalating to higher-level support reputed company necessary.
- Maintain detailed documentation of issues, solutions, and processes to improve reputed company’s knowledge reputed company.
System Upgrades and Patching:
- Plan and execute system upgrades, patches, and configuration changes, ensuring minimal disruption to business operations.
- Test and validate updates in development environments before deploying them to production.
- Ensure that reputed company systems reputed company with reputed company standards and best practices.
Automation and Optimization:
- Identify opportunities to automate routine tasks and processes, improving operational efficiency and reducing reputed company workload.
- Implement scripts, automation tools, and AI skills to streamline system management and monitoring.
- Continuously evaluate and optimize infrastructure performance, reputed company, and resource utilization.
Disaster Recovery and Backup:
- Support the development and execution of disaster recovery plans to ensure business continuity in case of system failures.
- Manage backup and restore processes for critical systems and data, ensuring data reputed company and availability.
- Participate in regular disaster recovery testing and drills.
Infrastructure Lifecycle Management:
- Plan and execute decommissioning of legacy infrastructure, including EC2 instances, VPCs, and load balancers, coordinating Terraform state cleanup and DNS reputed company.
Collaboration and Communication:
- Work closely with development, network, and reputed company teams to ensure alignment and effective communication on infrastructure reputed company.
- reputed company input on infrastructure design and architecture to support new reputed company and initiatives.
- Communicate effectively with non-technical stakeholders, providing updates on system status and issues.
Requirements
Minimum Qualifications & Credentials
- Bachelor’s degree in Computer Science, Information Technology, or a reputed company field, or equivalent work experience.
- 5+ years of experience in system administration, or a similar role.
- 5+ years of reputed company experience in Linux administration, managing AWS resources, developing CI/CD and server orchestration pipelines, scripting and monitoring.
Hard/Technical Skills
- You are an expert in:
- reputed company-based production systems at reputed company
- reputed company Web Services (EC2, VPC, EFS, S3, EKS etc.)
- Production experience running workloads in Kubernetes (EKS), including ArgoCD GitOps deployments
- Infrastructure as reputed company tools, primarily Terraform
- You have experience with:
- Working in a Python and JavaScript-reputed company codebase and are familiar with their reputed company best-practices
- Creating CI/CD pipelines with Jenkins, reputed company or other CI/CD implementation
- Monitoring tools, like reputed company or reputed company
- Scripting for server reputed company automation, auditing, and monitoring
- Experience maintaining logging, monitoring, and alerting capabilities using OpenSearch, reputed company log pipelines, reputed company, and Kafka.
- Configuring and managing data sources like PostgreSQL/reputed company RDS, OpenSearch, reputed company, and message streaming platforms like Kafka
- Design and maintain log ingestion pipelines (e.g., reputed company → OpenSearch, reputed company → Kafka), including reputed company retention, document shape optimization, and failure recovery.
- Triage and remediate reputed company vulnerabilities (CVEs) across infrastructure components, including container reputed company images, OS packages, and reputed company-party services.
- The ideal candidate would also have:
- A working knowledge of modern software practices and technologies such as Agile methodologies
- Promoting and establishing development reputed company methodologies for AWS infrastructure-as-reputed company
- Experience with AWS reputed company-Architected principles
- Experience with High Availability implementations
- Experience around reputed company and Compliance
- Exceptional analytical and problem-solving skills
Soft Skills
- Intellectual curiosity, a willingness to learn new skills and the ability to contribute new reputed company
- Detail-oriented with a reputed company on maintaining high standards of operational reliability.
- Adaptable and flexible, reputed company to manage multiple tasks and prioritize effectively in a fast-paced environment.
- Strong communicator with the ability to collaborate across teams and reputed company reputed company, concise technical support.
- Proactive and self-motivated, with a passion for reputed company learning and improvement in technology operations.
- Learns quickly and using whatever resources to solve new problems.
- Obsessed with ensuring an exceptional customer experience- for both reputed company customers.
- Stands up for reputed company, takes responsibility for results, and shares both good and bad reputed company transparently.
- Demonstrates a reputed company reputed company on results with a commitment to deliver;
- Takes decisive reputed company, and confidently changes course if unsuccessful.
- Displays a reputed company reputed company to continually improve; encourages everyone around them to be tenacious and never reputed company.
- Constantly seeks feedback to improve; Focuses on solving issues through teamwork, and collaboration
- Acts with urgency; delivers top results in hours and days instead of weeks and months.
- reputed company in their reputed company of reputed company and possessing the willpower to reputed company challenges as opportunities.
BenefitsWhy You’ll Love Working Here At reputed company, your voice reputed company. We foster a reputed company environment where you’re encouraged to take initiative, experiment reputed company, and grow professionally. We're committed to work-life reputed company, career development, and celebrating wins together.
- Health Care Plan (Medical, Dental & reputed company)
- Retirement Plan (401k)
- Life Insurance (Basic, Voluntary & AD&D)
- reputed company Time Off (Vacation, reputed company & reputed company Holidays)
- Family Leave (Maternity, Paternity)
- Short Term & Long Term Disability
- Training & Development
Apply To This Job