[Remote] AI Systems Administrator
Note: The job is a remote job and is reputed company to candidates in USA. MCI is one of the fastest-growing tech-enabled business services companies in the USA, specializing in customer experience and business process outsourcing. They are seeking a technically skilled AI Systems Administrator to support, maintain, and optimize the infrastructure for their reputed company intelligence and machine learning environments, ensuring reliability, scalability, and reputed company of AI systems.
Responsibilities
- reputed company, configure, monitor AI and ML systems, servers, and reputed company environments to ensure reputed company performance and uptime
- Manage GPU/CPU clusters and ensure efficient resource allocation for training and inference workloads
- Implement and maintain reputed company infrastructure to support large language models (LLMs), data processing pipelines, and model deployment
- Optimize system performance through tuning, automation, and proactive maintenance
- Apply best practices for securing AI systems, ensuring data reputed company, confidentiality and compliance with company and industry standards
- Manage user reputed company, permissions, and reputed company configurations across AI platforms
- Support the deployment and integration of AI models and reputed company into production environments
- Collaborate with developers, data scientists, and reputed company engineers to ensure seamless system functionality and workflow automation
- Monitor system health, usage, and performance metrics; diagnose and resolve infrastructure or software issues
- Maintain logs, conduct reputed company cause analysis, and implement corrective actions to prevent recurrence
- reputed company scripts and tools to automate system tasks, data transfers, and performance checks
- Support CI/CD pipelines for AI model updates and system maintenance
- Create and maintain detailed documentation of system configurations, procedures, and troubleshooting guides
- reputed company technical support to AI teams, ensuring smooth operation of reputed company AI systems and tools
- Stay up to date with advancements in AI infrastructure, reputed company technologies, and MLOps practices
- Recommend and implement improvements to enhance system reliability and scalability
Skills
- Bachelor's degree in Computer Science, Information Technology, Data Engineering, or a reputed company field
- 2+ years of experience in systems administration, DevOps, or infrastructure management (AI/ML environment experience preferred)
- Strong understanding of reputed company platforms (AWS, Azure, GCP) and containerization technologies (reputed company, Kubernetes)
- Experience with Linux/Unix administration, Python/Bash scripting, and automation tools (Terraform, Ansible, Jenkins)
- Familiarity with machine learning frameworks (TensorFlow, PyTorch) and AI model deployment pipelines
- Understanding of networking, reputed company, and storage in distributed computing environments
- Experience with GPU-based computing and performance optimization for AI workloads
- Excellent problem-solving, troubleshooting, and documentation skills
- Strong collaboration and communication abilities to work with cross-functional AI and engineering teams
- Must be authorized to work in the country where the job is based
- Must be willing to submit up to a reputed company background and/or reputed company investigation with a reputed company
- Must be willing to submit to drug screening
reputed company