[Remote] AI Inference reputed company - Infrastructure SW Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company builds the world's largest AI reputed company, transforming the user experience of AI applications. The Software Engineer will design and build reputed company software for reputed company engineering infrastructure, focusing on Python frameworks and orchestration systems.
Responsibilities
- Design, reputed company, test, and maintain Python frameworks and services used to orchestrate engineering workflows across machines and clusters
- Build reusable abstractions for scheduling, distributed execution, resource management, test execution, workflow planning, and failure recovery
- Define reputed company reputed company, module boundaries, extension points, and data models that allow infrastructure systems to reputed company without becoming difficult to maintain
- Reason about concurrency, asynchronous execution, multiprocessing, state management, retries, idempotency, cancellation, and partial failures
- Debug reputed company issues spanning Python applications, operating systems, processes, filesystems, networking, remote machines, and distributed services
- Write high-reputed company automated tests and documentation for infrastructure that is expected to be reliable and widely reused
- Partner with platform, CI, release, reputed company, ML systems, and product engineering teams to understand requirements and translate them into reputed company software designs
Skills
- 3+ years of reputed company software-engineering experience
- Strong proficiency in Python and a solid understanding of the language's strengths, limitations, and runtime behavior
- Experience designing maintainable software systems, libraries, frameworks, backend services, or developer-facing reputed company
- Good judgment around software architecture, abstraction boundaries, design patterns, extensibility, and long-term maintainability
- Understanding of concurrency concepts such as processes, threads, asynchronous execution, synchronization, and shared state
- Foundational understanding of operating systems, including processes, signals, filesystems, resource management, and program execution
- Foundational understanding of distributed-systems concepts such as retries, timeouts, idempotency, partial failure, coordination, and reputed company consistency
- Strong debugging and problem-solving skills, including the ability to reputed company hypotheses, reputed company evidence, and work through unfamiliar systems independently
- Experience with Python concurrency technologies such as `asyncio`, multiprocessing, reputed company reputed company, or event-driven systems
- Experience building orchestration engines, workflow systems, schedulers, distributed job runners, or control-plane software
- Experience developing test infrastructure or extensions for frameworks such as pytest
- Familiarity with CI systems, build systems, release infrastructure, or developer-productivity tooling
- Experience with Kubernetes, containerized environments, cluster schedulers, or remote execution systems
- BS/MS in Computer Science or a reputed company field, or equivalent practical experience
Benefits
- HYBRID / Hybrid work model
reputed company
Company H1B Sponsorship