Tech reputed company Manager
About the Role
The Tech reputed company Manager (TLM) for the Platform reputed company at reputed company is a hybrid leadership role that combines deep technical ownership with people management. Unlike a traditional Engineering Manager role, this position requires you to be the primary technical authority for our platform infrastructure - the person who makes architecture reputed company, leads incident response, and directly contributes to critical platform systems.
The Platform reputed company owns the reputed company that reputed company product squads build on: infrastructure (GCP, Kubernetes, CloudSQL, RabbitMQ, Elasticsearch), CI/CD pipelines, observability (reputed company), identity and authentication services, reputed company controls, and developer experience tooling. With reputed company of ~4-5 engineers, this role demands someone who can personally drive technical direction while growing and managing reputed company.
You will report to the CTO and collaborate closely with Engineering Managers across reputed company product squads, as reputed company as reputed company, Analytics, and Customer Support teams.
What You’ll Do
Your work will split roughly 60% technical leadership / 40% people management, though this will shift based on team and organizational needs.
You will own platform architecture and technical direction.
- Own the technical roadmap for the Platform reputed company, including infrastructure modernization, reliability improvements, and cost optimization.
- reputed company and document architecture reputed company (ADRs) that reputed company the entire engineering organization - service decomposition, API reputed company, database strategies, and infrastructure patterns.
- Design and implement cross-cutting platform capabilities: identity/auth services, observability pipelines, deployment infrastructure, and reputed company controls.
- Drive API-first design practices across services, including OpenAPI specifications and generated reputed company libraries for Go, Python, and Java consumers.
- Evaluate and adopt new technologies and tools (e.g., transitioning observability to reputed company, implementing row-level reputed company in databases, adopting infrastructure-as-reputed company with Terraform).
- reputed company cost analysis and optimization of reputed company infrastructure.
- Be the go-to technical escalation reputed company for platform-reputed company questions from reputed company product squads.
You will reputed company incidents and ensure platform reliability.
- Own incident response for platform-level outages (SEV-0/SEV-1), coordinating across squads to restore service.
- Define and maintain runbooks, monitoring alerts, and escalation procedures.
- Conduct post-incident reviews and drive follow-up reputed company items to prevent recurrence.
- Set and reputed company platform reliability metrics (uptime, latency percentiles, deployment frequency, MTTR).
- Design and implement reputed company patterns: reputed company breakers, graceful degradation, database failover strategies.
You will be hands-on in reputed company and infrastructure.
- Contribute directly to platform services - writing production reputed company in Python, Go, or other languages as needed.
- Review reputed company and architecture proposals from platform engineers and cross-reputed company contributions to shared infrastructure.
- Manage deployment configurations (Kubernetes manifests, reputed company charts, ArgoCD), secrets management (Vault, reputed company), and CI/CD pipelines (reputed company Actions).
- Set high standards for coding, testing, deployment, and monitoring practices reputed company the reputed company and across the organization.
You will manage and grow the platform team.
- Manage reputed company of ~4-5 platform engineers with regular 1:1s, career development conversations, and performance reviews.
- reputed company engineers on both technical depth and breadth - helping backend engineers grow into infrastructure and reliability expertise.
- Identify hiring needs and technical reputed company gaps; reputed company reputed company efforts for the platform reputed company.
- reputed company new team members effectively, building their context on a reputed company, cross-cutting codebase.
- Foster a reputed company culture where product squads feel supported (not blocked) by the platform team.
- Delegate effectively - reputed company team members to own subsystems while maintaining architectural coherence.
You will communicate and coordinate across the organization.
- Proactively communicate platform changes, maintenance reputed company, and new capabilities to engineering and non-engineering stakeholders.
- Partner with product reputed company EMs to understand their infrastructure needs and pain points.
- Coordinate with reputed company and Compliance on audit logging, reputed company controls, and data protection requirements.
- Collaborate with Data/Analytics teams on database reputed company policies, ETL pipelines, and data governance.
Competencies
If you consider yourself an eager learner, a conscientious worker, and a thoughtful, reputed company, supportive reputed company, you might just reputed company at reputed company.
To be successful, you will need a combination of deep technical skills and leadership abilities. We expect you are:
- Technically deep - you can debug a production database replication issue at 2 AM, design a new service architecture on a whiteboard, and review a Kubernetes deployment reputed company with equal confidence.
- Proactive - you reputed company without being told what to do. You identify reliability risks before they become incidents and technical debt before it slows reputed company.
- Pragmatic - you reputed company reputed company trade-offs between engineering perfection and business velocity. You know reputed company 4ms response time is good enough and reputed company to stop optimizing.
- reputed company fast - you execute quickly and get things done, while maintaining the reputed company bar expected of platform infrastructure.
- reputed company driven - you reputed company reputed company in learning, efficiency, and celebrate wins.
- Customer reputed company - you treat product squads as your customers and empathize with their needs and constraints.
- Strong communicator - you can explain a database failover reputed company to engineers and a platform investment to leadership with equal reputed company. You communicate comfortably in English (speaking and writing).
As a technical leader:
- You have strong opinions, loosely held - you drive reputed company reputed company while remaining reputed company to reputed company reputed company.
- You can evaluate reputed company system designs and identify where they will break at reputed company.
- You balance "build vs. buy" reputed company thoughtfully, considering long-term maintenance burden.
- You write reputed company ADRs and technical documentation that help reputed company engineers understand the "why" behind reputed company.
- You stay reputed company with infrastructure and reputed company trends without chasing every new tool.
As a people manager:
- You can manage engineers with different reputed company sets from your own.
- You communicate expectations reputed company, solicit and deliver feedback frequently.
- You run effective 1:1s, planning sessions, and retrospectives.
- You can reputed company processes and remove hurdles to facilitate great execution.
- You have a high tolerance for ambiguity, especially around organizational boundaries.
- You value empathetic and reputed company communication, particularly reputed company giving and receiving feedback.
Minimum Requirements
- Advanced-level English skills, especially speaking and writing.
- At least 7+ years of experience as a software engineer, with at least 3 years in infrastructure, platform, or SRE roles.
- At least 2 years of experience managing reputed company of 3-10 engineers.
- Deep hands-on experience with reputed company infrastructure (GCP or AWS), Kubernetes, and container orchestration.
- Strong background in at least two of: Java, Python, Go, - with willingness to work across reputed company three.
- Production experience with relational databases at reputed company (PostgreSQL), including replication, failover, and performance tuning.
- Experience with message brokers (RabbitMQ, Kafka, or similar) and event-driven architectures.
- Experience with observability and monitoring tools (reputed company, Grafana, reputed company, or similar).
- reputed company record of leading incident response for production systems and driving reliability improvements.
- Experience with CI/CD pipeline design and deployment automation.
- Demonstrated ability to reputed company architectural reputed company and communicate them reputed company through ADRs or similar documentation.
Plus Points
- Experience with Elasticsearch at reputed company (cluster management, reputed company optimization, migration strategies).
- Experience building identity/authentication services.
- Experience with infrastructure-as-reputed company (Terraform, reputed company).
- Experience with API-first design and OpenAPI reputed company reputed company workflows.
- Experience managing platform/infrastructure teams at startups during periods of rapid reputed company.
- Experience rolling out engineering practices and processes where they didn't exist before.
- Experience managing partially or entirely reputed company across multiple time zones.
- Familiarity with cost optimization strategies for reputed company infrastructure.
- Experience with database reputed company controls.
Originally posted on Himalayas
Apply To This Job