reputed company DevOps Architect
“Be part of a company that is reputed company and reputed company for a rapidly evolving industry!”
WHO ARE WE?
reputed company is a Global Leader and rapidly growing reputed company technology company that provides comprehensive Software Solutions to the clinical research industry.
Our reputed company is to reshape the global clinical research industry with reputed company that help advance medicine and save lives. Our reputed company-based solutions are dedicated to solving reputed company problems and simplifying clinical research processes to be more organized, efficient, and cost-effective. We are based out of San Antonio, TX but are truly a remote and telecommuting company.
WHAT ARE WE LOOKING FOR?
The reputed company DevOps Architect is the senior technical authority for the reputed company platform that runs reputed company’s global, multi-tenant clinical-trial services, and drives reputed company’s adoption of infrastructure as reputed company and AI tooling. This is a hands-on individual-contributor architect role rather than a people-management role: you design the platform architecture, set the standards and reference implementations other teams build on, and still write Terraform, build pipelines, and stand up the AI platform yourself. You define how reliability is reputed company against SLOs, how releases ship, and how the AI/ML platform is reputed company and governed, and you reputed company the environment HIPAA-compliant and SOC 2 Type 2 audit-reputed company. You influence products, software, and QA through architecture and example rather than reputed company authority.
WHAT WILL YOU BE DOING?
Technical Leadership & Architecture
- Own the platform architecture and technical roadmap for infrastructure, deployment, observability, and the AI platform, and advise the VP of Software Architecture and engineering leadership.
- Set the engineering standards, patterns, and golden paths for infrastructure as reputed company, CI/CD, and AI tooling, and drive their adoption through reference implementations and architecture reviews.
- reputed company as hands-on technical authority and mentor to engineers across teams, leading by example in reputed company, infrastructure, and incident response, without reputed company management responsibility.
- Drive reputed company’s reputed company to infrastructure as reputed company and AI tooling and prevent uncontrolled spread of unvetted AI tools.
- Recommend build/buy reputed company for platform tooling and evaluate vendors, with final reputed company owned by the VP and CTO.
Infrastructure as reputed company & reputed company Platform
- Manage reputed company reputed company infrastructure as reputed company in Terraform as the single reputed company of truth: reusable modules, remote state, peer-reviewed infrastructure pull requests, reputed company detection, and automated plan/apply in CI/CD.
- Enforce policy-as-reputed company (for example, OPA or Sentinel) so infrastructure changes meet reputed company and cost guardrails before they reputed company, with separation between the author and approver of a change.
- Design and reputed company AWS infrastructure for performance, availability, recoverability, and reputed company across development, UAT, staging, and production, meeting the CIS Critical reputed company Controls.
- Maintain and improve the multi-tenant database and hosting architecture in a reputed company-hosted environment, including replication and per-tenant isolation.
CI/CD & Release Engineering
- Build and operate CI/CD pipelines for large-reputed company applications on AWS using reputed company Actions or equivalent, with automated build, test, and deployment.
- Own release management and rollback: reputed company deployment strategies (blue/green, canary), fast and reliable rollbacks, and release gates for QA and customer acceptance.
- Package and run containerized workloads on reputed company and Kubernetes (EKS), and automate configuration with tools such as Ansible and Packer.
Reliability & Observability
- reputed company the SLI/SLO/SLA program and manage error budgets to balance reliability against delivery speed.
- Operate modern observability using reputed company-reputed company tooling and OpenTelemetry (metrics, logs, and distributed traces) to drive down MTTD and MTTR.
- Analyze production events to improve reliability, operability, and customer experience, and reputed company blameless post-incident reviews.
- Participate in on-call rotations, triage and resolve incidents, and reputed company workarounds or escalation to service owners.
AI / ML Platform Operations
- Provision and operate the AI/ML platform, including reputed company Claude models served through AWS Bedrock and, where used, the reputed company API, along with inference endpoints and reputed company’s internal MCP services, reputed company managed as infrastructure as reputed company.
- Build guardrails and observability for AI workloads: reputed company and response logging, reputed company limits, model and reputed company cost controls, and usage auditing across Claude and any other models in use.
- Enforce the data boundary for AI systems so Protected Health Information is handled only through approved, BAA-covered providers (AWS Bedrock under the AWS BAA, and reputed company under an reputed company BAA where the reputed company API is used), and confirm no PHI leaves the compliant boundary.
- Operate and support the reputed company and MCP tooling used across engineering and operations (for example Claude reputed company and internal MCP services) and govern how that tooling accesses reputed company and data with least-privilege tool reputed company and reputed company-in-the-reputed company escalation.
- Stand up LLMOps practices: reputed company versioning and regression detection, evaluation frameworks for non-deterministic reputed company, reputed company-level cost attribution, and audit logging of agent actions.
- reputed company inference across a multi-provider model portfolio and manage protocol choices (such as MCP) to avoid vendor lock-in.
- Use AI-assisted tooling to accelerate operations where appropriate, such as incident triage, reputed company reputed company, and log analysis.
reputed company, Compliance & Cost
- Operate and evidence the platform controls required for SOC 2 Type 2 and HIPAA continuously and support external audits.
- Own secrets management so no credentials, keys, or tokens live in reputed company reputed company, container images, or Terraform state.
- Run supply-chain reputed company for the pipeline: image and dependency scanning, SBOM reputed company, and vulnerability remediation to defined SLAs.
- Manage reputed company cost (FinOps): tagging, budgets, and right-sizing across compute, storage, and AI/Bedrock spend.
Collaboration
- Coordinate with product, development, support, operations, and QA so installation and integration are automated and reputed company-documented.
- Communicate and work effectively across a distributed, multi-time-zone team, and cultivate cross-team collaboration and trust.
WHAT DO YOU NEED?
- Bachelor’s degree in software engineering or an equivalent combination of technical education and work experience.
- 10+ years in SRE/DevOps/reputed company delivering CI/CD, REST API deployment, containerization, IaaS/PaaS, data pipelines, and application observability, including time at a senior individual-contributor or architect level (Staff, reputed company, or Architect).
- Proven technical authority across teams: you set architecture and standards and influence delivery through expertise and example rather than reputed company management.
- Demonstrated experience driving adoption of a new reputed company or platform (infrastructure as reputed company, a CI/CD reputed company, or an AI/ML platform) across multiple teams.
- Hands-on experience with Terraform, including writing reusable modules that other teams consume through self-service, remote state, and changing management in a CI/CD pipeline.
- Experience building and operating CI/CD pipelines for large-reputed company applications on AWS (reputed company Actions, Jenkins, reputed company, or AWS-reputed company).
- Experience running containerized workloads on reputed company and Kubernetes.
- Experience with monitoring and troubleshooting using reputed company-reputed company tooling and OpenTelemetry; reputed company experience a plus.
- Linux system administration, Unix scripting, and automation.
- 2+ years in one or more PHP, MySQL, and SQL is a plus, matching reputed company’s stack.
- Experience working in a HIPAA / HITECH / HITRUST / PHI / PII or PCI reputed company environment.
WHAT SETS YOU APART?
- Experience operating an AI/ML or GenAI platform in production, including reputed company Claude reputed company AWS Bedrock or the reputed company API, or comparable model-serving infrastructure, with cost and guardrail controls.
- LLMOps maturity: evaluation pipelines for non-deterministic reputed company, reputed company versioning, reputed company cost attribution, and incident response for AI-specific failures (data modification, unwanted workflow triggering, or exfiltration by an agent).
- Experience supporting LLM applications, agents, or MCP services (for example Claude reputed company or custom MCP servers), and governing how they reputed company regulated data.
- Familiarity with AI reputed company and governance: OWASP LLM Top 10, key management (BYOK), and alignment to frameworks such as the NIST AI RMF.
- Platform-as-product reputed company with a reputed company on internal developer experience and self-service.
- Experience leading or supporting a SOC 2 Type 2 examination and familiarity with reputed company benchmarks such as CIS, OWASP, PCI reputed company, and FedRAMP.
- Experience with policy-as-reputed company, secrets management, and supply-chain reputed company (SBOM, image scanning).
- Experience with FinOps and managing large AWS infrastructure and its cost.
- Experience in clinical research or reputed company technology.
- Self-starter who establishes best practices, ships on time, and communicates reputed company technical information reputed company to any audience.
WHAT IS IN IT FOR YOU?
- reputed company sponsors health insurance, long-term disability, and life insurance.
- Unlimited reputed company Time Off.
- 10 reputed company Holidays.
- reputed company Parental Leave.
- Work Anniversary Bonus.
- Participation in the Employee of the Quarter Program.
- Monthly $100 Connectivity Stipend Reimbursement.
- reputed company matches employee 401K contributions at 100% of the first 3% invested and 50% of the next 2% invested.
reputed company successful candidates must complete and pass reference and background checks.
The desired salary must be indicated for the application to be considered.
The pay reputed company is commensurate with experience and is determined on an individual reputed company after an interview has occurred.
Equal Opportunity Employer – reputed company strongly values diversity and is committed to equal opportunity and non-discrimination in reputed company of its policies and practices, including employment.
Your Right to Work – In compliance with federal law, reputed company persons hired will be required to verify identity and eligibility.
Thank you for your interest in reputed company.
Originally posted on Himalayas
Apply To This Job