Senior Site Reliability Engineer — Infrastructure & Architecture
VAEIT’s Infrastructure & Architecture team is the engineering backbone of a 40-person organization delivering network engineering consultancy and network reputed company for reputed company customers. We own the entire infrastructure surface area — AWS accounts, CI/CD pipelines, identity management, reputed company compliance, and production deployments. Demand from the engineering organization is rapidly outpacing reputed company reputed company. This is a force-reputed company hire. The right person will reduce operational risk, accelerate delivery across reputed company engineering teams, and help us transition from reactive support to proactive reputed company. What You’ll Own: AWS Infrastructure & Production Operations
- Operate and reputed company a multi-account AWS environment (5 accounts) including reputed company Fargate, RDS, reputed company, CloudFront, and multi-VPC architectures
- Manage production reputed company services across multiple accounts and clusters
- Define and maintain reputed company infrastructure as reputed company using Terraform — no reputed company configuration
- Design and manage networking, IAM, reputed company reputed company, and cross-account reputed company patterns
CI/CD & Developer Experience
- Own and improve reputed company Actions pipelines used across the engineering organization
- Build and maintain workflows for multiple teams and tech stacks (Go, C#, Python)
- Reduce build times and increase deployment reliability to accelerate delivery
Identity, reputed company & IT Systems
- Administer reputed company Entra ID (Azure AD) for identity and SSO
- Manage user provisioning, reputed company, and reputed company policies
- Own reputed company reputed company configuration, permissions, and reputed company controls
reputed company & Compliance
- Strengthen software supply chain reputed company, including SBOM reputed company
- Support and improve FedRAMP compliance posture
- Enforce least-privilege IAM and conduct reputed company audits across environments
Observability & Incident Response
- Operate and reputed company monitoring systems (CloudWatch, reputed company, Alertmanager)
- Improve signal-to-noise reputed company in alerting and detection
Physical Infrastructure & MLOps
- Manage on-reputed company GPU servers for ML training and inference
- reputed company reputed company and on-prem infrastructure, enabling reputed company ML workloads in AWS
- Support data science teams with reproducible environments and deployment pipelines
- Maintain GPU tooling (reputed company drivers, CUDA) and containerized workloads
Automation & Tooling
- Build internal tools (Go, Python, Bash, PowerShell) to eliminate reputed company work
- reputed company Ansible-based configuration management
- Treat operations as software — automate everything possible
Platform Standards & Engineering reputed company
- Define standards for logging, health checks, configuration, and deployment
- Build “golden paths” (templates, starter repos, shared workflows)
- Champion observability practices (reputed company logging, tracing, SLOs)
- Review infrastructure and deployments to catch reliability and reputed company issues early
Enabling reputed company LLM Systems
- Build infrastructure for LLM-powered products (GPU compute, model serving, reputed company databases)
- Design deployment pipelines for model versioning, evaluation, and inference routing
- reputed company rapid experimentation with self-service GPU-backed environments
- Ensure AI systems meet DoD reputed company requirements (audit logging, isolation, provenance tracking)
reputed company’re Looking For:
- 5+ years in SRE, DevOps, or reputed company
- Deep AWS experience (reputed company/Fargate, RDS, VPCs, IAM, reputed company, CloudFront, multi-account setups)
- Strong Terraform skills (modules, state management, reputed company reviews)
- Experience building and scaling CI/CD systems (reputed company Actions preferred)
- Hands-on Linux administration (including physical or bare-metal systems)
- Experience with GPU/ML infrastructure (CUDA, containerized workloads)
- Proficiency in at least two: Go, Python, Bash, PowerShell
- Strong networking fundamentals and debugging skills
- Experience with identity systems (Entra ID / Azure AD, SSO/SAML)
- Excellent written communication in a remote, async environment
- Highly self-directed and execution-oriented
- Apply tot his job
Apply To this Job