TalentReach is hiring a Senior Site Reliability Engineer for an innovative company building infrastructure for large-scale simulation workloads in autonomous systems testing. The role involves designing and operating multi-region AWS cloud infrastructure, managing Kubernetes clusters, leading incident response, implementing security governance, and improving developer experience. This is a hands-on position with broad ownership of complex technical challenges including GPU scheduling, Windows workloads on Kubernetes, and large-scale batch simulation.
Senior Site Reliability Engineer
Seattle, WA or Vancouver, BC (Hybrid)
Full-Time
•TalentReach is hiring on behalf of an innovative company building the infrastructure that powers large-scale simulation workloads for autonomous systems testing and validation.
•We are seeking a
What You'll Do
Infrastructure Ownership & Cloud Operations
•Design, build, and maintain multi-region AWS infrastructure using Terraform
•Operate and scale Amazon EKS clusters across production environments
•Manage autoscaling, node lifecycle management, and workload health
•Design and maintain networking infrastructure, including VPCs, DNS, load balancing, and cross-region connectivity
•Support infrastructure expansions, migrations, and deployment into new regions
•Contribute to GitOps deployment workflows using GitHub Actions, Helm, and Kustomize
Reliability Engineering & Incident Response
•Build and improve incident management processes, including severity definitions, escalation paths, and on-call procedures
•Lead incident response, debugging, and root cause analysis efforts
•Write postmortems and implement systemic reliability improvements
•Enhance observability through metrics, logging, tracing, and dashboards
•Support GPU-enabled and large-scale batch workloads running on Kubernetes
Security & Access Management
•Provide security-focused guidance on platform architecture decisions
•Manage cloud IAM governance, including roles, policies, and access boundaries
•Support audit readiness, compliance initiatives, and customer security requirements
•Assist with partner certifications and customer security questionnaires
Platform Tooling & Developer Experience
•Improve CI/CD pipelines and infrastructure validation processes
•Support engineering teams with infrastructure troubleshooting and performance optimization
•Build tooling and automation using Python and Bash
•Contribute wherever needed in a collaborative startup environment
What You'll Bring
•5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering roles
•Experience operating large-scale production systems across multiple cloud regions
•Strong Terraform expertise, including modules, state management, and multi-environment patterns
•Deep AWS experience, including VPC, IAM, EKS, S3, and CloudWatch
•Strong Kubernetes expertise, including cluster operations, autoscaling, RBAC, and Helm
•Experience with GitOps and CI/CD tools such as GitHub Actions, ArgoCD, or similar platforms
•Strong networking fundamentals, including CIDR, DNS, VPNs, load balancing, and cross-region networking
•Experience with observability platforms such as Prometheus and Grafana
•Proficiency with Python and Bash for automation and tooling
•Working knowledge of both Linux and Windows environments
•Strong ownership mindset with the ability to balance speed, reliability, and complexity
Preferred Qualifications
•Experience supporting Windows workloads on Kubernetes
•Familiarity with GPU scheduling and NVIDIA device plugin configuration
•Experience supporting simulation, machine learning, or rendering workloads
•Exposure to AWS Storage Gateway, Active Directory integrations, or AWS Transfer Family
•Familiarity with service mesh or service proxy architectures
•Experience with container-optimized operating systems such as Bottlerocket or Packer
•Experience optimizing cloud infrastructure costs at scale
TalentReach is committed to fostering an inclusive and equitable workplace. We welcome applicants of all identities and backgrounds and provide equal employment opportunities without regard to race, color, religion, sex, gender identity or expression, sexual orientation, national origin, age, disability, veteran status, or any other protected characteristic. We believe diversity strengthens our team and the communities we serve.
Resume Intelligence
Assess your experience against this job and tailor your bullet points instantly.
Sign in required
Please sign in to assess your resume alignment and generate custom copyable work bullets.