Summary
Design, build, and maintain the backend services, APIs, data-platform automation, and AI platform / agent systems behind our data products. This is a backend-heavy role involving distributed Python services, streaming LLM/agent runtimes on AWS, RAG pipelines, and the dbt/Airflow automation that powers a Snowflake-based lakehouse.
You will work inside the Developer Experience team, whose mandate is to build the tooling, paved paths, and automation that let internal data teams ship discoverable, governed, and consumable data products quickly and safely.
Platform Overview
- •Lakehouse and warehouse: Snowflake with Apache Iceberg tables in a medallion architecture (Bronze, Silver, Gold, Semantic Views)
- •Transformation: dbt (Cloud on Fusion) with data contracts, tag-driven governance, and automated project-evaluation checks
- •Orchestration: Airflow on AWS MWAA; DAGs and schedules authored and deployed through platform tooling and CI (metadata driven)
- •Developer tooling: an internal CLI and reusable utilities (dbt, Iceberg operations, Snowflake load, REST API based ingestions) that abstract the platform for data teams
- •Data quality and governance: Monte Carlo for data observability (monitors as code) and centralized governance (tagging, lineage, PII classification, cataloging)
- •AI layer: Agent/MCP/Skill deployment framework, RAG, semantic router/broker agents over data products and catalog metadata; conversational agents in Slack; model access via AWS Bedrock/AgentCore, centralized LLM Gateway, and Snowflake Cortex
- •Cloud and delivery: AWS (ECS Fargate, Lambda, API Gateway, DynamoDB, S3, IAM), Terraform IaC, GitHub Actions CI/CD, containerized builds, structured logging, metrics, and tracing
Responsibilities
Backend Services
- •Design, build, and maintain backend services in Python/Go to support Data Platform Services including Agent runtimes on AWS Bedrock AgentCore/ECS Fargate and event-driven Lambda handlers
- •Design and maintain abstraction layers for data domain customers via custom developed utilities
- •Maintain and service infrastructure using IaC (Terraform) for AWS accounts
AI and Agents
- •Build RAG pipelines and tool-calling AI agents for data products: retrieval, orchestration, grounding/citations, and evaluation
- •Integrate model providers with the LLM Gateway with response streaming, prompt caching, and structured outputs in an auditable way (OTeL)
- •Stand up eval harnesses, guardrails, and tracing so agent quality, latency, and cost are measurable and regressions are caught before release
Data Platform Automation
- •Build and maintain data-platform automation: dbt services, MWAA/Airflow orchestration, and tooling that makes data products discoverable, governed, and consumable
- •Extend the platform CLI and shared utilities that data teams use as their day-to-day interface to the platform
Identity and Security
- •Implement auth and identity: OAuth/OIDC flows (per-user 3LO, token vaulting, session binding), least-privilege IAM, and secrets management
- •Enforce multi-tenant isolation and per-user identity/RBAC across services and agents
Reliability and Delivery
- •Own service reliability and delivery: Terraform, GitHub Actions CI/CD, container builds, structured logging, metrics/tracing, alerting, and cost controls
- •Set technical direction: system and API design, code review, and mentoring
Required Skills
- •15+ years designing, building, and maintaining production backend services at scale
- •Expert-level Python, Go, or Java for server-side development; solid grasp of relevant frameworks and the WSGI/ASGI model
- •Service and API design: REST and/or gRPC, request/response and streaming patterns, pagination, versioning, idempotency, and backward-compatible contracts
- •Data layer: SQL and data modeling, query optimization and indexing, transactions, connection pooling; relational, warehouse (Snowflake), and NoSQL/key-value (DynamoDB) stores
- •Server-side patterns: caching strategies, background jobs/workers, queues and event-driven processing, rate limiting, retries/backoff, and timeouts
- •Performance and reliability: profiling, load handling, latency/throughput trade-offs, graceful degradation, and designing for failure
- •Observability: structured logging, metrics, distributed tracing, and debugging live production issues
- •OOP: encapsulation, abstraction, inheritance, composition, polymorphism; SOLID principles; design patterns applied pragmatically; strong domain modeling
- •Solid data structures and algorithms; ability to reason about time/space complexity
- •Concurrency and async programming (async/await, threading, event loops) and their failure modes
- •Testing (unit, integration, end-to-end) and testable design; Git and PR-based workflows; disciplined code review
- •Production AWS: ECS/containers, Lambda, IAM, API Gateway, DynamoDB
- •Infrastructure as code with Terraform; CI/CD (GitHub Actions or equivalent) and container builds
- •Distributed systems: horizontally scalable, resilient, and loosely coupled services; handling consistency, retries, idempotency, and partial failure
- •OAuth/OIDC, authn/authz, token handling, least-privilege access, multi-tenant isolation, secrets management
- •Building LLM applications in production (not research)
- •RAG: chunking, embeddings, vector search, hybrid search, reranking, grounding/citations, context-window management, retrieval evaluation
- •Agents: prompt engineering, tool use/function calling, structured outputs, single- and multi-step orchestration, prompt caching
- •Integration with model providers: Anthropic/Claude, AWS Bedrock/AgentCore, Snowflake Cortex; response streaming
- •AI quality and ops: eval harnesses, guardrails, tracing/observability, token/latency/cost optimization
- •AI security: prompt injection, data exfiltration, PII handling, per-user identity/RBAC enforcement
Nice to Have
- •dbt, Airflow/MWAA, Snowflake hands-on experience
- •MCP (Model Context Protocol) and/or applied RAG
- •Slack platform (Bolt, Socket Mode, Block Kit) or other real-time/conversational backends
- •Data observability (Monte Carlo) or data-quality/governance tooling
- •Fine-tuning/adaptation, semantic caching, or model routing/fallback
- •Experience adding AI capabilities to existing production systems
- •Iceberg/open table formats and lakehouse patterns
- •Tableau or other BI integration
Compensation
The typical base salary range for this position is $197,300 - $313,700 annually. In select cities within the San Francisco and New York City metropolitan area, the base salary range is $237,700 - $344,700 annually. The range represents base salary only and does not include bonus, equity, or benefits.