Senior Data Engineer

Gravie
GlobalRemote
Full Time
Posted Yesterday
Apache Kafka*

Role Overview

We're building a near real-time streaming data platform for operational data across the business. We seek a Senior Data Engineer to both build/extend and operate it: you'll build the infrastructure and you'll own the platform in production - latency and throughput SLOs, backpressure under spiky load, replay and backfill, dead-letter triage, and connector failure recovery. This is a hands-on role for someone who wants to build, extend, and operate a live production platform, not just design it and hand it off.

You are self-driven, calm under production pressure, and comfortable owning streaming systems in a regulated environment.

Responsibilities

You Will:

Take full-lifecycle ownership of the streaming platform - architecture, implementation, production operations.

● Run the platform in production: own latency/throughput SLOs, monitoring and alerting (e.g. Datadog) - including replays, per-source backfills, connector-failure recovery, dead-letter triage, and tuning for spiky, batch-driven claims load.

● Build and extend streaming pipelines that ingest CDC events from operational databases and SaaS sources into canonical, contract-validated form.

● Transform and enrich in Spark (Structured Streaming or dbt-on-Spark micro-batch), including cross-stream joins that correlate events into unified lifecycle entities.

● Enable secure, governed data access over PHI: classification, row/column controls, and access policy applied as data is served to consumers (e.g. ABAC).

● Make it reliable and observable: idempotent/replayable pipeline design, data-quality validation, runbooks, and observability the broader data team can rely on.

● Provision as code: define the platform (streaming, processing, storage, and catalog services on AWS) in CDK with CI/CD for data pipelines, and right-size for cost against the latency SLO.

● Partner across teams: work with upstream producers on source changes and contracts, with downstream consumers on access and data needs, and with stakeholders to turn requirements into what the platform delivers.

● Demonstrate commitment to our core competencies of being authentic, curious, creative, empathetic and outcome oriented.


Requirements

You Bring:

● 6+ years building and operating production data systems, including demonstrated ownership of streaming or event-driven pipelines - on-call, incident response, SLOs, runbooks, and recovery, not just development.

● Deep, production experience with Apache Kafka - partitioning, consumer groups, consumer-lag and broker-health troubleshooting, exactly-once/idempotent semantics, schema registry, and replay/backfill under load.

● Strong, hands-on Apache Spark experience (PySpark) for streaming and batch transformation in Production.

● AWS-native data engineering across streaming, processing, storage, and catalog services (e.g. MSK, EMR, Glue, S3, Athena), with infrastructure-as-code - AWS CDK (preferred) or Terraform - CI/CD for data pipelines, and cost awareness.

● Comfort debugging distributed data pipelines (consumer lag, data skew, backpressure, late/out-of-order events) with observability tooling (e.g. Datadog/Cloudwatch).

● Expert-level SQL and Python, and experience building and consuming REST APIs.

● An AI-forward engineering mindset, with demonstrated hands-on use of AI-assisted and agentic development tools, an opinion on where AI adds value (and where it doesn’t), and an understanding of agentic data consumption patterns and needs to act on trusted operational data—including context management, lineage, provenance, permissions, freshness, and low-latency access for agentic discovery.

● Change data capture and open table formats - CDC (e.g. Debezium) plus Iceberg or Delta Lake: schema evolution, partitioning, and table maintenance.

● Data contracts, schema governance, and cataloging - schema registries and compatibility rules with dead-letter handling; and familiarity with a technical metastore (e.g. Glue Data Catalog, Unity Catalog) and a governance/discovery catalog (e.g. Atlan, Alation, Collibra).

● Degree in Computer Science, Information Systems or another quantitative field, and comfort on the command line / a Unix-based OS (we are 100% Mac+Linux at Gravie).

● Health insurance domain knowledge - HIPAA Protected Health Information (PHI) and governing access to it.

● Excellent communication skills and demonstrated success in driving results through influence and collaboration.


Extra Credit:

● Experience with Apache Flink or other stateful stream processors - helpful context but not required.

● Knowledge of JVM-based languages like Kotlin or Java.

● Familiarity with serving data to AI and agentic consumers - exposing canonical data as low-latency context or inputs for automated/agentic workloads.

● Previous venture-backed start-up company experience.

About Gravie

Company

Job Details

Job TypeFull Time
Experience LevelMid Level
EducationBachelor's Degree
PostedSeptember 3, 2026 at 04:46 AM

Ready to Apply?

Don't miss out on this opportunity. Apply now and take the next step in your career.

Apply Now