Omada Health is on a mission to bend the curve of chronic disease. We rely on trusted, high-quality data to power intelligent products, personalized member experiences, and data-driven decision making.
As machine learning becomes increasingly central to our platform, we're investing in the data foundations that enable scalable model development, experimentation, and production inference.
Job overview:
We are seeking a Staff Software Engineer, Data Engineering to lead the design and development of the production data platform that powers machine learning across Omada.
In this role, you will partner closely with Data Scientists, Applied AI Engineers, Product Engineers, and fellow Data Engineers to identify, design and build trusted, reusable datasets foundations that serve as the foundation for feature engineering, model training, experimentation, and production inference.
Rather than building one-off pipelines for individual models, you'll create scalable data products and feature pipelines that enable multiple machine learning use cases while ensuring consistency, reliability, and governance across the ML lifecycle.
You will own the technical design of feature datasets—from ingesting raw behavioral, clinical, and operational data through transforming, validating, and publishing production-grade datasets that are reusable across modeling teams.
This role is ideal for someone who enjoys solving complex data problems, designing scalable distributed data systems, and enabling machine learning through well-engineered data foundations.
Key Responsibilities:
- Design, build, and maintain reusable feature datasets that support machine learning use cases including personalization, engagement, risk prediction, churn modeling, recommendation systems, and experimentation.
- Establish self-service foundations that streamline and democratize dataset creation across the data organization.
- Partner with Data Scientists to translate modeling requirements into production-ready feature pipelines, supporting the full model lifecycle from exploration to deployment.
- Identify source data, transformations, and historical windows needed for feature engineering. Help define and build shared, reusable feature definitions across models rather than one-off datasets.
- Balance features freshness, correctness, latency, and computational efficiency when designing data pipelines.
- Build datasets that support both historical model training and future production inference.
- Design and implement batch and streaming pipelines that transform raw healthcare, behavioral, product, and operational data into trusted ML-ready datasets.
- Build reliable data processing systems using Python, SQL, Spark, and modern cloud data platforms.
- Optimize large-scale distributed processing for performance, scalability, and cost.
- Design data pipelines that are modular, testable, observable, and easy to evolve as product requirements change.
- Ensure data quality through testing, anomaly detection, schema validation, and pipeline monitoring.
- Partner with platform teams to support near real-time feature generation where appropriate.
- Improve reproducibility by standardizing feature computation across experimentation and production.
- Support rapid experimentation without sacrificing long-term maintainability.
- Ensure data quality through testing, anomaly detection, schema validation, and pipeline monitoring.
- Establish engineering standards for correctness, documentation, and maintainability.
- Familiarity with feature stores or feature management platforms.
- Familiarity with model training pipelines and MLOps workflows.
Technical Leadership:
- Lead architecture and design discussions for large-scale ML data systems, driving adoption of reusable patterns and platform capabilities across Data Engineering.
- Influence technical direction across multiple engineering teams, embedding with Product, Engineering, and business stakeholders (Clinical, Finance, Growth, Enrollment) during early design phases to shape data capture requirements at the source.
- Translate ambiguous business requirements from Business domain SMEs into concrete technical specs, maintaining consistency of business logic and definitions across systems.
- Mentor engineers on distributed data processing, software engineering best practices, and scalable data modeling.
About you:
Experience
- 8+ years building large-scale production data platforms and distributed data pipelines.
- Experience designing reusable datasets that power machine learning, experimentation, or advanced analytics.
- Demonstrated experience partnering closely with Data Scientists to productionize feature engineering workflows.
- Experience leading cross-team technical initiatives and influencing engineering direction.
- Strong experience working with cloud-native data platforms such as AWS.
- Experience building production data systems using Databricks, Iceberg, Spark, Redshift, Snowflake, or similar technologies.
- Experience developing reliable batch and streaming data pipelines.
- Experience working with healthcare, behavioral, or other large-scale event data is a plus.
Technical Skills
- Expert SQL with strong data modeling skills.
- Strong programming skills in Python, Java, or Scala.
- Experience with Apache Spark or similar distributed compute frameworks.
- Experience with Airflow or similar orchestration platforms.
- Experience designing dimensional models, event models, and feature datasets.
- Experience implementing testing, CI/CD, observability, and production monitoring for data pipelines.
- Understanding of Feature Stores and ML data lifecycle concepts.
- Experience with Lakehouse Architecture such as Databricks, Iceberg is a strong plus.
- Experience with streaming technologies such as Kafka, Flink, or Spark Structured Streaming.
- Familiarity with NoSQL Databases (document & graph databases Nepture, Neo4j etc.)
- Understanding of software engineering best practices, distributed systems, and cloud-native architectures.
Communication Skills: An exceptional people leader who develops engineers into future technical leaders.