We are looking for a Senior Technical Product Manager to drive the next generation of our high-throughput, enterprise data ingestion and data management architecture. In this role, you will lead the strategic transformation of our Data Ingress Engine using technologies like Apache Spark, Databricks, and Medallion Lakehouse Architecture
You will also own the product roadmap for our core Master Data Management (MDM) and Data Remediation Engines, defining how millions of disparate clinical records are cleansed, matched into a single "Golden Patient Record," and safely reprocessed when errors occur.
If you thrive at the intersection of distributed systems, high-volume streaming/batch data pipelines, and complex data models, this role is for you.
Responsibilities:
- Lead the product vision and roadmap to decouple core data ingestion modules from operational API boundaries, turning them into highly scalable, standalone micro-services.
- Productize event-driven and streaming ingestion pipelines (e.g., Kafka/Spark Structured Streaming) capable of processing billions of high-velocity records.
- Establish clear API contracts, event interfaces, and integration specs for external sources feeding raw payloads (HL7 v2, C-CDA, FHIR JSON, Claims) into our platform.
- Define requirements for transitioning mass analytics pipelines onto a Databricks / Medallion Architecture using Apache Spark.
- Drive the design of optimized Silver and Gold layer datasets, ensuring deeply nested JSON formats (e.g., FHIR) are efficiently flattened and structured into high-performance columnar models (Delta Lake/Parquet) for OLAP analytics.
- Partner with engineering architects to balance real-time, low-latency micro-batches against cost-optimized, high-throughput daily batch workloads.
- Own the product requirements for our FHIR-native MDM Engine, defining rules for deterministic and probabilistic patient matching at scale.
- Productize Survivorship Logic to establish rules for creating and maintaining the "Golden Record" across fragmented data feeds.
- Design crosswalk management capabilities to track source-to-target entity linkages across millions of patient lives.
- Build out the end-to-end exception lifecycle for data validation failures—from isolation in Dead-Letter Queues (DLQs) to structured error reporting.
- Define requirements for Data Steward interfaces and APIs, allowing users to review validation exceptions, perform manual record merges/unmerges, and correct bad data payloads.
- Architect Idempotent Replay Mechanisms to re-ingest corrected payloads through Spark processing pipelines without generating duplicate data or corrupting state.
1. Data Ingress & Decoupled Architecture
2. Big Data & Medallion Analytics Strategy
3. Master Data Management (MDM) & Identity Resolution
4. Error Remediation & Exception Management
Requirements:
- 5+ years of Technical Product Management experience leading complex backend data platforms, distributed data engineering products, or big-data ingestion systems.
- Hands-on Distributed Computing Knowledge: Proven track record productizing solutions powered by Apache Spark (batch or streaming) and distributed data processing frameworks.
- Databricks / Lakehouse Expertise: Deep familiarity with Medallion Architecture (Bronze/Silver/Gold) and Delta Lake/Parquet performance patterns.
- OLAP / Columnar Strategy Mindset: Demonstrated experience optimizing complex, highly nested arrays and hierarchical data structures for analytical query engines.
- Identity & Governance: Direct experience productizing Master Data Management (MDM), entity resolution, or identity matching tools.
- Fault-Tolerant Systems: Experience with event-driven architectures, Dead-Letter Queues (DLQs), and data replay/remediation workflows.
- Technical Proficiency: Ability to comfortably write and execute SQL and Python to analyze raw datasets, validate engine output, and define logic rules
- Experience with healthcare data standards, including HL7 FHIR, C-CDA, or HL7 v2.
- Exposure to healthcare terminologies (SNOMED CT, LOINC, RxNorm, ICD-10).
- Familiarity with SQL-on-FHIR analytics standards.
This position is a replacement role, created to support Smile’s continued growth and commitment to operational excellence.

