Machinify, Inc.

Bending the healthcare cost curve with AI.

Senior Data Engineer – Analytics

Data EngineerData EngineerFull TimeRemoteTeam 51-200H1B No SponsorCompany SiteLinkedIn

Location

California

Posted

56 days ago

Salary

Not specified

Bachelor Degree4 yrs expEnglishAirflowAWSCloudKafkaPythonSparkSQL

Job Description

• Design and implement robust, production-grade pipelines using Python, Spark SQL, and Airflow to process high-volume file-based datasets (CSV, Parquet, JSON). • Lead efforts to canonicalize raw healthcare data (837 claims, EHR, partner data, flat files) into internal models. • Own the full lifecycle of core pipelines — from file ingestion to validated, queryable datasets — ensuring high reliability and performance. • Onboard new customers by integrating their raw data into internal pipelines and canonical models; collaborate with SMEs, Account Managers, and Product to ensure successful implementation and troubleshooting. • Build resilient, idempotent transformation logic with data quality checks, validation layers, and observability. • Refactor and scale existing pipelines to meet growing data and business needs. • Tune Spark jobs and optimize distributed processing performance. • Implement schema enforcement and versioning aligned with internal data standards. • Collaborate deeply with Data Analysts, Data Scientists, Product Managers, Engineering, Platform, SMEs, and AMs to ensure pipelines meet evolving business needs. • Monitor pipeline health, participate in on-call rotations, and proactively debug and resolve production data flow issues. • Contribute to the evolution of our data platform — driving toward mature patterns in observability, testing, and automation. • Build and enhance streaming pipelines (Kafka, SQS, or similar) where needed to support near-real-time data needs. • Help develop and champion internal best practices around pipeline development and data modeling.

Job Requirements

  • 4+ years of experience as a Data Engineer (or equivalent), building production-grade pipelines.
  • Strong expertise in Python, Spark SQL, and Airflow.
  • Experience processing large-scale file-based datasets (CSV, Parquet, JSON, etc) in production environments.
  • Experience mapping and standardizing raw external data into canonical models.
  • Familiarity with AWS (or any cloud), including file storage and distributed compute concepts.
  • Experience onboarding new customers and integrating external customer data with non-standard formats.
  • Ability to work across teams, manage priorities, and own complex data workflows with minimal supervision.
  • Strong written and verbal communication skills — able to explain technical concepts to non-engineering partners.
  • Comfortable designing pipelines from scratch and improving existing pipelines.
  • Experience working with large-scale or messy datasets (healthcare, financial, logs, etc).
  • Experience building or willingness to learn streaming pipelines using tools such as Kafka or SQS.
  • Bonus: Familiarity with healthcare data (837, 835, EHR, UB04, claims normalization).

Benefits

  • Real impact — your pipelines will directly support decision-making and claims payment outcomes from day one.
  • High visibility — partner with ML, Product, Analytics, Platform, Operations, and Customer teams on critical data initiatives.
  • Total ownership — you’ll drive the lifecycle of core datasets powering our platform.
  • Customer-facing impact — you will directly contribute to successful customer onboarding and data integration.

Related Categories

Related Job Pages

More Data Engineer Jobs

Manager, Ads Data Engineering

Netflix

Where you come to do the best work of your life. Follow @WeAreNetflix on Twitter, IG, Facebook, & Youtube for more

Data Engineer56 days ago
Full TimeRemoteTeam 10,001+Since 1997H1B Sponsor

Data Engineering Manager leading ads data initiatives at Netflix

HadoopSparkSQL
United States
$360K - $920K / year

Full Stack Software Engineer 5 – Data Architecture, Integrations

Netflix

Where you come to do the best work of your life. Follow @WeAreNetflix on Twitter, IG, Facebook, & Youtube for more

Data Engineer56 days ago
Full TimeRemoteTeam 10,001+Since 1997H1B Sponsor

Software Engineer focused on Data Architecture & Integrations at Netflix

ETLJavaJavaScriptNode.jsNoSQLPythonSQLTypeScriptGo
United States
$100K - $600K / year

Data Engineering Consultant

Eton Technologies

ERP | Cloud | Analytics | Integrations | IT Support

Data Engineer56 days ago
Full TimeRemoteTeam 51-200Since 2016H1B No Sponsor

Data Engineering Consultant designing modern data platforms for global clients

Amazon RedshiftAWSAzureCloudETLGoogle Cloud PlatformNumpyPandasPythonSQLTableau
United States
Data Engineer57 days ago
Full TimeRemoteTeam 1,001-5,000H1B No Sponsor

Data Engineer developing ETL/ELT pipelines for Nakupuna Solutions supporting NAVSUP

Amazon RedshiftAWSCloudETLPostgresPythonSQL
Pennsylvania
$125K - $140K / year