Senior Data Engineer

overt minds

📍 Remote🌐 Remote🕐 24d ago🔗 himalayas

Job Description

### This is a remote position. We are seeking a Senior Data Engineer to build scalable, cloud-native data platforms and enable end-to-end MLOps workflows. You will design ETL/ELT pipelines, manage data lakes/warehouses/feature stores, and ensure high-performance, secure, and cost-efficient pipelines for AI/ML and analytics. This role blends Data Engineering + MLOps to deliver production-ready, automated, and reliable ML workflows. **Responsibilities** * Data Pipelines: Design & optimize batch/streaming ETL/ELT pipelines at scale. * Platforms: Build/manage data lakes, warehouses, feature stores for ML/BI workloads. * MLOps: Enable model training, deployment, CI/CD, monitoring, retraining, versioning using SageMaker (AWS), Vertex AI (GCP), Azure ML. * Streaming: Implement real-time pipelines with Kafka, Spark Streaming, AWS Kinesis, GCP Pub/Sub, Azure Event Hubs. * Automation: Leverage Terraform, CloudFormation, ARM, Kubernetes for infra-as-code & scaling. * Quality & Governance: Ensure data lineage, metadata, observability, security, compliance, cost efficiency. * Collaboration: Work with Data Scientists & ML Engineers to productionize ML models across cloud environments. ### Requirements * 5+ years of hands on experience in Data Engineering, Big Data, or Cloud Data Platform roles, working on large scale production systems. * Strong command of Python and SQL, using them to build and optimize ETL/ELT pipelines. * Deep working knowledge of distributed data systems (e.g., Spark, Hive, Presto, Dask) for batch and real-time processing. * Proven track record with cloud-native platforms across AWS, GCP, or Azure — e.g., BigQuery, Redshift, EMR, Databricks — for data storage and analytics. * Experience designing and maintaining event driven and streaming architectures (Kafka, Pub/Sub, Flink). * Solid background in data modeling (star schema, OLAP cubes, graph databases) to support BI and analytics. * Practical exposure to data security, encryption, and compliance frameworks (e.g., GDPR, HIPAA). **Preferred Skills** * Direct experience enabling MLOps workflows building feature stores, managing versioned datasets, or integrating pipelines with ML platforms (SageMaker, Vertex AI, Azure ML). * Familiarity with real-time analytics systems such as Clickhouse or Apache Pinot. * Exposure to data observability tools (e.g., Monte Carlo, Databand) to monitor quality, lineage, and reliability. * Demonstrated ability to build scalable, resilient, and secure data systems that support mission critical applications. * Interest and experience in supporting AI/ML innovation with robust data infrastructure. * Strong mindset for automation, scalability, DevOps/MLOps practices, and engineering excellence. ### Benefits * Competitive compensation as per industry standards * Opportunity to work on enterprise‑scale AI/ML and analytics platforms * High‑impact role driving cloud‑native and MLOps transformation * Collaborative, engineering‑driven work culture * Strong growth path into Lead Data Engineer, ML Platform Engineer, or MLOps Architect roles Originally posted on [Himalayas](https://himalayas.app)