About this role
We are looking for a strong hands-on Data Engineer with ML engineering experience to build data pipelines and analytics products on complex, high-volume refinery datasets. This role demands deep technical skill on the lakehouse, comfort with messy time-series and process data, and the maturity to deliver in a client-facing project environment.
Key Responsibilities:
- Build and operate scalable data pipelines on Databricks / Spark to ingest, clean, and model refinery datasets — including process historian data, lab data (LIMS), planning data, maintenance, energy, and emissions data.
- Develop the medallion architecture (bronze / silver / gold) with strong attention to data quality, lineage, and reconciliation against source systems.
- Work with refinery process SMEs to translate operational questions (yield, energy, throughput, reliability, emissions) into analytical datasets and ML features.
- Develop, train, and productionise ML models for use cases such as yield optimisation, energy efficiency, predictive maintenance, soft sensors, anomaly detection, and APC performance monitoring.
- Implement MLOps practices — feature stores, experiment tracking (MLflow), model registry, CI/CD, monitoring, and drift detection.
- Collaborate with the enterprise architect on platform decisions and with the refinery SMEs on use-case design and validation.
- Build dashboards and analytical outputs in Power BI / Tableau / Databricks SQL for plant managers and corporate stakeholders.
- Document pipelines, models, and decisions to a standard suitable for client handover.
Must-Have Skills & Experience:
- 5–8 years of relevant experience in data engineering / ML engineering, with a strong project delivery track record.
- Expert-level Python and SQL; strong Spark / PySpark; proficient with Delta Lake.
- Strong hands-on experience with Databricks (Workflows, DLT, Unity Catalog, MLflow). Experience with Snowflake or Microsoft Fabric is a plus.
- Experience working with industrial time-series data — OSIsoft PI / AVEVA PI System (PI Web API, PI AF, PI Integrator), Honeywell PHD, or Aspen IP.21 — including handling high-frequency tag data, gaps, outliers, and time alignment.
- Solid understanding of ML — supervised/unsupervised learning, time-series forecasting, anomaly detection — with libraries such as scikit-learn, XGBoost, PyTorch/TensorFlow, and Prophet/statsmodels.
- MLOps experience: MLflow, model deployment patterns, monitoring, feature engineering at scale.
- Cloud proficiency on at least one of Azure, AWS, or GCP — including IAM, networking basics, and cost-aware design.
- Strong data quality discipline — testing (Great Expectations / Soda), reconciliation, observability.
- Education B.E/Btech.
Good-to-Have:
- Prior delivery on refinery / petrochemical / chemicals analytics — familiarity with concepts such as crude assays, CDU/VDU/FCC/HCU/Reformer operations, blending, energy balance, and yield accounting.
- Experience with Aspen suite (HYSYS, PIMS, GDOT, DMC3) data extraction and integration.
- Streaming experience — Kafka, Event Hubs, Kinesis, Spark Structured Streaming.
- Certifications: Databricks Data Engineer Professional, Databricks ML Professional, Azure Data Engineer / AWS Data Analytics Specialty.
- Exposure to Generative AI / LLM use cases on industrial data (RAG over P&IDs, SOPs, incident reports).
Engagement Conditions:
- Client-site presence required as per project schedule (minimum 4 days/week on-site, subject to client policy).
- Willingness to travel to refinery sites for data discovery, validation, and deployment workshops.
- Available to start within 2–4 weeks of selection.
What we offer:
- An opportunity to work on cutting-edge projects in the data engineering and machine learning space.
- A collaborative and innovative work environment.
- Opportunities for professional growth and development.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.