Location: REMOTE / Toronto, Ontario
This job allows you to work remotely.
We're looking for a Senior Data Engineer to join a team responsible for building the data infrastructure that powers large-scale machine learning and recommendation systems. You'll help design, develop, and optimize the data platform behind billions of intelligent decisions each day, supporting enterprise customers across a wide range of digital experiences.
In this role, you'll own the Spark-based data pipelines that feed production machine learning models, working closely with ML engineers, data scientists, and software engineers to build reliable, scalable, and high-performance data systems. You'll play a key role in developing the infrastructure that enables advanced AI-driven products and personalization capabilities, with the opportunity to influence architecture and solve complex data engineering challenges at scale.
What You'll Do
- Design, build, maintain, and optimize production data pipelines supporting AI and machine learning workloads at enterprise scale.
- Develop and scale Apache Spark batch processing pipelines, including cluster configuration, performance tuning, and optimization within Google Cloud.
- Build and maintain a centralized data platform that supports machine learning training, experimentation, and production inference.
- Partner closely with Machine Learning Engineers and Data Scientists to provide high-quality, reliable datasets for model development and evaluation.
- Monitor, troubleshoot, and improve the performance, reliability, and scalability of data processing systems.
- Collaborate with platform and distributed systems engineers to evolve the overall data architecture.
- Continuously improve the efficiency, reliability, and maintainability of the data infrastructure.
- Deliver data solutions and platform enhancements that provide measurable business impact.
Must Have Skills:
What We're Looking For
- 5+ years of professional experience in Data Engineering.
- Strong expertise with Apache Spark, including extensive experience using the PySpark DataFrame API to solve large-scale data processing challenges.
- Experience designing, optimizing, and operating distributed data processing pipelines.
- Strong Python development skills, including unit testing, version control, code reviews, and CI/CD practices.
- Experience working with modern data storage formats such as Parquet and Delta Lake.
- Experience building or consuming streaming data pipelines using technologies such as Kafka.
- Hands-on experience with cloud infrastructure, preferably Google Cloud Platform.
- Strong understanding of query optimization and data processing performance tuning.
- Experience working with modern software development practices, including Docker, Kubernetes, GitHub Actions, and automated deployment pipelines.
- Excellent collaboration skills with cross-functional engineering teams, including Machine Learning Engineers, Data Scientists, and Platform Engineers.
- Comfortable working in a fast-moving, highly collaborative environment where ownership and continuous improvement are valued.