Location: REMOTE / Toronto, Ontario
This job allows you to work remotely.
We're looking for a Senior Backend Engineer to help build and evolve the distributed systems that power large-scale AI and machine learning platforms. In this role, you'll lead the modernization of a high-throughput production system into a more scalable, resilient, and maintainable distributed architecture. You'll have the opportunity to solve complex systems engineering challenges while working alongside machine learning engineers and platform teams to deliver reliable, high-performance infrastructure.
This is an ideal opportunity for an engineer who enjoys designing distributed systems, improving system reliability, and taking ownership of services that operate at significant scale.
What You'll Do
- Lead the evolution of a core backend platform toward a modern distributed architecture while ensuring production stability throughout the transition.
- Design, build, and operate scalable, fault-tolerant backend services that support high-volume AI and machine learning workloads.
- Identify and resolve performance bottlenecks, reliability issues, and scalability challenges across distributed systems.
- Collaborate with Machine Learning Engineers, Product Managers, and other engineering teams to translate business and technical requirements into scalable production solutions.
- Own critical platform components throughout their lifecycle, including architecture, implementation, deployment, monitoring, and operational support.
- Promote engineering best practices around software architecture, testing, code quality, observability, and operational excellence.
- Mentor other engineers and help foster a culture of technical excellence, ownership, and continuous improvement.
Must Have Skills:
What We're Looking For
- 5+ years of professional software engineering experience, with a strong focus on backend and distributed systems development.
- Proven experience designing, building, and operating large-scale distributed systems that are highly available, fault tolerant, and capable of handling significant concurrency.
- Strong software engineering fundamentals, including system design, automated testing, version control, code reviews, and CI/CD.
- Strong programming experience in Python.
- Experience with distributed data processing technologies such as Apache Spark, Apache Flink, or Apache Beam. Hands-on Spark experience is an asset.
- Solid understanding of distributed systems concepts, including replication, partitioning, consensus, idempotency, back-pressure, and failure recovery.
- Experience building cloud-native applications on a major cloud platform such as Google Cloud Platform, AWS, or Microsoft Azure.
- Experience working with modern distributed systems technologies such as Kubernetes, Kafka, gRPC, and containerized microservices.
- Familiarity with modern software delivery practices, including Docker, Kubernetes, GitHub Actions, automated testing, and continuous deployment.
- Strong technical leadership skills with the ability to influence architecture, collaborate across teams, and mentor fellow engineers.
- Excellent analytical and problem-solving abilities, particularly when diagnosing and resolving complex production issues.
- A passion for learning new technologies, continuously improving engineering practices, and contributing to a collaborative, high-performing team.