Principal Machine Learning Developer: AI/ML Platform
Autodesk
Principal Machine Learning Operations Developer: AI/ML Platform
Location: Canada. Open to Toronto, Ontario or Remote Ontario.
About Autodesk
Autodesk makes software for people who make things. We are a global leader in 3D design, engineering, manufacturing, and entertainment software. Our customers use Autodesk software to design and make the physical and virtual worlds that we live in. If you've ever driven a high-performance car, admired a towering skyscraper, used a smartphone, or watched a great film or played an immersive game, chances are you've experienced what millions of Autodesk customers are doing with our software.
Position Overview
Autodesk, a global leader in 3D design, engineering, manufacturing, and entertainment software, is seeking a skilled Principal MLOps Developer to join our AI/ML Platform team. This role is pivotal in ensuring the smooth operationalization of machine learning models and the overall efficiency of our next-generation AI/ML platform used in the development of machine learning and generative AI solutions powering Autodesk’s suite of products and services. You will collaborate with research and product engineering from various domains including design, construction, manufacturing, and media & entertainment to deliver platform capabilities that support the full AI/ML development lifecycle.
As a principal-level contributor, you will help build innovative capabilities that enable faster, more secure development and deployment of machine learning and generative AI solutions. You will take ownership of critical platform components, provide architectural direction, and contribute to scalable systems for model training, inference, data processing, deployment automation, monitoring, governance, and operations.
Responsibilities
Operational Excellence: Drive the operational excellence and technical direction of our AI/ML Platform by implementing and optimizing MLOps practices across the full machine learning development lifecycle
Innovative System Design: Lead the design and engineering of software systems and platform services for the AI/ML Platform, contributing to scalable, secure, and reliable ML development and operations
Deployment Automation: Design and implement automated deployment pipelines for machine learning models and ML artifacts, ensuring seamless transitions from development to production
Workflow Automation: Develop comprehensive systems to automate and optimize laborious ML development and operational processes, integrating them into the platform to streamline operations
Scalable Infrastructure: Collaborate with cross-functional teams to design, implement, and maintain scalable infrastructure for model training, inference, data processing, and ML artifact management
ML Solution Deployment: Develop tools for building, deploying, and operating ML artifacts in production environments, facilitating a smooth transition from development to deployment
Big Data Management: Automate and orchestrate tasks related to managing large-scale data transformation, data processing, and data stores that support model training, validation, deployment, and operations
Scalable Services: Design and implement low-latency, scalable prediction and inference services to support the diverse needs of platform users and Autodesk product teams
Monitoring and Logging: Develop and maintain robust monitoring and logging systems to track model performance, system health, operational reliability, and overall platform efficiency
Collaboration with Data Engineers: Work closely with data engineers to ensure efficient data pipelines for model training, validation, deployment, and ongoing platform operations
Cross-Functional Collaboration: Collaborate across diverse teams, including machine learning researchers, data engineers, software developers, product managers, software architects, and operations teams, fostering a collaborative and cohesive work environment
Version Control and Model Governance: Implement version control systems for machine learning models and contribute to model governance practices
Governance and Trust: Contribute to the implementation of robust model governance practices, version control systems, and adherence to compliance standards. Uphold data privacy and ethical considerations, fostering trust in our AI/ML solutions
Security and Compliance: Enforce security best practices and compliance standards in all aspects of MLOps, ensuring data privacy and platform security
Continuous Improvement: Identify opportunities for process automation and optimization, and implement strategies to enhance the overall MLOps lifecycle
Architectural Leadership: Take ownership of critical components of the platform, providing architectural direction and contributing to the overall success of the AI/ML Platform
Basic Qualifications
- Educational Background: BS or MS in Computer Science, or equivalent practical experience
- Experience: 8+ years of experience in software development and engineering, with a solid record of delivering production systems and services
- Strong background in AI/ML with experience in deep learning, statistical modeling, and neural networks
- Expertise in AI/ML Technologies: Hands-on experience with AI/ML frameworks (such as TensorFlow, PyTorch) and familiarity with the lifecycle of AI/ML model development, from training to deployment
- Proficiency in Programming Languages: Strong coding skills in languages commonly used in AI/ML and system development, such as Python, Java, or Go
- Strong Analytical and Problem-Solving Skills: Ability to tackle complex technical challenges, analyze potential solutions, and implement the most effective ones
- Excellent Communication and Teamwork Abilities: Strong communication skills to effectively collaborate with cross-functional teams, along with the ability to work independently
- System Performance Optimization: Deep understanding of performance metrics and latency optimization techniques, with the ability to diagnose, tune, and enhance the efficiency of serving systems
- Commitment to Continuous Learning: A continuous learning mindset to stay updated with the latest trends and technologies in AI/ML, cloud computing, and software engineering
Preferred Qualifications
- GPU Computing: Exposure to leveraging GPU computing for AI/ML workloads, including experience with CUDA, OpenCL, or other GPU programming tools, to significantly enhance model training and inference performance
- Experience with Big Data Technologies: Experience with big data technologies and ecosystems (Hadoop, Spark, Kafka) for processing and analyzing large datasets in a distributed computing environment
- AI/ML Model Monitoring Tools: Familiarity with tools and frameworks for monitoring and managing the performance of AI/ML models in production (e.g., MLflow, Kubeflow, TensorBoard)
- Expertise in High-Performance Computing (HPC): Experience with HPC techniques and technologies for optimizing computational workloads, particularly in the context of AI/ML model training and inference
At Autodesk, we're building a diverse workplace and an inclusive culture to give more people the chance to imagine, design, and make a better world. Autodesk is proud to be an equal opportunity employer and considers all qualified applicants for employment without regard to race, color, religion, age, sex, sexual orientation, gender, gender identity, national origin, disability, veteran status or any other legally protected characteristic. We also consider for employment all qualified applicants regardless of criminal histories, consistent with applicable law.
Autodesk has always valued flexibility in how we work. We continue to provide employees flexibility to support their work preferences wherever possible and nearly all roles are hybrid or remote, unless otherwise indicated.