Text copied to clipboard!

Title

Text copied to clipboard!

MLOps Engineer

Description

Text copied to clipboard!
We are looking for an MLOps Engineer to design, build, deploy, and maintain scalable machine learning infrastructure that enables data science and engineering teams to deliver models to production efficiently and safely. This role sits at the intersection of machine learning, software engineering, cloud infrastructure, and platform operations. The ideal candidate is passionate about creating robust systems that support the full machine learning lifecycle, from data ingestion and experimentation to model deployment, monitoring, retraining, and governance. As an MLOps Engineer, you will work closely with data scientists, machine learning engineers, software developers, DevOps professionals, and product stakeholders to streamline workflows and improve the reliability of AI-powered applications. You will help establish best practices for version control, reproducibility, CI/CD pipelines, feature management, model registry usage, infrastructure as code, and observability across machine learning environments. Your work will directly influence how quickly and securely machine learning solutions can move from prototype to business value. In this position, you will be responsible for building automated pipelines for training, testing, validation, deployment, and monitoring of machine learning models across development, staging, and production environments. You will evaluate and implement tools for orchestration, containerization, experiment tracking, and performance monitoring. You will also contribute to system architecture decisions that improve scalability, fault tolerance, cost efficiency, and compliance with organizational standards. A successful candidate combines strong programming and cloud engineering skills with a practical understanding of machine learning workflows. You should be comfortable working with distributed systems, APIs, containers, and modern deployment practices, while also understanding the needs of model developers and analytics teams. Experience with security, access control, and data governance is highly valuable, especially in regulated or high-scale environments. This role offers the opportunity to shape the foundation of machine learning operations within a growing organization. You will help define technical standards, reduce operational friction, and improve the performance and trustworthiness of production AI systems. If you enjoy solving complex infrastructure challenges, enabling cross-functional teams, and building platforms that make machine learning repeatable and dependable, this role is an excellent fit.

Responsibilities

Text copied to clipboard!
  • Design and maintain end-to-end machine learning deployment pipelines
  • Build CI/CD workflows for model training, testing, and release
  • Manage containerized workloads and orchestration platforms for ML services
  • Implement model monitoring for performance, drift, latency, and failures
  • Collaborate with data scientists to productionize experiments and notebooks
  • Develop infrastructure as code for reproducible and scalable environments
  • Maintain model registries, artifact stores, and feature management systems
  • Improve observability, logging, alerting, and incident response processes

Requirements

Text copied to clipboard!
  • Bachelor's degree in computer science, engineering, or a related field
  • Experience with Python and software engineering best practices
  • Hands-on knowledge of cloud platforms such as AWS, Azure, or GCP
  • Experience with Docker, Kubernetes, and container-based deployments
  • Understanding of CI/CD tools and automated testing frameworks
  • Familiarity with ML lifecycle tools such as MLflow, Kubeflow, or Airflow
  • Knowledge of monitoring, logging, and infrastructure automation tools
  • Ability to work cross-functionally with engineering and data teams

Potential interview questions

Text copied to clipboard!
  • What experience do you have deploying machine learning models to production?
  • Which cloud platforms and MLOps tools have you used most extensively?
  • How do you monitor model performance and detect data drift?
  • Describe a pipeline you built for training and deploying ML models.
  • How do you ensure reproducibility across machine learning environments?
  • What is your experience with Kubernetes and container orchestration?
  • How do you balance model performance, reliability, and operational cost?
  • Have you worked in regulated environments with governance requirements?