Three lessons from our AI projects
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
Machine learning (ML) doesn't end when you develop a model—that's really just the beginning. Many organisations focus mainly on building a model, but forget that the real work starts once it's in production. A model is only valuable if it continues to perform well, and that requires ongoing maintenance and model monitoring.
MLOps, a combination of machine learning and operations, manages the entire lifecycle of a model, with monitoring playing a crucial role. How do you ensure a model remains reliable over the long term? How do you prevent it from becoming outdated or making inaccurate predictions? This blog explores the importance of monitoring, its benefits and common pitfalls, and offers practical tips for implementing it in Python so your models keep running smoothly, even when circumstances change.
Monitoring is essential to ensure that machine learning models continue to perform reliably in production. With various techniques and tools, you can continuously track your model's performance, detect changes in data and maintain the accuracy of its predictions. But what does this involve?
The four main pillars of ML monitoring are:
Monitoring not only helps you maintain reliable models, but also lets you respond immediately when anomalies or problems arise.
A good monitoring solution does not automatically solve all your problems. As with every step in the MLOps cycle, it is important to make the right trade-offs. Just as when developing your model, you need to choose the right metrics for monitoring. These metrics should fit your specific use case so you can take targeted action and ensure the model continues to contribute as effectively as possible to your objectives.
This applies not only to your model’s performance, but also to monitoring data drift. Before we look more closely at drift detection methods, a brief clarification: data drift refers to changes in the distribution of input data. Concept drift, by contrast, refers to changes in the relationship between input and output variables. Both types of drift carry risks, and effective monitoring can help you detect and address them. It is crucial, however, to identify the underlying cause of the drift so you can implement a targeted solution and keep your model robust.
Detecting drift means choosing the right statistical test for your dataset and the problem you are trying to solve. This requires a thorough understanding of your data and a considered decision. Here are some important questions to ask:
Choosing the right drift detection method therefore depends heavily on your goal and the characteristics of your dataset. Monitoring is not a 'plug-and-play' solution. It takes careful consideration to find the most suitable approach.
Once drift has been detected, the next step depends on the type of drift and its cause. It is always sensible to evaluate your model regularly and retrain it. A good experiment tracking platform helps you make an informed decision about whether to put the new model into production.
If retraining does not produce the desired result, you may need a more in-depth analysis. This could mean revisiting your feature engineering or your choice of algorithm. This manual work can be time-consuming, but is sometimes unavoidable if you want to improve your model.
Although manual steps are sometimes necessary, much of the monitoring of data and models can be automated. You can schedule drift detection checks, log results automatically and set triggers that, for example, start a retraining job as soon as drift is detected. But be careful not to automate too much: what happens if a new model is automatically deployed to production without sufficient checks? Manual approval may be necessary to avoid risks.
Alongside drift monitoring, it is important to log operational information. This means tracking how your model is used, including error messages, latency data and user data. Operational logging helps you identify bugs in the production environment and provides valuable insights for efficient maintenance.
Now that you know what drift detection involves, you may be wondering: how do I apply it in practice? It is a fair question. While drift detection makes sense in theory, putting it into practice can be more challenging. This is especially true when real-time (online) drift detection is needed. With online detection, data is processed as soon as it arrives, whereas with offline drift detection, data is analysed in batches.
Many use cases require online drift detection because batch processing is not always fast enough or feasible. This requires not only the right tests but also infrastructure robust enough for real-time data processing. This is often why drift detection is only implemented once a project or team has reached a higher level of MLOps maturity.
Cloud providers are responding to the need for drift detection and operational logging. For example, Azure offers an integrated Data Drift solution on its Azure Machine Learning platform, making it much easier to monitor data drift. Cloud solutions are particularly useful for operational logging.
Amazon CloudWatch (AWS) provides detailed insights into the operational performance of your ML models, such as error messages and latency data.
Azure Monitor offers similar functionality, including Log Analytics and Application Insights, allowing you to centralise logs and set up alerts for important events in your production code.
Alongside these cloud solutions, powerful open-source tools are also available. For drift detection in Python, there are three popular PIP packages you can use:
Although all three tools are useful, each has its own advantages and disadvantages. Evidently and NannyML are easier to use, while Alibi-detect offers more features but is more complex. Combining multiple tools can often be the best way to build a robust drift monitoring system.
At DSL, we have chosen to combine several of these open-source tools so we can offer a general-purpose, user-friendly solution for drift monitoring. By drawing on the strengths of each package, we can monitor tabular data easily while also detecting more complex drift in images or text. This approach offers flexibility and scalability across a range of use cases.
As noted earlier, an ML model, like the software it runs in, needs to be maintained as a "living" system. Models need to adapt regularly to their environment to keep performing well. This maintenance, however, brings both benefits and challenges.
Benefits of monitoring
Challenges and pitfalls
Monitoring is an essential part of successfully implementing and maintaining a mature MLOps cycle. By focusing on the right metrics, drift detection and automation, organisations can optimise model performance and ensure their models remain reliable, even in a rapidly changing environment. Tools such as Evidently AI, Alibi Detect, NannyML, and the monitoring options offered by Azure and AWS can be valuable here.
There are, however, practical challenges. Organisations need to reach a certain level of maturity and have robust infrastructure in place to use monitoring effectively. But the benefits, including improved performance, greater reliability and faster responses, are clear and compelling. Being aware of the potential pitfalls is crucial to successful implementation.
Want to know more about how your organisation can benefit from effective model monitoring and drift detection? Get in touch with us. We would be happy to help you take your MLOps infrastructure to the next level and get the most out of your AI solutions. MLOps makes that possible.
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
Many organisations have now run an AI pilot. The model works, the demo gets applause, and then nothing else happens. In our experience, most projects…
Artificial Intelligence is developing rapidly. New models appear almost every week, and more and more organisations are experimenting with AI. At the…
Want to be the first to hear about a new blog post?
Thanks for signing up!