Skip to content
Go to insights Blog

Model monitoring: maximise your models’ performance

Model monitoring: maximise your models’ performance
Written by
Data Science Lab
Published on
28 October 2024

Machine learning (ML) doesn't end when you develop a model—that's really just the beginning. Many organisations focus mainly on building a model, but forget that the real work starts once it's in production. A model is only valuable if it continues to perform well, and that requires ongoing maintenance and model monitoring. 

MLOps, a combination of machine learning and operations, manages the entire lifecycle of a model, with monitoring playing a crucial role. How do you ensure a model remains reliable over the long term? How do you prevent it from becoming outdated or making inaccurate predictions? This blog explores the importance of monitoring, its benefits and common pitfalls, and offers practical tips for implementing it in Python so your models keep running smoothly, even when circumstances change. 

What is monitoring? 

Monitoring is essential to ensure that machine learning models continue to perform reliably in production. With various techniques and tools, you can continuously track your model's performance, detect changes in data and maintain the accuracy of its predictions. But what does this involve? 

 
The four main pillars of ML monitoring are: 

  • Model Performance Monitoring 
    Monitoring your model's performance means measuring key metrics such as accuracy, precision, recall, F1-score, and R-squared. By tracking these statistics, you can act quickly if the model's performance starts to decline. 
  • Data Drift Detection 
    Models can become outdated when input data changes over time. Data drift detection helps you spot shifts in the data distribution, so you can intervene before model performance deteriorates significantly. 
  • Operational Monitoring 
    Besides model performance, it's important to track how the system operates. This includes operational metrics such as latency, availability and error messages. These give you insight into the robustness of the software surrounding your model. 
  • Detecting adversarial attacks 
    Attackers sometimes deliberately try to mislead models with manipulated data. Monitoring helps you detect early signs that unusual input data is being used to manipulate your model, such as images that people can still recognise but that cause a model to make a mistake. 

Monitoring not only helps you maintain reliable models, but also lets you respond immediately when anomalies or problems arise. 


The importance of choosing the right metrics

A good monitoring solution does not automatically solve all your problems. As with every step in the MLOps cycle, it is important to make the right trade-offs. Just as when developing your model, you need to choose the right metrics for monitoring. These metrics should fit your specific use case so you can take targeted action and ensure the model continues to contribute as effectively as possible to your objectives. 

This applies not only to your model’s performance, but also to monitoring data drift. Before we look more closely at drift detection methods, a brief clarification: data drift refers to changes in the distribution of input data. Concept drift, by contrast, refers to changes in the relationship between input and output variables. Both types of drift carry risks, and effective monitoring can help you detect and address them. It is crucial, however, to identify the underlying cause of the drift so you can implement a targeted solution and keep your model robust. 

How do you choose the right metrics for drift detection? 

Detecting drift means choosing the right statistical test for your dataset and the problem you are trying to solve. This requires a thorough understanding of your data and a considered decision. Here are some important questions to ask: 

  • What type of data are you using? 
    Is the data categorical or numerical? Numerical data may be better suited to tests such as the KS test or MMD, while Chi-square and Cramer’s V tests are more suitable for categorical data. 
  • How large is your dataset? 
    The sensitivity of some tests can vary depending on the size of the dataset. 
  • Is your dataset normally distributed? 
    This can affect which tests you choose. 
  • Do you want to focus on data drift or concept drift? 
    Do you want to monitor changes in the input data or in the relationship between input and output? 
  • Do you want online or offline drift detection? 
    Are you working with batch data or streaming data? Do you need to respond to changes in real time? 
  • Do you want univariate or multivariate drift detection? 
    Are you focusing on each feature separately (univariate), or do you want to compare multiple features at once (multivariate)? 
  • Which test do you choose for univariate drift? 
    Do you use a separate test for each feature, or choose one test for all features? 
  • Which dataset do you use for drift detection? 
    Do you use the training dataset as a reference, or choose a different, unprocessed dataset from the same period? 

Choosing the right drift detection method therefore depends heavily on your goal and the characteristics of your dataset. Monitoring is not a 'plug-and-play' solution. It takes careful consideration to find the most suitable approach. 

What do you do after detecting drift? 

Once drift has been detected, the next step depends on the type of drift and its cause. It is always sensible to evaluate your model regularly and retrain it. A good experiment tracking platform helps you make an informed decision about whether to put the new model into production. 

If retraining does not produce the desired result, you may need a more in-depth analysis. This could mean revisiting your feature engineering or your choice of algorithm. This manual work can be time-consuming, but is sometimes unavoidable if you want to improve your model. 

Automating monitoring and drift detection 

Although manual steps are sometimes necessary, much of the monitoring of data and models can be automated. You can schedule drift detection checks, log results automatically and set triggers that, for example, start a retraining job as soon as drift is detected. But be careful not to automate too much: what happens if a new model is automatically deployed to production without sufficient checks? Manual approval may be necessary to avoid risks. 

Operational logging 

Alongside drift monitoring, it is important to log operational information. This means tracking how your model is used, including error messages, latency data and user data. Operational logging helps you identify bugs in the production environment and provides valuable insights for efficient maintenance. 

Getting started with monitoring

Now that you know what drift detection involves, you may be wondering: how do I apply it in practice? It is a fair question. While drift detection makes sense in theory, putting it into practice can be more challenging. This is especially true when real-time (online) drift detection is needed. With online detection, data is processed as soon as it arrives, whereas with offline drift detection, data is analysed in batches. 

Many use cases require online drift detection because batch processing is not always fast enough or feasible. This requires not only the right tests but also infrastructure robust enough for real-time data processing. This is often why drift detection is only implemented once a project or team has reached a higher level of MLOps maturity. 

Cloud solutions for drift detection 

Cloud providers are responding to the need for drift detection and operational logging. For example, Azure offers an integrated Data Drift solution on its Azure Machine Learning platform, making it much easier to monitor data drift. Cloud solutions are particularly useful for operational logging. 

Amazon CloudWatch (AWS) provides detailed insights into the operational performance of your ML models, such as error messages and latency data. 

Azure Monitor offers similar functionality, including Log Analytics and Application Insights, allowing you to centralise logs and set up alerts for important events in your production code. 

Open-source tools for drift detection

Alongside these cloud solutions, powerful open-source tools are also available. For drift detection in Python, there are three popular PIP packages you can use: 

  • Evidently AI 
    This user-friendly tool focuses specifically on tabular data and supports various drift detection methods. 
  • Alibi-detect 
    A flexible, well-documented option that supports tabular data, images and text. Alibi-detect also includes a component for adversarial attack detection, helping to improve the security of your models. 
  • NannyML 
    Like Evidently, this tool is best suited to tabular data and is easy to integrate. It is slightly more user-friendly, but has limited support for other data types. 
     

Although all three tools are useful, each has its own advantages and disadvantages. Evidently and NannyML are easier to use, while Alibi-detect offers more features but is more complex. Combining multiple tools can often be the best way to build a robust drift monitoring system. 

Our solution at Data Science Lab

At DSL, we have chosen to combine several of these open-source tools so we can offer a general-purpose, user-friendly solution for drift monitoring. By drawing on the strengths of each package, we can monitor tabular data easily while also detecting more complex drift in images or text. This approach offers flexibility and scalability across a range of use cases. 

What are the advantages and disadvantages of monitoring?

As noted earlier, an ML model, like the software it runs in, needs to be maintained as a "living" system. Models need to adapt regularly to their environment to keep performing well. This maintenance, however, brings both benefits and challenges. 

Benefits of monitoring 

  1. Improved model performance 
    Continuous model monitoring can detect changes in performance early. This allows you to make adjustments quickly, helping the model continue to perform at its best. 
  1. Greater reliability 
    Monitoring helps maintain the integrity of the model. This increases confidence in the decisions it makes, which is essential for organisations that rely on AI solutions. 
  1. Faster response to change 
    Monitoring enables organisations to respond proactively to changes in data or model behaviour. Rather than reacting to problems after they arise, you can take action in advance, making operations more efficient. 
     

Challenges and pitfalls 

  1. Practical implementation 
    Although monitoring sounds straightforward in theory, it can be difficult in practice. It requires robust infrastructure for real-time data processing and careful selection of monitoring tests. Setting up an effective monitoring solution can be particularly challenging for organisations without the right expertise or resources. 
  1. Volume of data 
    Tracking too many metrics can be overwhelming. Without a clear strategy, it can be difficult to gain meaningful insights from the abundance of data, making the monitoring process unnecessarily complex. 
  1. Over-sensitivity to small changes 
    When a monitoring solution is too sensitive, it can trigger overreactions to small changes in data. This can consume unnecessary time and resources as models are frequently retrained, investigated or further developed when there is no real need. 

Conclusion

Monitoring is an essential part of successfully implementing and maintaining a mature MLOps cycle. By focusing on the right metrics, drift detection and automation, organisations can optimise model performance and ensure their models remain reliable, even in a rapidly changing environment. Tools such as Evidently AI, Alibi Detect, NannyML, and the monitoring options offered by Azure and AWS can be valuable here. 

There are, however, practical challenges. Organisations need to reach a certain level of maturity and have robust infrastructure in place to use monitoring effectively. But the benefits, including improved performance, greater reliability and faster responses, are clear and compelling. Being aware of the potential pitfalls is crucial to successful implementation. 

Ready for the next step?

Want to know more about how your organisation can benefit from effective model monitoring and drift detection? Get in touch with us. We would be happy to help you take your MLOps infrastructure to the next level and get the most out of your AI solutions. MLOps makes that possible. 

Blog

You may also find this interesting,

Sign up for our newsletter.

Want to be the first to hear about a new blog post?

Enter a valid email address.