Skip to content
Go to insights Blog

Working efficiently with MLflow 3.0

Working efficiently with MLflow 3.0
Written by
Data Science Lab
Published on
17 July 2025

What’s new and how can you apply it in practice?

Why MLflow?

Machine learning projects are becoming more complex. Data scientists and data engineers often work together on models in a single team. It is important that those models remain reproducible, scalable and easy to manage throughout the process. With multiple models, datasets, notebooks and scripts, it is easy to lose track without a clear structure. MLOps has become indispensable because it standardises and automates the work, and makes it scalable. MLflow helps you stay in control of your machine learning workflow. 

With MLflow, you can track experiments, register models and automatically see performance metrics, such as accuracy, precision and recall, directly in your interface. No extra code, but more control. 

MLflow 3.0 was recently released with several new features that address the latest trends, such as GenAI and LLMs, and the need for larger, faster and reproducible data science models. 

What’s new in MLflow 3.0?

MLflow 3.0 introduces several important improvements that help you work faster, keep a clearer overview and use it in a wider range of applications. 

Introducing the LoggedModel architecture.

The biggest development in deep learning is improved model tracking. Although MLflow 2.0 already made it technically possible to log multiple model checkpoints, this was not standardised, and there was often no direct link between a logged model and the context in which it was trained and evaluated. The updated model tracking provides extensive traceability across models, runs and evaluation metrics through the LoggedModel architecture. This architecture lets you log multiple model checkpoints and metadata within a single run and track their performance across different datasets. This is particularly useful in deep learning workflows, where you want to save and compare models at different points in training. You can log multiple checkpoints during training and find them alongside performance data for different datasets. The step parameter in the logging functions lets you log checkpoints at specific points in training. Each logged model receives a unique model ID that you can easily refer to later. You can then use performance metrics to select the best-performing model from the checkpoints. 

In short, you no longer need to organise everything manually. MLflow does it for you. Each logged model variant gets a unique ID linked to its context, training step and evaluation results. You can easily find which model was logged when and how it performed.

Support for LLMs and GenAI

MLflow now offers:

  • Prompt logging and output tracking;
  • You can log prompts and generated outputs, making LLM experiments easier to reproduce and compare. 
  • Evaluation with the GenAI Evaluation Suite;
  • With the GenAI Evaluation Suite, you can manage GenAI applications as you do traditional machine learning models. It helps you not only measure the quality of these applications, but also improve and maintain it throughout the lifecycle. 
  • A prompt registry.
    MLflow 3.0 offers a prompt registry to help you create better, reusable prompts. You can automatically improve your prompts using evaluation feedback and labelled datasets. A separate blog post on how to use LLM evaluation effectively with MLflow 3.0 is coming soon.

Use case: Image classification with MobileNetV2

Data and preprocessing

We use the Intel Image Classification 

dataset of landscape photos from Kaggle, categorised into different classes. After splitting the data into training, validation and test sets, we use Keras for data augmentation. 

Screenshot of code for a layer that applies random data transformations.

Below are examples of what the images look like after augmentation.

Nine versions of the same mountain photo after different edits.

Next, we preprocess the data using the same pipeline as MobileNetV2.

MobileNetV2 model and MLflow integration

We choose MobileNetV2 for its balance of performance and speed. The model is lightweight and pre-trained, and we freeze the base layers because we do not want to retrain all the weights.

Screenshot of Python code building a Keras model.

Instead of autolog, we deliberately use manual logging. This gives you more control over what is and is not logged, and over the information you store about your model. This is essential for scalable evaluation. We log checkpoints for each epoch, different metrics for each dataset and parameters for each model variant. A custom MLflow callback logs everything automatically after each epoch.

Screenshot of code showing the log_model_metrics function.

This function is used in a custom callback class that logs information about the model and its associated training run after each epoch during training.

Screenshot of Python code with a callback that logs to MLflow at each epoch.

Model training and run management

Now that the model has been compiled, we’re ready to train it. We also want to use MLflow to track training and make the model’s performance visible, so we first need to set up a tracking experiment. This is where you specify where all training runs can be found and give the experiment a name.

Screenshot of code configuring the MLflow folder and experiment.

The model is ready to train, with each run tracked in MLflow. This lets you carry out multiple runs within one experiment and compare them later.  One of our runs looks like this:

Screenshot of Python code logging metrics to MLflow.
Screenshot of code logging the final metrics and a confusion matrix to MLflow.

After running the model, we can use the new functionality in MLflow 3.0 to retrieve the logged models and their results with the code below:

Screenshot of code ranking model checkpoints by accuracy.

MLflow 3.0 makes it easy to filter by performance to select the best model. The ability to apply filters is a useful addition. You can set conditions when searching for the best and most suitable model, using search terms such as “accuracy > 0.9”.  This lets you select the best-performing model each time, then save and deploy it.

MLflow UI

For this case study, we looked at which model performed best based on the epoch. On the Run page in the MLflow UI, we can see the run and the logged models with the correct names.

Screenshot of an MLflow run with parameters and metrics.

The Models tab shows all saved models and their metrics.

Screenshot of the MLflow experiments overview.

The artifacts tab contains all the additional files, results and figures that can be saved for each model, such as information about the environment in which the model was run or the model’s architecture.

Screenshot of the MLflow interface showing a model run.

MLflow also lets you add files/figures manually within a run. You can find them in the run’s Artifacts tab, where they can be useful for logging additional information, such as a confusion matrix.

Screenshot of a confusion matrix in MLflow.

Next steps

This use case shows how MLflow 3.0 can help you lay the groundwork for a structured, reproducible machine learning workflow. But there is much more you can do with MLOps to improve your platform further.

Want to integrate MLflow further into your organisation? Consider:

  • CI/CD integration for automatic retraining & deployment;
  • Using stage transitions in the model registry ("Staging", "Production") to manage your model lifecycle more effectively;
  • Setting up alerts for model drift through monitoring tools.

What you need depends on where you are in your MLOps process. The maturity of your data science process, the scale at which you work and the complexity of your models all matter.

Key takeaways

  • LLMs & GenAI: Several new features for LLMs, including prompt logging and evaluation through the GenAI evaluation suite.
  • Deliberate evaluation: Explicit logging instead of autolog encourages more deliberate model evaluation.
  • With the extended LoggedModel architecture, the model is central to the entire MLflow lifecycle. This gives you control not only over experiments, but also over managing, versioning and deploying models to production, so you can set up the full development workflow flexibly and at scale.
  • UI: A central environment where all results for models and runs are clearly presented in a UI, making it quick and easy to compare models.
  • Selection: Easily filter for the best-performing model for each use case.

Want to get started with MLOps and deep learning in your organisation too?

At DSL, we're happy to think things through with you. Book an introductory call.

Blog

You may also find this interesting,

Sign up for our newsletter.

Want to be the first to hear about a new blog post?

Enter a valid email address.