Skip to content
Go to insights Blog

Predicting Formula 1 races

Predicting Formula 1 races
Written by
Data Science Lab
Published on
13 July 2021
As a Formula 1 fan, I spend every race weekend on the edge of my seat. I’m cheering on Max Verstappen, of course, but I’m also trying to beat my family in NU.nl’s GP game. Although I’m certain I know more about Formula 1 than anyone else in my family, I often let maths give me a hand.

How does the GP game work?

The GP game was developed by NU.nl and runs for one Formula 1 season. Each race weekend takes place in a different country and consists of a qualifying session and the race itself. For every race weekend, players must put together a team of four drivers and predict the top three. You can also score extra points by correctly predicting Max Verstappen’s position. In this blog, I’ll explain how I used mathematical models to optimise my team of four drivers. I’ll leave predictions for qualifying and the race aside this time, although they would make an interesting data science application.

To put together a team of four drivers, you have a total budget of 100 million. The team at NU.nl has set a price for each driver: the better the driver, the higher the price. Lewis Hamilton, for example, costs no less than 50 million, while Mick Schumacher – son of - costs ‘only’ 5 million. You earn points based on where the drivers in your team finish in qualifying and the race. For example, if Max Verstappen is in your team, takes pole position on Saturday and wins the race on Sunday, you earn (10 for pole position + 25 for winning the race) 35 points. The aim, of course, is to score as many points as possible with your chosen team.

Knapsack problem

The problem described above – choosing the optimal team of drivers – is a knapsack problem. This is a well-known mathematical problem centred on the following question:

‘Given a set of items I, where each item i has an associated weight c_i and value w_i, determine which items to put in the knapsack so that the total value is as high as possible without exceeding the maximum weight.’

Illustration of a backpack surrounded by weights from 1 to 20 kilos.

The aim is to maximise the value of the items in the knapsack; we call this the objective function. At the same time, the maximum weight must not be exceeded; we call this a constraint.

The GP game as a knapsack problem

We can now formulate the task of assembling a team of drivers as a knapsack problem. We have a set of drivers I, each with a cost c_i. The aim is to maximise the team’s total value. But we do not know the value each driver will deliver, so we need to come up with something ‘clever’ for that (I’ll return to this later). Alongside the budget of 100 million, the game imposes a few additional constraints that we need to account for:

  1. You can choose at most one driver from each team.
  2. You must choose exactly four drivers.
  3. The solution to the problem is binary; you either choose a driver (1) or you do not (0).

The GP game as an ILP problem

The next step is to formulate our mathematical problem as an integer linear programming problem. In our field, we know that a knapsack problem can be solved using linear programming. Linear programming is a method for solving optimisation problems in which the objective function and constraints are linear. That is also true of our problem. We call it an integer linear programming problem because the solution is binary (and therefore integer-valued). For each driver i, we define a decision variable x_i. This variable equals 1 if driver i is chosen in the solution, and 0 if we do not choose them. We can then formulate our problem as an integer linear programming problem (warning: mathematical formulas ahead!).

Mathematical formulas for an objective function with constraints.

The mathematical formulation of our problem looks quite complicated, but it is actually straightforward. Let’s start with the objective function. As mentioned earlier, we want to maximise the value of our team. For each driver i (20 in total), we therefore assign a value w_i. This value is based on the current championship standings and the results of the free practice sessions. That way, I account for a driver’s performance both over the season and during the race weekend. For example, Max Verstappen’s value for the race weekend in Austria (2 to 4 July 2021) can be calculated as follows:

Championship points on 3 July 2021: 156 championship points (KP)
Free practice 1 position: 1 (P1)
Free practice 2 position: 3 (P2)
Free practice 3 position: 1 (P3)
Max’s value: KP + (21 – P1) + (21 – P2) + (21 – P3) = 156 + (21 – 1) + (21 – 3) + (21 – 1) = 214

We can calculate the value of the selected team by adding up the values of its drivers. Note that this matches the objective function in our mathematical formulation above, since the variable x_i equals 0 if a driver is not in our team.

Next come the constraints. Constraint 1 states that the total cost of all the drivers we choose for our team must not exceed the budget of 100 million. We calculate the total cost by adding up the cost c_i of each driver i in our team. The same reasoning applies as with the objective function: if we do not choose a driver for our team, the variable x_i equals 0 and we do not include that cost. Constraint 2 ensures that we select exactly four drivers for our team. The third constraint ensures that we choose no more than one driver from each team. Suppose driver j is Max Verstappen (team Red Bull) and driver k is Sergio Perez (also team Red Bull). This constraint says that the sum of their decision variables can be at most 1. This means we cannot choose both drivers (because then x_j + x_k = 2).

Calculating the optimal solution

For the modelling, this time I’m choosing Excel rather than Python or R. Excel makes it easy to formulate and solve linear programming problems using Solver. Below is a screenshot of my Excel spreadsheet.

Screenshot of an Excel spreadsheet with the Solver window open.

First, we list all the data we need, such as each driver’s cost and value. We then use the data to specify our integer linear programming problem. We need to enter the right cells in the right places in Solver:

  • Objective function (‘Set Objective’)
    The objective function is in cell C26. This cell contains the following formula: SUMPRODUCT(I4:I23, J4:J23). We want to maximise the objective function.
  • Decision variables (‘By Changing Variable Cells’)
    The decision variables are in cells J4:J23.
  • Constraints (‘Subject to the Constraints’)
    Here we add all the constraints:
Constraint Cells Formula in Excel
You must choose 4 drivers C27 <= D27 SUM(J4:J23) <= 4
You cannot spend more than 100 million C28 <= D28 SUMPRODUCT(D4:D23,J4:J23) <= 100
You can choose only 1 driver per team C29:C38 <= D29:D38 For example, SUM(J4:J5) <= 1 for Mercedes.
Binary decision variables J4:J23 J4:J23 = binary

Next, we can tell Solver which algorithm to use to solve the problem. We choose Simplex LP because our objective function and constraints are linear. This gives us the following optimal solution:

Four cards featuring Formula 1 drivers, their value and their results this season.

This solution has a value of 472 points. I decide to trust my model completely for the race weekend in Austria and choose these drivers for my team. Fingers crossed…

Was this the best team?

Now that the race weekend in Austria is over, we can assess the result: did the model choose the best team? We can calculate this using the qualifying and race results. My team scored 69 points in total: 27 from qualifying and 42 from the race. We can now have the model calculate the optimal team again, this time using the GP game’s points allocation based on the qualifying and race results as the value. This gives us the following optimal team:

Four cards featuring Formula 1 drivers, their value and their results this season. This team scored 4 more points than the team I chose, bringing its total to 73. What is clear is that you should have picked Max Verstappen and Lando Norris for your team: together, they earned 49 points. Unfortunately, in hindsight, the model did not pick the optimal team for the race weekend in Austria. This was not down to the model itself or the simplex method, but to the fact that we estimated the drivers’ value before qualifying using the value function. Of course, to determine whether the model is statistically better than a human participant, we need to look at more than one race weekend. One thing is certain: my team scored the most points in our family league.

Areas for improvement

There are still several ways to improve my model, particularly how it determines the drivers’ value. For example, the calculation could include qualifying results or performance on similar circuits. It would also be worth investigating whether a weighted function is the best way to determine value. The value function could potentially be optimised using machine learning, or value could be treated as stochastic rather than deterministic. The latter in particular could add a great deal to the model, as luck, good or bad, can play a major role in Formula 1.

Applications of linear programming

Linear programming has all kinds of business applications. These include finding the optimal mix of products for a company to produce to maximise profit, drawing up a work schedule for hospital staff, finding the shortest route from A to B, or solving a transport problem. I believe operations research techniques such as linear programming are still underused in data science. Sometimes highly complex neural networks are trained when a much simpler model could answer the same question. In our work, we must therefore keep weighing complexity against effectiveness. Is my model effective enough to win the NU.nl F1 game? We’ll find out at the end of this Formula 1 season ;).

Blog

You may also find this interesting,

Sign up for our newsletter.

Want to be the first to hear about a new blog post?

Enter a valid email address.