Skip to content
To news Blog

Predicting Formula 1 races

Predicting Formula 1 races
Written by
Data Science Lab
Published on
13 July 2021
As a Formula 1 fan, I spend every race weekend on the edge of my seat. I’m cheering on Max Verstappen, of course, but I’m also trying to beat my family members at NU.nl’s GP-spel. Although I’m certain I know more about Formula 1 than anyone else in my family, I often let maths give me a helping hand.

How does GP-spel work?

The GP-spel was developed by NU.nl and runs for one Formula 1 season. Each weekend, a race takes place in a different country, with a qualifying session and the race itself. For every race weekend, participants must put together a team of four drivers and predict the top three. You can also score extra points by correctly predicting Max Verstappen’s position. In this blog, I’ll explain how I used mathematical models to optimise my team of four drivers. I won’t cover predicting the qualifying results and the race this time, although that would be an interesting data science application.

To put together a team of four drivers, you have a total budget of 100 million. The NU.nl team has set a cost for each driver: the better the driver, the more expensive they are. Lewis Hamilton costs a hefty 50 million, while Mick Schumacher – son of - costs ‘only’ 5 million. You earn points based on where the drivers in your team finish in qualifying and the race. For example, if Max Verstappen is in your team, takes pole position on Saturday and wins the race on Sunday, you earn (10 for pole position + 25 for winning the race) 35 points. The aim, of course, is to collect as many points as possible with your chosen team.

Knapsack problem

The problem described above — choosing the optimal team of drivers — is a knapsack problem. This is a well-known mathematical problem centred on the following question:

‘Given a set of items I, where each item i has an associated weight c_i and value w_i, determine which items to put in the knapsack to maximise the total value without exceeding the maximum weight.’

Illustration of a backpack surrounded by weights from 1 to 20 kilos.

The aim is to maximise the value of the items in the knapsack; we call this the objective function. At the same time, the maximum weight must not be exceeded; we call this a constraint.

The GP game as a knapsack problem

We can now formulate the task of putting together a team of drivers as a knapsack problem. We have a set of drivers I, each with a cost c_i. The aim is to maximise the total value of the team. However, we do not know how much value each driver will bring, so we need to come up with something ‘clever’ (I’ll return to this later). Besides the budget of 100 million, the game has a few additional constraints we need to account for:

  1. You can choose at most one driver from each team.
  2. You must choose exactly four drivers.
  3. The solution to the problem is binary: you either choose a driver (1) or you do not (0).

The GP game as an ILP problem

The next step is to formulate our mathematical problem as an integer linear programming problem. In our field, we know that a knapsack problem can be solved using linear programming. Linear programming is a method for solving optimisation problems in which the objective function and constraints are linear. That is also true of our problem. We call it an integer linear programming problem because the solution is binary (and therefore integer-valued). For each driver i, we define a decision variable x_i. This variable equals 1 if driver i is chosen for the solution and 0 if we do not choose them. We can then formulate our problem as an integer linear programming problem (warning: mathematical formulas ahead!).

Mathematical formulas for an objective function with constraints.

The mathematical formulation of our problem looks quite complicated, but it is actually straightforward. Let’s start with the objective function. As mentioned earlier, we want to maximise the value of our team. For each driver i (20 in total), we therefore calculate a value w_i. This value is based on the current championship standings and the free practice results. That way, I account for a driver’s performance both over the season and during the race weekend. For example, Max Verstappen’s value for the race weekend in Austria (2 to 4 July 2021) can be calculated as follows:

Championship points on 3 July 2021: 156 championship points (KP)
Position in free practice 1: 1 (P1)
Position in free practice 2: 3 (P2)
Position in free practice 3: 1 (P3)
Max’s value: KP + (21 – P1) + (21 – P2) + (21 – P3) = 156 + (21 – 1) + (21 – 3) + (21 – 1) = 214

We can calculate the value of the selected team by adding up the values of its drivers. Note that this matches the objective function in our mathematical formulation above, because the variable x_i is 0 when a driver is not in our team.

Next come the constraints. Constraint 1 states that the total cost of all the drivers we choose for our team must not exceed the budget of 100 million. We calculate the total cost by adding up the cost c_i of each driver i in our team. The same reasoning applies as for the objective function: if we do not choose a driver for our team, the variable x_i is 0, so their cost is not included. Constraint 2 ensures that we select exactly four drivers for our team. The third constraint ensures that we choose no more than one driver from each team. Suppose driver j is Max Verstappen (Red Bull team) and driver k is Sergio Perez (also Red Bull team). This constraint says that the sum of their decision variables can be at most 1. This means we cannot choose both drivers (as that would give x_j + x_k = 2).

Calculating the optimal solution

This time, I’m using Excel rather than Python or R to build the model. Excel makes it easy to formulate and solve linear programming problems using Solver. Below is a screenshot of my Excel sheet.

Screenshot of an Excel spreadsheet with the Solver window open.

First, we list all the data we need, such as each driver’s cost and value. We then use the data to specify our integer linear programming problem. We need to enter the right cells in the right places in Solver:

  • Objective function (‘Set Objective’)
    The objective function is in cell C26. This cell contains the following formula: SUMPRODUCT(I4:I23, J4:J23). We want to maximise the objective function.
  • Decision variables (‘By Changing Variable Cells’)
    The decision variables are in cells J4:J23.
  • Constraints (‘Subject to the Constraints’)
    Here we add all the constraints:
Constraint Cells Formula in Excel
You must choose 4 drivers C27 <= D27 SUM(J4:J23) <= 4
You cannot spend more than 100 million C28 <= D28 SUMPRODUCT(D4:D23,J4:J23) <= 100
You can choose only 1 driver per team C29:C38 <= D29:D38 For example, SUM(J4:J5) <= 1 for Mercedes.
Binary decision variables J4:J23 J4:J23 = binary

Next, we can specify in Solver which algorithm to use to solve the problem. We choose Simplex LP because our objective function and constraints are linear. This gives us the following optimal solution:

Four cards featuring Formula 1 drivers, their value and their results this season.

This solution has a value of 472 points. For the race weekend in Austria, I decide to trust my model completely and choose these drivers for my team. Fingers crossed…

Was this the best team?

Now that the race weekend in Austria is over, we can assess the result: did the model choose the best team? We can calculate this based on the qualifying and race results. My team scored 69 points in total: 27 for qualifying and 42 for the race. We can now have the model calculate the optimal team again, using the GP game’s points allocation based on the qualifying and race results as the value. This gives us the following optimal team:

Four cards featuring Formula 1 drivers, their value and their results this season. This team scored 4 more points than the team I chose, with 73 points. What is clear is that you should have picked Max Verstappen and Lando Norris for your team: together, they account for 49 points. Unfortunately, in hindsight, the model did not pick the optimal team for the race weekend in Austria. This was not due to the model itself or the simplex method, but to the fact that we estimated the drivers’ values before qualifying using the value function. Of course, to determine whether the model is statistically better than a human participant, we need to look at more than one race weekend. One thing is certain: I scored the most points in our family league with my team.

Areas for improvement

There are, of course, several ways to improve my model, particularly how it determines the drivers’ values. For example, the calculation could be expanded to include qualifying results or performance on similar circuits. It could also be interesting to investigate whether the weighted function is the best way to determine value. The value function could potentially be optimised using machine learning, or value could be treated as stochastic rather than deterministic. That last approach in particular could add a lot to the model, as luck—good and bad—can play a major role in Formula 1.

Applications of linear programming

Linear programming has all sorts of business applications. These include finding the optimal mix of products for a company to produce to maximise profit, creating a work schedule for hospital staff, finding the shortest route from A to B, or solving a transport problem. I believe operations research techniques, such as linear programming, are still underused in data science. Sometimes the most complex neural networks are trained when the same question could be answered with a much simpler model. In our work, we should therefore always weigh complexity against effectiveness. Is my model effective enough to win the NU.nl F1 game? We’ll find out at the end of this Formula 1 season ;).

Blog

You may also find this interesting,

Sign up for our newsletter.

Want to be the first to hear about a new blog post?

Enter a valid email address.