Three lessons from our AI projects
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
How does GP-spel work?
The GP-spel was developed by NU.nl and runs for one Formula 1 season. Each weekend, a race takes place in a different country, with a qualifying session and the race itself. For every race weekend, participants must put together a team of four drivers and predict the top three. You can also score extra points by correctly predicting Max Verstappen’s position. In this blog, I’ll explain how I used mathematical models to optimise my team of four drivers. I won’t cover predicting the qualifying results and the race this time, although that would be an interesting data science application.
To put together a team of four drivers, you have a total budget of 100 million. The NU.nl team has set a cost for each driver: the better the driver, the more expensive they are. Lewis Hamilton costs a hefty 50 million, while Mick Schumacher – son of - costs ‘only’ 5 million. You earn points based on where the drivers in your team finish in qualifying and the race. For example, if Max Verstappen is in your team, takes pole position on Saturday and wins the race on Sunday, you earn (10 for pole position + 25 for winning the race) 35 points. The aim, of course, is to collect as many points as possible with your chosen team.
The problem described above — choosing the optimal team of drivers — is a knapsack problem. This is a well-known mathematical problem centred on the following question:
‘Given a set of items I, where each item i has an associated weight c_i and value w_i, determine which items to put in the knapsack to maximise the total value without exceeding the maximum weight.’
We can now formulate the task of putting together a team of drivers as a knapsack problem. We have a set of drivers I, each with a cost c_i. The aim is to maximise the total value of the team. However, we do not know how much value each driver will bring, so we need to come up with something ‘clever’ (I’ll return to this later). Besides the budget of 100 million, the game has a few additional constraints we need to account for:
The next step is to formulate our mathematical problem as an integer linear programming problem. In our field, we know that a knapsack problem can be solved using linear programming. Linear programming is a method for solving optimisation problems in which the objective function and constraints are linear. That is also true of our problem. We call it an integer linear programming problem because the solution is binary (and therefore integer-valued). For each driver i, we define a decision variable x_i. This variable equals 1 if driver i is chosen for the solution and 0 if we do not choose them. We can then formulate our problem as an integer linear programming problem (warning: mathematical formulas ahead!).
Championship points on 3 July 2021: 156 championship points (KP)
Position in free practice 1: 1 (P1)
Position in free practice 2: 3 (P2)
Position in free practice 3: 1 (P3)
Max’s value: KP + (21 – P1) + (21 – P2) + (21 – P3) = 156 + (21 – 1) + (21 – 3) + (21 – 1) = 214
We can calculate the value of the selected team by adding up the values of its drivers. Note that this matches the objective function in our mathematical formulation above, because the variable x_i is 0 when a driver is not in our team.
Next come the constraints. Constraint 1 states that the total cost of all the drivers we choose for our team must not exceed the budget of 100 million. We calculate the total cost by adding up the cost c_i of each driver i in our team. The same reasoning applies as for the objective function: if we do not choose a driver for our team, the variable x_i is 0, so their cost is not included. Constraint 2 ensures that we select exactly four drivers for our team. The third constraint ensures that we choose no more than one driver from each team. Suppose driver j is Max Verstappen (Red Bull team) and driver k is Sergio Perez (also Red Bull team). This constraint says that the sum of their decision variables can be at most 1. This means we cannot choose both drivers (as that would give x_j + x_k = 2).
This time, I’m using Excel rather than Python or R to build the model. Excel makes it easy to formulate and solve linear programming problems using Solver. Below is a screenshot of my Excel sheet.
| Constraint | Cells | Formula in Excel |
| You must choose 4 drivers | C27 <= D27 | SUM(J4:J23) <= 4 |
| You cannot spend more than 100 million | C28 <= D28 | SUMPRODUCT(D4:D23,J4:J23) <= 100 |
| You can choose only 1 driver per team | C29:C38 <= D29:D38 | For example, SUM(J4:J5) <= 1 for Mercedes. |
| Binary decision variables | J4:J23 | J4:J23 = binary |
Next, we can specify in Solver which algorithm to use to solve the problem. We choose Simplex LP because our objective function and constraints are linear. This gives us the following optimal solution:
Now that the race weekend in Austria is over, we can assess the result: did the model choose the best team? We can calculate this based on the qualifying and race results. My team scored 69 points in total: 27 for qualifying and 42 for the race. We can now have the model calculate the optimal team again, using the GP game’s points allocation based on the qualifying and race results as the value. This gives us the following optimal team:
There are, of course, several ways to improve my model, particularly how it determines the drivers’ values. For example, the calculation could be expanded to include qualifying results or performance on similar circuits. It could also be interesting to investigate whether the weighted function is the best way to determine value. The value function could potentially be optimised using machine learning, or value could be treated as stochastic rather than deterministic. That last approach in particular could add a lot to the model, as luck—good and bad—can play a major role in Formula 1.
Linear programming has all sorts of business applications. These include finding the optimal mix of products for a company to produce to maximise profit, creating a work schedule for hospital staff, finding the shortest route from A to B, or solving a transport problem. I believe operations research techniques, such as linear programming, are still underused in data science. Sometimes the most complex neural networks are trained when the same question could be answered with a much simpler model. In our work, we should therefore always weigh complexity against effectiveness. Is my model effective enough to win the NU.nl F1 game? We’ll find out at the end of this Formula 1 season ;).
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
Many organisations have now run an AI pilot. The model works, the demo gets applause, and then nothing else happens. In our experience, most projects…
Artificial Intelligence is developing rapidly. New models appear almost every week, and more and more organisations are experimenting with AI. At the…
Want to be the first to hear about a new blog post?
Thanks for signing up!