Three lessons from our AI projects
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
How does the GP game work?
The GP game was developed by NU.nl and runs for one Formula 1 season. Each race weekend takes place in a different country and consists of a qualifying session and the race itself. For every race weekend, players must put together a team of four drivers and predict the top three. You can also score extra points by correctly predicting Max Verstappen’s position. In this blog, I’ll explain how I used mathematical models to optimise my team of four drivers. I’ll leave predictions for qualifying and the race aside this time, although they would make an interesting data science application.
To put together a team of four drivers, you have a total budget of 100 million. The team at NU.nl has set a price for each driver: the better the driver, the higher the price. Lewis Hamilton, for example, costs no less than 50 million, while Mick Schumacher – son of - costs ‘only’ 5 million. You earn points based on where the drivers in your team finish in qualifying and the race. For example, if Max Verstappen is in your team, takes pole position on Saturday and wins the race on Sunday, you earn (10 for pole position + 25 for winning the race) 35 points. The aim, of course, is to score as many points as possible with your chosen team.
The problem described above – choosing the optimal team of drivers – is a knapsack problem. This is a well-known mathematical problem centred on the following question:
‘Given a set of items I, where each item i has an associated weight c_i and value w_i, determine which items to put in the knapsack so that the total value is as high as possible without exceeding the maximum weight.’
We can now formulate the task of assembling a team of drivers as a knapsack problem. We have a set of drivers I, each with a cost c_i. The aim is to maximise the team’s total value. But we do not know the value each driver will deliver, so we need to come up with something ‘clever’ for that (I’ll return to this later). Alongside the budget of 100 million, the game imposes a few additional constraints that we need to account for:
The next step is to formulate our mathematical problem as an integer linear programming problem. In our field, we know that a knapsack problem can be solved using linear programming. Linear programming is a method for solving optimisation problems in which the objective function and constraints are linear. That is also true of our problem. We call it an integer linear programming problem because the solution is binary (and therefore integer-valued). For each driver i, we define a decision variable x_i. This variable equals 1 if driver i is chosen in the solution, and 0 if we do not choose them. We can then formulate our problem as an integer linear programming problem (warning: mathematical formulas ahead!).
Championship points on 3 July 2021: 156 championship points (KP)
Free practice 1 position: 1 (P1)
Free practice 2 position: 3 (P2)
Free practice 3 position: 1 (P3)
Max’s value: KP + (21 – P1) + (21 – P2) + (21 – P3) = 156 + (21 – 1) + (21 – 3) + (21 – 1) = 214
We can calculate the value of the selected team by adding up the values of its drivers. Note that this matches the objective function in our mathematical formulation above, since the variable x_i equals 0 if a driver is not in our team.
Next come the constraints. Constraint 1 states that the total cost of all the drivers we choose for our team must not exceed the budget of 100 million. We calculate the total cost by adding up the cost c_i of each driver i in our team. The same reasoning applies as with the objective function: if we do not choose a driver for our team, the variable x_i equals 0 and we do not include that cost. Constraint 2 ensures that we select exactly four drivers for our team. The third constraint ensures that we choose no more than one driver from each team. Suppose driver j is Max Verstappen (team Red Bull) and driver k is Sergio Perez (also team Red Bull). This constraint says that the sum of their decision variables can be at most 1. This means we cannot choose both drivers (because then x_j + x_k = 2).
For the modelling, this time I’m choosing Excel rather than Python or R. Excel makes it easy to formulate and solve linear programming problems using Solver. Below is a screenshot of my Excel spreadsheet.
| Constraint | Cells | Formula in Excel |
| You must choose 4 drivers | C27 <= D27 | SUM(J4:J23) <= 4 |
| You cannot spend more than 100 million | C28 <= D28 | SUMPRODUCT(D4:D23,J4:J23) <= 100 |
| You can choose only 1 driver per team | C29:C38 <= D29:D38 | For example, SUM(J4:J5) <= 1 for Mercedes. |
| Binary decision variables | J4:J23 | J4:J23 = binary |
Next, we can tell Solver which algorithm to use to solve the problem. We choose Simplex LP because our objective function and constraints are linear. This gives us the following optimal solution:
Now that the race weekend in Austria is over, we can assess the result: did the model choose the best team? We can calculate this using the qualifying and race results. My team scored 69 points in total: 27 from qualifying and 42 from the race. We can now have the model calculate the optimal team again, this time using the GP game’s points allocation based on the qualifying and race results as the value. This gives us the following optimal team:
There are still several ways to improve my model, particularly how it determines the drivers’ value. For example, the calculation could include qualifying results or performance on similar circuits. It would also be worth investigating whether a weighted function is the best way to determine value. The value function could potentially be optimised using machine learning, or value could be treated as stochastic rather than deterministic. The latter in particular could add a great deal to the model, as luck, good or bad, can play a major role in Formula 1.
Linear programming has all kinds of business applications. These include finding the optimal mix of products for a company to produce to maximise profit, drawing up a work schedule for hospital staff, finding the shortest route from A to B, or solving a transport problem. I believe operations research techniques such as linear programming are still underused in data science. Sometimes highly complex neural networks are trained when a much simpler model could answer the same question. In our work, we must therefore keep weighing complexity against effectiveness. Is my model effective enough to win the NU.nl F1 game? We’ll find out at the end of this Formula 1 season ;).
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
Many organisations have now run an AI pilot. The model works, the demo gets applause, and then nothing else happens. In our experience, most projects…
Artificial Intelligence is developing rapidly. New models appear almost every week, and more and more organisations are experimenting with AI. At the…
Want to be the first to hear about a new blog post?
Thanks for signing up!