Three lessons from our AI projects
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
A lease car’s residual value is important information for leasing companies because it has a significant bearing on the lease rates they offer customers. An accurate estimate of a new car’s residual value can help a leasing company set an appropriate lease rate and maximise its profit. Today, a wealth of information is available online, including general sentiment, news articles and posts on internet forums. These can indicate the popularity of particular cars or car brands. This large volume of text data can be combined with the vehicle data the leasing company already holds, such as make, engine capacity and number of doors, to predict residual value as accurately as possible using machine learning models.
The data you use first needs to be converted into numbers. This matters because raw text containing words and letters cannot be used as input for a machine learning model. For NLP, word embeddings often work well as a way to represent text. The best-known word embedding models are word2vec, fastText and GloVe. When choosing a machine learning model for an NLP problem, neural networks are often a good option. Long Short-Term Memory (LSTM) is a well-known method, but a Convolutional Neural Network (CNN) also performs well when making predictions from text.
Finally, choosing the right time window for the text data is important. To calculate a car’s residual value on 1 September 2021, we could include social media posts, car forum discussions or news articles from the past month in our model. But we do not know whether one month is enough; we may need at least 2, 3 or 4 months.
A common problem in organisations is that data is not standardised. This means that the same value within a field is not described consistently throughout the data. If a dataset contains both “Mercedes-Benz C Klasse” and “Mercedes Classe C”, we can quickly see that they refer to the same type of car, while a database or computer will generally treat them as two different things. This creates inconsistencies when modelling the data and generating insights. To build robust predictive models and generate accurate insights, we need the data to be consistent. We want to convert all “unclean” data to a standardised name. To do this, we first need to decide how we ideally want to describe the data and create a standardised list. This list contains all the make and model names, written as we would ideally want them to appear. If “Mercedes-Benz C Klasse” is a standard name on that list, we want to convert data points labelled, for example, “Mercedes C-Klasse”, “Mercedes C220”, “Benz Klasse C” or “Mercedes-Benz C220d Automaat” to that name. We want to minimise manual corrections to the data, so we look for a way to automate the standardisation process.
We could approach this process in several ways. One simple but effective option is to use fuzzy string matching. This lets you match the input data—in our case, the make and model names you want to convert—against the names in your standardised list. Fuzzy string matching is a method for matching two pieces of text that are similar, or partly similar, rather than exactly the same. One measure of the similarity between two pieces of text is the Levenshtein distance. In essence, it indicates how many characters need to be changed to make the two pieces of text identical.
One advantage of fuzzy string matching is that it is relatively straightforward to implement. Unlike conventional machine learning models, fuzzy matching does not require you to train a model on the data. You can match pieces of text directly. This also means that data points that occur infrequently can potentially be matched successfully, whereas machine learning models often need a lot of data before they can classify certain labels. One drawback of fuzzy matching is that texts need to bear some resemblance to each other before they can be matched. With examples such as “C-Klasse”, “C Classe” and “Klasse C”, we can be fairly confident of finding a successful match. But with examples such as “C220” and “C-Klasse”, finding a match quickly becomes much harder, even though they refer to the same car model.
The previous example is well suited to traditional machine learning models. Models such as a Random Forest are often successful at finding relationships that are not immediately obvious. With enough training data, a Random Forest could quickly learn that a “C220” is a “C-Class”. Alongside tree-based models such as Random Forest or XGBoost, neural networks such as LSTMs often perform well on this type of problem. Besides choosing the type of predictive model, we first need to decide how to represent the words in a text. One way is to use a TF-IDF matrix. This matrix includes every word in the dataset, with a numerical value for each word. That value reflects how ‘important’ a word is in a piece of text compared with all the texts in the dataset.
As mentioned earlier, one advantage of these predictive models is that they are often very accurate, provided there is enough data (which also makes this requirement a drawback). They are often less effective at identifying outcomes that rarely occur in the data or are entirely new. Regularly retraining the model is therefore often necessary.
Accidents can happen when you least expect them — we have all heard that often enough. Driving carries a considerable risk of accidents, and damage to the vehicle itself can be costly. For leased cars, liability for damage often rests with the driver. Checking for damage at the end of a lease term is a process that still involves many manual steps and has considerable potential for automation.
Object detection algorithms such as R-CNN and YOLO can use machine learning and cameras to identify new damage. Before a new lease car is delivered to its driver, cameras can inspect it to establish its condition when new. At the end of the lease, the cameras can inspect the car again and compare it with that initial condition. This makes it possible to determine quickly, with minimal manual work, whether the car was damaged during the lease. The findings can also provide an immediate estimate of the potential repair costs.
Many employers offer employees a fuel card to make it easy to pay for business mileage. Employees then do not need to keep track of their business mileage and claim it back afterwards. Unfortunately, fuel cards also carry a risk of fraud. Criminals can skim a fuel card and use it to buy fuel at considerable cost to the employer. Employees can also commit fraud themselves, for example by regularly lending their card to family members or friends.
To detect this type of fraud in time and avoid high costs, we can use machine learning methods to predict whether a transaction is fraudulent or legitimate. Fraud detection is a common challenge in data science. These solutions often focus on credit cards, but the underlying techniques can be used for any type of transaction, provided enough data is available. Algorithms such as XGBoost can detect potential fraud when sufficient labelled data is available. Even without labels, we can use unsupervised methods such as Isolation Forest and Random Cut Forest to identify anomalies (fraud) in transaction data.
Does your organisation handle a lot of repetitive questions? Do you want to start a data science project, or find out which data-driven solutions are right for you? Get in touch with no obligation to explore how we can achieve your data-driven goals. Together, we create the future.
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
Many organisations have now run an AI pilot. The model works, the demo gets applause, and then nothing else happens. In our experience, most projects…
Artificial Intelligence is developing rapidly. New models appear almost every week, and more and more organisations are experimenting with AI. At the…
Want to be the first to hear about a new blog post?
Thanks for signing up!