Three lessons from our AI projects
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
AI has considerable power, and with it comes great responsibility. Ethical questions about AI are no longer purely philosophical. They now have practical implications that governments and businesses need to examine. In response, the EU has set out seven pillars of Responsible AI[1]:
This blog focuses on the steps companies can take to develop and deploy Responsible AI. There is much more to say about ethics and AI beyond this discussion. We will not look at systems such as self-driving cars or drones, which face ethical dilemmas. Instead, we will look at everyday algorithms that our society encounters daily.
First, let’s discuss bias. In this context, it means a distortion of reality. Poor data leads to poor models and outcomes. Suppose an algorithm is used to recognise dogs in photos, but the model repeatedly misses photos that clearly show a dog or classifies a blueberry muffin as a chihuahua [2], that gives you reason to think the model is not working as well as it should. Perhaps the input data was labelled incorrectly or was not representative, the model was configured incorrectly, or there were problems with overfitting or underfittin
The consequences of a misclassification in these examples are not far-reaching. But what if a facial recognition model is used to identify criminals? If the model wrongly identifies someone as a criminal, the consequences could unfairly change that person’s life. The problem may stem from how the algorithm is optimised, but a more insidious cause could be bias embedded in the training data. Machines learn only from what you show them – they do not account for nuanced contextual or cultural factors. An algorithm that assesses whether a female applicant might be suitable for a role may give her a lower score if the model was trained on historical data in which men are overrepresented. Conversely, datasets in which minority groups are overrepresented can lead an algorithm to label people from those groups as ‘higher risk’. This could create a situation in which a group is more likely to end up in prison or receive a longer sentence (which causes these groups to be overrepresented in new data, which increases the risk of these groups ending up in prison more often, which…).
Sources of potential bias in machine learning (ML) systems include [3]:
It is important for organisations to think carefully about the data used to train a model and any potential sources of bias in the dataset. The data must be representative. Ideally, bias should be prevented altogether. Continuing to monitor the algorithm for bias can help minimise its influence on decisions. Data scientists develop the algorithms, but the responsibility cannot rest entirely on their shoulders. Everyone involved in an AI project has a role in actively preventing bias so that decisions are sound and ethical.
A previously published blog post about explainability and explainable AI (XAI) and its importance describes how XAI provides insight into how decisions are made and which features play a role. This ‘explainability’ can help prevent unfair outcomes and identify errors in the system early. Explainability is related to transparency. While explainability mainly concerns a model’s output, transparency is more about the processes leading up to training and deploying an algorithm: the input data, analytical choices and how the algorithm works.
As an end user, you usually have little visibility into the training data, where it came from, how it was cleaned or exactly which features the model was trained on. You ‘just have to trust’ that the algorithm works in your interests. Being transparent with the public and allowing end users to provide feedback may help build trust in algorithms. The municipality of Amsterdam offers an example with its algorithm register. The idea is that automated public services should respect the same principles as the municipality’s other services – including being open and accountable, treating people equally and not restricting their freedom or control over decisions that affect them.
Full transparency is not always possible, for example when algorithms are, or must remain, confidential. Consider fraud detection by the Dutch Tax and Customs Administration. If it disclosed its algorithm, fraudsters could complete their tax returns in a way that lets them avoid detection. For commercial companies, releasing an algorithm is not always desirable either. Google has proposed using ‘model cards’ to give both experts and non-experts insight into models. Model cards can include information about a model’s performance, the data used, its limitations, ethical risks and the strategies used to address them. A Model Card Toolkit is available that works with scikit-learn to create reports like these. This approach could allow companies that do not want to reveal their entire algorithm to be more transparent about the algorithms they use. Even so, transparency is like ‘looking through a window’: you only see what the window lets you see.
Under Article 22 of the GDPR, as a rule, an algorithm alone may not make the final decision when that decision could have a significant impact on a person’s life:
“The data subject shall have the right not to be subject to a decision based solely on automated processing, including profiling, which produces legal effects concerning him or her or similarly significantly affects him or her.”
In practice, people often find it difficult to override an algorithm. After all, they may assume the computer knows better. It is therefore important that automatic approval is not enabled when a person needs to review an algorithm’s decision.
We live in a society where algorithms continually influence what we do. Algorithms also increasingly make decisions that affect our lives, for example when we apply for a loan or mortgage. People, not the algorithm itself, bear the greatest responsibility for the ethical implications of those decisions. That is why it is important to consider which data you use to train a model and to prevent potential bias from leading to unfair decisions. Within an organisation, responsibility for this should not rest entirely with data scientists; there should be a dialogue with everyone involved in the project. Being transparent about the choices made when developing an AI system makes it possible to get feedback and gives end users more confidence in the system. It is up to the organisation to document all the considerations behind those choices. People must also always be able to reverse a decision prepared by an AI system. There is a great deal of trust in AI, and we must work to ensure that trust is justified.
Would you like to know more about this topic or see examples of how we approach it in our projects? Get in touch for more information, with no obligation.
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
Many organisations have now run an AI pilot. The model works, the demo gets applause, and then nothing else happens. In our experience, most projects…
Artificial Intelligence is developing rapidly. New models appear almost every week, and more and more organisations are experimenting with AI. At the…
Want to be the first to hear about a new blog post?
Thanks for signing up!