Three lessons from our AI projects
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
To avoid this scenario, it is important to identify which of the thousands of mutations could pose a serious threat, so that vaccine development can respond as quickly as possible. Artificial intelligence (AI) can help. Researchers at MIT have developed a new way to model ‘viral escapes’ using models originally built to analyse language: the familiar Natural Language Processing (NLP) models.
The idea behind using an NLP model in this case is as follows: viruses mutate in ways that follow the biological rules of protein structure, while also favouring their own survival. For example, SARS-CoV-2 mutates in the spike protein shown in the image below. People who have recently had COVID-19 or been vaccinated currently have antibodies that fit the spike protein. As a result, SARS-CoV-2 virus cells can no longer make contact with human cells, and someone becomes infected. The virus therefore aims to mutate its spike proteins quickly so that the antibodies no longer fit, but the receptors on human cells still do. It needs to mutate in a way that lets it escape the human immune system (the antibodies) without dying or losing its ability to reproduce. Similarly, for an NLP model, a sentence must not only have the right meaning (semantics), but also be grammatically correct (syntax). Using those same two principles, the researchers have creatively adapted NLP models to detect changes in the genetic code of viruses.
The example below illustrates how the NLP model estimates which virus mutations could lead to viral escape. The first sentence represents the virus before it mutates. The second sentence (from the left) shows a small mutation. Its meaning has barely changed, and the sentence is grammatically correct. With this mutation, the new virus still resembles the original closely enough for the immune system to recognise and attack it. No new antibodies are needed. The third sentence is grammatically incorrect. In the language of the virus, such a mutation would be considered unsuccessful. The last sentence is where the danger lies. It is grammatically correct and has the right semantics. These are the exceptions that the NLP model estimates could lead to viral escape. The researchers call the search for these exceptions 'constrained semantic change search' (CSCS).
Of course, the actual NLP model was not trained on sentences, but on the building blocks of various spike proteins from coronaviruses: amino acid sequences. The training set contained just under 1,000 sequences of the SARS-CoV-2 spike protein and a further 3,000 spike protein amino acid sequences from other types of coronavirus. The NLP model is shown in the image below. Internally, the model constructs a semantic representation, also known as an “embedding”, for a given amino acid sequence. The model’s output shows how well an amino acid fits the “grammar” of the sequence. In the image, the amino acid that fits the sequence best grammatically is marked with a capital A.
Of the 891 different coronavirus spike protein amino acid sequences the researchers examined with the model, one came from a strain that reinfected someone who had recovered from Covid-19 the previous year. This sequence was quickly ranked highly by the CSCS. Only three other sequences in the set showed both greater semantic change and greater so-called grammaticality. The researchers also fed some of the new variants into their algorithm and found that both the South African and British strains scored "quite high" for their likelihood of escape.
What now? If new mutations are tracked closely and the algorithm is used to assess the threat they pose, researchers can test suspicious strains in the laboratory as quickly as possible and adjust the vaccines accordingly. The test works as follows. Suspicious strains are brought together with antibodies. If the antibodies do not bind to the spike proteins, the current antibodies no longer offer protection. It is still unclear how much time vaccine developers would actually save with an AI-based approach like this. What we do know is that, in a pandemic on this scale, every second counts.
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
Many organisations have now run an AI pilot. The model works, the demo gets applause, and then nothing else happens. In our experience, most projects…
Artificial Intelligence is developing rapidly. New models appear almost every week, and more and more organisations are experimenting with AI. At the…
Want to be the first to hear about a new blog post?
Thanks for signing up!