Skip to content
Go to insights Blog

The power of embeddings

The power of embeddings
Written by
Data Science Lab
Published on
28 June 2021
Natural Language Processing (NLP, also known as text mining) is an umbrella term for techniques that extract insights from text data. These techniques are used in practice with varying degrees of success. In recent years, rapid developments in neural networks (particularly by tech giants such as Google) have led to several major breakthroughs. How do they work? How useful are these new techniques in practice? And do they work with Dutch-language text data? That is too much to cover in one blog post, so this will be a two-part series. In this first part, we take a closer look at embeddings: a way of representing text data that underpins recent developments. The second part looks at practical applications and specific possibilities for Dutch-language text.

NLP is not a recent development; it has been around for a long time. Conceptually, it dates back to the 1950s, and in practice to around the 1990s. Topics such as Named Entity Recognition, Part-of-Speech tagging and stemming & lemmatisation are standard components of NLP toolkits. The main problems that text data is used to address include semantic search, sentiment analysis, topic modelling, classification and clustering. Common techniques include bag-of-words, tf-idf, Naïve Bayes and LDA. What we have seen in recent years, however, is a shift towards NLP techniques based on neural networks. This began with the introduction of word2vec[1] in 2013: a kind of dictionary that lets you translate a word into a sequence of numbers that captures its meaning (word embeddings). These embeddings were widely used in recurrent neural networks. The principle of attention[2] made neural networks better able to learn to retain long-term relationships within passages of text. The arrival of the transformer[3] improved this considerably, and today the focus is on pre-trained language models from the BERT family[4]. This has finally made transfer learning a reality in NLP, as pre-trained convolutional neural networks had already done in computer vision. In chronological order, the most important academic developments are therefore:

Timeline from word embeddings in 2013 to pre-trained language models in 2018.

But as mentioned, the rapid progress began with word2vec and word embeddings. Why is this so important? Because it radically changes how text is represented.

Representations of text data

A computer cannot calculate with text, only with numbers. If we want to apply algorithms to text data, we first need to ‘translate’ that text into numbers. In other words, we need to find a numerical representation.

Diagram showing individual letters passing through a matrix to a neural network.

Bag-of-words

The bag-of-words (BoW) representation has been widely used for years. The approach is simple. Suppose you have a set of documents (pieces of text). You then:

  • Identify which words occur
  • Use them to build a vocabulary
  • For each document: count how often each word in the vocabulary occurs
  • Each document is now a long sequence of numbers
  • Your ‘corpus’ (text dataset) is now a matrix of numbers
  • B. Stop words are often filtered out of the vocabulary

Photo of a woman with a dog, with detection boxes around both.

This matrix of numbers can now be used with all kinds of algorithms for classification, clustering, topic modelling, etc. tf/idf is also often used to give distinctive words a higher weight in this matrix and irrelevant words a lower one, but this technique is not limited to BoW. So what is the problem with this representation? BoW’s simplicity is its strength, but it also has a number of limitations. To understand text, we need to:

  1. Understand the meaning of individual words
  2. Understand how words relate to one another

How does BoW perform on these points?

  • Understanding the meaning of individual words
    • A word is represented only by an index referring to the vocabulary. No meaning is attached to it.
  • Understanding how words relate to one another
    • The order of the words in the text is ignored.

Two examples to make this clear:

  • BoW sees no similarity between ‘animal’ and ‘creature’, even though they mean almost the same thing
  • By contrast, BoW sees these sentences as exactly the same:
    • “Not as I expected, this film is great”
    • “As I expected, this film is not great”

Word embeddings

The central idea behind word embeddings is John Firth’s distributional hypothesis[5]:

“You shall know a word by the company it keeps”.

In other words, you can infer the meaning of a word from the context in which it occurs. That is exactly how the word2vec algorithm works: it examines pieces of text to see which words occur alongside which context words.

Four cards featuring Formula 1 drivers, their value and their results this season.

In the example of The quick brown fox, this may not seem particularly useful. But if you do it for millions or even billions of sentences, patterns start to emerge. That gives you a large amount of data on which to train a model (a neural network) to predict the context words given a word. This may seem like a pointless, unpromising task, but it has a useful by-product: a matrix containing a set of weights (numbers) for each word. And there is something special about these numbers. The meaning of the word is actually encoded in them — and you can extract it! How? By comparing the numbers: they can be close together (81 and 82) or far apart (2 and 81), which you can interpret as similar or dissimilar. And because each word is represented by several numbers, you can look at different kinds of similarity. An example is the now-canonical “king – man + woman = queen”:

Let’s look at the numbers we get (which we call embeddings) for the four words “king”, “man”, “woman” and “queen” in the image below.

Here’s what stands out:

  • Some numbers/columns (shown in colour below) are close together when we compare “king” and “man”. That suggests a semantic similarity between the two words (which there is)
  • The same goes for “woman” and “queen”
  • There is a similar relationship between “king” and “queen”, but in different numbers/columns

Diagram comparing the word vectors for king, man, woman and queen.

Semantically, if you take the word “king” and replace the male aspect with the female one, you arrive at “queen”. And that is exactly (well, almost exactly) what happens when you do the same calculation with the word embeddings for these words! This suggests that some numbers in an embedding capture gender, while others capture ‘royalty’. Now imagine that we have embeddings with just three numbers (in reality, there are hundreds or even thousands) and treat those numbers as points in 3D space. We can then visualise the relationships between words:

Diagram with coloured dots and connecting lines.

If you’d like to discover a few more interesting and surprising connections, I recommend this blog: https://graceavery.com/word2vec-fish-music-bass/

Word embeddings as transfer learning

So how do you get word embeddings? Generally, you don’t create them yourself; you reuse existing ones. Researchers at Google trained their word2vec algorithm on enormous datasets and made the embeddings publicly available in various languages. There are also other algorithms for creating word embeddings. The best known are fastText from Facebook and GloVe from Stanford. You can download and use all of them. The first form of transfer learning in NLP.

From word embeddings to text embeddings

So, we’ve found a new way to represent words, with their meaning more or less encoded in numbers. How can we assess this representation?

  • Understanding the meaning of individual words
    • This seems to have worked very well! The word embeddings for “animal” and “beast” will be very similar
  • Understanding the relationships between words
    • Erm, the relationships…?

Well, as the name suggests, word embeddings represent words, but not yet sentences or passages of text. With BoW, we could at least represent passages of text, so where do we go from here? One common approach was to use recurrent neural networks (e.g. LSTM, GRU) alongside word embeddings. The text is split into words and converted into a sequence of word embeddings, which are read in step by step. At each step, the network’s internal ‘memory’ is updated, so you end up with an embedding for the whole text. This embedding is then used for the task at hand (prediction, text generation, etc.).

Diagram of an LSTM network classifying the sentence 'best movie ever' as positive.

From here, we can broadly distinguish two approaches:

  1. ‘direct’ text embedding techniques
  2. improvements focused on sensitivity to context

Direct text embeddings

There are several techniques for obtaining text embeddings. These are generally used to embed sentences or paragraphs. Embedding entire pages of text in one go does not yet work well. The simplest method (though not necessarily the worst) is to use average word embeddings: take the word embeddings of every word in your text and calculate their average. The drawback is that this ignores word order and includes irrelevant words, but you can address the latter by weighting the words with tf/idf first and then calculating a weighted average. A series of other approaches followed. The doc2vec[6] algorithm is essentially an extension of word2vec, in which the text is provided as context alongside the individual words. Words and texts are thus encoded in the same embedding space and can be compared with one another. With skip-thought vectors[7], the word2vec algorithm is scaled up by replacing words with sentences.

So we move from: “the meaning of a word is: the words it appears between”, to: “the meaning of a sentence is: the sentences it appears between”. This is done using RNN’s (GRU) with word embeddings in an encoder-decoder model.

Context sensitivity

Another line of research focused more on helping neural networks use word embeddings to handle context better. The concept of attention began as a technique for focusing on particular parts of a text during a given task. At certain points, a specific selection of the text (the context) is relevant within the larger whole. With the arrival of the transformer architecture, combinations of attention mechanisms replaced earlier architectures based on recurrent and convolutional neural networks. What followed was a series of large-scale language models based on transformers. The best-known models, from the BERT family, provide pre-trained, context-sensitive word embeddings. From there, it is a relatively small step to text embeddings. Sentence-BERT[8] (better known by the term sentence transformers) is a practical implementation of this approach. We have reached a significant point: transfer learning for NLP is finally a reality!

Diagram of pre-training on a large text collection and fine-tuning on a small one.

Wrap up

Now that we have arrived at text embeddings, we have a way to create a powerful representation of a piece of text that captures both the meaning of the words and the relationships between them (the sentence structure). So do we just need to download a pre-trained model, use it to generate embeddings for our text, and we’re ready to go? Depending on the task, that is true to some extent. At least in theory.

In part 2, we’ll see how well these promises hold up in practice…

References

[1] [Mikolov et al. 2013] https://arxiv.org/abs/1301.3781
[2] [Bahdanau et al. 2014] https://arxiv.org/abs/1409.0473
[3] [Vaswani et al. 2017] https://arxiv.org/abs/1706.03762
[4] [Devlin et al. 2018] https://arxiv.org/abs/1810.04805
[5] https://en.wikipedia.org/wiki/Distributional_semantics
[6] [Le, Mikolov 2014] https://arxiv.org/abs/1405.4053
[7] [Kiros et al. 2015] https://arxiv.org/abs/1506.06726
[8] [Reimers, Gurevych 2019] https://arxiv.org/abs/1908.10084
Blog

You may also find this interesting,

Sign up for our newsletter.

Want to be the first to hear about a new blog post?

Enter a valid email address.