Natural Language Processing (NLP, also known as text mining) is an umbrella term for techniques that extract insights from text data. These techniques are used in practice with varying degrees of success. In recent years, rapid developments in neural networks (particularly by tech giants such as Google) have led to several major breakthroughs. How do they work? How useful are these new techniques in practice? And do they work with Dutch-language text data? That is too much to cover in one blog post, so this will be a two-part series. In this first part, we take a closer look at
embeddings: a way of representing text data that underpins recent developments. The second part looks at practical applications and specific possibilities for Dutch-language text.
NLP is not a recent development; it has been around for a long time. Conceptually, it dates back to the 1950s, and in practice to around the 1990s. Topics such as
Named Entity Recognition,
Part-of-Speech tagging and
stemming & lemmatisation are standard components of NLP toolkits. The main problems that text data is used to address include
semantic search,
sentiment analysis, topic modelling, classification and clustering. Common techniques include
bag-of-words,
tf-idf,
Naïve Bayes and
LDA. What we have seen in recent years, however, is a shift towards NLP techniques based on neural networks. This began with the introduction of word2vec
[1] in 2013: a kind of dictionary that lets you translate a word into a sequence of numbers that captures its meaning (word embeddings). These embeddings were widely used in recurrent neural networks. The principle of
attention[2] made neural networks better able to learn to retain long-term relationships within passages of text. The arrival of the
transformer[3] improved this considerably, and today the focus is on pre-trained language models from the BERT family
[4]. This has finally made
transfer learning a reality in NLP, as pre-trained convolutional neural networks had already done in computer vision. In chronological order, the most important academic developments are therefore:

But as mentioned, the rapid progress began with word2vec and word embeddings. Why is this so important? Because it radically changes how text is represented.
Representations of text data
A computer cannot calculate with text, only with numbers. If we want to apply algorithms to text data, we first need to ‘translate’ that text into numbers. In other words, we need to find a numerical representation.

Bag-of-words
The bag-of-words (BoW) representation has been widely used for years. The approach is simple. Suppose you have a set of documents (pieces of text). You then:
- Identify which words occur
- Use them to build a vocabulary
- For each document: count how often each word in the vocabulary occurs
- Each document is now a long sequence of numbers
- Your ‘corpus’ (text dataset) is now a matrix of numbers
- B. Stop words are often filtered out of the vocabulary

This matrix of numbers can now be used with all kinds of algorithms for classification, clustering, topic modelling, etc.
tf/idf is also often used to give distinctive words a higher weight in this matrix and irrelevant words a lower one, but this technique is not limited to BoW. So what is the problem with this representation? BoW’s simplicity is its strength, but it also has a number of limitations. To understand text, we need to:
- Understand the meaning of individual words
- Understand how words relate to one another
How does BoW perform on these points?
- Understanding the meaning of individual words
- A word is represented only by an index referring to the vocabulary. No meaning is attached to it.
- Understanding how words relate to one another
- The order of the words in the text is ignored.
Two examples to make this clear:
- BoW sees no similarity between ‘animal’ and ‘creature’, even though they mean almost the same thing
- By contrast, BoW sees these sentences as exactly the same:
- “Not as I expected, this film is great”
- “As I expected, this film is not great”
Word embeddings
The central idea behind word embeddings is John Firth’s distributional hypothesis[5]:
“You shall know a word by the company it keeps”.
In other words, you can infer the meaning of a word from the context in which it occurs. That is exactly how the word2vec algorithm works: it examines pieces of text to see which words occur alongside which context words.

In the example of
The quick brown fox, this may not seem particularly useful. But if you do it for millions or even billions of sentences, patterns start to emerge. That gives you a large amount of data on which to train a model (a neural network) to predict the context words given a word. This may seem like a pointless, unpromising task, but it has a useful by-product: a matrix containing a set of weights (numbers) for each word. And there is something special about these numbers. The meaning of the word is actually encoded in them — and you can extract it! How? By comparing the numbers: they can be close together (81 and 82) or far apart (2 and 81), which you can interpret as similar or dissimilar. And because each word is represented by several numbers, you can look at different kinds of similarity. An example is the now-canonical “king – man + woman = queen”:
Let’s look at the numbers we get (which we call embeddings) for the four words “king”, “man”, “woman” and “queen” in the image below.
Here’s what stands out:
- Some numbers/columns (shown in colour below) are close together when we compare “king” and “man”. That suggests a semantic similarity between the two words (which there is)
- The same goes for “woman” and “queen”
- There is a similar relationship between “king” and “queen”, but in different numbers/columns

Semantically, if you take the word “king” and replace the male aspect with the female one, you arrive at “queen”. And that is exactly (well, almost exactly) what happens when you do the same calculation with the word embeddings for these words! This suggests that some numbers in an embedding capture gender, while others capture ‘royalty’. Now imagine that we have embeddings with just three numbers (in reality, there are hundreds or even thousands) and treat those numbers as points in 3D space. We can then visualise the relationships between words:

If you’d like to discover a few more interesting and surprising connections, I recommend this blog:
https://graceavery.com/word2vec-fish-music-bass/
Word embeddings as transfer learning
So how do you get word embeddings? Generally, you don’t create them yourself; you reuse existing ones. Researchers at Google trained their word2vec algorithm on enormous datasets and made the embeddings publicly available in various languages. There are also other algorithms for creating word embeddings. The best known are fastText from Facebook and GloVe from Stanford. You can download and use all of them. The first form of transfer learning in NLP.
From word embeddings to text embeddings
So, we’ve found a new way to represent words, with their meaning more or less encoded in numbers. How can we assess this representation?
- Understanding the meaning of individual words
- This seems to have worked very well! The word embeddings for “animal” and “beast” will be very similar
- Understanding the relationships between words
Well, as the name suggests, word embeddings represent words, but not yet sentences or passages of text. With BoW, we could at least represent passages of text, so where do we go from here? One common approach was to use recurrent neural networks (e.g. LSTM, GRU) alongside word embeddings. The text is split into words and converted into a sequence of word embeddings, which are read in step by step. At each step, the network’s internal ‘memory’ is updated, so you end up with an embedding for the whole text. This embedding is then used for the task at hand (prediction, text generation, etc.).

From here, we can broadly distinguish two approaches:
- ‘direct’ text embedding techniques
- improvements focused on sensitivity to context
Direct text embeddings
There are several techniques for obtaining text embeddings. These are generally used to embed sentences or paragraphs. Embedding entire pages of text in one go does not yet work well. The simplest method (though not necessarily the worst) is to use average word embeddings: take the word embeddings of every word in your text and calculate their average. The drawback is that this ignores word order and includes irrelevant words, but you can address the latter by weighting the words with tf/idf first and then calculating a weighted average. A series of other approaches followed. The doc2vec[6] algorithm is essentially an extension of word2vec, in which the text is provided as context alongside the individual words. Words and texts are thus encoded in the same embedding space and can be compared with one another. With skip-thought vectors[7], the word2vec algorithm is scaled up by replacing words with sentences.
So we move from: “the meaning of a word is: the words it appears between”, to: “the meaning of a sentence is: the sentences it appears between”. This is done using RNN’s (GRU) with word embeddings in an encoder-decoder model.
Context sensitivity
Another line of research focused more on helping neural networks use word embeddings to handle context better. The concept of attention began as a technique for focusing on particular parts of a text during a given task. At certain points, a specific selection of the text (the context) is relevant within the larger whole. With the arrival of the transformer architecture, combinations of attention mechanisms replaced earlier architectures based on recurrent and convolutional neural networks. What followed was a series of large-scale language models based on transformers. The best-known models, from the BERT family, provide pre-trained, context-sensitive word embeddings. From there, it is a relatively small step to text embeddings. Sentence-BERT[8] (better known by the term sentence transformers) is a practical implementation of this approach. We have reached a significant point: transfer learning for NLP is finally a reality!

Wrap up
Now that we have arrived at text embeddings, we have a way to create a powerful representation of a piece of text that captures both the meaning of the words and the relationships between them (the sentence structure). So do we just need to download a pre-trained model, use it to generate embeddings for our text, and we’re ready to go? Depending on the task, that is true to some extent. At least in theory.
In part 2, we’ll see how well these promises hold up in practice…
References
[1] [Mikolov et al. 2013]
https://arxiv.org/abs/1301.3781
[2] [Bahdanau et al. 2014]
https://arxiv.org/abs/1409.0473
[3] [Vaswani et al. 2017]
https://arxiv.org/abs/1706.03762
[4] [Devlin et al. 2018]
https://arxiv.org/abs/1810.04805
[5]
https://en.wikipedia.org/wiki/Distributional_semantics
[6] [Le, Mikolov 2014]
https://arxiv.org/abs/1405.4053
[7] [Kiros et al. 2015]
https://arxiv.org/abs/1506.06726
[8] [Reimers, Gurevych 2019]
https://arxiv.org/abs/1908.10084