An AI Christmas card as a gift from Data Science Lab
Last year at Data Science Lab, we created an AI Christmas card using generative AI. It was widely used and is still one of our most visited pages. So…
DeepSeek has emerged as a formidable competitor in the world of Large Language Models (LLMs). With stock prices fluctuating and political tensions rising, DeepSeek has been hard to ignore in recent weeks. But what makes this model stand out? Is it genuinely revolutionary, or simply a budget-friendly alternative to market leaders such as OpenAI’s ChatGPT? In this blog, we look at DeepSeek’s technical architecture and analyse why its models are so efficient and fast.
While market leaders such as OpenAI, Meta and Google are investing tens of billions in developing LLMs, DeepSeek claims to have trained its latest model for just a few million. Yet its performance comes surprisingly close to that of these established players. DeepSeek has achieved an impressive combination: the power and accuracy of leading AI models, such as OpenAI-o1, with remarkable efficiency and scalability. We have not seen this balance between performance and cost savings on this scale before.
Figure 1: Benchmark performance of DeepSeek-R1. (DeepSeek-R1)
To understand why DeepSeek is so fast and efficient, we need to look at its underlying architecture. Like most advanced LLMs, DeepSeek is based on the Transformer architecture (Vaswani et al., 2017). But while conventional AI models activate every neuron/parameter in the network for every input, DeepSeek takes a different approach with a Mixture-of-Experts (MoE) architecture. Only a small selection of specialised subnetworks - the ‘experts’- is activated for each input. Which experts are selected depends on their relevance to the current input.
This mechanism drastically reduces the computing power required and increases the model’s speed without compromising performance. DeepSeek-V3, for example, contains 671 billion parameters but activates only 37 billion per token – just 18% of the total model.
DeepSeekMoE is the driving force behind DeepSeek’s impressive efficiency. It is DeepSeek’s own MoE model. DeepSeekMoE consists of multiple experts and a gating mechanism that determines which experts are activated for each input (see Figure 2).
Traditional MoE models face two major problems:
DeepSeekMoE addresses these limitations with two key innovations:
These innovations enable DeepSeekMoE to achieve performance comparable to dense models while using significantly less computing power. DeepSeekMoE 16B, for example, performs as well as LLaMA2 7B while using 40% fewer computational resources. This makes the model not only powerful, but also more scalable and efficient than many traditional alternatives.
Figure 2: Visualisation of DeepSeekMoe (subfigure c). The number of parameters and computational costs are the same across all three architectures. (DeepSeekMoE)
DeepSeek uses an advanced training process with gradient-based optimisation. Unlike conventional neural networks, only the activated experts in DeepSeek receive updates, allowing them to specialise quickly in specific tasks. In other words, experts learn only from input relevant to them.
The experts specialise through:
Because input with specific features or characteristics is consistently routed to the same expert, that expert becomes specialised in a particular task. For example, one expert may specialise in programming code, another in legal texts and a third in everyday conversations.
A common risk during this process is that some experts are overused while others remain underused. This reduces the model’s efficiency and ability to generalise. To address this, DeepSeek uses load balancing techniques, such as auxiliary losses (DeepSeekMoE), which penalise overusing an expert, and bias-adjusted routing (DeepSeek-V3), which gives each expert a dynamic bias term based on its load.
In short, experts in the DeepSeek models improve through:
Ultimately, the specialised experts lead to:
DeepSeek has further refined its architecture in DeepSeek-V2 and DeepSeek-V3. These models introduce new features such as:
Multi-Head Latent Attention is a new attention mechanism that is faster than traditional Multi-Head Attention (MHA). Conventional Transformer models typically use Multi-Head Attention (MHA) (Vaswani et al., 2017), creating a bottleneck that slows down inference. To address this, DeepSeek designed the MLA mechanism for the DeepSeek-V2 and DeepSeek-V3 models (see Figure 3).
A more efficient alternative to auxiliary loss functions. Many MoE models, including DeepSeekMoE, require auxiliary loss functions to prevent some experts from being overloaded while others remain unused. DeepSeek-V3 addresses this by introducing the bias-adjusted routing strategy, which improves efficiency without compromising load balancing.
DeepSeek uses FP8 mixed precision arithmetic to reduce memory usage and computation time. Calculations are split between lower precision and, where needed, higher precision, balancing efficiency with numerical stability.
With Multi-Token Prediction, the model predicts multiple future tokens at once, speeding up both training and inference. Traditional LLM’s predict one token at a time.
An infrastructure designed to address GPU bottlenecks and enable efficient training at an extremely large scale. DeepSeek-V3 is trained using DualPipe pipeline parallelism. Combined with an optimised cross-node communication system, this makes DeepSeek-V3 one of the most efficient large-scale MoE models to date.
In short, the differences between these three model variants are as follows:
DeepSeek-V3 is therefore the most advanced model, combining these innovations for maximum performance and efficiency (see Figure 4).
Figure 3: Simplified visualisation of several attention mechanisms, with Multi-head Latent Attention (MLA) on the far right. By compressing keys and values together into a latent vector, MLA makes inference considerably faster. (DeepSeek-V2)
Figure 4: The basic architecture of DeepSeek-V2 and DeepSeek-V3. (DeepSeek-V2)
DeepSeek-R1 was the model that marked the real breakthrough. It not only shook up the AI community but also had a broad impact on society. From plunging share prices to heated political debates about the future of AI, DeepSeek-R1 proved it was more than just another new model: it was a technological milestone.
This model uses Reinforcement Learning (RL), enabling it to reason and make decisions independently. This not only improves its performance but also shows what is possible when you let the model think for itself.
In terms of architecture, DeepSeek R1 is identical to the DeepSeek-V2 and DeepSeek-V3 variants. It uses the “only-pretrained” versions of these models, but then undergoes its own post-training process.
Put simply, the developers built a test-verification harness around the model and applied an RL process to it. The model learns by being rewarded for good outcomes. During this process, they used a simple reward function (the RL equivalent of a loss function). The function is a weighted sum of two components:
This post-training process leads to:
Thanks to this process, DeepSeek-R1 is not only accurate but also more reliable at generating structured and checked outputs.
As a result, the final model has:
While DeepSeek offers considerable potential, it also brings challenges and risks:
1. Open-source security risks
DeepSeek is open-source. That is both a strength and a risk. It provides transparency and allows for custom-built solutions, but it also makes the model more vulnerable to misuse for malicious purposes, such as generating disinformation, spam or unethical automation.
2. Ecosystem and developer support
Compared with proprietary models (such as OpenAI’s GPT-4), DeepSeek has a smaller ecosystem of integrations, tools and community support. This could make it harder for DeepSeek to compete with larger models in the long term. Adoption and real-world testing will determine whether the model can effectively compete with its larger rivals over time.
3. Bias and language limitations
The model is heavily optimised for English and Chinese, which may affect its performance in other languages.
Several articles soon appeared demonstrating interference and censorship by the Chinese government (including CNN and The Guardian). There are also significant privacy concerns, including data leaks, which have led several countries to impose restrictions on the use of DeepSeek (primarily in professional settings) (including Wiz, NowSecure and BBC).
In short, DeepSeek offers considerable potential as an open-source alternative to the closed-source LLMs that have led the market so far, but its long-term success depends on continuous improvements in accuracy and safety, and on the growth of its ecosystem.
DeepSeek shows that open-source AI models can compete with proprietary models. It offers a new perspective on efficient, scalable AI while also prompting discussion about transparency, security and ethics in AI development.
With DeepSeek-V3, the company shows that open-source models can do more than hold their own: they may even shape the future of AI. Its combination of strong performance, scalability and efficiency makes DeepSeek revolutionary. Whether you’re a researcher, developer or entrepreneur – DeepSeek is changing how we think about Large Language Models. We could even say that DeepSeek has set a new standard.
DeepSeek’s open-source nature and efficiency also make it attractive for business applications, not just research. We’ll cover that in a later blog post. Are you already wondering how your organisation could benefit from these developments? We’d be happy to help. Arrange a meeting with us and explore the possibilities.
Last year at Data Science Lab, we created an AI Christmas card using generative AI. It was widely used and is still one of our most visited pages. So…
Generative AI (GenAI) is developing rapidly. We’re seeing major improvements in quality, not only in text but especially in images and video…
Since the launch of tools such as ChatGPT, interest in generative AI (GenAI) has grown enormously. The applications seem endless: from content…
Want to be the first to hear about a new blog post?
Thanks for signing up!