Skip to content
Go to insights Blog

Super-resolution

Super-resolution
Written by
Data Science Lab
Published on
17 August 2021

If you watch series such as CSI, you’ll probably recognise scenes like this: detectives are on the culprit’s trail and study security camera footage expectantly. At first glance, the footage doesn’t seem to show what they’re looking for, but then they spot a barely visible detail. They zoom in, enhance the image quality and… the detail that emerges reveals important information, leading to a breakthrough in the investigation.

In reality, unfortunately, you can’t simply zoom in and conjure up details. Images consist of a limited number of pixels, and a sharper image requires more pixels. If those pixels weren’t captured, they can’t be recovered. Today, however, Artificial Intelligence can get us closer by predicting those additional pixels: what would the image most likely have looked like? The process of creating a high-resolution image from a low-resolution image is called super-resolution.

A beach photo before and after sharpening.


Super-resolution has useful applications in various fields, for example:
  • Medical imaging (source): constraints such as limited time or patient movement mean that the resolution of MRI scans is not always as high as desired. Super-resolution is therefore used to improve the quality of MRI scans, which can help doctors make a diagnosis.

Two MRI images of a knee side by side, grainy on the left and sharp on the right.

  • Satellite imaging (source): satellite images are taken from a great distance, so the detail they show is limited. Objects such as cars, for example, may be only 10 pixels across. Super-resolution can help detect objects more accurately.
  • Security (source): because high-resolution cameras cannot be installed everywhere, security cameras often produce low-resolution images. Improving the quality of camera footage can make details – such as number plates – easier to see.

Traditional methods of adding pixels to an image generate new pixels based on the surrounding ones. Unfortunately, super-resolution is not that simple: these interpolation techniques often result in blurry or pixelated images. Advances in deep learning methods have driven rapid progress in super-resolution in recent years. Using a large dataset of images, neural networks can learn to add detail to images. How does that work? In this blog, we introduce supervised deep learning methods for super-resolution.

Data preparation

The first step in supervised learning is to create a dataset for training a model. The aim of super-resolution models is to learn a mapping between low-resolution (LR) and high-resolution (HR) images: in other words, how to turn an LR image into an HR image. We can give the model an LR image as input and an HR image as ground truth. This allows it to learn a function that produces an image as close as possible to the original from an LR image. We therefore need LR–HR pairs to train the model. A dataset of LR–HR pairs is often created by taking a set of HR images and applying a function that reduces their quality. This can be done by removing pixels, applying blur and adding noise. The result is a high-quality and a low-quality version of each image, which we can feed into the model.

Diagram showing a sharp photo being reduced in size and then restored.

Model frameworks

Now that we know what kind of dataset we need, we can look at the different types of models available for super-resolution. Super-resolution is an ill-posed problem: an LR image can correspond to multiple HR images, so there is no single correct solution. That is why the way pixels are added, known as upsampling , is an important part of the method. Broadly speaking, there are four model frameworks with different approaches to upsampling:

  • Pre-upsampling: in this method, the LR image is enlarged first (using conventional interpolation techniques) to generate a rough HR image. Convolutional Neural Networks (CNNs) are then used to learn a mapping between this rough HR image and the original HR image. The advantage is that the neural network only has to learn this one mapping, because the image is enlarged using conventional interpolation methods. The disadvantage is that this can cause blurring.

Diagram of a network that first enlarges an image and then refines it.

  • Post-upsampling: the LR images first pass through the convolutional layers, with upsampling performed only in the final layer. The advantage of this approach is that the mapping function is learned in a lower-dimensional space, making the computations less complex.

Grid of GAN-generated images, from blurry to recognisable.

  • Iterative up-and-down sampling: this alternates between upsampling and downsampling. Models that use this framework are often better at finding deep relationships between LR-HR pairs and therefore often produce higher-quality images.

Diagram of a network that enlarges a photo of a butterfly step by step.

  • Progressive-upsampling: this framework uses multiple CNNs to generate an increasingly larger image in small steps. Breaking a difficult task into several simpler ones reduces its complexity, which can improve the results.

Grid of simple line drawings of sheep.

  Alongside the upsampling method, the type of deep learning network is another important component of super-resolution models. Discussing the different networks is beyond the scope of this blog, but you can consult this source for more information.

Learning strategies

To train a super-resolution model, you need to be able to measure the difference between the generated HR image and the original HR image. This difference is used to calculate the reconstruction error for the HR image and thus optimise the model. Loss functions are used for this calculation. Some examples of these loss functions are:

  • Pixel loss: pixel loss is the simplest loss function: each pixel in the generated image is compared directly with the corresponding pixel in the original image. Research has shown that this does not fully reflect reconstruction quality, so other functions are used as well.
  • Content loss: this function ensures visual similarity between the original and generated images by comparing their content rather than individual pixels. Using this loss function therefore improves perceptual quality and produces more realistic images.
  • Adversarial loss: this is used in GAN-related architectures. Generative Adversarial Networks (GANs) consist of two neural networks – the generator and the discriminator – that compete with each other (want to know more about GANs? Read about them in this blog post). In super-resolution, the generator tries to produce images that the discriminator believes are real. The discriminator tries to distinguish the original HR images from the generated ones. This training process ultimately produces a generator that is good at creating images similar to the originals.
  • Total variation loss: total variation loss calculates the absolute difference between neighbouring pixels and measures the amount of noise in an image. This function is used to reduce noise in the generated images.

Evaluation

To evaluate the results of trained super-resolution models, we measure the quality of the generated images. There are several ways to do this, including objective and subjective methods. Because subjective methods often take considerable time and money, especially when the dataset is large, objective methods are used more often in super-resolution.

  • Peak signal-to-noise ratio (PSNR): this quantitative metric is based on individual pixels and gives the ratio between signal and noise.
  • Structural similarity index metric (SSIM): this is also an objective method. It measures structural similarities between images by comparing their contrast, light intensity and structural details.
  • Opinion scoring: a subjective method in which people are asked to assess specific criteria such as sharpness, colour or natural appearance.

Discussion

Now that we’ve discussed several ways to train super-resolution models using deep learning, you might be thinking: wow, super-resolution lets us reveal all sorts of details we couldn’t otherwise see! In practice, it’s a little more complicated. Some time ago, the following photo circulated on social media:

Two images side by side: a highly enlarged, pixelated portrait and the sharper result.

On the left, we see a low-resolution image, clearly of Obama. On the right, we see the result produced by a super-resolution model trained on faces: a white man. This sparked debate because, like many other applications of AI, it shows the danger of bias: this super-resolution model produced white faces far more often than faces of other skin colours. What happened here? Rather than ‘improving’ the low-resolution image, the model generates an entirely new face that looks the same as the input image at low resolution. The bias in the dataset used to train the model is reflected in its results. So it’s important to realise that super-resolution does not produce a true reconstruction: we’re making a guess based on information learned from a dataset. The information is still ‘made up’, so we need to be careful about the conclusions we draw from it and think carefully about which purposes we can and cannot use it for.

References

Blog

You may also find this interesting,

Sign up for our newsletter.

Want to be the first to hear about a new blog post?

Enter a valid email address.