Skip to content
Go to insights data engineering

PySpark basics: upserting data on Databricks

PySpark basics: upserting data on Databricks
Written by
Data Science Lab
Published on
10 June 2025

In this blog, Tim (Data Engineer) explains how data engineers and data scientists can move from pandas to PySpark. Although the two libraries have similar syntax, there are important conceptual differences, especially when processing large volumes of data through distributed computing.


Tim explains how Apache Spark, the engine behind PySpark, works and how it uses a driver and executors to process data in parallel. In PySpark, you work with DataFrames divided into partitions, which are essential for scalability. He also covers the difference between transformations and actions, the importance of lazy evaluation, and why some operations, such as wide transformations, are much more resource-intensive than others. Finally, he gives a practical example of a common task in data pipelines: upserting data into a bronze table using PySpark in Databricks.


Key insights:

  • PySpark is the Python interface for Apache Spark.
  • Spark is optimised for big data and uses distributed computing.
  • Lazy evaluation allows Spark to determine the most efficient execution strategy.
  • PySpark is ideal when your data no longer fits on a single machine.


👉 Read the full article on Medium: From pandas to PySpark - Tim Winter

Blog

You may also find this interesting,

Sign up for our newsletter.

Want to be the first to hear about a new blog post?

Enter a valid email address.