Managed DevOps Pool - Efficient and secure DevOps pipelines?
A Managed DevOps Pool lets you manage the infrastructure you need to run a DevOps pipeline without having to maintain servers or machines – Azure…
In this blog, Tim (Data Engineer) explains how data engineers and data scientists can move from pandas to PySpark. Although the two libraries have similar syntax, there are important conceptual differences, especially when processing large volumes of data through distributed computing.
Tim explains how Apache Spark, the engine behind PySpark, works and how it uses a driver and executors to process data in parallel. In PySpark, you work with DataFrames divided into partitions, which are essential for scalability. He also covers the difference between transformations and actions, the importance of lazy evaluation, and why some operations, such as wide transformations, are much more resource-intensive than others. Finally, he gives a practical example of a common task in data pipelines: upserting data into a bronze table using PySpark in Databricks.
Key insights:
👉 Read the full article on Medium: From pandas to PySpark - Tim Winter
A Managed DevOps Pool lets you manage the infrastructure you need to run a DevOps pipeline without having to maintain servers or machines – Azure…
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
Many organisations have now run an AI pilot. The model works, the demo gets applause, and then nothing else happens. In our experience, most projects…
Want to be the first to hear about a new blog post?
Thanks for signing up!