Skip to content
Go to insights Blog

Introducing Skippa

Introducing Skippa
Written by
Data Science Lab
Published on
30 December 2021

Summary

Every data scientist is probably familiar with pandas and scikit-learn. The usual workflow starts with cleaning data in pandas, followed by further pre-processing using pandas or scikit-learn transformers such as StandardScaler and OneHotEncoder. Then you move on to a machine learning algorithm (scikit-learn).

There are a few problems with this workflow:
1. The development phase of your workflow is quite complex and requires a lot of code ?
2. It is difficult to reproduce the workflow for predictions during deployment ?
3. Existing solutions to reduce these problems are not good enough (yet) ?

Skippa is a package designed to:

  • ✨ drastically simplify development
  • ? package all data cleaning and pre-processing together with the algorithm in a single pipeline file
  • ? reuse the pandas and scikit-learn interfaces you already know"

Skippa helps you define data cleaning and pre-processing transformations easily. Here is roughly how it works:

from skippa import Skippa, columns from sklearn.linear_model import LogisticRegression X, y = get_training_data(...) pipeline = ( Skippa() .impute(columns(dtype_include='object'), strategy='most_frequent') .impute(columns(dtype_include='number'), strategy='median') .scale(columns(dtype_include='number'), type='standard') .onehot(columns(['category1', 'category2'])) .model(LogisticRegression()) ) pipeline.fit(X, y) predictions = pipeline.predict_proba(X)

☝️Skippa does not claim to solve every problem, cover every feature you might ever need or offer a highly scalable solution, but it should provide a substantial simplification for > 80% of typical pandas/sklearn-based machine learning projects. 

Links

Read the rest of the blog post here > Introduction Skippa [ENG]

Blog

You may also find this interesting,

Sign up for our newsletter.

Want to be the first to hear about a new blog post?

Enter a valid email address.