1
0 Comments

Feedback wanted: make data pipelining and transformations easier

I and my partner are currently working on an idea, could be explained as - " a visual self-service tool for data pipeline and data preparation, with intelligence".

We want to tackle two problems. First is the amount of time that is spent on getting the data ready vs. actual modelling/analysis. Secondly, the resources needed (Data engineers, scientist, etc.) to set up and operate a "data workflow" needed to leverage analysis/ML on a continuous basis.

The motivation of this startup mainly comes from a background as CTO at a "Uber for trucks" startup, where consistent clean data flows for models was essential. But was really costly and time-consuming to set up and operate.

Q: you as a founder at a data-intensive startup (maybe you have ML embedded even), how do you handle this today? Is it a pain in the ass?

Intended solution:
We think overall it can be easier and more autonomous than today.
We're imagining a drag-and-drop interface for data pipelining and a GUI for data transformation (to make people less dependent on data engineers). At the same time, not disconnected from a Data Scientist workflow and flexible (i.e. fully integrable with Python).

To speed up the process, we are developing a model/framework to be applied to the data itself. That learns from the data, the user and the high-level objectives. To overtime, automate more and more of the mundane data tasks, ensure data quality and enhance the data scientist.

Q: Which vertical would you go after?

on December 21, 2020