Kestra is an infinitely scalable orchestration and scheduling platform, creating, running, scheduling, and monitoring millions of complex pipelines.
Cost / License
- Free
- Open Source (Apache-2.0)
Platforms
- Mac
- Windows
- Linux
- Online
- Self-Hosted




Kestra is an infinitely scalable orchestration and scheduling platform, creating, running, scheduling, and monitoring millions of complex pipelines.




CocoIndex is Python-native data transformation for any engineer, designed for AI workloads, with a smart incremental engine for always-fresh, explainable data.

Orbital automates integration between data sources (APIs, Databases, Queues and Functions). BFF's, API Composition and ETL pipelines that adapt as your specs change.


A visual data pipeline studio that runs on your laptop. Drag sources, transforms, validators, and sinks onto a canvas. Wire them together. Press Run. Duckle compiles the graph to SQL and executes it through a real columnar engine, with live previews, generated SQL on every node...




Cloud-based platform for managing complex data integrations using templates and pre-configured functions, enhancing diverse data projects.

A low code Machine Learning personalized ranking service for articles, listings, search results, recommendations that boosts user engagement. A friendly Learn-to-Rank engine.

Context Data is an enterprise data infrastructure built to accelerate the development of data pipelines for Generative AI applications. The platform automates the process of setting up internal data processing and transformation flows using an easy-to-use connectivity framework...

Data Engineering is needlessly complex today. What if we can abstract away a lot of the tedious parts of this and let companies focus on just the data they want to store, and the questions they have on their data?




Preparing for a data engineer interview and are overwhelmed by all the tools and concepts? Enroll now for free and learn how to ace the data engineering interview.
Quadratic is a Web-based spreadsheet application with Python, SQL, and Formulas that runs in the browser and as a native app (via Electron).

A next-generation data discovery and observability tool for startups and enterprises that helps to efficiently democratize data, powers collaboration of data science and data engineering teams, significantly reduces time to data discovery, cuts on data downtime and offers a...




Open Standard for Metadata. A Single place to Discover, Collaborate and Get your data right.




TABLUM.IO is a data management tool that specializes in data staging and preparation, specifically for raw and unstructured data from files, feeds, and API responses.




Shipyard’s your go-to cloud-based DataOps platform for secure data extraction, transformation, reverse ETL, workflow orchestration, monitoring, alerting, and more.

The concept behind Dataplane is to make it quicker and easier to construct a data mesh with robust data pipelines and automated workflows for businesses and teams of all sizes. In addition to being more user friendly, there has been an emphasis on scaling, resilience...




AI-first ETL/ELT tool in Rust. Build data pipelines visually, then drop into SQL or Python — with a built-in AI assistant that composes pipelines from plain English.




Open metadata and governance for enterprises - automatically capturing, managing and exchanging metadata between tools and platforms, no matter the vendor.



Mage.ai is an open-source data pipeline tool designed to simplify the process of building, running, and maintaining machine learning and data workflows. With an intuitive interface, it allows users to create powerful pipelines for ETL (Extract, Transform, Load) processes, data...



Dagster is a cloud-native data pipeline orchestrator for the whole development lifecycle, with integrated lineage and observability, a declarative programming model, and best-in-class testability.



Lunapad is an open-source data notebook that helps you move from exploration to production without changing tools. Analyze data with SQL and Python, collaborate through living documentation, build dashboards and data apps, and reuse your work instead of starting over.




High-performance synthetic data generator written in Rust. Produces GDPR/HIPAA-compliant test data for PostgreSQL & MySQL.

Neum AI is a best-in-class framework to manage the creation and synchronization of vector embeddings at large scale.

Native macOS data IDE with Claude AI in the SQL editor, visual ETL pipelines, and scheduled jobs. Replaces DBeaver, Data Loader, and Fivetran on Mac.




Real-time data quality screening API — PASS / WARN / BLOCK in <10ms. Schema drift, null spikes, type mismatches, outlier detection. Python SDK.



