Kestra is an infinitely scalable orchestration and scheduling platform, creating, running, scheduling, and monitoring millions of complex pipelines.
Cost / License
- Free
- Open Source (Apache-2.0)
Platforms
- Mac
- Windows
- Linux
- Online
- Self-Hosted




Kestra is an infinitely scalable orchestration and scheduling platform, creating, running, scheduling, and monitoring millions of complex pipelines.




Orbital automates integration between data sources (APIs, Databases, Queues and Functions). BFF's, API Composition and ETL pipelines that adapt as your specs change.


MINEO is the platform to explore your data, build, deploy and share data apps based on supercharged Python Notebooks powered by Code, No-code & AI.




Cloud-native microservice data processing for streaming and batch pipelines, optimized for Cloud Foundry and Kubernetes, with flexible integration options.

Cloud-based platform for managing complex data integrations using templates and pre-configured functions, enhancing diverse data projects.

Google Cloud Data Fusion is a managed service for creating and managing data pipelines. This cloud-native solution enables users to build and control data pipelines quickly. It offers a web interface for constructing scalable data integration solutions, allowing users to clean...
A visual data pipeline studio that runs on your laptop. Drag sources, transforms, validators, and sinks onto a canvas. Wire them together. Press Run. Duckle compiles the graph to SQL and executes it through a real columnar engine, with live previews, generated SQL on every node...




Anakin.io is a web scraping and structured data extraction API built for developers and AI teams. It converts any website into clean Markdown or JSON with a single API call - handling JavaScript rendering, anti-bot bypass, proxy rotation, and authenticated scraping (behind login...




This next-gen AI connects to any system, like APIs, websockets, or databases, and ingests data into various targets, including vector databases and email.


Orchestration and durable functions replace queues and simplify step workflows, enabling developers to create reliable systems effortlessly.




PipeRider is an open-source data reliability toolkit for identifying data quality issues across pipelines.




Operational software for an AI-powered world - connecting systems, unifying data + enabling custom applications and AI on existing infrastructure.




Palantir Foundry allows you to build apps using data from external databases or apps you create in their platform. Imported or inputted data is stored in the "Ontology" which is a relational/graph/linked database hybrid that allows you to represent and relate data...
OpenSnowcat is an open-source fork of Snowplow under Apache 2.0 License, fully compatible with Snowplow and Segment SDKs.

AI-first ETL/ELT tool in Rust. Build data pipelines visually, then drop into SQL or Python — with a built-in AI assistant that composes pipelines from plain English.





Turn any website into clean, LLM-ready data. Open-source web crawler with stealth mode, distributed crawling, real-time WebSocket progress & Markdown output. Power your AI apps with GcrawlAI.

Oneprofile enables businesses to sync customer profiles and events between their tools so the right data is always in the right place.

Rawbbit is a self-hosted, open-source event tracking and raw-storage pipeline for game analytics.

The operating system for small businesses. We replace every separate tool you pay for — CRM, billing, email, content, contracts, ads, scheduling, automations — in one AI-native interface.


Stitch is a cloud-first, open source platform for rapidly moving data. A simple, powerful ETL service, Stitch connects to all your data sources – from databases like MySQL and MongoDB, to SaaS applications like Salesforce and Zendesk – and replicates that data to a destination...
Baselight unlocks the power of data, combining openness, community, and AI to make high-quality structured data accessible to all.

Real-time data quality screening API — PASS / WARN / BLOCK in <10ms. Schema drift, null spikes, type mismatches, outlier detection. Python SDK.



