Track Awesome Data Engineering Updates Weekly

A curated list of data engineering tools for software developers

🏠 Home · 🔍 Search · 🔥 Feed · 📮 Subscribe · ❤️ Sponsor · 😺 igorbarinov/awesome-data-engineering · ⭐ 8.1K · 🏷️ Big Data

[ Daily / Weekly / Overview ]

Jan 12 - Jan 18, 2026

Testing / Data Profiler

Snowflake Emulator (⭐4) - A Snowflake-compatible emulator for local development and testing.

Jan 05 - Jan 11, 2026

Testing / Data Profiler

daffy (⭐25) - Decorator-first DataFrame contracts/validation (columns/dtypes/constraints) at function boundaries. Supports Pandas/Polars/PyArrow/Modin.

Dec 08 - Dec 14, 2025

Data Lake Management

FlightPath Data - FlightPath is a gateway to a data lake's bronze layer, protecting it from invalid external data file feeds as a trusted publisher.

Profiling / Data Profiler

YData Profiling - A general-purpose open-source data profiler for high-level analysis of a dataset.

Desbordante (⭐459) - An open-source data profiler specifically focused on discovery and validation of complex patterns in data.

Nov 03 - Nov 09, 2025

Stream Processing

Pathway (⭐56k) - Performant open-source Python ETL framework with Rust runtime, supporting 300+ data sources.

Testing / Data Profiler

GreatExpectation - Open Source data validation framework to manage data quality. Users can define and document “expectations” rules about how data should look and behave.

Sep 29 - Oct 05, 2025

Workflow

SQLMesh - An open-source data transformation framework for managing, testing, and deploying SQL and Python-based data pipelines with version control, environment isolation, and automatic dependency resolution.

Sep 22 - Sep 28, 2025

Data Ingestion

db2lake (⭐2) - Lightweight Node.js ETL framework for databases → data lakes/warehouses.

Sep 15 - Sep 21, 2025

Community / Books

Learn AI Data Engineering in a Month of Lunches - A fast, friendly guide to integrating large language models into your data workflows.

Sep 01 - Sep 07, 2025

Charts and Dashboards

QueryGPT (⭐26) - Natural language database query interface with automatic chart generation, supporting Chinese and English queries.

Aug 18 - Aug 24, 2025

Community / Books

Architecting an Apache Iceberg Lakehouse - A guide to designing an Apache Iceberg lakehouse from scratch.

Aug 04 - Aug 10, 2025

Testing / Data Profiler

Spark Playground - Write, run, and test PySpark code on Spark Playground's online compiler. Access real-world sample datasets & solve interview questions to enhance your PySpark skills for data engineering roles.

Jul 14 - Jul 20, 2025

Data Ingestion

Estuary Flow - No/low-code data pipeline platform that handles both batch and real-time data ingestion.

Jun 23 - Jun 29, 2025

Data Lake Management

Gravitino (⭐2.6k) - An open-source, unified metadata management for data lakes, data warehouses, and external catalogs.

Apr 21 - Apr 27, 2025

Charts and Dashboards

Seaborn - A Python visualization library based on matplotlib. It provides a high-level interface for drawing attractive statistical graphics.

Apr 14 - Apr 20, 2025

Data Lake Management

Ilum - A modular Data Lakehouse platform that simplifies the management and monitoring of Apache Spark clusters across Kubernetes and Hadoop environments.

Community / Books

Best Data Science Books - This blog offers a curated list of top data science books, categorized by topics and learning stages, to aid readers in building foundational knowledge and staying updated with industry trends.

Mar 17 - Mar 23, 2025

Stream Processing

CocoIndex (⭐5.6k) - An open source ETL framework to build fresh index for AI.

Batch Processing

Substation (⭐389) - A cloud native data pipeline and transformation toolkit written in Go.

Testing / Data Profiler

RunSQL - Free online SQL playground for MySQL, PostgreSQL, and SQL Server. Create database structures, run queries, and share results instantly.

Community / Books

Snowflake Data Engineering - A practical introduction to data engineering on the Snowflake cloud data platform.

Feb 24 - Mar 02, 2025

Data Ingestion

CsvPath Framework - A delimited data preboarding framework that fills the gap between MFT and the data lake.

Oct 21 - Oct 27, 2024

Workflow

Hamilton (⭐2.3k) - A lightweight library to define data transformations as a directed-acyclic graph (DAG). If you like dbt for SQL transforms, you will like Hamilton for Python processing.

Sep 02 - Sep 08, 2024

Data Ingestion

Artie - Real-time data ingestion tool leveraging change data capture.

Jul 29 - Aug 04, 2024

Workflow

Mage - Open-source data pipeline tool for transforming and integrating data.

Jul 22 - Jul 28, 2024

Stream Processing

SwimOS - A framework for building real-time streaming data processing applications that supports a wide range of ingestion sources.

Jul 15 - Jul 21, 2024

Testing / Data Profiler

DataKitchen - Open Source Data Observability for end-to-end Data Journey Observability, data profiling, anomaly detection, and auto-created data quality validation tests.

Jun 24 - Jun 30, 2024

Data Ingestion

Google Sheets ETL (⭐21) - Live import all your Google Sheets to your data warehouse.

Workflow

Kestra (⭐26k) - A versatile open source orchestrator and scheduler built on Java, designed to handle a broad range of workflows with a language-agnostic, API-first architecture.

May 27 - Jun 02, 2024

Data Ingestion

AWS Data Wrangler (⭐4.1k) - Utility belt to handle data on AWS.

File System

JuiceFS (⭐13k) - A high-performance Cloud-Native file system driven by object storage for large-scale data storage.

Workflow

CronQ - An application cron-like system. Used w/Luigi. Deprecated.

Kestra - Scalable, event-driven, language-agnostic orchestration and scheduling platform to manage millions of workflows declaratively in code.

May 20 - May 26, 2024

Data Ingestion

Meltano - CLI & code-first ELT.
- Singer SDK - The fastest way to build custom data extractors and loaders compliant with the Singer Spec.

Apr 08 - Apr 14, 2024

Data Comparison

datacompy (⭐623) - A Python library that facilitates the comparison of two DataFrames in Pandas, Polars, Spark and more. The library goes beyond basic equality checks by providing detailed insights into discrepancies at both row and column levels.

Mar 25 - Mar 31, 2024

Data Ingestion

dlt - A fast&simple pipeline building library for Python data devs, runs in notebooks, cloud functions, airflow, etc.

Testing / Data Profiler

DQOps (⭐177) - An open-source data quality platform for the whole data platform lifecycle from profiling new data sources to applying full automation of data quality monitoring.

Mar 18 - Mar 24, 2024

Data Lake Management

Project Nessie (⭐1.4k) - A Transactional Catalog for Data Lakes with Git-like semantics. Works with Apache Iceberg tables.

Community / Podcasts

The Data Stack Show - A show where they talk to data engineers, analysts, and data scientists about their experience around building and maintaining data infrastructure, delivering data and data products, and driving better outcomes across their businesses with data.

Feb 26 - Mar 03, 2024

Workflow

SuprSend - Create automated workflows and logic using API's for your notification service. Add templates, batching, preferences, inapp inbox with workflows to trigger notifications directly from your data warehouse.

Feb 19 - Feb 25, 2024

Workflow

Multiwoven (⭐1.6k) - The open-source reverse ETL, data activation platform for modern data teams.

Feb 05 - Feb 11, 2024

Profiling / Data Profiler

Data Profiler (⭐1.5k) - The DataProfiler is a Python library designed to make data analysis, monitoring, and sensitive data detection easy.

Jan 29 - Feb 04, 2024

Workflow

Prefect - An orchestration and observability platform. With it, developers can rapidly build and scale resilient code, and triage disruptions effortlessly.

Jan 08 - Jan 14, 2024

Data Ingestion

Sling - CLI data integration tool specialized in moving data between databases, as well as storage systems.

Dec 04 - Dec 10, 2023

Workflow

PACE (⭐37) - An open source framework that allows you to enforce agreements on how data should be accessed, used, and transformed, regardless of the data platform (Snowflake, BigQuery, DataBricks, etc.)

Nov 27 - Dec 03, 2023

Databases

Relational
- RQLite (⭐17k) - Replicated SQLite using the Raft consensus protocol.
- MySQL - The world's most popular open source database.
  - TiDB (⭐40k) - A distributed NewSQL database compatible with MySQL protocol.
  - Percona XtraBackup - A free, open source, complete online backup solution for all versions of Percona Server, MySQL® and MariaDB®.
  - mysql_utils (⭐882) - Pinterest MySQL Management Tools.
- MariaDB - An enhanced, drop-in replacement for MySQL.
- PostgreSQL - The world's most advanced open source database.
- Amazon RDS - Makes it easy to set up, operate, and scale a relational database in the cloud.
- Crate.IO - Scalable SQL database with the NOSQL goodies.

Key-Value
- Redis - An open source, BSD licensed, advanced key-value cache and store.
- Riak - A distributed database designed to deliver maximum data availability by distributing data across multiple servers.
- AWS DynamoDB - A fast and flexible NoSQL database service for all applications that need consistent, single-digit millisecond latency at any scale.
- HyperDex (⭐1.4k) - A scalable, searchable key-value store. Deprecated.
- SSDB - A high performance NoSQL database supporting many data structures, an alternative to Redis.
- Kyoto Tycoon (⭐281) - A lightweight network server on top of the Kyoto Cabinet key-value database, built for high-performance and concurrency.
- IonDB (⭐595) - A key-value store for microcontroller and IoT applications.

Column
- Cassandra - The right choice when you need scalability and high availability without compromising performance.
  - Cassandra Calculator - This simple form allows you to try out different values for your Apache Cassandra cluster and see what the impact is for your application.
  - CCM (⭐1.2k) - A script to easily create and destroy an Apache Cassandra cluster on localhost.
  - ScyllaDB (⭐15k) - NoSQL data store using the seastar framework, compatible with Apache Cassandra.
- HBase - The Hadoop database, a distributed, scalable, big data store.
- AWS Redshift - A fast, fully managed, petabyte-scale data warehouse that makes it simple and cost-effective to analyze all your data using your existing business intelligence tools.
- FiloDB (⭐1.5k) - Distributed. Columnar. Versioned. Streaming. SQL.
- Vertica - Distributed, MPP columnar database with extensive analytics SQL.
- ClickHouse - Distributed columnar DBMS for OLAP. SQL.

Document
- MongoDB - An open-source, document database designed for ease of development and scaling.
  - Percona Server for MongoDB - Percona Server for MongoDB® is a free, enhanced, fully compatible, open source, drop-in replacement for the MongoDB® Community Edition that includes enterprise-grade features and functionality.
  - MemDB (⭐593) - Distributed Transactional In-Memory Database (based on MongoDB).
- Elasticsearch - Search & Analyze Data in Real Time.
- Couchbase - The highest performing NoSQL distributed database.
- RethinkDB - The open-source database for the realtime web.
- RavenDB - Fully Transactional NoSQL Document Database.

Graph
- Neo4j - The world's leading graph database.
- OrientDB - 2nd Generation Distributed Graph Database with the flexibility of Documents in one product with an Open Source commercial friendly license.
- ArangoDB - A distributed free and open-source database with a flexible data model for documents, graphs, and key-values.
- Titan - A scalable graph database optimized for storing and querying graphs containing hundreds of billions of vertices and edges distributed across a multi-machine cluster.
- FlockDB (⭐3.3k) - A distributed, fault-tolerant graph database by Twitter. Deprecated.

Distributed
- DAtomic - The fully transactional, cloud-ready, distributed database.
- Apache Geode - An open source, distributed, in-memory database for scale-out applications.
- Gaffer (⭐1.8k) - A large-scale graph database.

Timeseries
- InfluxDB (⭐31k) - Scalable datastore for metrics, events, and real-time analytics.
- OpenTSDB (⭐5.1k) - A scalable, distributed Time Series Database.
- QuestDB - A relational column-oriented database designed for real-time analytics on time series and event data.
- kairosdb (⭐1.8k) - Fast scalable time series database.
- Heroic (⭐846) - A scalable time series database based on Cassandra and Elasticsearch, by Spotify.
- Druid (⭐14k) - Column oriented distributed data store ideal for powering interactive applications.
- Riak-TS - Riak TS is the only enterprise-grade NoSQL time series database optimized specifically for IoT and Time Series data.
- Akumuli (⭐842) - A numeric time-series database. It can be used to capture, store and process time-series data in real-time. The word "akumuli" can be translated from esperanto as "accumulate".
- Rhombus - A time-series object store for Cassandra that handles all the complexity of building wide row indexes.
- Dalmatiner DB (⭐690) - Fast distributed metrics database.
- Blueflood (⭐597) - A distributed system designed to ingest and process time series data.
- Timely (⭐385) - A time series database application that provides secure access to time series data based on Accumulo and Grafana.

Other
- Tarantool (⭐3.6k) - An in-memory database and application server.
- GreenPlum - The Greenplum Database (GPDB) - An advanced, fully featured, open source data warehouse. It provides powerful and rapid analytics on petabyte scale data volumes.
- cayley (⭐15k) - An open-source graph database. Google.
- Snappydata (⭐1k) - OLTP + OLAP Database built on Apache Spark.
- TimescaleDB - Built as an extension on top of PostgreSQL, TimescaleDB is a time-series SQL database providing fast analytics, scalability, with automated data management on a proven storage engine.
- DuckDB - A fast in-process analytical database that has zero external dependencies, runs on Linux/macOS/Windows, offers a rich SQL dialect, and is free and extensible.

Data Ingestion

Kafka - Publish-subscribe messaging rethought as a distributed commit log.
- BottledWater (⭐6) - Change data capture from PostgreSQL into Kafka. Deprecated.
- kafkat (⭐502) - Simplified command-line administration for Kafka brokers.
- kafkacat (⭐5.7k) - Generic command line non-JVM Apache Kafka producer and consumer.
- pg-kafka (⭐112) - A PostgreSQL extension to produce messages to Apache Kafka.
- librdkafka (⭐880) - The Apache Kafka C/C++ library.
- kafka-docker (⭐7k) - Kafka in Docker.
- kafka-manager (⭐12k) - A tool for managing Apache Kafka.
- kafka-node (⭐2.7k) - Node.js client for Apache Kafka 0.8.
- Secor (⭐1.9k) - Pinterest's Kafka to S3 distributed consumer.
- Kafka-logger (⭐45) - Kafka-winston logger for Node.js from Uber.
- Kroxylicious (⭐241) - A Kafka Proxy, solving problems like encrypting your Kafka data at rest.

AWS Kinesis - A fully managed, cloud-based service for real-time data processing over large, distributed data streams.

RabbitMQ - Robust messaging for applications.

FluentD - An open source data collector for unified logging layer.

Embulk - An open source bulk data loader that helps data transfer between various databases, storages, file formats, and cloud services.

Apache Sqoop - A tool designed for efficiently transferring bulk data between Apache Hadoop and structured datastores such as relational databases.

Heka (⭐3.5k) - Data Acquisition and Processing Made Easy. Deprecated.

Gobblin (⭐2.3k) - Universal data ingestion framework for Hadoop from LinkedIn.

Nakadi - An open source event messaging platform that provides a REST API on top of Kafka-like queues.

Pravega - Provides a new storage abstraction - a stream - for continuous and unbounded data.

Apache Pulsar - An open-source distributed pub-sub messaging system.

Airbyte - Open-source data integration for modern data teams.

File System

HDFS - A distributed file system designed to run on commodity hardware.
- Snakebite (⭐859) - A pure python HDFS client.

AWS S3 - Object storage built to retrieve any amount of data from anywhere.
- smart_open (⭐3.4k) - Utils for streaming large files (S3, HDFS, gzip, bz2).

Alluxio - A memory-centric distributed storage system enabling reliable data sharing at memory-speed across cluster frameworks, such as Spark and MapReduce.

CEPH - A unified, distributed storage system designed for excellent performance, reliability, and scalability.

OrangeFS - Orange File System is a branch of the Parallel Virtual File System.

SnackFS (⭐14) - A bite-sized, lightweight HDFS compatible file system built over Cassandra.

GlusterFS - Gluster Filesystem.

XtreemFS - Fault-tolerant distributed file system for all storage needs.

SeaweedFS (⭐29k) - Seaweed-FS is a simple and highly scalable distributed file system. There are two objectives: to store billions of files! to serve the files fast! Instead of supporting full POSIX file system semantics, Seaweed-FS choose to implement only a key~file mapping. Similar to the word "NoSQL", you can call it as "NoFS".

S3QL (⭐1.2k) - A file system that stores all its data online using storage services like Google Storage, Amazon S3, or OpenStack.

LizardFS - Software Defined Storage is a distributed, parallel, scalable, fault-tolerant, Geo-Redundant and highly available file system.

Serialization format

Apache Avro - Apache Avro™ is a data serialization system.

Apache Parquet - A columnar storage format available to any project in the Hadoop ecosystem, regardless of the choice of data processing framework, data model or programming language.
- Snappy (⭐6.5k) - A fast compressor/decompressor. Used with Parquet.
- PigZ - A parallel implementation of gzip for modern multi-processor, multi-core machines.

Apache ORC - The smallest, fastest columnar storage for Hadoop workloads.

Apache Thrift - The Apache Thrift software framework, for scalable cross-language services development.

ProtoBuf (⭐70k) - Protocol Buffers - Google's data interchange format.

SequenceFile - A flat file consisting of binary key/value pairs. It is extensively used in MapReduce as input/output formats.

Kryo (⭐6.5k) - A fast and efficient object graph serialization framework for Java.

Stream Processing

Apache Beam - A unified programming model that implements both batch and streaming data processing jobs that run on many execution engines.

Spark Streaming - Makes it easy to build scalable fault-tolerant streaming applications.

Apache Flink - A streaming dataflow engine that provides data distribution, communication, and fault tolerance for distributed computations over data streams.

Apache Storm - A free and open source distributed realtime computation system.

Apache Samza - A distributed stream processing framework.

Apache NiFi - An easy to use, powerful, and reliable system to process and distribute data.

Apache Hudi - An open source framework for managing storage for real time processing, one of the most interesting feature is the Upsert.

VoltDB - An ACID-compliant RDBMS which uses a shared nothing architecture.

PipelineDB (⭐2.7k) - The Streaming SQL Database.

Spring Cloud Dataflow - Streaming and tasks execution between Spring Boot apps.

Bonobo - A data-processing toolkit for python 3.5+.

Robinhood's Faust (⭐1.9k) - Forever scalable event processing & in-memory durable K/V store as a library with asyncio & static typing.

HStreamDB (⭐728) - The streaming database built for IoT data storage and real-time processing.

Kuiper (⭐1.7k) - An edge lightweight IoT data analytics/streaming software implemented by Golang, and it can be run at all kinds of resource-constrained edge devices.

Zilla (⭐672) - - An API gateway built for event-driven architectures and streaming that supports standard protocols such as HTTP, SSE, gRPC, MQTT, and the native Kafka protocol.

Batch Processing

Hadoop MapReduce - A software framework for easily writing applications which process vast amounts of data (multi-terabyte data-sets) - in-parallel on large clusters (thousands of nodes) - of commodity hardware in a reliable, fault-tolerant manner.

Spark - A multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters.
- Spark Packages - A community index of packages for Apache Spark.
- Deep Spark (⭐197) - Connecting Apache Spark with different data stores. Deprecated.
- Spark RDD API Examples - Examples by Zhen He.
- Livy - The REST Spark Server.
- Delight (⭐346) - A free & cross platform monitoring tool (Spark UI / Spark History Server alternative).

AWS EMR - A web service that makes it easy to quickly and cost-effectively process vast amounts of data.

Data Mechanics - A cloud-based platform deployed on Kubernetes making Apache Spark more developer-friendly and cost-effective.

Tez - An application framework which allows for a complex directed-acyclic-graph of tasks for processing data.

Bistro (⭐8) - A light-weight engine for general-purpose data processing including both batch and stream analytics. It is based on a novel unique data model, which represents data via functions and processes data via columns operations as opposed to having only set operations in conventional approaches like MapReduce or SQL.

Batch ML
- H2O - Fast scalable machine learning API for smarter applications.
- Mahout - An environment for quickly creating scalable performant machine learning applications.
- Spark MLlib - Spark's scalable machine learning library consisting of common learning algorithms and utilities, including classification, regression, clustering, collaborative filtering, dimensionality reduction, as well as underlying optimization primitives.

Batch Graph
- GraphLab Create - A machine learning platform that enables data scientists and app developers to easily create intelligent apps at scale.
- Giraph - An iterative graph processing system built for high scalability.
- Spark GraphX - Apache Spark's API for graphs and graph-parallel computation.

Batch SQL
- Presto - A distributed SQL query engine designed to query large data sets distributed over one or more heterogeneous data sources.
- Hive - Data warehouse software facilitates querying and managing large datasets residing in distributed storage.
  - Hivemall (⭐314) - Scalable machine learning library for Hive/Hadoop.
  - PyHive (⭐1.7k) - Python interface to Hive and Presto.
- Drill - Schema-free SQL Query Engine for Hadoop, NoSQL and Cloud Storage.

Charts and Dashboards

Highcharts - A charting library written in pure JavaScript, offering an easy way of adding interactive charts to your web site or web application.

ZingChart - Fast JavaScript charts for any data set.

C3.js - D3-based reusable chart library.

D3.js - A JavaScript library for manipulating documents based on data.
- D3Plus - D3's simpler, easier to use cousin. Mostly predefined templates that you can just plug data in.

SmoothieCharts - A JavaScript Charting Library for Streaming Data.

PyXley (⭐2.3k) - Python helpers for building dashboards using Flask and React.

Plotly (⭐24k) - Flask, JS, and CSS boilerplate for interactive, web-based visualization apps in Python.

Apache Superset (⭐70k) - A modern, enterprise-ready business intelligence web application.

Redash - Make Your Company Data Driven. Connect to any data source, easily visualize and share your data.

Metabase (⭐45k) - The easy, open source way for everyone in your company to ask questions and learn from data.

PyQtGraph - A pure-python graphics and GUI library built on PyQt4 / PySide and numpy. It is intended for use in mathematics / scientific / engineering applications.

Workflow

Luigi (⭐19k) - A Python module that helps you build complex pipelines of batch jobs.

Cascading - Java based application development platform.

Airflow (⭐44k) - A system to programmatically author, schedule, and monitor data pipelines.

Azkaban - A batch workflow job scheduler created at LinkedIn to run Hadoop jobs. Azkaban resolves the ordering through job dependencies and provides an easy-to-use web user interface to maintain and track your workflows.

Oozie - A workflow scheduler system to manage Apache Hadoop jobs.

Pinball (⭐1k) - DAG based workflow manager. Job flows are defined programmatically in Python. Support output passing between jobs.

Dagster (⭐15k) - An open-source Python library for building data applications.

Kedro - A framework that makes it easy to build robust and scalable data pipelines by providing uniform project templates, data abstraction, configuration and pipeline assembly.

Dataform - An open-source framework and web based IDE to manage datasets and their dependencies. SQLX extends your existing SQL warehouse dialect to add features that support dependency management, testing, documentation and more.

Census - A reverse-ETL tool that let you sync data from your cloud data warehouse to SaaS applications like Salesforce, Marketo, HubSpot, Zendesk, etc. No engineering favors required—just SQL.

dbt - A command line tool that enables data analysts and engineers to transform data in their warehouses more effectively.

RudderStack (⭐4.3k) - A warehouse-first Customer Data Platform that enables you to collect data from every application, website and SaaS platform, and then activate it in your warehouse and business tools.

Data Lake Management

lakeFS (⭐5.1k) - An open source platform that delivers resilience and manageability to object-storage based data lakes.

ELK Elastic Logstash Kibana

docker-logstash (⭐237) - A highly configurable Logstash (1.4.4) - Docker image running Elasticsearch (1.7.0) - and Kibana (3.1.2).

elasticsearch-jdbc (⭐2.8k) - JDBC importer for Elasticsearch.

ZomboDB (⭐4.7k) - PostgreSQL Extension that allows creating an index backed by Elasticsearch.

Docker

Gockerize (⭐667) - Package golang service into minimal Docker containers.

Flocker (⭐3.4k) - Easily manage Docker containers & their data.

Rancher - RancherOS is a 20mb Linux distro that runs the entire OS as Docker containers.

Kontena - Application Containers for Masses.

Weave (⭐6.6k) - Weaving Docker containers into applications.

Zodiac (⭐200) - A lightweight tool for easy deployment and rollback of dockerized applications.

cAdvisor (⭐19k) - Analyzes resource usage and performance characteristics of running containers.

Micro S3 persistence (⭐14) - Docker microservice for saving/restoring volume data to S3.

Rocker-compose (⭐409) - Docker composition tool with idempotency features for deploying apps composed of multiple containers. Deprecated.

Nomad (⭐16k) - A cluster manager, designed for both long-lived services and short-lived batch processing workloads.

ImageLayers - Visualize Docker images and the layers that compose them.

Nov 20 - Nov 26, 2023

Testing / Data Profiler

Grai (⭐313) - A data catalog tool that integrates into your CI system exposing downstream impact testing of data changes. These tests prevent data changes which might break data pipelines or BI dashboards from making it to production.

Feb 08 - Feb 14, 2021

Community / Conferences

Data Council - The first technical conference that bridges the gap between data scientists, data engineers and data analysts.

Feb 04 - Feb 10, 2019

Datasets / Realtime

Twitter Realtime - The Streaming APIs give developers low latency access to Twitter's global stream of Tweet data.

Datasets / Data Dumps

GitHub Archive - GitHub's public timeline since 2011, updated every hour.

Aug 21 - Aug 27, 2017

Community / Podcasts

Data Engineering Podcast - The show about modern data infrastructure.

Apr 10 - Apr 16, 2017

Community / Forums

/r/dataengineering - News, tips, and background on Data Engineering.

/r/etl - Subreddit focused on ETL.

Mar 20 - Mar 26, 2017

Datasets / Realtime

Reddit - Real-time data is available including comments, submissions and links posted to reddit.

Datasets / Data Dumps

Common Crawl - Open source repository of web crawl data.

Wikipedia - Wikipedia's complete copy of all wikis, in the form of Wikitext source and metadata embedded in XML. A number of raw database tables in SQL form are also available.

Sep 07 - Sep 13, 2015

Datasets / Realtime

Eventsim (⭐534) - Event data simulator. Generates a stream of pseudo-random events from a set of users, designed to simulate web traffic.

Jul 20 - Jul 26, 2015

Monitoring / Prometheus

Prometheus.io (⭐62k) - An open-source service monitoring system and time series database.

HAProxy Exporter (⭐624) - Simple server that scrapes HAProxy stats and exports them via HTTP for Prometheus consumption.