Skip to main content
Acacia
Fund
Start Here
News
Compliance
Markets
Data
|
Study
Review
Diagnostic
Journeys
Console
A-Z
Search
Mode
Beginner
Advanced
Expert
×
Reading Flow
Focus mode
Reading guide
Information Density
Compact
Standard
Comfortable
Presentation
Preferred way to absorb content.
Visual
Balanced
Verbal
Learning Profile
Difficulty
Calibrate
Interests
Choose
Home
/
Data Engineering Research
Data Engineering Research
Data Engineering
2027 Data Engineering Predictions: AI-Augmented Pipelines, Real-Time Universal Catalogs, and the Death of Batch
Data Engineering
Data Platform as a Product: UX Patterns for Internal Developer Platforms
Data Engineering
The Rise of the Analytics Engineer: dbt, SQLMesh, and the Modern Data Stack
Data Engineering
Cost Optimization in Data Pipelines: Engineering for Efficiency at Petabyte Scale
Data Engineering
Schema Registry Patterns: Avro, Protobuf, and JSON Schema in Production
Data Engineering
Data Products: Designing APIs for the Internal Data Platform
Data Engineering
Data Mesh in Practice: Implementing Domain Ownership Without Chaos
Data Engineering
Terraform for Data Infrastructure: Infrastructure as Code for the Data Platform
Data Engineering
Kubernetes for Data Engineering: Running Data Pipelines on K8s
Data Engineering
Feature Stores at Scale: Feast vs Tecton in Production Deployments
Data Engineering
ML Pipeline Orchestration: From Notebook to Production with Feast and MLflow
Data Engineering
Data Observability: Monitoring, Lineage, and Incident Response for Pipelines
Data Engineering
Building a Data Platform on a Budget: The Open Source Stack in 2026
Data Engineering
SQLMesh: The SQL-First Data Transformation Framework Challenging dbt
Data Engineering
dbt Mesh: Decentralizing Data Transformation at Enterprise Scale
Data Engineering
Delta Lake vs Apache Iceberg vs Apache Hudi: Lakehouse Format Shootout
Data Engineering
Debezium and CDC: Capturing Database Changes at Scale
Data Engineering
Apache Iceberg Deep Dive: Table Formats for the Lakehouse Era
Data Engineering
Real-Time Streaming with Apache Kafka: From Pub/Sub to Event-Driven Architecture
Data Engineering
Data Contracts: Schema as API for the Analytics Team
Data Engineering
Airflow vs Prefect vs Dagster: Choosing the Right Orchestrator in 2026
Data Engineering
Data Quality at Scale: Great Expectations Beyond Unit Tests for Data
Data Engineering
Dagster 2.0: Next-Gen Data Pipeline Orchestration for the Modern Data Platform
Data Engineering
Data Pipeline Patterns for High-Throughput Genomics: Orchestrating Bioinformatics Workflows with Dagster
Data Engineering
Reproducible ML Pipelines in Computational Biology: MLOps for CRISPR Target Discovery
Data Engineering
EY Canada published a cybersecurity report and most citations were hallucinated -- 🛡️ AML 2026-06-01
Data Engineering
Fellegi & Sunter (1969) - A Theory for Record Linkage
Data Engineering
Wang & Strong (1996) - Beyond Accuracy: What Data Quality Means to Data Consumers
Data Engineering
Dean & Ghemawat (2004) - MapReduce: Simplified Data Processing on Large Clusters
Data Engineering
Chang et al (2006) - Bigtable: A Distributed Storage System for Structured Data
Data Engineering
Akidau et al (2015) - The Dataflow Model: Balancing Correctness, Latency, and Cost
Data Engineering
Armbrust et al (2021) - Lakehouse: A New Generation of Open Platforms
Data Engineering
Kreps, Narkhede & Rao (2011) - Kafka: A Distributed Messaging System for Log Processing
Data Engineering
Zaharia et al (2012) - Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing
Data Engineering
Apache Iceberg - Open Table Format Specification
Data Engineering
Stonebraker et al (2018) - What Goes Around Comes Around: Data Management Cycles
Data Engineering
The Lakestream Paradigm: How Streaming-First Lakehouse Architecture Is Replacing Batch ETL in 2026
Data Engineering
SQLite in Production: Optimizing WAL Mode, Concurrency, and VFS Layers
Data Engineering
Rune 1.1: adds Python, an Emacs editor, a symbol index and is now free
Data Engineering
Choose DuckDB rather than SQLite
Data Engineering
CosmosEscape: Taking over Every Database in Azure Cosmos DB
Data Engineering
The development pipeline is a production system
Data Engineering
Show HN: A local merge queue for parallel Claude Code agents
Home
Paths
Study
Me