;
;

DZone Big Data Zone

Recent posts in Big Data on DZone.com

From ETL, ELT, and EtLT to Agent: What Is Changing in Enterprise Data Engineering?

For the past two decades, most enterprise data engineering systems have been built on one default assumption: People understand the system. The sys...
Posted on 8 September 2026 | 6:00 pm

Handling Large API Responses Without Freezing the Client: A Practical Architecture With Temporal, Kafka, and RAG

A large API response becomes a client problem long before it becomes a network problem. A browser can receive hundreds of megabytes and still becom...
Posted on 4 September 2026 | 5:00 pm

Your Spark Job Isn't Slow Because of Bad Code. It's Slow Because of the Wrong Join

I learned this lesson the hard way. We had a critical data pipeline running for over 3 hours every single day. The logic was perfectly clean. The o...
Posted on 3 September 2026 | 3:00 pm

Designing Replay-Safe CDC Pipelines With Kafka, Debezium, and Recovery Contracts

Change data capture (CDC) pipelines look straightforward on paper: capture database changes, publish them to Kafka, and update downstream systems. ...
Posted on 1 September 2026 | 6:00 pm

Real-Time Supply Chain Event Streaming With Kafka and Neo4j

In a previous article, we built a static supply chain graph in Neo4j using Apache Spark, with suppliers, warehouses, distribution centers, and reta...
Posted on 18 August 2026 | 6:00 pm

Orchestrating Small Language Models Without Losing Events or Context

Reliable orchestration for small language models depends less on model sophistication than on the durability of event flow and state. Under the ass...
Posted on 13 August 2026 | 2:00 pm

A Practical Pipeline for Identifying Sensitive Columns Before Test Data Masking

I work as a data analyst at a legal services company. Part of my work involves protecting sensitive data during the Test Data Management (TDM) proc...
Posted on 10 August 2026 | 1:00 pm

Supply Chain Resilience Analysis With Apache Spark and Neo4j

Supply chains are graphs. Suppliers feed into warehouses, warehouses feed into distribution centers, and distribution centers feed into retailers. ...
Posted on 10 August 2026 | 12:00 pm

How We Cut PyFlink Pipeline p99 Latency from 3-5 Seconds to ~500ms

The Problem: Our p99 Was 3-5 Seconds Our PyFlink pipeline was missing its latency SLO by seconds. The pipeline itself was straightforward: consume ...
Posted on 7 August 2026 | 4:00 pm

TensorFlow vs PyTorch: The Real Difference Isn’t Accuracy

A few days ago, I set out to build a simple image classification model using convolutional neural networks (CNNs). The task itself wasn’t particula...
Posted on 5 August 2026 | 5:00 pm

Why LLM Pipelines Fail in Production and How Temporal and Kafka Fix Them

A production LLM pipeline is rarely just a prompt and a response. It typically combines retrieval, prompt rendering, model inference, output shapin...
Posted on 5 August 2026 | 3:00 pm

Why Enterprise AI Agents Fail: A Runtime Data Governance Pattern for Reliable Answers

The Failure You Have Probably Already Seen An enterprise AI agent is deployed against production data. It answers the first ten questions confident...
Posted on 3 August 2026 | 2:00 pm

Compliance Reporting Without Losing the Spreadsheet or the Control

Compliance-reporting teams keep spreadsheets in the loop for a practical reason: a workbook lets domain experts inspect assumptions, formulas, sour...
Posted on 14 July 2026 | 5:00 pm

AWS Glue ETL Design Principles for Production PySpark Pipelines

AWS Glue makes it easy to get a PySpark pipeline running quickly. It is significantly harder to build one that stays maintainable as logic grows, p...
Posted on 14 July 2026 | 2:00 pm

Top 10 Best Places to Prepare for Your Next Data Engineer Interview

Landing a data engineering role means clearing a gauntlet that no other software discipline has to face all at once: airtight SQL, production-grade...
Posted on 10 July 2026 | 12:00 pm

Building Production-Grade Delta Lake Pipelines With Apache Spark on Databricks

Why Delta Lake? Apache Parquet on cloud storage was a great first step for data lakes — but it left engineers dealing with a painful set of problem...
Posted on 8 July 2026 | 2:00 pm

Azure Databricks for Scalable MLOps and Feature Engineering With Apache Spark, Delta Lake, and MLflow

Raw data doesn't win model competitions. Features do. And when your raw data is tens of billions of rows sitting across multiple sources, you can't...
Posted on 6 July 2026 | 2:00 pm

From Polling to PubSub: Building an Asynchronous OPC UA Stack in Python

Industrial control systems are generating more data than ever before, but the Python tooling used to process this telemetry often encounters severe...
Posted on 3 July 2026 | 6:00 pm

Real-Time AI Feature Engineering With Spark Structured Streaming and Databricks Feature Store

The Feature Engineering Problem Feature engineering is where most ML projects silently fail in production. Not because the model is wrong — but bec...
Posted on 2 July 2026 | 6:00 pm

Dead Letter Queue Patterns in Apache Flink: Handling Poison Messages Without Stopping Your Stream

Streaming systems usually fail in one of two ways: Loudly, when infrastructure breaks Quietly, when one bad record keeps replaying until th...
Posted on 2 July 2026 | 1:00 pm