;
;

DZone Data Zone

Recent posts in Data on DZone.com

From ETL, ELT, and EtLT to Agent: What Is Changing in Enterprise Data Engineering?

For the past two decades, most enterprise data engineering systems have been built on one default assumption: People understand the system. The sys...
Posted on 8 September 2026 | 6:00 pm

Select AI and Vector Search on a Legacy Oracle Schema: What It Actually Takes

Who this is for: DBAs and developers sitting on an Oracle schema that's been in production for a decade or more, who keep hearing that Select AI an...
Posted on 7 September 2026 | 6:00 pm

Prompting AI for Analytics: The Missing Optimization Layer Between Your Question and the Model

The answer was right. The question cost four times what it needed to. Every analytics team using AI models runs into the same quiet cost: wasted to...
Posted on 4 September 2026 | 7:00 pm

Your Spark Job Isn't Slow Because of Bad Code. It's Slow Because of the Wrong Join

I learned this lesson the hard way. We had a critical data pipeline running for over 3 hours every single day. The logic was perfectly clean. The o...
Posted on 3 September 2026 | 3:00 pm

Best Practices for Handling Bad Data in Stream Processing Platforms

Today, stream processing platforms facilitate the real-time analysis of data flowing continuously from Internet of Things (IOT) devices, financial ...
Posted on 2 September 2026 | 5:00 pm

Ampere PMU Profiler: A Guide to Microarchitecture Profiling

Executive Summary The Ampere® PMU Profiler (APP) is a Python-based tool designed to provide deep insight into the microarchitectural behavior of ap...
Posted on 1 September 2026 | 6:01 pm

Designing Replay-Safe CDC Pipelines With Kafka, Debezium, and Recovery Contracts

Change data capture (CDC) pipelines look straightforward on paper: capture database changes, publish them to Kafka, and update downstream systems. ...
Posted on 1 September 2026 | 6:00 pm

Stop Hardcoding Database Checks: Building a Metadata-Driven Data Quality Framework

In high-volume data platforms, hardcoding validation logic into individual processing pipelines creates significant operational drag. As an enterpr...
Posted on 1 September 2026 | 3:00 pm

Evolve or Automate: What It Actually Means to Be an AI-Native Data Engineer

The Moment It Gets Real At some point in the last year, every data engineer had the same experience. You opened a copilot tool, typed a rough descr...
Posted on 1 September 2026 | 12:00 pm

Designing a Dynamic Multi-Hierarchy Security Model for Analytics and Decision Support Systems

A simple access check uncovered something alarming: several dashboards still showed employee compensation based on an organizational hierarchy that...
Posted on 31 August 2026 | 7:00 pm

When "Roughly Right" Looks Like a Liability: Engineering Financial-Grade Data Pipelines

Analytics teams do not get too upset about small errors. If a product dashboard is off by half a percent on a Tuesday, nobody files a ticket. If yo...
Posted on 31 August 2026 | 2:00 pm

Understanding RabbitMQ Exchange Types in Spring Boot

In this blog, you will take a closer look at the different exchange types that can be used in RabbitMQ. All are demonstrated by means of examples...
Posted on 26 August 2026 | 1:00 pm

Designing Rayfall: One Expression Language for a Columnar Database

Columnar engines naturally organize computation around vectors to make effective use of single instruction, multiple data (SIMD) instructions. This...
Posted on 25 August 2026 | 5:00 pm

Stop Paying Your AI Agent to Do the Same Job Twice

If you have wired an AI agent into a real production workflow, you have probably hit this wall; the agent is genuinely good at the task, but it is ...
Posted on 21 August 2026 | 3:00 pm

Building Data Pipelines: Here's What Palantir Foundry Did That Surprised Me.

Senior data engineers are trained to be skeptical of proprietary platforms. When I entered a Palantir Foundry training bootcamp, I expected to find...
Posted on 18 August 2026 | 3:00 pm

Vector Database Indexing Explained: Why It Matters More Than the Embeddings Themselves

Most conversations about vector databases start and end with embeddings. Discussions typically center around how they're generated, which model pro...
Posted on 18 August 2026 | 1:00 pm

The Embedding Model You Choose Matters More Than Your LLM

The Uncomfortable Truth You’ve spent days prompt-engineering your LLM. You’ve benchmarked Claude against GPT. You’ve debated whether to use Mixtral...
Posted on 17 August 2026 | 7:00 pm

Audit-Ready by Design: Building Lineage, Point-in-Time Reconstruction, and Immutability Into Data Architecture

Compliance Checkbox vs. Architectural Constraint Most data platforms treat audit-readiness as a downstream concern. The pipelines are built, the wa...
Posted on 17 August 2026 | 6:00 pm

Enterprise AI Data Engineering With Snowflake Cortex and RAG

Where the Data Actually Lives Every enterprise I have worked with hits the same wall. Mountains of data. Warehouses, ticketing systems, PDFs, old e...
Posted on 13 August 2026 | 7:00 pm

From Microservices to Agent Services: The Next Architectural Shift

The evolution from monolithic applications to microservices transformed enterprise software by decomposing business capabilities into independently...
Posted on 12 August 2026 | 6:00 pm