Data Pipelines

For much of the past decade, companies moved data into their analytics systems with pipelines written as code: Spark or Python jobs that pulled records out of operational databases, reshaped them, and loaded the results into a warehouse. The approach works, but it is brittle. A single change upstream, such as a renamed or newly added column, can stop a job and send an engineer digging through custom code to find out why.

That brittleness has become more expensive. Companies now want the same data to feed dashboards, machine-learning models and AI agents, and the condition of that data is increasingly the limiting factor. In a July 2024 survey of 1,203 data management leaders, Gartner found that 63 percent of organizations either do not have, or are unsure whether they have, the right data management practices for AI. The research firm predicted that through 2026, organizations will abandon 60 percent of AI projects that are not supported by AI-ready data.

Spending on the tools that move and prepare that data is rising with it. Grand View Research estimates the global data integration market at $17.9 billion in 2025 and projects that it could reach $47 billion by 2033, a compound annual growth rate of 12.7 percent.

The Industry Is Settling on a Common Pattern

Within that market, a recognizable design has taken hold. Raw data lands in inexpensive cloud object storage in an open table format, most often Apache Iceberg, which adds database features such as transactions, controlled schema changes, and "time travel" queries of earlier versions to what are otherwise plain files. Transformations are written in SQL and managed with dbt, an open-source tool that treats data models like software, with version control, dependency tracking, and automated tests. And the data is refined in stages.

Databricks, which popularized the term, describes this "medallion architecture" as a series of layers that denote data quality: bronze for raw data as it was ingested, silver for data that has been cleaned and validated, and gold for highly refined tables that drive analytics, dashboards, and machine learning. The company's own documentation calls it "a recommended best practice but not a requirement."

The commercial signals behind the pattern are hard to miss. Databricks acquired Tabular, the startup founded by Iceberg's original creators, in June 2024, for a price Bloomberg reported at nearly $2 billion. In December 2024, Amazon Web Services launched S3 Tables, managed Iceberg storage built into its S3 object store, claiming up to three times faster query performance than Iceberg tables kept in general-purpose S3 buckets. In August 2026, AWS Glue 6.0 added full support for version 3 of the Iceberg specification, alongside a 30 percent price cut.

The SQL side of the stack has consolidated too. Fivetran and dbt Labs completed their merger on June 1, 2026, and the combined company says more than 100,000 data teams use dbt. A January 2026 survey of 252 senior data and IT leaders in North America and Europe, commissioned by the Iceberg-tooling vendor Ryft, found that 58 percent were already using Iceberg for business-critical analytics. The survey also found that most of them relied on custom scripts and internal tools to keep their Iceberg environments optimized and governed.

Putting Numbers on the Switch

What is less clear is how much the switch is worth. Vendors publish favorable benchmarks and practitioners publish migration stories, but side-by-side comparisons under the same conditions are scarcer.

A paper titled "Medallion Architecture for Cloud Data Integration Using dbt on AWS" attempts such a comparison. The study was authored by Balakrishna Pothineni, Ashok Gadi Parthi, Ram Sekhar Bodala, Nitin Saksena, Aswathnarayan Muthukrishnan Kirubakaran, Abhirup Mazumder, Bikesh Kumar, and Sumit Saha. The paper was presented at the 2025 International Conference on Computer and Applications (ICCA), held in Bahrain on December 22–24, 2025, and appears in the conference proceedings published by IEEE.

Rather than evaluating dbt, Iceberg, or medallion architecture independently, the researchers examined how the technologies could operate as an integrated data platform.

The architecture begins with change records captured from PostgreSQL through AWS Database Migration Service. Those records land in Amazon S3 as Iceberg tables in the bronze layer. dbt models clean, deduplicate, and validate the data into a silver layer before producing business-oriented gold tables. Apache Airflow orchestrates the workflow, while Athena and Redshift provide analytical access. Automated tests check conditions such as duplicate transaction IDs, missing values, and orphaned customer records.

As lead author, Balakrishna Pothineni played a central role in shaping the research around a practical question: whether an integrated dbt, Apache Iceberg, and medallion architecture could deliver measurable improvements over conventional Glue-based ETL. His work focused on translating that architectural concept into a testable framework and evaluating it across performance, data quality, cost, and engineering-efficiency measures. That research direction is reflected in the study's controlled comparison of the two approaches under the same simulated workload.

"Most of these tools already existed. What we kept seeing was that every team stitched them together differently, and the pipeline ended up as something only one or two people really understood," Pothineni said. "We wanted to treat the whole thing as one design and then actually measure it."

The layering itself is part of that argument. Because bronze tables keep data exactly as it arrived, a faulty transformation can be fixed and rerun without going back to the source systems. In the authors' account, the contribution lies not in the individual technologies but in the development and quantitative evaluation of an operational framework showing how those technologies interact within an AWS data environment.

Pothineni described the role of the bronze layer in practical terms:

"The bronze layer is your insurance policy. If a transformation goes wrong, you still have the data as it arrived, and you can rebuild everything downstream from it."

What the Tests Found

The authors evaluated the proposed approach through a controlled, large-scale laboratory simulation designed to approximate demanding enterprise data workloads rather than through a live production deployment. The experimental environment used a synthetic retail dataset scaled to one petabyte, drawing on Kaggle e-commerce benchmarks and samples from TPC-DS, a widely used decision-support benchmark. The team also stress-tested incremental data processing under approximately 10,000 concurrent write events. As a baseline, the study used a conventional AWS implementation based on Glue ETL jobs written in PySpark, allowing the alternative architecture to be evaluated under consistent experimental conditions.

The results reported in the paper show substantial performance differences. For a one-terabyte workload, processing time decreased from approximately 100–120 minutes with the Glue ETL baseline to 35–50 minutes using dbt and Apache Iceberg. Athena query latency declined from 45–60 seconds to 15–25 seconds, while the proportion of records containing data-quality errors fell from 4.2–6 percent to 0.5–1 percent. In the simulated one-petabyte scenario, estimated monthly operating costs decreased from $1,800–$2,500 to $900–$1,200.

The study also examined engineering productivity and data-access efficiency. Pipeline development was reduced from an estimated 12–15 days to three to five days, while the modeled deployment cadence improved from weekly releases to hourly or daily deployments. The authors further report that Iceberg compaction and metadata pruning reduced the volume of data scanned by approximately 60 percent, while converting Parquet-based tables to Iceberg produced query-performance improvements of roughly two to three times.

By evaluating the architecture against a defined baseline, at terabyte- and petabyte-scale data volumes, and under high-concurrency conditions, the study provides quantitative evidence of the performance, data-quality, cost, and engineering-efficiency tradeoffs associated with integrating dbt and Apache Iceberg into an AWS-based analytical environment. Because the evaluation was conducted in a controlled simulation, the reported results should be interpreted as experimental evidence rather than claims of equivalent performance in every production environment. The paper also does not separate values measured in its own runs from the 2025 benchmark data it says it incorporated, or state which charges its monthly cost figure covers. Even so, the study offers a documented starting point for assessing an architectural approach that organizations could further validate against their own workloads and operational constraints.

One result stands out because it reflects a change in behavior rather than speed. With Glue, the authors report two to four hours of downtime when a table's schema changed; with Iceberg, which records each version of a table's structure, they report none. Abhirup Mazumder contributed to the study's examination of cloud data reliability and schema evolution, an important concern when integrating changing source data across analytical pipelines. The research evaluated how Iceberg's table-management capabilities could reduce the operational disruption traditionally associated with schema changes.

Mazumder connected that capability to a common data-engineering problem:

"Schema changes are where a lot of pipelines quietly break. Once the table format tracks every version, adding a column stops being an outage and becomes an ordinary change."

A Real-World Migration Points the Same Way

The closest independent comparison comes from a public body. In October 2024, the UK Ministry of Justice's data engineering team published an account on its engineering blog of moving a transaction data lake from Glue PySpark to Athena, Iceberg and dbt, broadly the same shift the paper models. The team reported that runtime for its longest jobs fell by 75 percent, that converting CSV files to Parquet with Athena was 99 percent cheaper than with Glue PySpark, and that refreshes moved from weekly to daily.

The ministry's data was far smaller than the paper's simulation: its largest dataset was under 500 gigabytes. The team also cautioned that the cost advantage "may diminish for data lakes operating at the petabyte-scale," the scale the paper simulates.

Where the Approach Could Apply

Retail provides the paper's immediate use case. Customer measures based on purchase recency, frequency, and spending can be incrementally refreshed rather than reconstructed through large weekly batches. Nitin Saksena contributed to the study's cloud data platform architecture, including the design principles behind its medallion-based data layers. His contribution connects the architecture's incremental data-processing capabilities with the need to make current, business-ready information available more quickly for downstream analytics and decision-making.

Saksena emphasized what more frequent data refreshes can mean in practice:

"In retail, a customer metric that's a week old is often answering last week's question. Moving from weekly batches toward daily or hourly refreshes changes what the business can actually do with that data."

Financial services introduce a different requirement: reproducibility. A bank may need to determine what a number looked like on a particular date and how it was calculated. Iceberg snapshots combined with dbt lineage could help reconstruct both the underlying data and transformation logic. Sumit Saha contributed to the study's data governance and customer-analytics perspective, examining how auditable, point-in-time data can support regulatory reporting as well as data-driven customer applications in financial services. His contribution connects the architecture's lineage and historical-data capabilities with the broader need for traceability across integrated cloud data platforms.

Saha highlighted why that traceability matters in financial services:

"In banking you're regularly asked what a number looked like on a specific date and how it was calculated," Saha said. "Snapshots and lineage don't make that question go away, but they make it a lot easier to answer."

High-volume operators such as telecommunications carriers and transportation companies, which capture large streams of changing records, could use the same change-capture and incremental-merge design to keep analytical tables current without reprocessing full histories. Public-sector analytics teams, as the Ministry of Justice case shows, are already using variations of the pattern.

From Controlled Research to Real-World Applications

The study was designed as a controlled evaluation rather than a production deployment, giving the researchers a consistent environment in which to compare the two architectures. Its synthetic retail workload allowed the team to examine processing performance, query latency, data quality, engineering effort, and cost at terabyte and simulated petabyte scale. The results therefore provide experimental evidence under defined conditions rather than a prediction of what every organization will experience in production.

The evaluation also helped identify where a hybrid approach remains useful. While dbt provides a SQL-centered framework for many transformation workloads, deeply nested and semi-structured data can require more flexible processing. The researchers therefore used Spark alongside dbt for those cases rather than treating the technologies as mutually exclusive. Ram Sekhar Bodala contributed to the study's data-integration and hybrid-processing strategy, examining how dbt transformations and Spark workloads could work together within the same cloud data architecture. His contribution helped define where SQL-based transformation is effective and where more flexible distributed processing remains necessary.

Bodala explained the rationale for combining the two approaches:

"SQL covers most of what business teams need, but not everything. Deeply nested event data is still awkward in dbt. For that we used Spark alongside it, so it's a hybrid, not a wholesale replacement."

The architecture also has implications beyond conventional reporting and analytics. Aswathnarayan Muthukrishnan Kirubakaran brought a machine-learning perspective to the research, connecting the cloud data-engineering framework with the requirements of downstream ML workloads. His contribution focused on how current, consistently transformed, and quality-checked data can provide a stronger foundation for machine-learning applications. The bronze, silver, and gold structure evaluated by the researchers provides a framework in which raw information can be retained while progressively validated datasets are prepared for analytics and machine-learning consumption.

Muthukrishnan Kirubakaran also identified production validation as an important next step for evaluating the architecture beyond controlled ML and data-processing workloads:

"This was a controlled simulation, and we want to be clear about that. Production systems have messier data, competing workloads and budgets that don't look like a lab. The real test is whether the gains hold up there."

What Comes Next

The authors identify tighter integration with OpenLineage, an open standard for tracking how data moves between systems, and extending the design to Microsoft Azure and Snowflake as next steps, along with AI-assisted pipeline optimization.

The comparison itself may also evolve. Glue 6.0 introduces Spark Declarative Pipelines, allowing engineers to define desired outcomes while the platform manages aspects of execution. That brings Glue closer to the declarative development model associated with dbt and suggests that the distinction between traditional code-based ETL and SQL-based ELT may continue to narrow.

For Pothineni, broader testing is part of the research's next phase. He sees independent evaluation across different data volumes, partitioning strategies, and query patterns as a way to build a clearer picture of where the architecture delivers the greatest operational and cost advantages.

Pothineni framed different results from future implementations as useful evidence rather than a contradiction of the study:

"If someone runs this on their own petabyte workload and gets different numbers, that's useful too. Cost depends heavily on how a team partitions, compacts and queries its data. We'd like to see more people publish those results."

For data teams, the broader question is increasingly not whether SQL models, layered architectures and open table formats can work together, but how their benefits change across workloads and scale. The Ministry of Justice's experience and the paper's petabyte-scale simulation point toward the same next step: evaluating these architectures against large production workloads and publishing enough configuration, performance, and cost information for others to compare and build on the results.