Apache Spark RDD Transformations Pipeline
RDDs (Resilient Distributed Datasets) are Spark’s fundamental abstraction — an immutable, partitioned collection of records that can be operated on in parallel.
Transformations like map, filter, and flatMap are lazy: they build a DAG of dependencies without executing until an action (like reduce or collect) triggers computation.
This pipeline model enables Spark to optimize execution by fusing transformations, minimizing data shuffles, and recovering from failures using lineage information.
Step through the pipeline below: the first four steps only build the DAG, and nothing runs until collect() asks for a result. Click a source record to change its value, click any other record to follow its key through every column, and switch map-side combine off to see how many more records the shuffle has to move.
Six partitions of key–value records on three machines flow through mapValues, filter and reduceByKey: the DAG stays lazy until collect(), then each partition maps, filters, combines and sorts locally, the records shuffle to the partition their key hashes to, and the reduced sums arrive at the driver.