Apache Spark — The answer to a slow job is in the plan and the event log
explain is the receipt showing what Spark decided not to read
In one line
Spark does not run the code you wrote as it is. The optimizer rewrites the plan so that it pushes conditions down to the file reader, cuts away columns that are not used, and skips partition directories that are not needed. The result is printed in the physical plan of explain as PushedFilters, ReadSchema and PartitionFilters, and the final plan, which AQE rewrites once more during execution, is visible only after the action has finished.
Why you need to read the plan
Two queries that produce the same result, one takes 3 seconds and the other 3 minutes. The code looks almost the same. The difference usually comes from how much was read. For example, wrapping one condition in a function made the whole Parquet file get read, or a single select("*") line lifted twenty unneeded columns from disk.
You cannot see such differences however much you read the code. You have to look at what plan Spark turned the code into. That receipt is explain. There is one important attitude here: do not trust the optimization; verify it. The optimizer does it only when it can, and it raises no warning when it could not.
How it works
According to the SQL EXPLAIN reference, EXTENDED shows four plans: the parsed logical plan extracted from the query, the analyzed logical plan with names and types resolved, the optimized logical plan after the optimization rules, and the physical plan that will actually run. PySpark's explain selects these by mode. The default (simple) shows only the physical plan, extended shows both logical and physical, codegen shows the generated code, cost shows the logical plan with statistics when statistics are available, and formatted shows the physical plan split into an overview and per-node details.
If you compare the optimized logical plan with the physical plan, you can see what the optimizer did. Most of it is gathered in the single line of the node that reads the file.
+- FileScan parquet [order_id#0,qty#3,day#7]
PartitionFilters: [isnotnull(day#7), (day#7 = 2026-01-03)]
PushedFilters: [IsNotNull(qty), GreaterThanOrEqual(qty,5)]
ReadSchema: struct<order_id:string,qty:int>
Filter pushdown (PushedFilters). It means that where(qty >= 5) does not exist only as a Filter node but was handed to the file reader. In the Parquet configuration, the default of spark.sql.parquet.filterPushdown is true. What can be skipped with the pushed condition depends on the statistics the file has. The section on leveraging statistics gives the counts and the minimum and maximum values in Parquet metadata as an example of statistics that Spark reads directly from the data source. With the minimum and maximum values alone, a chunk that cannot satisfy the condition does not have to be opened.
Column pruning (ReadSchema). Only the columns the query uses to the end are read. Parquet stores data grouped by column, so if you read only two columns, the bytes of the other columns are not lifted from disk. CSV is row by row, so this gain is much smaller. This is the reason to convert a file to Parquet once.
Partition pruning (PartitionFilters). This is a condition on a column whose value is in the directory name, such as day=2026-01-03/. Partition discovery in Parquet automatically extracts the partition columns from such paths and puts them into the schema. A condition on that column is handled at the directory listing stage, before any file is opened, so dates that do not match cost 0.
Merging consecutive filters. Even if you split a condition into several calls like where(a).where(b), they are merged into a single Filter in the optimized logical plan. So splitting conditions for readability does not hurt performance.
The moment pushdown breaks
Pushdown works best when the condition is a comparison of a bare column with a constant. What the reader understands is a simple shape such as "qty is 5 or more". If you wrap the column in a function, as in upper(status) = 'PAID', it is no longer a shape that can be handed to the reader, and the condition drops out of PushedFilters. The result is the same, and there is no error or warning. Only the amount read grows. In this lab you check exactly this difference in the plan.
Partition columns are the same. If you put day = '2026-01-03' on a date partition, it goes into PartitionFilters, but if you filter by processing the date string, pruning may not happen. This is where the habit comes from: put filter conditions on the column in the shape it is stored.
AQE: a plan that changes during execution
According to the performance tuning documentation, adaptive query execution (AQE) is a technique that re-optimizes the plan during execution using runtime statistics, and it has been on by default since Spark 3.2.0. This has an important consequence for reading plans: an explain printed before an action is not the final plan.
If you print it before an action, you see AdaptiveSparkPlan isFinalPlan=false at the top. Only after a shuffle stage has actually finished and its size is known does AQE merge shuffle partitions or change the join strategy. If you call an action once and look at the plan of the same DataFrame again, nodes such as AQEShuffleRead appear together with isFinalPlan=true. To check the plan that was executed, you have to look at this final plan. In the SQL tab of the web UI documentation too, you can expand the four plans for each query with Details and see metrics such as how many rows passed through each operator.
What it looks like in the field
"I put a filter on, so why does it read everything?" The most common causes are a condition wrapped in a function, a comparison between different types (a string column and a numeric constant), and reading a CSV as it is. All three show up in the PushedFilters and ReadSchema lines of the plan.
"I do the select later, so it doesn't matter" is mostly right. This is because the optimizer keeps only the columns that are used to the end. However, if you cache() in the middle or pass the entire row to a Python UDF, the columns needed at that point increase. Always check with ReadSchema.
Without statistics, plans become unstable. The section on leveraging statistics says that when statistics are missing or inaccurate, Spark cannot choose a good plan, and it advises looking at estimates with explain(mode="cost") and at statistics during execution with isRuntime=true in the SQL UI.
What really matters in practice
- Do not trust optimization; verify it with explain. No warning appears when it fails.
- The single line of the read node is the key. Look first at the three fields PushedFilters, ReadSchema and PartitionFilters.
- Put conditions on the column in the shape it is stored. If you wrap it in a function, pushdown silently disappears.
- A condition on a partition column skips directories. It is the cheapest pruning, so make a frequently filtered column the partition key.
- Look at AQE's final plan after the action. A plan with isFinalPlan=false is a trailer.
What you will do in the next lab
You convert the order CSV to Parquet with a date column added, and then print the extended and formatted plans of explain. You confirm how a condition and column selection appear in PushedFilters and ReadSchema, and compare how pushdown disappears when you wrap the same condition in upper(). From the Parquet written split by date, you pick one day and see the condition appear in PartitionFilters, then run a per-customer aggregation as an action, read the AQE final plan and the event log, and count how many of the 200 shuffle partitions had tasks that actually read the shuffle. Finally, you confirm in the optimized logical plan that consecutive filters are merged into one and that a constant expression such as 1 + 2 is precomputed.