Prepare for PySpark and Apache Spark interview questions grouped by experience level.
PySpark Interview Question & Answers
0-2 Years
Apache Spark is a genuinely distributed computing framework for actually processing large volumes of data across a cluster of machines, built to solve the genuine problem of processing data at a scale too large for a single machine to actually handle efficiently on its own.
Spark genuinely processes data primarily in memory, making it meaningfully faster for many workloads, especially anything genuinely iterative, like machine learning. MapReduce genuinely writes intermediate results to disk between steps, which is genuinely more reliable for a very large, single-pass job but genuinely slower for anything requiring multiple passes over the exact same data.
PySpark is the genuine Python API for Apache Spark, letting a Python developer actually write Spark applications using genuinely familiar Python syntax, while the actual underlying computation still runs on Spark's own core engine, which is genuinely written in Scala.
The driver program genuinely coordinates the overall application and creates the SparkContext. Executors genuinely run on worker nodes, actually performing the real computation and storing data. The cluster manager, like YARN or Kubernetes, genuinely allocates resources across the cluster for the driver and executors to actually use.
A SparkSession is the genuine entry point for actually working with Spark, providing access to DataFrame and SQL functionality. Every genuine PySpark application starts by creating one, and it's used throughout the application to actually read data, run queries, and coordinate execution.
PySpark lets a team genuinely build on existing Python expertise and its genuinely rich ecosystem of data science libraries, at a small, genuine performance cost compared to Scala for certain operations, since Python code genuinely has to communicate with Spark's underlying JVM through a genuine bridge.
An RDD is Spark's genuinely original core data structure, representing an genuinely immutable, distributed collection of objects that can actually be processed in parallel across a cluster, with Spark automatically genuinely handling how that data is actually partitioned and distributed.
Once an RDD is genuinely created, its actual data can't be modified directly. Instead, a genuine transformation applied to an RDD produces a genuinely new RDD, leaving the original completely unchanged, similar in spirit to how an immutable value works in a genuinely functional programming language.
Resilient refers to an RDD's genuine ability to automatically recover from a lost partition, since Spark genuinely tracks the sequence of transformations, called lineage, used to actually create that RDD, letting it genuinely recompute a lost partition from that lineage rather than needing to actually restart the entire job.
A transformation, like map() or filter(), genuinely produces a new RDD from an existing one, but doesn't actually trigger any real computation immediately. An action, like count() or collect(), genuinely triggers Spark to actually execute the accumulated transformations and produce a real, concrete result.
Lazy evaluation means Spark genuinely doesn't actually execute a transformation the moment it's called, but instead genuinely waits until an action is actually triggered. It matters because it lets Spark genuinely optimize the entire chain of transformations together before actually running anything, rather than executing each step in isolation.
sc.parallelize([1, 2, 3, 4, 5]) genuinely creates an RDD from the given Python list, distributing it across the cluster's own available partitions, letting Spark actually process it in parallel.
A DataFrame organizes data into genuinely named columns, similar to a table in a relational database, and it's genuinely built on top of RDDs internally. Unlike a raw RDD, a DataFrame lets Spark's own Catalyst optimizer genuinely understand the actual structure of the data, enabling genuinely significant automatic performance optimization.
spark.createDataFrame([(1, 'Anu'), (2, 'Raj')], ['id', 'name']) genuinely creates a DataFrame with the specified column names, letting you actually work with that data using DataFrame's own genuinely structured API.
spark.read.csv('file.csv', header=True, inferSchema=True) genuinely reads the file into a DataFrame, using the first row as genuine column headers and automatically inferring each genuine column's data type from the actual file contents.
df.select('name', 'age') genuinely returns a new DataFrame containing only the specified columns, letting you actually work with just the genuinely relevant subset of a larger dataset's own columns.
df.filter(df.age > 18) genuinely returns a new DataFrame containing only the rows where the age column's value actually satisfies that specific condition, similar in spirit to a WHERE clause in SQL.
df.show() genuinely prints a formatted, readable preview of the DataFrame's own first 20 rows by default, and df.show(n) genuinely lets you specify a different, genuinely specific number of rows to actually display instead.
The DataFrame API lets you actually manipulate data using genuine Python method calls, like df.filter() and df.select(). Spark SQL lets you actually write standard SQL queries directly against a registered DataFrame, and both genuinely compile down to the exact same underlying execution plan.
df.createOrReplaceTempView('people') genuinely registers the DataFrame under that name, letting you actually run spark.sql('SELECT * FROM people') afterward to actually query it using standard SQL syntax.
A schema defines the genuine column names and data types for a DataFrame. It matters because Spark genuinely uses it to actually validate data and optimize query execution, and an incorrect or genuinely mismatched schema can cause an actual error or an unexpected, silently incorrect result during processing.
df.withColumn('age_plus_one', df.age + 1) genuinely returns a new DataFrame with an genuinely additional column computed from the existing age column, without modifying the genuinely original DataFrame at all, since DataFrames themselves are genuinely immutable.
df.groupBy('department').agg({'salary': 'avg'}) genuinely groups rows by the department column and computes the genuine average salary within each group, similar in spirit to a GROUP BY clause combined with an aggregate function in SQL.
df.count() genuinely triggers an action, returning the actual total number of rows in the DataFrame as a single integer. df.show() genuinely triggers a different action, printing a formatted preview of the DataFrame's own actual row contents rather than just a count.
map() genuinely applies a function to every element. filter() genuinely keeps only elements satisfying a condition. groupBy() genuinely groups data by a key. join() genuinely combines two datasets based on a genuinely matching key. None of these genuinely trigger actual computation on their own.
collect() genuinely retrieves every element back to the driver program. count() genuinely returns the total number of elements. take(n) genuinely returns the first n elements. Each of these genuinely triggers Spark to actually execute the accumulated transformations.
collect() genuinely brings every single row back to the driver program's own memory, and if the dataset is genuinely too large to fit in the driver's available memory, this can genuinely cause the driver to actually run out of memory and crash, an especially real risk when working with a genuinely large, production-scale dataset.
Spark genuinely builds a DAG representing the sequence of transformations leading to a genuine action, capturing the actual dependencies between different stages of computation. This DAG lets Spark's own optimizer genuinely plan the most efficient way to actually execute the entire chain of operations together.
Seeing the entire chain of transformations at once, before any of them actually run, lets Spark's optimizer combine adjacent steps, reorder operations for efficiency, and decide the best overall execution strategy for the whole chain together, rather than committing to a suboptimal choice for one step in isolation without any visibility into what comes after it.
A narrow transformation, like map() or filter(), genuinely operates on data within a single partition, with no data actually needing to move between partitions. A wide transformation, like groupBy() or a join, genuinely requires shuffling data across partitions, which is a meaningfully more expensive operation.
A wide transformation genuinely triggers a shuffle, moving data across the network between executors, which is genuinely one of the most expensive operations in Spark. Minimizing genuinely unnecessary wide transformations, or restructuring logic to actually reduce how much data needs to shuffle, meaningfully improves overall job performance.
df.orderBy('age') genuinely sorts the DataFrame's rows by the age column in ascending order by default, and df.orderBy(df.age.desc()) genuinely sorts in descending order instead.
df.withColumnRenamed('old_name', 'new_name') genuinely returns a new DataFrame with the specified column renamed, without modifying the genuinely original DataFrame at all, since DataFrames are genuinely immutable.
df.drop('column_name') genuinely returns a new DataFrame with the specified column removed entirely, leaving every other genuine column in the DataFrame unchanged.
df.select('column_name').distinct().count() genuinely returns the number of distinct values found in that specific column, first genuinely narrowing to just that column, then removing duplicates, then counting what actually remains.
df.printSchema() genuinely prints a readable, tree-formatted view of the DataFrame's own column names and types. df.dtypes genuinely returns that exact same information as a plain Python list of tuples instead, useful when you actually need to programmatically inspect the schema rather than just visually reading it.
3-6 Years
Persistence, using cache() or persist(), genuinely keeps an RDD's computed data in memory (or on disk) after it's actually first computed, so a subsequent action reusing that exact same RDD doesn't have to genuinely recompute it entirely from scratch through its own lineage again.
cache() is genuinely a shorthand for persist() using the default storage level, MEMORY_ONLY. persist() lets you actually specify a genuinely different storage level explicitly, like MEMORY_AND_DISK, giving more genuinely fine-grained control over exactly how and where the data should actually be stored.
Partitioning divides a genuine dataset into smaller chunks distributed across the cluster, letting Spark process them in parallel. Its genuine impact on performance matters because too few partitions genuinely underutilize available cluster resources, while too many can genuinely introduce excessive overhead managing all those genuinely small tasks.
RDDs are genuinely useful when you need low-level control over data that doesn't fit cleanly into a structured, tabular format, like genuinely unstructured or complex nested data, or when implementing a genuinely custom algorithm that doesn't map naturally onto DataFrame's own higher-level API.
repartition() genuinely reshuffles data across a specified number of partitions, which can genuinely increase or decrease the partition count, but always triggers a full shuffle. coalesce() genuinely reduces the number of partitions without a full shuffle, making it genuinely more efficient specifically when only decreasing the partition count.
Inner join genuinely keeps only rows matching in both DataFrames. Left outer join genuinely keeps every row from the left DataFrame, filling in null for genuinely non-matching rows from the right. Right outer and full outer joins genuinely extend that same logic in their respective corresponding directions.
A broadcast join sends a genuinely small DataFrame's complete data to every executor, avoiding the genuine need to shuffle the much larger DataFrame across the network to actually perform the join. It solves the genuine problem of an expensive shuffle join being unnecessary when one side of the join is genuinely small enough to fit comfortably in memory.
A window function computes a genuine value across a set of rows related to the current row, without collapsing them into a single output row the way a plain aggregation would. A genuinely practical use case is calculating each employee's genuine rank within their own department by salary.
df.na.drop() genuinely removes rows containing a null value. df.na.fill(0) genuinely replaces a null value with a specified default. df.filter(df.column.isNotNull()) genuinely filters based on whether a genuinely specific column's value is actually null or not.
The explain() output genuinely shows the actual physical execution plan Spark chose, and it will explicitly genuinely show a BroadcastHashJoin or a genuine SortMergeJoin (a common shuffle-based join strategy) in its plan, letting you actually verify which specific join strategy Spark actually selected for a given query.
Spark SQL lets you actually query structured data using standard SQL syntax, and it genuinely benefits from Spark's own Catalyst optimizer, which can automatically genuinely optimize a query's execution plan in a way the more genuinely low-level RDD API doesn't automatically provide.
A temporary view, created with createOrReplaceTempView(), is genuinely scoped to the current SparkSession and disappears once that session ends. A global temporary view, created with createGlobalTempView(), genuinely persists across multiple SparkSessions within the exact same Spark application.
spark.sql('SELECT department, AVG(salary) FROM employees GROUP BY department') genuinely runs that SQL query against the registered employees view, returning the genuine result as a new DataFrame.
A UDF lets you genuinely define custom logic in Python (or Scala) and apply it as if it were genuinely a built-in SQL function, needed when the genuinely required transformation isn't already covered by one of Spark's own built-in functions.
A Python UDF genuinely requires data to be serialized and passed between the JVM (where Spark's own engine runs) and a genuinely separate Python process for every single row, introducing real, genuine overhead that a built-in function, executed entirely within the JVM, genuinely avoids altogether.
The driver genuinely runs the main application code, builds the execution plan, and coordinates the overall job. Executors genuinely run on worker nodes, actually performing the real computation on their own genuine assigned data partitions and reporting results back to the driver.
A cluster manager genuinely allocates resources, CPU and memory, across the cluster for Spark's driver and executors to actually use. Common options include YARN, genuinely common in Hadoop-based environments, Kubernetes, and Spark's own built-in standalone cluster manager.
A job corresponds to one genuine action triggered in the application. Spark genuinely breaks a job into several stages, separated by a genuine shuffle boundary. Each stage is further genuinely broken into tasks, with each task genuinely processing one single partition of data on one specific executor.
SparkContext is the genuinely original entry point for interacting with a Spark cluster's own core RDD functionality. SparkSession, introduced later, genuinely wraps SparkContext and provides a genuinely unified entry point also covering DataFrame and SQL functionality, and it's what a modern PySpark application typically genuinely uses directly.
Without caching, Spark genuinely recomputes a DataFrame's entire lineage every single time it's actually accessed by a genuinely separate action. Caching genuinely stores the computed result after the first access, letting every subsequent access reuse it directly rather than genuinely paying that same computation cost repeatedly.
A broadcast variable genuinely sends a read-only value to every executor once, efficiently, rather than genuinely shipping a fresh copy of that value along with every single individual task that actually needs it, which would genuinely waste network bandwidth for a value used repeatedly across many tasks.
Data skew happens when data is genuinely unevenly distributed across partitions, causing one or a genuinely small number of tasks to process dramatically more data than the others. It hurts performance because the overall job genuinely can't finish until that specific, genuinely overloaded task actually completes, leaving other executors sitting comparatively idle in the meantime.
The Spark UI provides a genuine web-based dashboard showing details about a running (or completed) Spark application, including its stages, tasks, and their actual execution time, letting you actually identify a genuinely slow stage or an unevenly distributed task.
6-8 Years
The Spark UI's stage view genuinely shows the actual task duration distribution within a stage, and a genuinely wide spread, where most tasks finish quickly but a handful take dramatically longer, is a genuine, telltale sign of data skew affecting a specific subset of partitions.
Salting the genuinely skewed key, appending a genuine random value to spread its records across more partitions before actually performing the aggregation or join, then combining the results afterward, is a genuinely common technique to actually reduce the impact of that specific data skew.
spark.sql.shuffle.partitions genuinely controls the number of partitions used during a shuffle operation. spark.executor.memory and spark.executor.cores genuinely control the resources allocated to each executor, all genuinely tunable based on the actual specific workload's own characteristics.
A genuinely common guideline targets roughly 100 to 200 megabytes of data per partition, and roughly two to three times the total number of available CPU cores across the cluster, though the actual, genuinely right number depends on the specific workload's own characteristics and needs to be genuinely tuned empirically.
Predicate pushdown lets Spark genuinely apply a filter condition as early as possible, ideally directly at the actual data source itself, like a Parquet file's own metadata, rather than reading every single row into memory first and only genuinely filtering afterward, meaningfully reducing the amount of data actually read and processed.
Parquet genuinely stores data column by column rather than row by row, letting Spark read only the specific columns a genuine query actually needs, and it also genuinely supports built-in compression and predicate pushdown, both of which meaningfully improve read performance compared to a genuinely plain, row-oriented CSV file.
Catalyst is Spark SQL's genuine query optimization engine, analyzing a query's logical plan and automatically applying genuine optimizations, like predicate pushdown and reordering a join, before actually generating the final, optimized physical execution plan.
Tungsten is Spark's genuine execution engine optimization, improving memory management and CPU efficiency by working with genuinely raw binary data directly rather than through Java's own standard object representation, meaningfully reducing memory overhead and improving overall execution speed.
A regular Python UDF genuinely processes one row at a time, incurring real serialization overhead for every single row. A Pandas UDF genuinely processes an entire batch of rows at once, using Apache Arrow for genuinely efficient data transfer between the JVM and Python, making it meaningfully faster than a genuinely row-by-row UDF.
Reduce the amount of data genuinely needing to shuffle, filtering early, using a broadcast join for a genuinely small table, and tuning the number of shuffle partitions to genuinely match the actual data volume, rather than leaving it at Spark's own generic default value.
AQE lets Spark genuinely re-optimize a query's execution plan at runtime, based on actual, real statistics gathered during execution, rather than relying purely on estimates made before actually running the query. It solves the genuine problem of an initial plan turning out to be genuinely suboptimal once real data characteristics, like an unexpected data skew, actually become apparent.
Dynamic partition pruning lets Spark skip reading a partition that a join condition rules out at runtime, based on the actual values found in the smaller side of the join, rather than reading every partition and filtering afterward. It works alongside Adaptive Query Execution's broader ability to adjust a query's plan based on real, actual runtime information rather than only relying on estimates made in advance.
8-10 Years
Structured Streaming lets you write a genuinely streaming computation using the exact same DataFrame API used for batch processing, treating an incoming stream as a genuinely continuously growing, unbounded table. This means the exact same code and mental model genuinely apply to both batch and streaming workloads.
Structured Streaming genuinely processes incoming data in small, discrete batches at a defined interval, rather than truly processing each genuinely individual record the instant it actually arrives, striking a practical genuine balance between low latency and processing efficiency.
Watermarking defines how genuinely late an event is allowed to arrive and still actually be included in a windowed aggregation. It solves the genuine problem of unboundedly retaining state for every possible late-arriving event forever, letting Spark genuinely discard state for a window once it's confident no further, relevant late data will actually arrive.
Append mode genuinely outputs only new rows added since the previous batch. Update mode genuinely outputs any row that's changed. Complete mode genuinely outputs the entire, complete result table every single time, which is only genuinely practical for a bounded aggregation, like a total count, rather than a genuinely unbounded result set.
Exactly-once semantics guarantees each genuine record is processed and reflected in the output exactly one time, even if a genuine failure and retry occurs partway through processing. It matters because a system genuinely relying on that data, like a financial reporting pipeline, can't tolerate either a genuinely missed or a duplicated record.
A checkpoint saves the streaming query's current processing progress and internal state to durable storage. If the streaming application genuinely fails and restarts, it reads that checkpoint to actually resume from where it left off, rather than reprocessing everything from the beginning or genuinely losing track of what had already been processed.
spark.readStream.format('kafka').option('subscribe', 'topic_name').load() genuinely reads a continuous stream of messages directly from the specified Kafka topic, letting you actually apply the exact same DataFrame transformations you'd use for batch processing to that genuinely continuous stream of incoming data.
DStreams genuinely represent a stream as a sequence of small, discrete RDDs, requiring you to actually think in terms of that lower-level RDD API. Structured Streaming genuinely, instead, uses the higher-level DataFrame API throughout, and is now the genuinely recommended, actively-developed approach for new streaming applications in Spark.
In cluster mode, the driver genuinely runs on one of the cluster's own worker nodes. In client mode, the driver genuinely runs on the machine that actually submitted the job, outside the cluster itself, which is often genuinely more convenient for interactive development but genuinely less resilient for a long-running production job.
Spark's own native Kubernetes support lets you submit a job directly using spark-submit with a Kubernetes master URL, and Kubernetes then genuinely schedules the driver and executor pods, letting Spark genuinely run within an existing, genuinely container-orchestrated infrastructure rather than requiring a genuinely dedicated, separate cluster manager like YARN.
Balance the total genuine number of executors and cores per executor against the cluster's own actual total available resources, and against other jobs genuinely sharing that same cluster, avoiding over-allocating so much that a single job genuinely starves every other job also needing to run concurrently.
Dynamic allocation lets Spark genuinely adjust the number of executors a running application actually uses, based on the current genuine workload, scaling up during a heavy processing phase and scaling back down during a genuinely lighter one, rather than requiring a genuinely fixed, statically-sized allocation for the entire job's duration.
A tool like Ganglia or a genuinely integrated cloud monitoring service tracks cluster-wide resource utilization over time, and integrating Spark's own metrics with a centralized monitoring system, like Prometheus, lets you actually track trends and set up alerting across genuinely many jobs, beyond just inspecting one at a time.
I'd weigh the actual, genuine data volume, whether it genuinely exceeds what a single machine can comfortably hold and process in memory, against the real, added operational complexity Spark introduces. A genuinely small dataset that comfortably fits on one machine often doesn't justify Spark's own real distributed-systems overhead.
10+ Years
I'd weigh the actual, genuine data volume and processing complexity against the real operational overhead Spark introduces, cluster management, tuning, monitoring. A genuinely modest data volume that comfortably fits a simpler tool often doesn't justify Spark's own real, additional complexity.
Migrate incrementally, starting with the pipelines that would genuinely benefit most, or that are genuinely easiest to migrate first, validating output against the genuinely existing pipeline's own results carefully before actually cutting over the downstream consumers relying on that specific pipeline.
I check whether it genuinely avoids an unnecessary wide transformation, whether partitioning is genuinely sensible given the actual expected data volume, and whether it accounts for genuine data skew rather than assuming perfectly even distribution across every single key.
Enforce genuine resource quotas at the cluster manager level so one team's job can't genuinely starve another's, and require genuinely standard logging and monitoring integration as part of a shared job submission template, rather than leaving that up to each individual team's own inconsistent practice.
I'd weigh the genuine time saved on cluster management, tuning, and upgrades that a managed service provides against the real, ongoing cost premium it typically carries. For most organizations, a managed service genuinely wins unless there's a genuinely specific, compelling reason to self-manage instead.
I'd check whether production data volume genuinely differs meaningfully from the development test data, since a genuinely small test dataset can hide a real performance issue, like a wide transformation or data skew, that only actually surfaces once the real, much larger production data volume is genuinely involved.
Track job duration, executor resource utilization, and shuffle read/write volume over time, alerting on meaningful deviation from an established baseline. A slowly, steadily growing job duration trend is often a genuine early warning sign well before it actually causes a real, missed processing deadline.
Treat the output schema as a genuine contract with every downstream consumer. Adding a genuinely new, additional column is generally safe. Changing or removing an existing column needs a documented migration plan and direct communication with every team genuinely consuming that output before actual removal.
I'd check whether the actual input data volume has genuinely grown recently, since a job that fit comfortably in memory before can genuinely exceed available memory as its input naturally grows over time, and I'd also check for a genuinely new data skew pattern that wasn't present in earlier, smaller runs.
I'd load test with a genuinely realistic, larger dataset volume, monitoring actual resource utilization under that load, and identify whether the coming bottleneck is genuinely likely to be memory, CPU, or shuffle I/O, since each of those calls for a meaningfully different scaling response.
This is a judgment question interviewers use to see how you reason under genuine uncertainty, not to test a specific textbook fact. A strong answer names the actual constraint that forced the decision, the realistic options that were genuinely on the table, why you picked one knowing it wasn't guaranteed to be right, and what you'd do differently with what you know now.
I'd walk through one of their actual jobs together against a realistically large dataset, showing the Spark UI's own stage timing directly and pointing out concretely where the real cost is genuinely coming from, rather than just telling them the code is inefficient in the abstract.
I wouldn't lead with the risk in the abstract. I'd point to a specific, real, already-experienced out-of-memory driver crash caused by exactly that pattern, and show concretely how an alternative, like take(n) or writing output directly to storage instead, would have genuinely avoided that exact same specific problem.
I'd bring the actual, concrete data volume and expected growth trajectory into the discussion, rather than a general, abstract preference for one tool over the other. Most disagreements like this genuinely resolve once both sides are looking at the exact same concrete numbers together.
I'd translate the work into terms leadership already tracks: the compute cost of a job running longer than it needs to, and the delay that imposes on any downstream report or product depending on its output. Framed as cost and delivery-time reduction with a concrete number attached, it competes far better for prioritization than framed as a purely technical improvement.




