Prepare for Apache Kafka interview questions grouped by experience level.
Kafka Interview Question & Answers
0-2 Years
Kafka is a distributed event streaming platform, letting different systems actually publish and subscribe to a genuinely continuous stream of data reliably and at a large scale. It was built to solve the genuine problem of tightly coupled, point-to-point integration between many different systems becoming unmanageable as an organization's own number of systems grows.
In pub-sub, a genuine producer publishes a message without knowing who will actually consume it, and a genuine consumer subscribes to receive messages without knowing who actually produced them. Kafka genuinely implements this through topics, where producers publish messages and consumers subscribe to actually read them.
A traditional message queue typically genuinely removes a message once it's actually consumed. Kafka instead genuinely retains messages for a configured period, letting genuinely multiple, independent consumers each read the exact same message at their own pace, and even genuinely re-read older messages later if actually needed.
Real-time data pipelines feeding data between genuinely different systems. Event-driven microservices architecture, where a service publishes an event other services actually react to. Log aggregation, collecting genuine log data from many different sources into one centralized place.
Kafka genuinely runs as a cluster of multiple servers, called brokers, working together, distributing data and load across them rather than relying on a single, genuinely centralized server, which improves both scalability and fault tolerance.
An event is a genuine, single piece of data representing something that happened, like an order being placed, published to a Kafka topic. Kafka genuinely treats each event as an immutable record, once written, it's genuinely never modified.
A topic is a genuine named category or feed to which a producer actually publishes messages and from which a consumer actually reads them, conceptually genuinely similar to a table in a database or a folder organizing genuinely related files.
A partition is a genuine, ordered subset of a topic's own data, and a topic is genuinely split into multiple partitions to actually let Kafka distribute that topic's data across multiple brokers, enabling genuinely parallel reads and writes that a single, undivided topic couldn't otherwise support.
An offset is a genuinely unique, sequential identifier assigned to each message within a specific partition, letting a consumer actually track exactly which message it has read up to, and where to genuinely resume reading from next time.
Kafka genuinely guarantees messages within one single partition are actually written and read in the exact same order they were originally produced. This guarantee genuinely doesn't extend across different partitions of the exact same topic, only within one, genuinely individual partition.
If the message genuinely has a key, Kafka uses a hash of that key to genuinely, consistently determine which partition it goes to, ensuring every message with the exact same key always lands on the exact same partition. Without a key, Kafka genuinely distributes messages across partitions in a genuinely round-robin fashion.
Retention determines how long Kafka genuinely keeps a message before actually deleting it, configured either by a genuine time period or a maximum total size. It lets Kafka act as a genuinely durable, replayable log for a defined window, rather than only genuinely, briefly holding a message until it's consumed.
Unlike a traditional queue where consuming a message typically removes it, Kafka genuinely leaves the message in the topic exactly as it was, governed only by the topic's own retention policy. This means another consumer, or even the same consumer reading again from an earlier offset, can genuinely still read that exact same message later.
A producer is a genuine client application that actually publishes (writes) a message to a specific Kafka topic, and Kafka then genuinely handles distributing and storing that message across the appropriate partition and broker.
A message key lets a producer actually control which partition a message is written to, since Kafka genuinely hashes the key to consistently route every message sharing the exact same key to the exact same partition, which matters when a genuine, specific message order needs to be preserved for related messages.
Without a key, Kafka genuinely distributes messages across the topic's own available partitions in a genuinely round-robin fashion, spreading load evenly, though without any genuine guaranteed ordering relationship between any two specific messages.
acks controls how many genuine broker replicas must actually confirm receiving a message before the producer considers it genuinely successfully sent. A genuinely higher acks setting provides stronger durability guarantees at the cost of slightly higher genuine latency.
Serialization converts a genuine in-memory object, like a Java object or a Python dictionary, into a genuine byte format Kafka can actually store and transmit, since Kafka itself genuinely works with raw bytes rather than a genuinely specific programming language's own native data types.
A producer can genuinely group several messages together into one single batch before actually sending them to a broker, reducing the genuine number of separate network requests needed, which meaningfully improves overall throughput compared to sending every genuinely single message as its own, separate request.
A consumer is a genuine client application that actually subscribes to one or more Kafka topics and reads (consumes) the messages published to them, tracking its own genuine progress through each partition using an offset.
A consumer group is a genuine set of consumers working together to actually read from a topic, with each individual partition genuinely consumed by only one consumer within that group at a time. It solves the genuine problem of letting multiple consumer instances share the total workload of reading a genuinely large topic in parallel.
Some consumers will genuinely be assigned more than one partition each, since Kafka distributes the topic's own total partitions across whatever number of consumers are genuinely currently active within that specific consumer group.
The genuinely excess consumers, beyond the actual number of partitions, will simply sit idle, receiving genuinely no partitions to actually consume from at all, since Kafka can only assign a specific partition to genuinely one consumer at a time within that group.
Committing an offset genuinely records how far a consumer has actually progressed through a partition, so if that consumer genuinely restarts or crashes, it knows exactly where to actually resume reading from, rather than genuinely reprocessing every message from the very beginning again.
A consumer configured to genuinely start from the earliest offset reads every message still retained in the topic, including genuinely older ones. A consumer configured to start from the latest offset genuinely only receives messages published after it actually started, missing anything published genuinely earlier.
A broker is a genuinely single server within a Kafka cluster, responsible for actually storing data and serving read and write requests from producers and consumers. A genuine Kafka cluster typically consists of several brokers working together.
A cluster is a genuinely group of brokers working together, distributing data and load across them. Running as a cluster lets Kafka genuinely scale beyond what a single server could handle, and provides genuine fault tolerance, since the cluster can keep functioning even if one specific broker actually fails.
The controller is the genuine broker responsible for managing cluster-wide administrative tasks, like tracking which broker is the leader for each partition and detecting when a broker joins or leaves the cluster. Only one broker in the cluster genuinely acts as the controller at any given time, and if it fails, another broker is genuinely elected to take over that role.
ZooKeeper genuinely managed cluster metadata, like which broker is the genuine leader for a specific partition, and coordinated broker membership in the cluster, in the older, traditional Kafka architecture, before genuinely more recent versions introduced an alternative that removes this genuine external dependency.
Each partition genuinely has one broker acting as its leader, handling every read and write for that specific partition. Other brokers genuinely hold follower replicas, passively replicating the leader's own data, ready to genuinely take over if the leader broker actually fails.
Replication genuinely keeps multiple copies of a partition's own data across different brokers. It solves the genuine problem of data loss or unavailability if a single broker genuinely fails, since a genuine follower replica on a different, healthy broker can actually take over as the new leader.
The replication factor genuinely specifies how many total copies (the leader plus its followers) of a partition's own data Kafka actually maintains across the cluster, with a genuinely higher replication factor providing stronger fault tolerance at the cost of using more actual storage.
kafka-topics.sh --create --topic my-topic --bootstrap-server localhost:9092 --partitions 3 --replication-factor 2 genuinely creates a new topic with the specified name, number of partitions, and replication factor.
kafka-topics.sh --list --bootstrap-server localhost:9092 genuinely returns the names of every topic currently present in that specific cluster, letting you actually confirm a topic exists or see what's already been created.
kafka-console-producer.sh --topic my-topic --bootstrap-server localhost:9092 genuinely opens an interactive prompt where any text you actually type and press enter on is immediately published as a genuine message to that specific topic.
kafka-console-consumer.sh --topic my-topic --bootstrap-server localhost:9092 --from-beginning genuinely reads and prints every message currently retained in that topic, starting from its genuinely earliest available offset.
It genuinely specifies the address of at least one broker in the cluster the tool should actually connect to, which the client then uses to actually discover the rest of the cluster's own brokers and metadata, rather than needing every single broker's address listed explicitly.
3-6 Years
acks=0 genuinely doesn't wait for any confirmation at all, fastest but genuinely least durable. acks=1 genuinely waits for just the partition leader to confirm. acks=all (or -1) genuinely waits for every in-sync replica to confirm, providing the genuinely strongest durability guarantee at the cost of higher latency.
An idempotent producer genuinely ensures a message isn't accidentally written more than once, even if the producer genuinely retries sending it after a network issue, by having Kafka genuinely track and deduplicate a sequence number per producer, avoiding the genuine risk of a duplicate message from a simple retry.
linger.ms genuinely controls how long a producer waits to actually accumulate more messages into a batch before sending it, even if the batch isn't genuinely full yet. A genuinely higher value improves throughput through larger batches, at the cost of slightly higher genuine per-message latency.
A producer can genuinely be configured to automatically retry sending a message that genuinely failed due to a transient issue, like a temporary network blip, rather than the message simply genuinely being lost. Combined with idempotency, this provides genuinely strong delivery reliability without risking a duplicate.
It genuinely controls how many unacknowledged requests a producer can have outstanding at once. If set genuinely too high alongside retries enabled, a genuine retried message could arrive out of order relative to a genuinely later message, which is exactly why Kafka's idempotent producer feature genuinely handles this specific ordering concern automatically.
Automatic commit genuinely, periodically commits the consumer's own current offset in the background at a defined interval. Manual commit gives the application genuinely explicit control over exactly when an offset is actually committed, typically done only after a message has genuinely been fully, successfully processed.
If the offset was genuinely already automatically committed before the message finished processing, and the consumer then crashes, that specific message could genuinely be lost, since a restarted consumer would resume from the genuinely already-committed offset, skipping the message that was never actually fully processed.
Rebalancing genuinely redistributes a topic's own partitions among the currently active consumers within a consumer group, occurring whenever a genuine consumer joins or leaves the group, like when a consumer crashes or a genuinely new instance is added to help share the load.
Frequent rebalancing genuinely, temporarily pauses processing across the entire consumer group while partitions are actually reassigned, and if it happens genuinely too often, overall consumer throughput can meaningfully suffer, since the group spends real time rebalancing rather than actually processing messages.
At-most-once genuinely risks losing a message but never genuinely processes a duplicate. At-least-once genuinely guarantees a message is never lost but might genuinely be processed more than once. Exactly-once genuinely guarantees a message is both never lost and never genuinely duplicated, the strongest but most complex guarantee to actually achieve.
An ISR is a genuine follower replica that's actually kept sufficiently up to date with the partition leader, within a configured time threshold. Only a replica genuinely considered in-sync is eligible to actually be elected as the new leader if the current leader fails.
Kafka genuinely elects a new leader from among that partition's own current in-sync replicas, letting the partition continue actually accepting reads and writes with genuinely minimal disruption, rather than the entire partition becoming unavailable until the original, failed broker actually recovers.
min.insync.replicas genuinely specifies the minimum number of in-sync replicas that must actually acknowledge a write for it to genuinely succeed when a producer uses acks=all. If genuinely fewer replicas than that minimum are currently in-sync, the write is genuinely rejected rather than silently accepted with weaker actual durability.
Unclean leader election genuinely allows a replica that isn't currently fully in-sync to actually become the new leader if genuinely no in-sync replica is available at all. It trades genuine data consistency, since that replica might be genuinely missing some recent messages, for continued availability rather than the partition becoming completely unavailable.
Each genuinely additional replica means one more complete, actual copy of that partition's own data stored on a genuinely separate broker, increasing storage and network usage for the ongoing replication traffic, in exchange for the cluster being able to genuinely tolerate that many more broker failures without any actual data loss.
A serializer converts a genuine in-memory object into a byte array Kafka can actually store and transmit, needed because Kafka itself genuinely operates on raw bytes and has genuinely no built-in understanding of a specific programming language's own native object types.
A String serializer genuinely converts plain text into bytes with no defined structure beyond that raw text. An Avro serializer genuinely encodes data according to a defined schema, providing genuine structure, type safety, and efficient compact encoding that plain string serialization genuinely doesn't offer on its own.
A Schema Registry stores and manages the genuine schemas used to serialize and deserialize messages, like Avro schemas, letting a producer and consumer agree on a genuinely consistent, versioned data structure, and letting a schema genuinely evolve over time in a controlled, backward-compatible way.
It provides a genuine, enforced contract for a message's own structure, catching a genuinely incompatible schema change early, before it actually breaks a consumer relying on the older, expected structure, which becomes genuinely critical once many independent teams are producing and consuming the exact same shared topics.
Kafka Streams is a genuine Java library for actually building applications that process data directly from Kafka topics in real time, transforming, aggregating, or joining streams of data, without needing a genuinely separate, external stream processing cluster.
A stream represents a genuinely unbounded sequence of individual events. A table represents the genuine current, latest state derived from that stream, conceptually similar to a database table being genuinely continuously updated as new events actually arrive.
A stateless operation, like filtering or mapping a single record, genuinely processes each record independently with no memory of any prior record. A stateful operation, like an aggregation or a join, genuinely needs to remember information across multiple records to actually produce its correct result.
Kafka Streams genuinely backs its internal state up to a dedicated, internal Kafka topic called a changelog topic, letting it actually recover that state correctly if the stream processing application itself genuinely crashes and needs to actually restart from where it left off.
6-8 Years
I'd weigh the genuine expected throughput needed and the genuine maximum number of consumers you'd ever want to actually run in parallel within a single consumer group, since the partition count genuinely sets a hard ceiling on how many consumers within one group can actually process that topic simultaneously.
Too many partitions genuinely increases the metadata overhead each broker has to actually manage, can genuinely slow down leader election during a failure, and increases the genuine number of open file handles and memory overhead required, even though each individual partition might genuinely carry very little actual data.
Checking consumer lag metrics directly reveals exactly how far behind each specific partition's consumption actually is. From there, I'd check whether the genuine bottleneck is the consumer's own processing logic being too slow, or genuinely insufficient consumer instances to actually keep pace with the topic's real incoming message rate.
Consumer lag measures the genuine difference between the latest message offset actually produced to a partition and the offset a specific consumer has actually processed up to. Monitoring it matters because a genuinely, continuously growing lag signals the consumer can't actually keep up with the incoming rate of messages.
Increasing batch.size and linger.ms lets the producer accumulate genuinely larger batches before sending, reducing the genuine number of separate network requests needed overall, meaningfully improving throughput at the cost of a slightly longer genuine delay before each individual message actually gets sent.
Compression, like gzip or snappy, genuinely reduces the actual size of data sent over the network and stored on disk, at the cost of some additional genuine CPU overhead required to actually compress and decompress that data on both the producer and consumer side.
Producer-side compression compresses a batch of messages before they're ever sent over the network, reducing bandwidth on the way to the broker as well as the broker's own storage footprint. If compression is only configured on the broker, the uncompressed data still travels over the network first, so producer-side compression is generally the more effective place to actually enable it.
Kafka Connect provides a genuinely standardized, configuration-driven way to actually move data between Kafka and an external system, like a database or a cloud storage service, without needing to actually write genuinely custom producer or consumer code for every single integration.
A source connector genuinely pulls data from an external system and publishes it into a Kafka topic. A sink connector genuinely reads data from a Kafka topic and writes it out into an external system, like a database or a data warehouse.
Standalone mode genuinely runs a single Connect worker process, appropriate for genuinely simple testing or a small deployment. Distributed mode genuinely runs multiple worker processes together as a cluster, providing genuine fault tolerance and the ability to actually scale a connector's own workload across multiple workers.
A connector genuinely splits its overall work into one or more tasks, and Kafka Connect distributes these genuine tasks across the available workers, letting a single connector's own configured data movement job actually run in parallel across multiple worker processes.
A genuinely well-designed connector, paired with a Schema Registry enforcing compatibility rules, handles a genuinely backward-compatible schema evolution automatically, though a genuinely breaking schema change usually requires updating the connector's own configuration or coordinating a planned migration with downstream consumers.
8-10 Years
Exactly-once processing guarantees a message is genuinely both never lost and never duplicated, even across a producer retry or a consumer failure and restart. Kafka genuinely achieves this through idempotent producers combined with transactions, atomically tying a message's own production and its offset commit together as one single, genuine unit.
A transaction lets an application genuinely, atomically write to multiple partitions (or topics) and commit its own consumer offset together as one single unit, ensuring that if any part of that genuine operation fails, none of it takes effect, avoiding a genuinely partial, inconsistent result.
KRaft is Kafka's genuinely newer, built-in consensus protocol for managing cluster metadata, removing the genuine need for a separate ZooKeeper cluster entirely. It solves the genuine problem of operating and scaling two genuinely separate distributed systems, ZooKeeper and Kafka itself, together as one cohesive unit.
Log compaction genuinely retains only the most recent message for each unique key within a topic, removing an older, genuinely superseded message with the exact same key, rather than deleting data purely based on its genuine age or the topic's overall total size.
A topic representing genuine current account balances, where only the genuinely most recent balance for each specific account actually matters, fits log compaction well, letting the topic act like a genuinely compact, continuously updated snapshot rather than retaining every single historical balance update forever.
Using a Schema Registry with a genuinely enforced backward-compatibility rule ensures a genuinely new schema version can still actually be read correctly by a consumer expecting an older version of that same schema, letting producers and consumers genuinely upgrade independently rather than needing a genuinely coordinated, simultaneous deployment.
Kafka Streams is genuinely embedded directly within your own application as a library, tightly coupled to Kafka itself. Flink runs as a genuinely separate, dedicated cluster and supports genuinely additional data sources beyond just Kafka, offering more flexibility for a genuinely complex, multi-source processing pipeline at the cost of additional operational complexity.
Track under-replicated partition count, consumer lag, broker disk usage, and request latency over time, alerting on meaningful deviation from an established baseline. A genuinely growing under-replicated partition count is often an early warning sign of a genuine broker health problem before it actually causes real data unavailability.
An under-replicated partition has genuinely fewer in-sync replicas than its own configured replication factor specifies, indicating a genuine follower broker has fallen behind or become unavailable. It matters because it genuinely reduces that partition's own actual fault tolerance until the replica catches back up.
MirrorMaker (or a genuinely similar cross-cluster replication tool) continuously replicates data from a genuinely primary cluster to a secondary cluster in a different datacenter, letting the organization actually fail over to that secondary cluster if the primary genuinely becomes unavailable.
Replication genuinely introduces some lag, meaning the secondary cluster's data is never genuinely perfectly, instantaneously in sync with the primary, and offset values on the replicated topic can genuinely differ between clusters, requiring careful, deliberate handling if consumers genuinely need to fail over seamlessly.
I'd load test with a genuinely realistic message volume and size, monitor broker CPU, disk I/O, and network throughput under that load, and add genuinely additional brokers (and rebalance existing partitions across them) before the coming increase actually starts degrading real, current performance.
Enabling TLS encryption genuinely protects data in transit between clients and brokers, and configuring ACLs (Access Control Lists) genuinely restricts which specific authenticated user or application is actually allowed to produce to or consume from a genuinely specific topic.
10+ Years
I'd weigh the actual, genuine need for high throughput, durable message replay, and multiple genuinely independent consumers reading the exact same data stream, against the real operational complexity Kafka introduces. A genuinely simple, low-volume messaging need often doesn't justify Kafka's own real, additional overhead.
I'd migrate incrementally, starting with the genuinely highest-value, most painful existing integration first, running the old and new integration genuinely in parallel during a transition period, and validating that the Kafka-based version genuinely produces the exact same, correct downstream result before actually retiring the older integration.
I check whether the genuine partitioning strategy fits the actual expected access pattern, whether the delivery semantics needed, at-least-once versus exactly-once, are genuinely correctly implemented, and whether the design accounts for a genuine consumer or broker failure rather than assuming everything simply always works.
Enforce genuinely hard requirements, like mandatory schema registration, through the platform's own tooling and access controls directly, rather than relying on manual, ad hoc review. For conventions that genuinely resist full automation, I'd document the handful of decisions that actually matter most.
I'd weigh the genuine time saved on cluster operations, upgrades, and scaling that a managed service provides against the real, ongoing cost premium it typically carries. For most organizations, a managed service genuinely wins unless there's a genuinely specific, compelling reason to self-manage instead.
I'd check whether the actual test load genuinely resembles production message volume and size, since a consumer that keeps up fine during a genuinely small-scale test can genuinely fall behind once real production traffic volume is genuinely involved, especially if the consumer's own processing logic includes a genuinely slow, external call.
Track consumer lag, broker health metrics, and end-to-end pipeline latency over time, alerting on meaningful deviation from an established baseline rather than relying only on a hard, static threshold. A genuinely, slowly growing consumer lag trend is often an early warning sign well before it actually causes a real, visible business impact.
Treat the schema as a genuine contract with every consuming team. Adding a genuinely new, optional field is generally safe. Removing or renaming an existing field needs a documented migration plan and direct communication with every team genuinely consuming that topic before actual removal, enforced through the Schema Registry's own compatibility rules.
I'd check whether the consumer application itself has actually crashed or is stuck, whether a genuinely recent deployment caused the issue, and consider a genuinely quick mitigation, like restarting the consumer or temporarily scaling it up, while properly investigating the actual, real root cause.
I'd load test with a genuinely realistic message volume and size, verify both broker-side capacity, disk, network, and consumer-side processing capacity can genuinely handle the expected peak, and identify whether the coming bottleneck is genuinely likely to be the broker cluster itself or the downstream consumer applications.
This is a judgment question interviewers use to see how you reason under genuine uncertainty, not to test a specific textbook fact. A strong answer names the actual constraint that forced the decision, the realistic options that were genuinely on the table, why you picked one knowing it wasn't guaranteed to be right, and what you'd do differently with what you know now.
I'd walk through an actual, real scenario together, showing concretely what genuinely happens to their specific application if a message is genuinely processed twice, rather than explaining delivery semantics as an abstract concept in isolation. Seeing the actual, concrete consequence tends to build that awareness far more effectively.
I wouldn't lead with event-driven architecture as an abstract best practice. I'd point to a specific, real, already-experienced incident caused by a tightly coupled, synchronous dependency, and show concretely how Kafka's own decoupled, asynchronous model would have genuinely prevented that exact same specific problem.
I'd bring the actual, concrete question of how genuinely different consumers actually need to access that data into the discussion, rather than a general, abstract preference for one topic structure over the other. Grounding the discussion in the specific, real consumption pattern resolves it faster than an abstract debate.
I'd translate the investment into terms leadership already tracks: the cost of a specific past incident traced back to a schema break or an undetected consumer lag spike, and the ongoing risk of a similar incident recurring. Framed as risk reduction with a concrete, already-incurred cost behind it, it competes far better for prioritization than framed as a general infrastructure improvement.




