Your AI Is Waiting on Your Data | Jun Rao (Co-Founder, Confluent)
Listen to full episode:
Summary: Hosts Anant and Ed sit down with Jun Rao, Co-Founder of Confluent and one of the original creators of Kafka, to trace how streaming data evolved from solving LinkedIn’s internal data silos to becoming a foundation for modern enterprise AI. The conversation focuses on why real-time data matters more than ever, how open source shaped Kafka’s adoption, and where agentic AI depends on fresh contextual data.
Chapters:
00:00 — Welcome back, the security callback, and IBM's Lightwell project
05:29 — Why quantum deserves its own episode
08:08 — Introducing Jun Rao: from Kafka to IBM and enterprise scale
11:47 — LinkedIn's data silos and the messaging systems that couldn't scale
19:09 — Choosing Scala and unifying LinkedIn's pipeline into real-time streams
22:19 — The commit log, Kafka Streams, and Flink
26:07 — What changed after Kafka moved to Apache
29:24 — AI-assisted contributions and open source governance
35:36 — From batch training to runtime context: how generative AI changes streaming
41:36 — Trust Bank: real-time AI escalation and the verbosity problem
47:06 — The enterprise data problem and unifying data at rest and in motion
51:14 — Hybrid, multi-cloud, and what's next for real-time agents
Sound Bites:
“The real data is in the log.”
“Human beings are still the end owner of the code.”

