Your AI Is Waiting on Your Data | Jun Rao (Co-Founder, Confluent)

 

Listen to full episode:

Summary: Hosts Anant and Ed sit down with Jun Rao, Co-Founder of Confluent and one of the original creators of Kafka, to trace how streaming data evolved from solving LinkedIn’s internal data silos to becoming a foundation for modern enterprise AI. The conversation focuses on why real-time data matters more than ever, how open source shaped Kafka’s adoption, and where agentic AI depends on fresh contextual data.

Chapters:

00:00 — Welcome back, the security callback, and IBM's Lightwell project

05:29 — Why quantum deserves its own episode

08:08 — Introducing Jun Rao: from Kafka to IBM and enterprise scale

11:47 — LinkedIn's data silos and the messaging systems that couldn't scale

19:09 — Choosing Scala and unifying LinkedIn's pipeline into real-time streams

22:19 — The commit log, Kafka Streams, and Flink

26:07 — What changed after Kafka moved to Apache

29:24 — AI-assisted contributions and open source governance

35:36 — From batch training to runtime context: how generative AI changes streaming

41:36 — Trust Bank: real-time AI escalation and the verbosity problem

47:06 — The enterprise data problem and unifying data at rest and in motion

51:14 — Hybrid, multi-cloud, and what's next for real-time agents

Sound Bites:

“The real data is in the log.”

“Human beings are still the end owner of the code.”

Previous
Previous

The Bigger Model Myth | Vanja Josifovski (CEO, Kumo AI)

Next
Next

Agentic Coding is Eating Itself | Jonathan Ellis (CEO, Brokk)