The Data Problem No One Solved | Peter Staar (Principal Research Scientist)
Listen to full episode:
Summary: This episode centers on two connected shifts: how AI is changing the economics of web content, and how unstructured data is being handled inside modern AI systems. Anant and Ed are joined by Peter Staar (Principal Research Scientist at IBM), one of the creators behind Docling, for a practical conversation about document conversion, structured extraction, and why open source matters more than ever. They discuss the tension between summarization and publisher traffic, the technical challenge of making document pipelines stable across fast-moving models, and Peter’s view that community is becoming the new moat in software.
Chapters:
00:00 - New year opening, AI jokes, and the episode topic: unstructured data and agents
04:41 - How AI summaries are disrupting web traffic economics
08:17 - Robots.txt, licensing, fair use, and the business response
13:00 - Introducing Peter Staar and the Docling project
15:13 - Peter's path from physics and HPC to knowledge graphs
21:08 - Docling's document model: stable abstractions for developers
22:32 - Specialist models vs. VLM pipelines, and the real parsing burden
25:53 - Docling's two pillars: conversion and structured extraction
29:40 - Open source as moat: community, benchmarking, and real-world evals
39:06 - Docling as "the pandas for documents"
44:41 - Why framework matters more than model: integrations and enterprise advice
48:40 - Naming Docling, reflections, and the future of software work
Sound Bites:
“Delivering code is not the same as running the code successfully.”
“The real problem was really: how do we get content or context into the LLMs?”
“If a kid in the basement can do this alone, you might have to rethink what you’re doing.”

