The Data Problem No One Solved | Peter Staar (Principal Research Scientist)

 

Listen to full episode:

Summary: This episode centers on two connected shifts: how AI is changing the economics of web content, and how unstructured data is being handled inside modern AI systems. Anant and Ed are joined by Peter Staar (Principal Research Scientist at IBM), one of the creators behind Docling, for a practical conversation about document conversion, structured extraction, and why open source matters more than ever. They discuss the tension between summarization and publisher traffic, the technical challenge of making document pipelines stable across fast-moving models, and Peter’s view that community is becoming the new moat in software.

Chapters:

00:00 - New year opening, AI jokes, and the episode topic: unstructured data and agents

04:41 - How AI summaries are disrupting web traffic economics

08:17 - Robots.txt, licensing, fair use, and the business response

13:00 - Introducing Peter Staar and the Docling project

15:13 - Peter's path from physics and HPC to knowledge graphs

21:08 - Docling's document model: stable abstractions for developers

22:32 - Specialist models vs. VLM pipelines, and the real parsing burden

25:53 - Docling's two pillars: conversion and structured extraction

29:40 - Open source as moat: community, benchmarking, and real-world evals

39:06 - Docling as "the pandas for documents"

44:41 - Why framework matters more than model: integrations and enterprise advice

48:40 - Naming Docling, reflections, and the future of software work

Sound Bites:

“Delivering code is not the same as running the code successfully.”

“The real problem was really: how do we get content or context into the LLMs?”

“If a kid in the basement can do this alone, you might have to rethink what you’re doing.”

Previous
Previous

Inside Moltbook | Kate Blair (Director, IBM Research)

Next
Next

What’s Next for CPUs, GPUs, & Quantum | Alessandro Curioni & Sarah Sheldon