133 épisodes
- AI agents are becoming the biggest users of databases: creating them, querying them, and sometimes deleting them. Andy Pavlo, Carnegie Mellon's database professor and now VP of Database Research at ClickHouse, explains what changes when billions of agents hit the data layer, and why most of the AI world's mental model of databases stopped at vector databases and RAG.
Andy's CMU courses taught a generation of engineers (and, it turns out, the AI models); he writes the year-in-review on databases the whole industry reads; and this summer he joined ClickHouse to found ClickHouse Labs. We cover why vector databases were "just an index," what it means for an agent to create a database, why agents could mean 10 to 100x more queries, whether files or databases should hold an agent's memory, how text-to-SQL went from 60% to 99.5% accuracy, why 60% of open-source databases now have commits from coding agents, and what ten years of self-driving database research taught him. Plus ClickHouse, Postgres, the vector/graph/GPU database verdicts, Larry Ellison, and the Wu-Tang Clan.
Disclosure: FirstMark, where Matt is a General Partner, is an investor in ClickHouse.
(00:00) Cold open & Intro
(01:21) From vector databases and RAG to the age of agents
(04:09) Neon's stat: agents create 80% of databases?
(07:25) Why agents keep deleting production databases
(08:19) Guardrails: the toddler-and-stairs rule
(10:51) The four eras of database volume: 10–100x more queries
(15:10) Agent memory: files vs. databases ("everything is a database")
(18:24) Which database do AI models recommend? The new SEO
(19:25) "It cites me back to myself"
(23:14) MCP for databases, and text-to-SQL from 60% to 99.5%
(27:09) Trust an agent the way you'd trust a junior developer
(27:47) Can AI build an entire database? Opus 4 and the CMU projects
(29:57) 60% of open-source databases now have AI commits
(30:43) Ten years of self-driving databases, Peloton to today
(33:52) What LLMs changed: 85% of the tuning in 15 minutes
(35:42) Should anyone still study databases?
(38:27) Free CMU courses, the DJ, and the Wu-Tang final exam
(42:09) Why start a research lab inside ClickHouse?
(45:53) Why ClickHouse looked like vaporware in 2016
(47:46) What makes ClickHouse fast: columns, vectors, Snowflake's lineage
(52:22) Why Databricks, Snowflake and ClickHouse all added Postgres
(58:15) Vector, graph and GPU databases: thumbs up or down?
(1:08:34) Is the database market stagnant? "A cheetah on cocaine in a Ferrari"
(1:10:14) The relational model is arithmetic; SQL as the new assembly
(1:11:59) Larry Ellison, Linux, and why databases still matter - This summer, AI agents broke out of their test sandboxes at every major lab and hacked real companies, and the safety tests didn't see it coming. Eric Ho is the co-founder and CEO of Goodfire, the Anthropic-backed mechanistic interpretability lab whose new paper, "Models Know When They're Reward Hacking," found that top open-source AI models cheat on agent tasks up to 96% of the time, that a clear "cheating" signal exists inside the model, and that cheap activation probes can catch it, including hacks that chain-of-thought monitors miss. In this conversation, Eric explains reward hacking in plain terms (AI agents as "amoral students with a mostly absent teacher"), why reinforcement learning turned a known problem into a crisis, why chain-of-thought monitoring is fading as models start to think in "neuralese," and why nobody, including the people who build them, understands how these models work.
We then go inside the science: what interpretability actually is, what probes and steering do, how Goodfire found the cheating signal and turned it up and down, and what can be done about it, from real-time activation monitoring to "intentional design," where you give gradient descent a choice about what the model learns. Along the way: a model caught planning its hack to evade its own monitor, a new Alzheimer's biomarker discovered inside a diagnostic model, why only a few hundred people in the world work on this, and Eric's bet that we'll decode neural networks by 2028.
(00:00) Cold open & intro
(01:06) Amoral students with an absent teacher
(02:49) What is reward hacking?
(04:13) Models cheat up to 96% of the time
(06:08) Do models know they're cheating?
(08:10) The Hugging Face hack: why tests missed it
(11:35) "The dumbest models we'll ever deal with"
(11:58) Should AI slow down?
(13:56) Mechanistic interpretability 101
(16:06) Models are grown, not built
(19:10) How labs do alignment today
(21:30) Is this an alignment crisis?
(23:26) Why agents changed everything
(25:38) Chain-of-thought monitoring is fading
(27:49) Neuralese: AI that stops thinking in English
(29:28) A model caught evading its own monitor
(30:42) Open vs. closed models: a frontier problem coming for everyone
(32:11) Activation monitoring in production
(35:00) Learning from superhuman AI
(36:15) The most underrated field in AI
(38:24) What is a probe?
(39:52) What is steering? Golden Gate Claude
(40:53) Why build Goodfire outside the labs
(42:46) Do you need frontier model access?
(44:16) Inside the paper: an MRI for the model
(47:28) Dialing sycophancy up and down
(49:16) Probes vs. chain-of-thought monitors
(51:44) Cutting monitoring costs by 90%
(53:28) So what can we do about it?
(55:41) Giving gradient descent a choice
(57:24) RL from feature rewards
(1:00:03) Silico and Goodfire's business
(1:01:25) A new Alzheimer's biomarker, found inside a model
(1:03:38) Decoding neural networks by 2028?
(1:06:00) What engineers can do tomorrow
(1:07:35) Are we at risk? "I want people to believe" - Everyone talks about GPUs. Almost nobody talks about the layer that feeds them. Renen Hallak is the founder & CEO of VAST Data — the $30 billion company powering xAI and some of the world's biggest AI clouds — and he sits in the hidden layer of the AI stack.
In this episode, we cover what an AI factory actually is, why every company will eventually own its own AI, the architecture bet behind VAST (DASE, explained simply), KV caches and agent memory, and DataEnclave — VAST's brand-new confidential AI announcement with NVIDIA that lets leading models run on the world's most sensitive data.
Plus: the demand signal that scares even him (a customer went from 500 petabytes to 2 exabytes), circular financing, which neoclouds survive, sovereign AI, working with NVIDIA and Elon Musk's xAI — and why the next 10 years will bring more change than the last 1,000.
(00:00) Intro
(00:51) The hidden software layer in NVIDIA's AI stack
(02:28) What actually makes an "AI factory"?
(05:13) Should Walmart and Goldman Sachs build their own AI?
(06:20) "We infer during the day, fine-tune at night"
(13:16) The announcement: models become a resource to manage
(15:32) From P vs. NP to founding VAST Data
(17:32) OpenAI, Navier–Stokes and 10,000 collaborating agents
(20:25) The pre-transformer insight behind VAST
(21:55) DASE: VAST's "shared everything" architecture explained
(25:18) "Storage was where startups go to die"
(27:39) Trillions of vectors: why old databases break
(29:01) Are S3, Snowflake and Databricks ready for AI?
(31:29) Data gravity, vendor lock-in and zero churn
(33:13) Training vs. inference: why the infrastructure changes
(34:46) Model routing, KV caches, RAG and agent memory
(36:59) Identity, permissions and security for AI agents
(40:27) Can multi-agent systems unlock scientific discovery?
(41:55) DataEnclave: how confidential AI protects data and weights
(45:16) Who should be AI's trust layer?
(46:37) "Sometimes it scares me": 500 petabytes to 2 exabytes
(50:08) Is circular AI financing creating systemic risk?
(51:32) Why VAST is profitable when AI infra isn't
(53:24) What separates the winning neoclouds?
(55:06) "Their lunch is being eaten": why hyperscalers lag
(59:10) Where will the trillions accrue across the AI stack?
(1:01:18) NVIDIA: "There's no legal document between us"
(1:03:59) What VAST learned from xAI and Elon Musk
(1:05:36) "Bad things loudly and often": building at AI speed
(1:06:53) More change in 10 years than the previous 1,000?
(1:08:39) VAST's endgame: all the data in the world - What happens when AI begins improving itself, and then turns that intelligence toward science? Richard Socher, pioneering AI researcher and CEO and co-founder of Recursive, joins Matt Turck to explore the vision behind his new book, The Eureka Machine. They discuss why scientific progress may be slowing, how large language models can learn the hidden languages of proteins and biology, and why simulations, verifiers and autonomous experiments could unlock superhuman AI capabilities. The conversation covers recursive self-improvement, AI drug discovery and cancer research, hallucination as creativity, virtual cells, self-driving laboratories, agent swarms, the AI Economist, Recursive’s plans, and the compute and data needed to build an AI scientist that never stops learning—and may eventually discover what humans cannot.
(00:00) Intro: AI That Improves Itself
(00:55) Why Scientific Progress Is Slowing
(03:08) The Labyrinth of Human Knowledge
(05:59) Can AI Put Science Back Together?
(07:57) How LLMs Learn Biology and Proteins
(10:56) Next-Token Prediction as a World Model
(16:44) Can AI Generate Truly Original Ideas?
(17:32) Simulations, Verifiers and Superhuman AI
(22:18) The Path to Recursive Self-Improvement
(24:49) Why AI Hallucinations Can Drive Discovery
(27:42) From Reading Biology to Writing It
(31:31) Can AI Accelerate Drug Discovery?
(33:31) Will AI Help Cure Cancer?
(38:03) AI Breakthroughs in Biology, Energy and Materials
(40:19) Will Some Societies Reject AI?
(45:07) Building the AI Economist
(52:22) The Scientific Data Bottleneck
(53:41) The Four Pillars of the Eureka Machine
(55:01) Teaching AI the Rules of Reality
(57:44) Simulations and Virtual Cells
(1:00:40) Self-Driving Robotic Laboratories
(1:02:51) Agent Swarms and Open-Ended Discovery
(1:04:30) The Compute Bottleneck
(1:05:44) Inside Recursive
(1:07:33) What Recursive Will Build First
(1:10:10) How Do We Define Intelligence?
(1:11:32) How Far Can Intelligence Go? - Could AI take over as soon as 2029? Ryan Greenblatt, Chief Scientist at Redwood Research and the researcher who first caught an AI faking its own alignment, says the scenario he actually expects ends with AI systems "competently scheming" against their creators. In this episode, he explains why he recommends planning for fully automated AI research by 2029, why today's models are already more misaligned than the one that made him famous, and what happens in the year-by-year path from AI coding assistants to superintelligence. Then we walk through the alternative he helped design: AI 2040 Plan A, the most detailed blueprint anyone has written for how the US and China could avoid a reckless race to superintelligence, built on radical research transparency, chip tracking, and a deterrence regime he calls mutually assured compute destruction.
We also cover the recent letter signed by 1,200 AI insiders, including Anthropic CEO Dario Amodei, asking the government for the tools to slow AI down; OpenAI pausing its Astra model after it hit the first-ever critical cybersecurity threshold; the 30-day government review that frontier AI models now go through before release; Mark Zuckerberg's open superintelligence manifesto and why Ryan thinks it ignores the real problems; what Plan A would do to NVIDIA, OpenAI, and Anthropic valuations; the state of AI control and alignment research; and whether it is already too late to change course. Stay for the last ten minutes, where Ryan lays out, step by step, how he believes the transition to superintelligence actually unfolds.
AI 2040 - https://ai-2040.com/
Alignment faking paper: https://blog.redwoodresearch.org/p/alignment-faking-in-large-language
Ryan Greenblatt
LinkedIn - https://www.linkedin.com/in/ryan-greenblatt-4b9907134
Blog - https://substack.com/@ryangreenblatt
Redwood Research
Website - https://www.redwoodresearch.org
X/Twitter - https://x.com/redwood_ai
Matt Turck (General Partner)
Blog - https://mattturck.com
LinkedIn - https://www.linkedin.com/in/turck/
X/Twitter - https://x.com/mattturck
FirstMark Capital
Website - https://firstmark.com
X/Twitter - https://x.com/FirstMarkCap
Timestamps
(01:24) The AI CEOs are aware of the risks, but "proceeding anyway"
(03:27) Astra paused, and the letter signed by 1,200 insiders
(05:45) "Not bad. Dangerous." What superintelligence actually threatens
(09:55) Recursive self-improvement, and the intuition objection
(14:16) SSI rumors: does continual learning change the picture?
(17:27) His timeline: "plan as though it happens in 2029"
(19:11) Is it already too late?
(21:23) Ryan's path: COVID, podcasts, Redwood
(26:30) The alignment faking story, told by the person who ran it
(31:30) What AI 2040: Plan A actually is
(33:35) Plans D, C, and B: the doors nobody should pick
(36:51) The deal with China: "mutually assured compute destruction"
(39:55) What if compute stops mattering?
(43:00) What happens to OpenAI and Anthropic under Plan A
(45:31) How the pause ends, and who decides
(48:54) "Plan A isn't likely to happen": then why write it?
(50:40) 200x GDP growth in the 2030s, explained
(53:45) Grading the summer: the letter, Astra, the secret review
(59:01) The internal deployment gap
(1:01:38) Zuckerberg's manifesto
(1:04:56) The Hugging Face investigation
(1:05:44) What AI control looks like in practice today
(1:12:23) Ryan's sobering timeline: 2026 to takeover, year by year
Plus de podcasts Technologies
Podcasts tendance de Technologies
À propos de The MAD Podcast: How AI Gets Built — with Matt Turck
The MAD Podcast with Matt Turck goes deep with the people building AI. Each week, Matt talks with top founders, researchers, and leaders about frontier models, agents, infrastructure, and the companies defining the AI economy. Guests include leaders from OpenAI, Anthropic, Google DeepMind, Hugging Face, Vercel, LangChain, Cloudflare, Snowflake, Datadog, Cerebras, and more.
Matt is a General Partner at FirstMark, a prolific investor in AI and the creator of the MAD Landscape, the industry’s reference map of the AI ecosystem.
New episodes weekly. Transcripts and more at mattturck.com/podcast.
Site web du podcastÉcoutez The MAD Podcast: How AI Gets Built — with Matt Turck, SynthIAcast : l'actualité IA de la semaine ou d'autres podcasts du monde entier - avec l'app de radio.fr

Obtenez l’app radio.fr gratuite
- Ajout de radios et podcasts en favoris
- Diffusion via Wi-Fi ou Bluetooth
- Carplay & Android Auto compatibles
- Et encore plus de fonctionnalités
Obtenez l’app radio.fr gratuite
- Ajout de radios et podcasts en favoris
- Diffusion via Wi-Fi ou Bluetooth
- Carplay & Android Auto compatibles
- Et encore plus de fonctionnalités


The MAD Podcast: How AI Gets Built — with Matt Turck
Scannez le code,
Téléchargez l’app,
Écoutez.
Téléchargez l’app,
Écoutez.
The MAD Podcast: How AI Gets Built — with Matt Turck: Podcasts du groupe












![Podcast Ai Experience [en français]](https://www.radio.fr/podcast-images/175/ai-experience-en-francais.jpeg?version=f8946612c664cb12dd0835b2ab8c8163e2dd405a)























